The AI Daily Brief: Artificial Intelligence News and Analysis - Why AI Users Are Raving About GLM 5.2
Episode Date: June 22, 2026GLM 5.2 is looking like the first open-weight model in a while that might survive contact with real-world usage, especially for coding and web design. NLW looks at why builders are comparing it to the... DeepSeek R1 moment, where the hype is justified, where the cost story is more complicated, and what it means for enterprise AI stacks that can no longer assume a simple OpenAI-versus-Anthropic race.Register for our new enterprise-grade AI training programs: http://training.besuper.ai/Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at kpmg.com/us/SophisticatedSection - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/Outsystems - Stop wondering how AI will change your business and start building the agents that will lead it - http://outsystems.com/Scrunch - The AI customer experience platform - https://scrunch.com/Zenflow Work - Agents for knowledge work - https://zenflow.free/Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/MissionCloud - Eliminate AWS complexity with end-to-end cloud and AI services https://www.missioncloud.com/AssemblyAI - The best way to build Voice AI apps - https://www.assemblyai.com/briefRobots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Our Newsletter is BACK: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? sponsors@aidailybrief.ai
Transcript
Discussion (0)
Today on the AI Daily Brief, why AI power users are raving about GLM 5.2.
Before that in the headlines, Trump talks anthropic and Fable 5 return rumors swirl.
The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, quick announcements before we dive in.
First of all, thank you to today's sponsors, KPMG, Scrunch, Mission Cloud, and Out Systems.
To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts.
And if you want to learn more about sponsoring the show, send us a note at sponsors at
AIDDailybrief.aI.
For those of you who are looking for deeper training programs, we announced last week that we
have upgraded Enterprise Claw and the Executive Catchup Program to be more enterprise grade
in collaboration with Super Intelligent. You can learn all about that at training.bysuper.
And specifically, the new Executive Agent Leadership Program, formerly known as Enterprise Claw,
is registering its next cohort, which will begin next week.
so if you're interested in that, again, go check it out at training.btsuper.a.i.
Last note today, we're in this kind of weird period where there's so much headline news
that I don't just want to be doing the Fable 5 update story every day for the main episode,
but the consequence of that is that the normally 5-minute headlines is extending to more like
10 or even 12 or 13 minutes. That won't be the case forever, but for now we got a little bit
of a weird balance. And so with that, let's dive into the slightly extended headlines.
The theme of this headlines episode is separating out,
fact from innuendo in the attempt to understand where things actually are in this very confusing
moment with AI. We're going to start with some comments that seem to some to shed light on the whole
Fable 5 mytho situation, and by the end of the headlines, see where it leaves us relative to whether
we might be getting Fable 5 back this week. Now, over the weekend, many folks thought that they
figured out some new, old information that seemed to make the Fable Ban make a little bit more sense.
Specifically, they dug up reporting from the economist from June 14th, in which the economist wrote,
On June 11th, Mark Warner, the vice chair of the Senate Intelligence Committee, said that General
Joshua Rudd, who leads the National Security Agency and the Pentagon Cyber Command,
had told him that mythos, quote, broke into almost all of our classified systems, not in weeks,
but in hours.
Now, June 11th was the same Thursday that Amazon CEO, Andy Jassy, informed the administration of
the jailbreak that became the center of the story.
Once the quote resurfaced, ex-commentators were quick to jump on it, commented chubby, summing up the feelings of many,
wow, that changes the whole Fable Five story completely.
University professor Pedroes Domingos, who is typically not a fan of the current administration,
commented mythos broke into almost all of the NSA's classified systems in hours, per its director.
It would have been irresponsible to not impose export controls on it, and on Fable with its pathetically inadequate guardrails.
Now, on the one hand, part of why this is resonant is that it happens.
has a feel of truthiness to it, in that it would make way more sense if the White House was
already keyed up about mythos slash fable being too powerful from some other evidence that they'd
seen with this weird jailbreak report, just providing enough pretext for them to do what they
had wanted to do in the first place, which is to disallow the model at this time. And yet for those
who are reading this line from Mark Warner, as some literal breach of the NSA which demanded some
response, the reporter behind the story Shashank Joshi added some additional context. While he
he said that the quote attributed to Mark Warner was accurate, he added, it would be a mistake
to read the quote literally, I think. It surely depends on using mythos alongside other tools
under very particular conditions. I quoted it to give a sense of mythos potency, but it was a mistake
not to have added caveats. In other words, this was not the director of the NSA reporting some
terrifying breach. It was them reporting how powerful they had found mythos in their specific
controlled tests.
AI policy commentator Peter Wildeford gave an example of what he thinks is a more plausible
scenario for what happened.
He wrote,
Senator Warner claimed that he was told by the head of the NSA and cyber command that
Mythos was breaking into classified systems in hours.
This is an important claim to understand better.
I thought Mythos was very good at cybersecurity, but break into classified systems in hours
good?
NSA classified networks are physically disconnected from the internet entirely, with specialized
hardware controlling what data can even cross between them. More plausible readings of what actually
happened then. One, this was a simulated exercise against replica systems, not the real NSA network.
Two, Mythos was given the relevant code and architecture docks up front rather than breaking
and blind. Three, it tore through poorly secured internal IT that got described as classified systems.
Four, Mythos was operated with significant additional tooling and human expertise. Of course, Peter
concludes, none of this means that Mythos underlying cyber capability isn't alarming, an AI that
compresses weeks of expert security research into ours is a genuine threat to systems that are
connected to networks as we've seen. Peter's point then was not to say that his scenarios were exactly
what happened, but that they were all in his estimation more plausible explanations than what
people assumed was some massive mythos breach of the NSA. The CyberSect guru added even more
context in a blog post breaking down the story. They noted additional reporting that confirmed
this breach occurred during a red team exercise run by the NSA, i.e. this was not some outside
attack or breach. It was in a specific controlled environment where they were trying to run
adversarial tests. Now, CyberSeg Guru also cast at least a little bit of skepticism on the
source, pointing out that NSA Director Rudd was appointed in a heavily contested vote this
March, with those who opposed his confirmation citing Rudd's background as a special operations
officer with no relevant experience in signals intelligence or cyber warfare.
Right, CyberSecur, that doesn't make his claim false, but it's relevant context for a statement
about a cyber incident made by the agency's own director.
He's a relatively new appointee in a technical domain that wasn't his original specialty,
testifying about his own agency's capabilities.
And yet, even with all of this, I think it is fair to say that the fable ban wasn't solely
or cleanly about the Amazon jailbreak and clearly wasn't just about personality differences
between Anthropic and the White House.
Interestingly, in an interview with the Axio show on Saturday, President Trump spoke
about the issue at length.
He said, we have a situation with Anthropic.
We didn't like what they're doing.
So far, I think they've responded very responsive.
to our request. When asked if he regards Anthropic and Dario Amade personally as a national security threat,
Trump responded, not now, but a week ago, maybe. Referring to the G7 summit, Trump added,
I was with him yesterday. He made a speech. I made a little speech. Seems like a nice guy, smart guy.
He responded to us very quickly because, you know, it's tremendous liability. People get put in
prison immediately for that. You can't play games with that. He responded very responsibly, I thought,
so far. Asked about the possibility of shutting down Anthropic, Trump commented, I don't want
to do that. You know, we're beating China. I was with President Xiwi talked about it. We're beating
China by a lot. Trump also explicitly ruled out the Defense Production Act to control AI, stating,
I don't think we have to do that. So far it's been very responsible. Summing up, Trump commented,
I think the good far outweighs the bad. We are going to find the bad and we're going to stop it.
I think the biggest takeaway is that right now, everything is heightened. People are completely
geared up. Everyone is looking for any tea leave that they can read to understand when Fable might be coming
back, what the new relationship between the White House and AI companies is going to be,
all of which is to say it's a good time to be extremely careful about the sourcing of reporting
and to try to separate what we know from what we think.
Now, one other story from the weekend that was real and does also seem to have some big
implications for the AI race was another high-profile departure at DeepMind as Nobel laureate
John Jumper left for Anthropic.
Jumper announced the move in an ex post on Friday, thanking Demis Sassabas for taking a chance on him
nine years ago and hiring him to lead the AlphaFold team shortly after he completed his PhD.
That work, of course, resulted in an AI model that predicts the 3D structure of proteins
based on their amino acid sequence, massively accelerating the field of biochemistry and drug
discovery.
For that work, Jumper shared the Nobel Prize in chemistry with Hizabas in 2024.
Honoring his colleague, Hesabas thanked Jumper for his collaboration, commenting,
what we achieved with AlphaFold changed the world and showing the field what was possible
with AI for science and medicine, lighting the way for how AI can benefit.
humanity. Now, from the outside, lots of folks were left to wonder what the heck is going on at
DeepMind to trigger an exodus of elite talent. Lassan on X wrote, Google is in freefall. This is the
second VP of engineering that left Google DeepMind this week. First, Nome Shazir, transformer and
mixture of experts pioneer. Today, Nobel laureate John Jumper, who basically built AlphaFold 1 through 3,
and most recently also worked on AI coding at DeepMind. Now, speaking of that, some suspected
that being assigned to lead AI coding efforts
rather than continue his work on AI for science
may have contributed to Jumper's exit.
But still, to have two very, very high-profile leaders
of DeepMind head one to Anthropic
and one to Open AI in a single week
doesn't look great from the outside.
A few minutes after Jumper made his announcement,
Leo at Sintwaved added some background
about plummeting morale in DeepMind.
They wrote,
After the release of Fable 5 and with GPD 5.6 looming,
the mood behind the scenes at Google DeepMind
is increasingly one of frustration
and broad discontent over the labs perceived fall into a distant third or even fourth place.
A well-connected DeepMind employee told me,
I can't blame Nome for walking. He won't be the last big name to go either.
Leo added that staff were demoralized by ZAI's GLM 5.2 overtaking Gemini 3.1 Pro
on the artificial analysis intelligence index.
In addition, the release of Gemini 3.5 Flash in Gemini Omni earlier this year was received with
little fanfare, and DeepMind has now gone four months without a flagship model release.
Another source at DeepMind told Leo that Gemini 3.5,
Pro is, quote, not the step change we need to be truly competitive in the race to AGI.
That model is reportedly slated to be released next Tuesday, June 30th.
Leo added,
The consensus seems to be that leadership at Google has all but conceded the race to Anthropic
and Open AI, and that only a big shakeup will propel them back to the heights of mid to late
2025.
Another deep mind source commented,
we no longer have a frontier model in text, image, video, voice, or even vision.
If we can't release a real frontier model after over four months of work with all these
resources, what are we doing?
Now, Googler Logan Kilpatrick did offer some pushback, responding,
everyone I know is hopeful and locked in,
lots of things in the pipeline that will hopefully pay off short and long term.
And once again, I will caution this is all behind-the-scenes sources and reporting,
meaning you have to take it with at least a little bit of a grain of salt.
I think in general that we tend to make too much of any individual career move.
For example, there was another story this weekend that Barrett Zulf was out at OpenAI
just five months after rejoining.
And this is a guy that has now absolutely ping-ponged between OpenAI and thinking machines labs and then back to OpenAI.
And while, of course, any high-profile departure could be an indication of something going on in a lab,
humans are complex creatures with lots and lots of reasons and motivations behind their decisions that we on the outside aren't going to be privy to.
What is true, however, and what is worth noting about the Google story, is first, that two very high-profile leaders does start to make a pattern,
and that, too, the drop-off in where Google fits, relative to at least the coding and enterprise
side of the AI race, is in 2026 absolutely notable.
Now, we haven't seen 3.5 Pro yet, and Google has many strengths outside just where they sit
at the state of the art.
But you do have to think that the stakes for Google DeepMind with every new model release
have raised significantly.
Now, really, as we head into this week, the biggest rumors that people care about is when
we're going to be getting Fable 5 back. And as much hay as the press made about Trump saying a week
ago that Dario Ananthropic were national security threats, others actually saw the interview as the first
step to a resolution. Dan McCatrier wrote, if you listen to Trump, he's quite conciliatory. He doesn't
want to kill the goose that lays the golden egg. Trump knows AI is the foundation of America's future.
Claude Fable back next week. Bet on it. Now beyond that, we did get even more substantial rumors about
what comes next. Andrew Curran, who's one of the best follows for actual
AI news and tends to have good sources when he reports something that hasn't been reported
yet, wrote, a new, more capable version of Mythos has emerged from training. I don't know whether
it will be called Mythos 5.1 or Myth06, or if Anthropic will keep it internal to accelerate
further development, but it has arrived. Then Andrew points out something important that we haven't
discussed enough. He continues, stopping models like Fable 5 or Mythos 5 from being served to the public
does nothing to slow down development. In fact, it probably speeds it up slightly by freeing up
resources. There are also no rules preventing the labs from continuing to advance capabilities
while any current model is under embargo, or from keeping progress quiet until they choose to release
it. None of them can afford to pause or slow down. We need only look at how capable GLM 5.2 is
as proof of this. To protect their business models, the frontier labs must continually train
increasingly capable systems to stay ahead of open source and each other. The current continues to
rage beneath the ice, and we continue to race towards our destination. Now, in addition to a potential
mythos 5.1 or 6 emerging in the labs, some found evidence that Sonnet 5 might be nearing release.
Leo at synthwaived again wrote, the slug Claude Sonnet 5 has appeared on an anthropic partner
provider. Gonna be a busy week. Chubby responded, so we get Claude Sonnet 5 instead of Fable 5 soon.
Looks like a busy week, probably GPD 5.6 and Sonnet 5. But hey, keep them coming.
Leo responded, I suspect it'll actually be Fable 5 plus Sonnet 5 plus 5, but let's see.
Now, regarding GPD 5.6, some are reporting that there are.
already seeing the model show up in Codex, implying that we're getting pretty close.
A French ex-user called Mirichil posted a playable demo of a Pokemon game supposedly
one-shotted by GPT5.6. Meanwhile, within OpenAI, Codex lead Tebow has begun the vague posting.
He wrote, we built the Codex app with models that were okayish at front-end. Wait to see
what we can do when we finally improve front-end capabilities significantly in our models.
That day will be something. Scientist Darya Anutmas, who typically gets early access to models,
joined in the vague posting writing, people were flabbergasted by Fable 5, rightly so.
But those who think this will remain the best AI for a long time will soon be proven wrong.
When some thought he was just stating the obvious, Anoumez urged them to read between the lines,
adding, read the words long time versus soon. I didn't say eventually.
Now, I think, Andrew Curran's visual metaphor of the current raging under the ice is a good one.
And what's important to note with all these rumors is that even if we are in line for a big week right now,
There is so much that could happen that could change that path.
Still, if you want to let yourself get excited about anything,
my fellow builders out there I have no doubt will be very excited to see
that the way that the OpenAI team seems to be teasing the next models
is them being better at front end.
We should be so lucky.
For now that that's going to do it for this extended headlines.
Next up, the main episode.
One of the most important AI questions right now isn't who's using AI.
It's who's using it well.
KPMG in the University of Texas at Austin just analyzed 1.4 million real workplace AI interactions
and found something surprising. The highest impact users aren't better prompt engineers. They treat
AI like a reasoning partner. They frame problems, guide thinking, iterate, and push for better answers.
And the good news? These behaviors are teachable at scale. If you're trying to move from AI access to real
capability, KPMG's research on sophisticated AI collaboration is worth your time. Learn more at
KPMG.com slash us-s-sophisticated. That's KPMG.com slash us-sophisticated.
Quick question. When was the last time you actually visited a website to research something?
If you're like me, AI pretty much does that work for you now. That, of course, raises a new question
for brands. If AI is doing the discovering, researching, and deciding, who or what is your website
really for? That shift in user behavior, the rise of AI bots becoming your most important new
visitors, is what my sponsor scrunch is taking head on.
Scrunch is the AI customer experience platform that helps marketing teams understand how AI agents experience their site,
where they show up in AI answers, where they don't, and what's preventing them from being retrieved, trusted, or recommended.
And it's not just visibility. Scrunch shows you the content gaps, citation gaps, and technical blockers that matter,
and helps you fix them so your brand is found and chosen in AI answers.
Now, for our listener, Scrunch is providing a free website audit that uncovers how AI sees your site,
where there's gaps, and how you're showing up in AI versus the competition.
run your site through it at scrunch.com slash AI Daily.
The average enterprise is spending $11.5 million on AI this year,
and most of them can't prove a single dollar came back.
What does AI actually look like when it produces ROI?
Ask the healthcare company that just made their payment processing
320 times faster, or the law firm whose document research went from three months to 10 minutes,
or the contact center who reduced wait times by 99%.
These are real Mission Cloud customers with real results.
Mission Cloud is a CDW company and an AWS premiere to your partner.
They're the AI-first outcomes obsessed to AWS experts who build AI solutions that drive your business forward.
Whether you're flooded with AI ambitions but no idea where to start, or six months into a deployment that's going sideways, they've seen it and they've fixed it.
Stop burning your budgets on AI that doesn't produce results.
Start at mission cloud.com.
This episode of the AI Daily Brief is brought to you by OutSystems, a leading agendic systems platform built for the enterprise.
organizations all over the world are building, orchestrating, and governing agentic system on the OutSystems platform and with good reason.
OutSystems open and unified platform allows teams to architect, deliver, and scale governed agentic systems with agility.
Teams of any size and technical depth can use OutSystems to build, deploy, and manage AI apps and agents quickly and cost effectively without compromising reliability and security.
Without systems, you can rapidly launch ideas from concept to completion.
It's the leading Agendic Systems platform that is unified, agile, and enterprise.
proven, allowing you to accelerate growth, reduce operational friction, and deliver real enterprise
impact with AI.
Out Systems.
Build your agentic future.
Welcome back to the AI Daily Brief.
Last week, in the wake of Fable 5 going offline, one of the major topics of conversation
on this show was the new models and new model approaches that were rushing in to fill the
gap, not only trying to win people's usage, but also having a side effect of making people
think differently about how to construct their AI stack.
Now, part of what has made this Fable 5 moments, so it was a lot of it.
resonant and important among businesses, is that already the changes in the cost paradigm,
based on the shift to agentic AI, and magnified by the broader compute shortage, were already
causing companies to look around and ask whether there would be different approaches than just
firing up the most-d-of-the-art model for every single AI use case. Now, in those conversations
last week, we mentioned the first impressions of GLM 5.2, and they were good, but we have now had a
weekend pass where people actually got their hands on the thing. And the statutes of the statutes.
of the model and people's belief about its implications has done nothing but grow. So today we're going
to talk a little bit about those second impressions of GLM 5.2 and explore whether it's something that you
should actively consider. Now, the analogy that everyone is plumbing for is the deep seek R1 moment.
Euchenjinjin writes, looking at my timeline, it feels like GLM 5.2 is having its deep seek R1 moment.
I never thought an open source model could break into the top three coding models this soon.
Gen Zhu writes,
GLM 5.2 feels like a turning point that's as significant as Deepseek R1.
The Fable Saga and GLM 5.2 release happening at the same time just changed the adoption calculations,
and things will only accelerate from here, with Deepseek's massive funding,
Kimi Minimax, Tencent, Quint, all lining up releases in the coming months.
Berkov writes, the GLM 5.2 moment is the new Deepseek R1 moment.
The last remaining mode is gone unless the U.S. labs pull something unseen before from the sleeve.
So what are we referring to when we talk about the deep seek moment?
Many of you might remember that there was this crazy thing that happened in January of 2025,
where all of a sudden there was this new model and this new application called Deepseek
that had raced at the top of the Apple App Store,
and that for many casual users felt distinctly better than whatever they were getting with ChatGPT.
Now, what had happened was that Deepseek, a Chinese lab spun out from or connected to a hedge fund,
believe it or not, had plopped a reasoning model inside a free app.
This was not the first reasoning model that was available.
OpenAIs 01 had been announced back in the previous September,
and it started to be rolled out more broadly in December of 24,
but it was behind a paywall,
whereas DeepSeek was putting its R1 reasoning model right there for free use.
Now, anyone who remembers the shift from non-reasoning models to reasoning models
will remember just what a huge difference it was,
and so all of a sudden all these people were having that experience in real time.
Pair that with some reports from Deep Seek themselves,
which ended up being a little misleading
about how little they had spent to train that model
and the market absolutely freaked out.
Nvidia had the single biggest daily loss
in terms of pure numbers
with DeepSeek peeling off $589 billion
from its market cap in a single day.
Now, of course, this ended up being the market getting way ahead of itself.
The Deep Seek phenomenon eventually receded,
but it did force American labs to think differently
about how fast they got reasoning models
into the free versions of their applications.
And yet, Teepzig has kind of a weird legacy.
Every time a new Chinese open weight model comes out, it scores super high on benchmarks,
everyone talks about how it's closed the gap with the Western Labs, and then a couple weeks
later no one's using it.
And honestly, a couple weeks later is being generous.
What usually happens is that these models don't really survive first contact with the real
world of usage and fade almost instantly.
Now, GLM 5.2 had some of these hallmarks, which is not to say that Chinese models have
been irrelevant.
In fact, over the course of the last year, they have been increased.
increasingly integrated into the stack, especially for startups and younger and smaller companies
that don't necessarily have as many big constraints in which models they can use.
And on top of that, as the overall state of the art increases, being just a few months behind
the state of the art still means lots of use cases that are viable.
In other words, what it means to be three or six months behind the state of the art now has a lot
more viable use cases than what it meant to be three or six months behind the state of
the art a year ago.
Now, coming back to GLM 5.2, at first blush, it did the same thing that these Chinese
open weight models always do.
impressive benchmarks, lots of excitement, but the vibes have been very clearly different.
Turns out it's not just random Twitter hypebees that are talking about GLM 5.2,
but some very respected figures in the industry.
Versal's CEO, Guillermo Rosh, writes, genuinely impressed, almost shocked at how good GLM 5.2 is at coding.
This changes things.
I'm Rolan writes,
GLM52 is not just another open model.
I played with it for a few hours, and for the first time, an open or public model felt
meaningfully close to frontier lab quality across real tasks. Not perfect, not fully benchmarked,
but very different. In another post, he wrote, this is not another AI slot model. Don't ignore it.
It feels like a chat GPT moment for public open models. However, Inamard does have a caveat about
cost, which will come back to in a moment. Now, even more than just random people talking about
GLM 5.2, I think a lot of folks started paying attention after Design Arena wrote a log post on X
about how GLM52 beat Fable 5 at website design. Now, this is a lot of people.
This is one of those benchmarks that one might have been tempted to be skeptical of when it was
first announced.
How could GLM 52 possibly be ahead of Fable 5 on design?
And how could it do so for a significantly lower price point?
Now importantly, they do caveat that GLM 52 doesn't surpass Fable 5 in everything.
It's behind Fable 5 on game development, data visualization, and 3D design, and it's down
all the way at fourth place on UI component.
But when it comes to websites themselves, that's where it ranks first.
The design arena pointed to three different model behaviors that they suggested made the difference.
The first they wrote is that the outputs seem to indicate a beautiful set of starting templates.
And while they point out that all of the models use a starting point of web design templates,
GLM 5.2s seem to avoid some of the most infamous anti-patterns like the purple gradients
that were all over early AI web design.
Now, what they found is that when you compare the outputs in GLM 52 to something like Fable 5,
they're much, much more concentrated.
That concentration means that on average,
they might be better, although they are going to be less diverse.
The second model behavior is that GLM 5.2 avoids common error cases.
Specifically, it seems to be really good at using certain dependencies
such as chart J.S and 3JS.
Design Arena writes,
while other models often fail to effectively use these libraries,
GLM 5.2 calls and uses them naturally.
It also uses tailwind CSS in 91% of sessions,
as compared to Opus 48, which only uses tailwind CSS,
in 57% of sessions.
The third model behavior is more intricate, detailed outputs.
Now, they do point out, though, that this complexity comes with a cost, that cost being
longer generation times as the model outputs more tokens.
In fact, GLM 5.2 websites produced 25% more characters and lines of code in their testing,
and had an average generation time that was about double Claude Fable 5.
Now, one thing that is worth pointing out that you start to see here with GLM 5.2 is that I think
that in general, people's assumption are that these Chinese open weight models,
are going to be much, much less expensive to run, and in this case, it's not as clear-cut.
YouTube and AI entrepreneur Theo writes,
I see a lot of people hyped about GLM 5.2, rightfully so.
Having an open-weight model surpassed GPD-54 in every Gemini model is dope.
That said, it's not cheap.
Both Opus 48 and GPT-55 set to medium are cheaper and smarter than GLM 5.2.
It also uses way more output tokens.
The tokens are cheaper, but the volume of them means you spend more time waiting for results.
still dope just trying to make sure people set their expectations properly.
Now, it's fascinating, especially for those of you who listened to the episode this weekend about local models with Newfar,
is that a lot of people are acting like the only way to use GLM 5.2 is running it locally yourself.
Going back to Intermar Golan, he writes,
The catch of GPD 5.2 is that running it properly is still expensive.
You probably need something like 8 Nvidia H200 GPUs, which means roughly 400K to buy or around 20K a month to rent.
Shutterstock founder John Oranger also did this sort of math.
talking about how many blackwells you need to run this under what settings. But of course, in
reality, most people are just going to use this model via one of the routing tools like OpenRouter,
or in one of the open source harnesses that they can maintain for themselves. OpenCode, for example,
tweeted, GLM 5.2 is a hit, been out for three days and it's already sixth on our leaderboard.
And I would strongly suggest for those of you who want to try this, that rather than trying to
hack at some very complex physical infrastructure, at least to start, you go try it via some service
like OpenRouter. Still, you can absolutely feel the Overton window shifting on how fast people
think we'll get a Fable Class model from China. Elon Musk actually got into a bit of a debate
with ex-user TOR taxes about this. After TOR suggested that we'd see a full Chinese mythos by
November or December of this year, Elon Musk argued that it would actually be Q1. When the founder
of ZAI responded that long, Elon Musk responded again. On benchmarks, yes, but as measured by true
usefulness, even Q1 would be very impressive. Anthropic is right.
focused on maximizing useful intelligence, which does not show up in benchmarks, but definitely
shows up in revenue. And yet, as Boxes Aaron Levy points out, the fact that open weight
models are being discussed credibly at this level of capability should be a huge update for many.
The implications of open models getting to frontier performance ensures that you can always
have sovereign AI, have the ability to post-train for your specific workflows, cost-optimized
for various workloads, and actually afford to do much more with AI, which opens up meaningfully
different applications. Huge win for the applied AI layer. So for those of you who are running
businesses and are trying to figure out what to do with this, my recommendation is not that
you race out and try to buy expensive hardware. Now, if you do want to start experimenting with
local AI, you can absolutely go check out the episode I did with Newfar. But I think that the bigger
thing here is that even assuming we get Fable back this week, I think the idea that AI was fully
down to a two-horse race between Open AI and Anthropic, with a little asterisk for Google if they can
get their mojo back has been broken over the course of the past six weeks or so. The combination of
workloads getting so much more intense, meaning so much more costly, plus the double-edged
sword of models getting powerful enough that they're subject to government review and restriction,
while also meaning that models just behind them are probably going to be viable for a lot of use
cases, means just this incredible potential flowering of new diverse model architectures and setups
inside companies that can optimize for different priorities, whether it's speed, cost, performance,
or something else. I don't think that on average, most companies need to be trying to race and
shift off of their core subscriptions, whatever they may be. But I do think that having some part
of the organization have some license in sandbox to experiment with some of these alternative
model architectures is probably time and money well spent right now. Certainly as we see more
evidence and case studies of how companies are putting this together, I will bring them to you.
For now, though, that's going to do it for today's AI Daily Brief.
Appreciate you listening or watching.
As always, until next time, peace.
