The AI Daily Brief: Artificial Intelligence News and Analysis - Why Fable 5.1 Is Worth the Upgrade

Episode Date: September 2, 2026

Fable 5.1 is the new state of the art—but its high token usage and restrictive limits mean the real question isn’t whether to switch, but where it belongs in your personal model stack. NLW examine...s its biggest capability gains, early user reactions, and how to decide when a frontier model is worth the cost. In the headlines: OpenAI’s Astra crosses a critical cybersecurity threshold, concerns grow around opaque model reasoning, Gemini 3.8 Flash targets coding, and World Labs unveils its Atlas world model.NEXT COHORT - Executive Agent Leadership - Returns in September -- Learn how to use agents - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://training.besuper.ai/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://kpmg.com/us/Sophisticated⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Harbor - Invest in the AI ecosystem. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.harborcapital.com/aidaily⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Hyperagent - Hire a team of always-on agents. New users get $100 in free credits. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠hyperagent.com/aidailybrief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.rackspace.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Section - Section turns AI investment into workforce transformation and ROI - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.sectionai.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠AssemblyAI - The best way to build Voice AI apps - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.assemblyai.com/brief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The AI Daily Brief helps you understand the most important news and discussions in AI. Newsletter: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Interested in sponsoring the show? sponsors@aidailybrief.ai

Transcript
Discussion (0)
Starting point is 00:00:00 Anthropic has released its latest models, Fable 5.1 and Mythos 5.1. On the benchmarks, they are undeniably state-of-the-art, outperforming everything else that exists on pretty much every category. Anthropic also claims that they've made major advances in the cost, so that for many tasks, including long-running agentic tasks, Fable 5.1 should cost as much as 25 or even 40% less than the comparative task in Fable 5. Initial responses are pretty good, Although users are getting pretty varied mileage in terms of just how much the costs actually are and how far you can even get with Fable 5.1 given usage limits. Still, the question comes up, as it will now forever with every new model,
Starting point is 00:00:40 is this one good enough that it's worth switching to? Except I think that that's no longer the right question. Instead, the question should be, what can I use this model for? How does it fit in to my overall model stack? What can I do to take most advantage of it while recognizing whatever tradeoffs it comes with? That's what we're getting into in today's episode, so let's dive in. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in.
Starting point is 00:01:12 First of all, thank you to today's sponsors, KPMG, Blitzy, robots and pencils, and hyperagent. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors at aidailydief.aI. Also, as I've been mentioning recently, our next set of executive agent leadership, programs at Super Intelligent are coming up just after Labor Day. You can find out about those at training.bysupor.aI. Again, you can find out all about that at training.bysup. We have kind of a dramatic set of headlines today.
Starting point is 00:01:44 The first up is an update about OpenAI's forthcoming Astra. In a Tuesday blog post, OpenAI said that they now believe that Astra meets the critical cybersecurity capability threshold under their preparedness framework. In layman's terms, that means that the model is capable of finding and exploiting previously unknown security flaws without human guidance. In their previous assessment at the beginning of August, OpenAI believed that it was possible Astra would reach the threshold but weren't sure yet. Essentially, this is the same concern that saw Anthropic keep Mythos under lock and key earlier this year. Sharing some details on how they assessed Astros capabilities,
Starting point is 00:02:17 OpenAI shared that the model achieved a perfect 100% score on exploit bench. This benchmark evaluates a model's ability to develop exploits based on known vulnerabilities. Open AI then took it a step further and developed their own internal version of the benchmark, consisting of 20 high-severity vulnerabilities that were recently disclosed. The idea was to test whether the model was actually capable of creating novel exploits from scratch by using tests that couldn't be in the training data. OpenAI wrote, on this data set, Astra achieves much higher arbitrary code execution rates than GPT-5-6 sole, using far fewer output tokens. During the evaluation, the model even discovered and used two zero-day vulnerabilities as part of an exploit chain. Now, to put some numbers,
Starting point is 00:02:57 members around this comparison, Astra managed a 30% score on their internal version of exploit bench with 40,000 tokens used, as opposed to GPT56 Seoul, which wasn't capable of any significant results until it spent around 110,000 tokens. But if we extrapolate out to other capabilities, this could mean the model is much more token efficient for running agents across the board. In further testing with expert partners, OpenAI found that Astra was able to design and execute full exploit chains to gain route access to a hardened operating system and execute commands on a hardened browser. As a result, OpenAI will deploy a series of new safeguards for Astra's release. The model itself has received additional training to refuse cybersecurity tasks,
Starting point is 00:03:37 resulting in a 91.5% refusal rate up from 59% for GPT56 sold. OpenAI is also adding more classifiers to detect cyber abuse and attempted jail breaks, and in addition, open AI will now be flagging certain accounts as higher risk and applying more stringent model behavior guardrails to those accounts. OpenAI says that they believe that Astra is more likely to respect security boundaries than previous models, but they're still implementing additional chain of thought monitoring to detect and stop misaligned actions early. In an unusually serious post on X, that even used like correct grammar and punctuation, Sam Altman added, there is an obvious tension here. On one hand, Astra is very good and we are excited to see what people will build with
Starting point is 00:04:18 We are proud of our work. On the other hand, we are clearly in a phase of development where we believe caution is warranted, and we are pacing our progress to ensure that we can meet the safety standards required by new capability levels. Astra has been done with training for a while now and is a significant step forward in both capabilities and alignment. For the models after that, we have been slowing things as needed to ensure that we can do sufficient work on safety and alignment.
Starting point is 00:04:41 Hinting at the mood inside OpenAI, he continued, We've been living with the tension between being excited and anxious about progress for some time, and it is still discordant for us. We know it is much more discordant for other people, and yet we believe strongly that the world needs to understand where AI is going and how models perform in the real world. More importantly, we believe the world will need aligned AI to manage the future phases of this transition. An iterative loop where society and this technology evolved together is what will lead to the highest chance of getting this right. So we hope you enjoy our new model, and we hope the world continues to take what's happening in AI extremely seriously.
Starting point is 00:05:13 Now, sources suggest that Astro could be coming as soon as this week, which would be perfect timing given that I'm traveling and theoretically I'm doing preload episodes. But believe it or not, that is not the only discourse going on about Astra. In a late-night scoop on Tuesday, the information revealed a technical breakthrough that makes Astra much better at reasoning, and, according to some, potentially much more dangerous. The technique is called recurrent depth, which uses a looped transformer. Functionally, this means the model can process the same text string multiple times to improve its response before generating an output. While sources say the technique improved performance and reduced
Starting point is 00:05:45 cost, the big downside is a lack of observability. Part of the reasoning process now takes place inside the model without generating an output. This means chain of thought will be partially obscured and unable to be read or understood by humans. Open AI sources said that they've used the technique in a limited way in Astra to ensure that reasoning can still be adequately monitored. However writes the information, AI researchers quote, worry that some AI developers may not impose the same kind of limits OpenAI did if they adopt the same technique for their own models, and that unfettered use of the technique could potentially lead to runaway AI whose actions can be hard to oversee. Now, folks working in AI safety have already been concerned about
Starting point is 00:06:22 agent observability getting more difficult. In their analysis of the hugging face attack, Meeter noted that logs were impossible for a human to piece together and required AI analysis to get the full picture. Duarkech Patel, in his dramatic and controversial retelling of the attack earlier this week, commented, I don't think this is the final warning shot we'll get, but it's probably the last one that I'll personally be able to understand. Following the report, Nathan Calvin of Encode AI posted, really huge and extremely concerning story from the information tonight. Looks like OpenAI utilized a breakthrough in Neurolease for Astra that could destroy chain of thought monitorability. It seems quite likely that if OpenAI discover this architecture
Starting point is 00:06:57 and found performance or efficiency gains, that other companies are likely to find it soon, too, if they haven't already, and may not choose to prioritize monitorability at the expense of efficiency. If some folks do, it may be difficult to avoid a race to the bottom. Ryan Greenblad of Redwood Research, who was one of the lead researchers on the meter investigation of the Hugging Face incident, wrote, My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space. This would very likely destroy the usefulness of chain of thought for monitoring and oversight.
Starting point is 00:07:28 Former Open AI researcher Stephen Adler said, If this is true, Open AI seems to be violating one of the few red lines that exists in the AI industry. Absolutely do not train your models like this. What is going on? Still, a number of folks tried to jump in and calm down sentiment a little bit. Amir Afradi from the information again jumped in to reinforce the notion that their reports suggest that OpenAI is putting limits on this technique and trying to make sure chain of thought is visible, but it's concerned that other AI developers may not. And OpenAI chief scientist Jacob Pachaki wrote, I want to prevent a race into unmonorability kicked off by confused reporting. The depth of the computation graph for our present frontier models,
Starting point is 00:08:03 including Astra, is within a factor or two of GPD4. OpenAI has worked to preserve and utilize chain of thought monitoring since our very first reasoning models. We care deeply about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program. So you know, another uncontroversial release coming up. Speaking of releases, the Wall Street journal reports? That Gemini 3.8 Flash is on the way and could fix one of the longest
Starting point is 00:08:38 standing problems for Google's AI, which is coding. Now, try as they might, Google has never produced a state-of-the-art coding model, and at this point they have fallen drastically behind in this critical capability. Yet the journal reports that during testing within the company, engineers preferred their forthcoming 3.8 Flash model to Anthropics Opus. Now, the model is expected to be released this week, possibly today, so we'll soon see whether it lives up to the hype, but the article also covered what's been happening behind the scenes for the Gemini Pro series. horses said that all internal candidates to be released as 3.5 Pro were scrapped because they weren't sufficiently better than the Flash models. However, researchers are pleased with the performance
Starting point is 00:09:11 of Gemini 4 during pre-training evals. The model is still in post-training, meaning there's more time before it's ready, but perhaps some good news for those who want to see more competition than just OpenAI and Anthropic. Now, one model which, were it not for our main topic of Fable 5.1, could have easily been the entire main topic for today, is World Labs' newly released model Atlas. They describe it as the world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space and time. World Labs Ben Mildenhall writes, Atlas is an auto-regression diffusion model built from the ground up for the task of next
Starting point is 00:09:48 frame prediction. It is simultaneously a world-class method for camera-controlled video generation, novel view synthesis, and sparse 3D reconstruction. World Labs co-founder Fei-Lee writes, Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating spacetime by reframing videos, natively outputing 3D spaces from one or more input images, composing multiple posed images into a consistent 3D world and more. This is the best camera-conditioned world model ever,
Starting point is 00:10:17 opening doors to many possible use cases from VFX to robotics. Now this is one that you really have to go see, but it's controllable video in real-world environments like nothing you've ever seen. explaining an example of a use case Justin Ryan writes, Atlas is an AI model that can reconstruct moving 3D scenes from as few as three cameras. Creators can record a real moment, then view it from camera angles that were never filmed. A16Z's Martin Casado writes, Think of it as a video model with full camera control,
Starting point is 00:10:44 and the scene remains nearly 3D consistent, built on a fully internal base model. There are many use cases from video editing to 3D construction to robotics. Peter Yang summed up the feeling of more than a few when he wrote, Yeah, Fable 5.1 is really cool, but this is bonkers. And for Elvis on X, it's more than just Fable 5.1 that this is cooler than. He writes, Omni Models are the next frontier. And simply put, this is the most exciting release I've seen this year. Alas, for now, for most people, when it comes to our day-to-day use cases,
Starting point is 00:11:13 the bigger topic is indeed Fable 5.1. So with that, we will close the headlines and move on to the main episode. Hello, everyone. One big change around AI is we've shifted our thinking from how we rank our pages to how do we become the source that AI trusts enough to answer with? At KPMG, they're seeing this firsthand. AI generated results now surface answers directly, often without a single click. That's why they are increasingly focused on generative engine optimization or GEO, structuring content so AI systems can retrieve it, understand it, and cite it as trusted
Starting point is 00:11:48 authority. This is not just an SEO evolution, but a visibility mandate. And indeed, the GEO mandate from KPMG is simple. If AI is shaping decisions, your expertise needs to show up inside the answer. Read all about it at KPMG.com slash us slash geo. Again, that is KPMG.com slash us slash geo. Here's why most legacy modernization projects fail. The AI doing the work can't understand codebases at scale. It sees a small slice of context, examine syntax, and misses years of decisions distributed across the global application ecosystem. Blitzy solves this the way it solves everything. Grounded in your code before any migration begins, Blitzie's agents reverse engineer the entire legacy system into a persistent knowledge graph, every dependency, every
Starting point is 00:12:33 constraint, every piece of tribal knowledge that used to live in one engineer's head. From that understanding, Blitzie autonomously executes language migrations, framework upgrades, and monolith to microservices transformations, all validated end to end. One Blitsey customer modernized a $10 million monolithic insurance stack in 16 weeks against a 137 week baseline with coding agents. That's 9x compression. Retire technical debt while accelerating your roadmap. See how at blitzy.com. That's B-L-I-TZY.com. At this point, it's no longer a question of whether companies are actively using AI. Using it well, on the other hand, is a whole different story.
Starting point is 00:13:08 Robots and Pencils, though, is a company that I can point to that is actually built for this time. They're an applied AI engineering firm working directly with clients on problems that matter to the business, not experiments that live in a slide deck. Every engagement starts by working backwards from the outcome a client actually needs. If you're trying to tell real AI engineering apart from noise in this space, that's the difference maker. Head to Robots and Pencils.com. This episode of the AI Daily Brief is brought to you by Hyperagent, where you run fleets of agents your team can manage together.
Starting point is 00:13:36 Forget local agents and chat workflows waiting on your laptop to be prompted. Hyperagent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing agents turn competitor moves into landing pages. Sales agents enrich leads, draft emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at Hyperagent.
Starting point is 00:14:03 Get $100 in credits at hyperagent.com slash AI Daily Brief. Welcome back to the AI Daily Brief. Today is one of my favorite types of days around these parts at the AI Daily Brief, and that is a new model day. On Tuesday, Anthropic released Claude Fable 5.1 and Mythos 5.1. And what's interesting is not just how the capabilities have improved, but the other aspects that Anthropic chose to focus on with this launch. Still, let's start with the capabilities. From here on out, though, the question around every single state-of-the-art advance will be given how powerful our existing models are, are the capabilities jumps or some other new feature worth making the switch to? Still, let's talk about capabilities first, because if they aren't a big upgrade, the rest of the conversation is
Starting point is 00:14:49 kind of pointless. In short, Fable 5.1 is the new state-of-the-art, unambiguously. 5-1 scored 55.8 on Terminal Bench 4.0, which tests Aigentic coding, and that goes all the way to 60.9% for Mythos 5.1. That's up from 42% for Fable 5 and 52.3% for Opus 5, and way above GPt 56 sole at 37.3%. There was a similar jump on Cursor Bench 3.2.0, with Fable 51 scoring 73.4% against Fable 55 score of 70.5%. GpT 56S scored 67.2%, so again, a pretty significant gap. Fable 51 has also got a new state-of-the-art score on GDPVal AA, beating Fable 5 by 130 Elo points and Opus 5, which was the previous state-of-the-art by 29 points. GBT 5-6 sole was already 12 points behind Fable, and is now over 140 points behind Fable 5.1.
Starting point is 00:15:39 For business tasks, Fable 5.1 scored 31.4% on automation bench, which is a huge jump from the 17.1% score that Fable 5 achieved, and 19.6% for GPt 56 sole. computer use, which is obviously a key part of agentic capabilities, is also up with Fable 51 coming in meaningfully above previous models as well. Still, it's very clear from the announcement that Anthropic was concerned not just with an improvement in capability, but also an improvement in cost. The charts that the team was most keen to share on social media were the charts that not just showed the score comparison, but a graph of score matched against mean cost per task.
Starting point is 00:16:15 Across agentic scientific research, agentic terminal coding, multidisciplinary reasoning, and broader agentic coding, not only did Fable 5-1 score higher at each effort level from low to max, but each of their mean costs per task were lower at each comparable level. In other words, at a low, medium, or high effort setting with Fable 5.1, you were going to get a better score and at a lower cost than the low, medium, or high effort setting on Fable 5. And right up top in the blog post, it is clear that price is a major focus. Anthropic writes, Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads,
Starting point is 00:16:49 wherever usage is billed by token. This is because we're reducing our pricing on cash reads, where the model reads inputs that have already been processed and stored. For highly agentic work, the savings will often be much larger, up to approximately 45%. In other words, as we have been discussing, the question around new model releases is no longer just about capabilities jumps, but also about efficiency increases, and you can see that even a purest company like Anthropic is not immune to that new reality. Then again, it's one thing for a company to make claims about its benchmark scores and costs for its own models, it's another thing when they get tested in the wild. So when it comes to the artificial analysis intelligence index, Fable 51 is undisputedly at the top
Starting point is 00:17:26 of the benchmarks. It jumped from an overall score of 62 with Fable 5 to 66 for Fable 5.1. That puts it ahead of Opus 5 as well, which was the previous leader at 63. Anthropic now has the top three models on the index, all slightly ahead of GPT56 sold. However, A.A. find that the cost per task was brutal. In fact, artificial analysis found that Fable 51 was actually a little more expensive than Fable 5, even with the cut to cash-read pricing. The model cost $3.76 per task compared to $3.14 for Fable 5.5. Artificial analysis blamed much higher token consumption, with Fable 5.1 using 70% more tokens across the benchmark run. They acknowledged that the reduction in cash costs did save an average of $1.40 per task, largely concentrated in the agenic benchmarks, but that was
Starting point is 00:18:15 enough to offset a much more token-hungry model. Notably, testing Fable 5.1 on extra high rather than max, produced a 28% reduction in cost, with only a 1 point overall drop in performance, scoring a 65 overall. Public LLM evaluation platform VALS AI also found Fable 5.1 in the lead of its charts, and asked, has Anthropic solidified itself as the frontier leader? Fable 5.1 debuts at number one on the VALS index, and with Opus 5 and Fable 5 right behind it, Anthropic now holds the top three spots. Now, the Ark Prize found results a little bit more in line with Anthropics' promises. 5-1 scored a 90% on Arc AGI2 and a 97.5% on Arc AGI1, and they reported that its average cost per task was about 32% lower than Fable 5s, driven by better token efficiency.
Starting point is 00:19:00 Unfortunately, they couldn't really get clear Arc AGI3 results, as they write, The requests were frequently misclassified by Anthropic as reverse engineering attempts, preventing us from completing testing before release. One other benchmark that Anthropic was very keen to highlight was the big, jump on agentic scientific research on the terminal bench science benchmark, where Fable 5 won at 52.6% scored almost double the previous high of Opus 5 at 29%. There have been a lot of indications recently that Anthropic wants to spend more time and more focus in the areas of medicine and biology and scientific research more broadly, and given the top billing of the Agentic Scientific
Starting point is 00:19:35 Research benchmark, this seems to be more evidence of that. Another big thing that Anthropic was pitching in the announcement blog was the fact that the guardrails were much improved between Fable 5 and 51, which has specific implications for something like biology and medical question, where they claim that they've reduced the fallback rate, i.e. the times when the model switches from Fable 51 to instead an Opus model by about 85%. Indeed, what's interesting about the announcement overall is how much the focus is not strictly on the capabilities. Axio Senior AI reporter Madison Mills writes, so it's not enough to release a new model anymore. Now we're getting new models, new safeguards, cybersecurity, and intervention changes, cost cuts, and new enterprise
Starting point is 00:20:13 IP protections all in one release. And what she's referring to is that right up at the top of that blog post, in addition to Price, Anthropic is also selling that they have a new Enterprise Frontier safeguard system or EFS, which allows them to offer enterprises zero data retention. One of the absolute biggest blockers to Fable 5 usage has been that enterprises simply weren't willing or weren't able to deal with the 30-day retention policies that came when Fable 5 came back online after being shut down by the U.S. government. And so this is a major, major upgrade, although the EFS system is not going to be rolled out all at once, with Anthropics saying that it will be made available to enterprise customers in phases beginning later this fall. Still, they say until EFS is available, eligible customers will be
Starting point is 00:20:51 able to use Fable 5-1 with zero data retention. Finally, in addition to price and data retention, they also pitch this improvement in their safeguards, specifically improvements to reduce false positives where the system flags benign content. In addition to the 85% reduction that I just mentioned for biology and basic medicine questions, they also said that the new safeguards block 60% fewer false positives than before in cybersecurity as well. Now, on the cybersecurity front, it sounds like they've conceptually refocused things, making it so that Fable 51 can be used to discover vulnerabilities without being able to develop exploits for them. So what were people's first impressions? The one other thing that a lot of folks from the Anthropic team were pitching was Fable
Starting point is 00:21:29 51 sounding less Claude-like, i.e. having less of the hallmarks of an AI writer and less of some of the patronizing tone that people have been annoyed with with recent iterations of Claude. Claude Code creator Boris Churny posted, We heard your feedback and are actively working on reducing ClaudeSpeak, solid progress with 5-1, more to come. But what did users outside Anthropic find? Professor Ethan Malik did find that overall the model is a meaningful advance, although perhaps not as much of an advance on that ClaudeSpeak as we might like.
Starting point is 00:21:55 He wrote, Had early access to Claude Fable 5.1. It's a real advance in long-run work that requires judgment and taste, but less of an advance in the clodish. Part of the way he tested it was creating a retro game. And interestingly, a lot of people seem to be looking for games as the way to test things. Bridgemind AI shared a video of a Mario Kart clone saying Fable 5.1 one shot at this Mario Kart game, one of the best results I've had so far and I am super impressed with the game development capabilities.
Starting point is 00:22:22 Alex Albert from Claude showed how he used Fable 5.1 to generate videos through code. For those of you not watching, the video is a walkthrough of the type that you might see in a real estate listing. Alex says, for this one, I gave it a picture of a property lot. It designed a house for the lot, rendered it, and produced a cinematic walkthrough. Meng, too, found that Fable 5 was really good at advanced JavaScript for more visual and interactive sites. He wrote, It's faster, understands complex design instructions better, and recreates references with surgical precision.
Starting point is 00:22:51 With this much power, it's hard to settle for static sites, especially when so many AI sites look generic. That said he did point out that it's not all of a sudden perfect, that it can still create generic AI illustrations if you don't specify the images, that it still has some difficulty with 3D subjects like people and dogs, that you still had to deploy taste, fixing overlapping elements, negative space, and scroll behavior, and that because it works faster, he went through tokens very, very quickly. That token burning effect is something that we'll come back to in just a minute. On frontend design, Mia writes, I've asked Claude Fable 5.1 to create 100 HTML files. The rules
Starting point is 00:23:23 were simple. Look stunning, zero repeat designs go full creative mode. All 100 files created in one single prompt. These are the best results I've had with this type of experiment beating any other model. It's really good on the front end, and there's almost no broken files. It's truly impressive. Entrepreneur Matthew Miller wrote, Fable 51 is the best model I have ever used. I have thrown everything at it since it dropped. Every single task done to perfection.
Starting point is 00:23:46 The one-shot capabilities are unlike anything I have seen. You ask once and it just delivers. But the thing that actually blew me away is security. I can hand Fable 5.1 security tasks, and it does not fall back or refuse. It found and patched vulnerabilities in my codebase that Fable 5 refused to. This is the fastest I have ever felt AI advance, GPT Astra and GROC 4.7 are both about to release. The world is about to change. Now, every time there's a new model, you can always count on every to have one of the most
Starting point is 00:24:13 comprehensive reviews. This is of course their vibe check series, and their summation of Fable 5.1 is Anthropic is Soback again. CEO Dan Shipper wrote, It's the strongest coding model we've used, but now it's fast, token efficient, and crucially actually speaks like a normal person. The team at every found that it was a monster at coding. Dan said that Kieran Klassen rebuilt a working version of one of the their products from one prompt, and Fable 51 added useful details that he hadn't requested. On writing, Dan said, it had clearer pros, fewer AI tells, and it takes an edit without arguing. They found that on agentic tasks that it used about half the tokens is Opus 5, and delivered
Starting point is 00:24:48 things in about 60% of the time. Previously, Dan said, the big knock on Anthropic was that they built a super genius in a data center that was almost unusable. It was too slow, argued back, and talked to technical gibberish. They've managed to solve those problems and more with Fable 5.1. And what's even more important than that is that I think that Dan, landed on the usage pattern that many power users might. He wrote, I still use chat GPT for work more day to day, but I use way more tokens in Fable 5.1.
Starting point is 00:25:14 I send it off at the beginning of the day to do big programming projects like end-to-end MVP builds and check in every once in a while. This has sort of been a power user's division of labor for some time at this point. The GPT-5-6 models in Codex for interactive tasks where you are co-working with the AI, and the Fable models for long-running tasks that don't require as much interaction. Will Brown from Prime Intellect agreed saying, God, this model is nuts. They really just made it smarter and better at coding. It can just do things. They made it reasonable and not slop. This is so cool. You can give it way more work and it just does it. The code is pretty good. It explains the important stuff well, follows instructions, catches its own mistakes. The most AGI pilling model for me in
Starting point is 00:25:51 several weeks at least. Now, to the extent that there are critiques so far, it is absolutely about how token hungry the model can be and how quickly that runs up against subscription usage limits. Steve Jabbs writes, Fabel 51 and about 12 subagents equals one hour of usage on the 20x Klaude Max plan. Chubby writes, literally unusable, the rate limits are absurd. And oh, by the way, Fable's automatic continuation is bugged and doesn't even work. Issue Agrawal writes, Fable 5.1 is unusable. It's so expensive that you can barely get more than 30 minutes of usage out of it. And weekly limits will also be lower in two weeks. This is not a model for extended work. Even people not prone to hyperbole like former investor Jeffrey Emanuel wrote,
Starting point is 00:26:30 something definitely seems screwy with the Fable 5.1 usage. Probably a caching bug in the new Claude Code, if I had to guess. I managed to blow through all of my 28 max 20x accounts today, at least the five-hour usage limit, just doing audits of a bunch of my projects. First time ever. Entrepreneur Adam B. Levine dug in and suggested that he might have found the problem. Pro tip, he writes, Fable 51 was burning a lot of credits, and turns out it decided every sub-agent should be a Fable 5.1, ignoring our longstanding rule to the contrary. In another 20, he said, seems like 5-1 is super trigger-happy with big workflows that use like 10 Fable 5.1 sub-agents that then eat even a 20x limit if you're running more than one agent or it's a bigger project.
Starting point is 00:27:10 Basically, if you just let it go on the default settings, it's going to use 5.1 to spin up the sub-agents that it uses to do work, and that could burn through things very quickly. Already people started jumping in with their own cost optimization approaches, but I sort of think that Jan Velek has it right when he says, subscriptions will end API pricing is awaiting us. I think at this point that is pretty inevitable. But I also think that people always do this thing when they judge costs in the very first hours of even having a model
Starting point is 00:27:35 before people have really figured out how to use it and before all the norms settle. So I wouldn't be surprised if your mileage actually goes a bit farther than some of the responses that you're seeing. The question, though, is, especially if there are strict usage limits and you're going to find yourself on API pricing which is pretty expensive soon.
Starting point is 00:27:50 Matvee writes, serious question. What can you do with Fable 5.1 that you can't do with opus, soul, Kimmy, composer, or GROC? give me your actual use cases. Tell me what I'm missing. On the one hand, I think this is the right type of question for people to be asking. In the same way that pretty much every enterprise right now is trying to figure out a multi-model architecture that allows them to connect the right task with the right level of capability, most individuals are going to have something similar, or perhaps they
Starting point is 00:28:17 don't have any sort of automated router, but they just understand and have designed systems so that they know which model and setting to use for different types of requests, so that they're not just burning everything on the most state-of-the-art, most expensive model on the highest settings. At the same time, there's this idea that's been around for a while, that the models are so good now that for many use cases no one can really tell the difference between them, and to even consider using the most expensive state-of-the-art models, you must be deluding yourself into thinking that there's actually a difference. I reject that pretty wholesale. The idea that just because multiple models can successfully complete a task means that they're all interchangeable with one another,
Starting point is 00:28:55 is akin to saying that if two people can complete the same work task, it doesn't matter which one does because the task got done. Now, certainly there are going to be tasks for which that is the case. And those are precisely the tasks that you should be optimizing using cheaper models for. But when it comes to a lot of high-end important work, I still find that as capable as all of these models are, there are still massive differences between them. One thing I strongly advocate for is to have a standing slate of personal benchmarks for new model testing. They don't know it to be anyone else's tasks, they can just be the things that matter to you. And you might find that for your particular tasks, models that other people are complaining about work great and models that
Starting point is 00:29:34 other people love don't work so well. For me, that personal benchmark lists includes a few things. It's basically some combination of research, writing, strategic and critical thinking, and building, which includes both an interface design and an architecture component. And what you'll notice is that, especially when it comes to something like writing or strategic thinking, a lot of preference is going to be subjective. In other words, Anthropic can't show me some benchmark for iterating on NLW's mad ideas for new businesses. That's something that I have to see how Fable 5 versus Opus versus Soul handle in practice. And even in this era of generally capable models, I still find massive differences in things like that. The reminder here is that for all of us, the question
Starting point is 00:30:13 when a new model comes out is no longer should I switch to that model. Instead, it's how does that model fit into my personal model architecture. For what uses, is that model better and worth whatever financial or other types of costs that come with it? The best users, in other words, are going to figure out how to get the most out of new models, rather than just clunking around from one to the next, with some old idea that you have to pick just one. Now, for one last qualification on that, I will note that if you are not in a financial position where you can be blithely shifting between models, a lot of these considerations get different. And for that, the advice that they are all pretty generally capable is accurate. It certainly is the case that it has never been a better time
Starting point is 00:30:52 to be locked into just one ecosystem because they are also individually capable, even if they do have different tradeoffs. Still, now the fun part begins where you get to go test and try these things. I'm excited to spend some time this Labor Day weekend testing things out, and I will, of course, report back next week. For now, though, that is going to do it for today's AI Daily Brief. Appreciate you listening or watching, as always. Until next time, peace.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.