The AI Daily Brief: Artificial Intelligence News and Analysis - AI Model Month Is Off to a Blistering Start

Episode Date: September 9, 2026

September’s model boom brings Gemini 3.8 Flash, Meta’s MuSpark 1.3, the Muse personal agent, and ChatGPT Images 2.5. NLW explores why faster, cheaper, more specialized AI makes model selection cri...tical. In the headlines: OpenAI’s disputed Navier-Stokes breakthrough, a Claude usage-limits lawsuit, ElevenLabs’ IPO preparations, and Cognition’s $48 billion valuation.Multiplayer AI Sprint - ⁠⁠https://multiplayerai.ai/⁠⁠Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://kpmg.com/us/Sophisticated⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Harbor - Invest in the AI ecosystem. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.harborcapital.com/aidaily⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Hyperagent - Hire a team of always-on agents. New users get $100 in free credits. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠hyperagent.com/aidailybrief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.rackspace.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Section - Section turns AI investment into workforce transformation and ROI - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.sectionai.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The AI Daily Brief helps you understand the most important news and discussions in AI. Newsletter: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Interested in sponsoring the show? sponsors@aidailybrief.ai

Transcript
Discussion (0)
Starting point is 00:00:00 Throughout the summer, the big thing we've been exploring at the AI Daily Brief is all about the move from a single model paradigm where you pick the best model overall and that's the one you stick with to a more complex model architecture where we are both as individuals and as teams, able to navigate nimbly between different models and even different harnesses to get the most out of AI based on whatever particular use case we might have. And what's more, this summer, we got really clear on the fact that getting the most out of AI is not just a question. a model or harness capability, but also a question of efficiency and cost, especially as we move to more complex agenic workloads. And so it's fitting that the beginning of September has been just a cavalcade of new models, from Fable 5-1 to GPT-6 Astra to the models that we're looking at today, including Muse Spark 1.3 and ChatGBTGBT Images 2.5, all of these add up to way more diversity in the tools we have access to for you to design the perfect AI stack for your actual life and work.
Starting point is 00:00:58 The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Section, and HyperAgin. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors at AIdailybrief.aI. And finally, if you haven't yet, you can check out our latest free self-directed training program. it is called the multiplayer AI sprint for teams,
Starting point is 00:01:33 and basically the idea is to shepherd you through a process of figuring out how to build agents that don't just help you, but actually sit at the intersection of work that is shared across your teams. I'm pretty convinced that this is the next big paradigm for AI inside companies, and so I wanted to build a sprint that could help you guys fully embrace that. There's, of course, a link to that on the AIDailybrief.ai website, but you can also find it at multiplayera.ai. We kick out today with a story that very easily could have been the main episode,
Starting point is 00:01:59 given how much drama is surrounding it. On Tuesday, OpenAI published a solution to the Navier-Stokes problem, one of the seven problems selected for the Millennium Prize in the year 2000. The Wall Street Journal characterized these problems as the, quote, holy grail of math, and that's fairly accurate. Each Millennium Prize problem has a million-dollar reward attached, and only one has been solved in the 26 years since the prize was established. The other problems include the most famous unsolved problems in math,
Starting point is 00:02:28 such as the Riemann hypothesis, N.P. You know, the things we all talk about when we get together for dinner. Now, for the purposes of this particular episode, I'm actually not going to get into the details of the problem itself, or debates around whether it has any significant real-world applications. I'll read OpenAI's description of the problem just to give you a flavor. They write, the Navier-Stokes equations use Newton's second law of motion, F-Equels MA, to describe how fluids move. Importantly, they treat a fluid as a continuous medium rather than tracking individual molecules. These equations are used for aircraft design, weather forecasting, and the study of blood flow. A fundamental open question for these dynamical equations has been whether the continuum approximation of the fluid can break down.
Starting point is 00:03:09 Specifically, can the Navier-Stokes equations for a three-dimensional, incompressible fluid with constant density develop a singularity even when the motion starts smoothly. Here, a singularity means the dynamics lead to speeds in the fluid growing without bound within a finite amount of time. The development of a singularity would have to deepen despite the presence of viscosity, which tends to smooth out motion. Because a real fluid cannot move infinitely fast, this would mark a breakdown in how the equations model the fluid. To continue modeling the system, one would then need to track the behavior of each particle individually. So that's the problem they're addressing here, and I think again for context, the important note is that this
Starting point is 00:03:43 represents a huge step up from something like the Erdos problems that made news last year. Open AI claims to have solved this problem, and they did so using an internal model that is significantly more capable than GPT-6 Astra. Noam Brown said that the result cost several million dollars to find, and it seems to have taken a week or two. However, the big controversy surrounded exactly how OpenAI had arrived at this result. Shortly after the result was published, New York University professor Tristan Buckmaster published his version of the events.
Starting point is 00:04:12 According to Buckmaster, he and an anthropic employee, named Levant Alpogee, had been working on the Navier-Stokes problem together for more than a year. This was an outside project for Levant and the pair had used a range of different AI models, including GPD 56 sole in the Codex harness. Crucially, Levant and Buckmaster did not find a solution to Navier-Stokes, but they did find novel solutions to related problems that could be viewed as a stepping stone to the Millennium Prize problem. They were also using extremely novel methodology that few in the mathematics world were pursuing.
Starting point is 00:04:44 Rumors of their work spread through AI circles in recent weeks, incorrectly claiming that Anthropic had solved a Millennium Prize problem. Buckmaster says he reached out to OpenAI last week to clarify the situation. According to Buckmaster's telling, Sebastian Bubeck from OpenAI informed him on Sunday that they had solved Navier Stokes and wanted to discuss publication. Buckmaster wrote, I asked whether the model had been trained on or had access to our sessions in Codex, into which we had been putting all of our drafts for the whole of the project.
Starting point is 00:05:13 I was told the model did not look up user data. I asked again about training and I did not get an answer. He said he was offered two proposals, for either OpenAI to publish separately or Buckmaster to join the publication if he agreed to remove Levinn from the authorship because of his ties to Anthropic. Buckmaster declined both options and threatened to go public if OpenAI published. Continuing his account, Buckmaster wrote, The reply was, why would you ruin your career? I replied that I am an academic and asked why he thought going public would ruin my career. The reply was, if you don't want me to be nice, then I don't have to be nice.
Starting point is 00:05:47 Open AI leaders responded with a series of statements. Sebastian Bubeck called the allegations false and inflammatory, then revealed part of his text message chain, which he claims contradicts Buckmaster's version of events. Sam Altman gave a series of explanations for how this went sideways, but claimed that OpenAI's approach was different than that of Buckmaster and Levin. The OpenAI account claimed, we, the researchers, and the agents,
Starting point is 00:06:09 did not see any of their work through any means until they released it publicly. In particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from the usage of our products helped improve our models. After Levin characterized this statement as quote-unquote coming clean, OpenAI chief research officer Mark Chen responded, two things to distinguish. Did any human or agent look at user data as part of the Navier-Stokes effort? No. Do we use user feedback and de-identified data to improve chat GPT and codex in a holistic way? Yes, and so does every LLM company. Now, as for the controversy, there are two distinct strains of conversation.
Starting point is 00:06:48 Firstly, academics are up in arms over what they see as unethical behavior. Assistant Professor Talia Ringer of Illinois University wrote, Rushing to get a result after you hear someone else has a result is messed up. That is AI scooping culture and goes against every academic norm that exists in reasonable fields like mathematics. This is how AI culture rots entire fields. Thomas Wolfe, the co-founder of Hugging Face, suggested this might just be a preview of accelerated AI science, commenting, hope this is not a glimpse of the future we'll get in science research, with these dominating players playing marketing games hurtful for the real scientific community.
Starting point is 00:07:21 The second and likely far more relevant criticism, at least for the AI Daily Brief audience, was questions of trust in OpenAI. From Buckmasters account, we can assume they were using some consumer version of Codex, but it's unclear whether they agreed to share data to improve OpenAI's models. For some, it's a wake-up call for anyone who is using AI to work on proprietary tasks. former deep mind employee Susan Zhang wrote, Everyone getting sniped by the personal drama, but miss the more interesting unanswered question.
Starting point is 00:07:49 Can these labs see all your work and scoop you when the stakes are high enough? Seeking clarification from OpenAI leaders? Mathematician Tryon Zylorus asked. Important question. If I opt out from training, then paste a trade secret using my paid subscription, do you de-identify my personal details but keep the trade secret and may add it to your training data? At the time of recording, that has not received a response.
Starting point is 00:08:09 So taking a step back, there are a few reasons that this whole episode is having such resonance. First is honestly the voyeurism of it. People love drama. Fighting against that is like trying to fight the tides. But the question of ethics around advanced AI and what these companies can do with data is a question that while it has been present, basically since the beginning of LLMs, has gotten a lot louder in consideration more recently. Especially as model leadership starts to consolidate around a couple of companies, it brings up a lot of uncomfortable questions for people. Now, some of those questions go to economic incentives and where the AI companies ultimately land. One of the things that there is a lot
Starting point is 00:08:45 more chatter about right now is questions of whether these companies will ultimately not view themselves just as selling the inputs to innovation, but also as selling the outputs of innovation. In other words, does it make more sense for OpenAI and Anthropic, to sell existing scientists and labs and companies, the ability to do novel drug discovery, or does it make more sense to do that drug discovery yourself and get the money on the other side of the patents? Given all the chatter that we had, around whether Dario had actually said that Anthropic was going to be the only company in the world at some point, those questions feel a little bit more pertinent now than they might have in the past. And finally, the fact that the story reveals that OpenAI has a much more powerful model that they're already using internally,
Starting point is 00:09:25 certainly captured some notice as well. Unfortunately, as is so often the case, no one looks particularly good coming out of this. The New York Times, Mike Isaac wrote, Not lost on me that the two labs asking the public to trust them as stewards of responsible AI leadership at existential stakes or having a slap fight on Twitter about who gets credit over a math problem. Like I said, this could have been a whole main episode, but since we're a little compressed on time, let's quickly rip through a couple of other stories before we get to our main, which is all about all sorts of new models that we haven't had a chance to talk about yet.
Starting point is 00:09:54 First up, speaking about Anthropic, there is a new class action lawsuit against the company filed on behalf of Claude Mac subscribers, with the claim that Anthropic used deceptive marketing and opaque fine print to underserved customers. In particular, the lawsuit claims that the $100 a month 5x plan and the $200 a month 20x plan don't actually deliver five and 20 times the usage of a $20 a month pro plan. Plaintiffs allege the way five-hour and weekly usage limits are calculated mean the actual usage is far lower than the advertised multiples. Now, usually this type of case wouldn't be all that interesting to me, but I think it's sort of representative of the type of thing that we're going to see a lot more of, as anthropic
Starting point is 00:10:31 and open AI become increasingly interwoven with just the normal way of doing business. The lawyers running the case themselves noted that what made it compelling to them was how ubiquitous and necessary a top-tier AI subscription has become. They said that they were frequently hearing from workers who felt they needed to pay high-price subscription costs to remain relevant in the job market, but felt they weren't getting what they were paid for. I'm not particularly sure I think this goes anywhere, but it is certainly representative of the level of scrutiny that the top AI labs are going to face going forward. Over in markets, it looks like it's not just open AI and anthropic that are thinking about IPO. 11 Labs has also hired a chief financial officer to help the
Starting point is 00:11:07 startup head for a public listing. On Tuesday, 11 Labs announced that Ethan Tandowski had joined the executive team, having most recently served as the CFO of Adyen, a Dutch fintech firm that went public in 2018. Eleven Labs co-founder Maddie Stanizuski said in a press release, Ethan brings a strong track record of scaling financial operations and high-growth environments and valuable experience as CFO of a public company. And that appears to be Eleven Labs' ambition as well. The information reports that they are beginning to explore a possible IPO. The company said that they are on track to reach 600 million in annualized revenue by the end of the year, up from 350 million at the end of last year, and sources said that the company has reached
Starting point is 00:11:43 profitability and is now generating more than half of their revenue from large enterprise customers. Cognition's recently rumored fundraising round has completed. The company raised $2 billion in new funds, catapulting them to a $48 billion valuation. Cognition last raised funds in May at $26 billion, meaning they've almost doubled their valuation in three months. In that time period, cognition has gone from a $492 million. revenue run rate to almost 900 million at present. Beyond the numbers, the raise suggests that cognition will continue to operate as an independent agent lab. Following SpaceX's acquisition of cursor for $60 billion, there were rumors that they would pursue cognition as well.
Starting point is 00:12:19 CEO Scott Wu strongly denied the chatter at the time, and this fundraising round certainly gives cognition more runway to continue building their coding agent Devin and pursuing their thesis. In an announcement post, they wrote, we're still at the dawn of the self-driving software era. In this next chapter, agents will become proactive by default, software will improve itself, and even resource allocation will become intelligent as compute budgets self-allocate towards the highest-impact use cases. Human engineers will increasingly act as architects, setting goals and priorities while agents take on more of the work to achieve them. Cognition wrote that independence is core to this strategy, saying, we can choose and combine the models best suited to the work,
Starting point is 00:12:55 including our own, rather than tie customers to one provider. You got to think that after watching OpenAI cut off access to cursor customers because of SpaceX's acquisition, cognition sees the value of staying independent even more acutely. For now, though, that is going to do it for today's AI Daily Brief headlines. Next up, the main episode. If you're leading AI inside an enterprise, you already know that the gap right now isn't capability but execution. That's why KPMG's You Can with AI is back with a new season featuring conversations
Starting point is 00:13:27 with leaders like Sorosia Chatterjee of Emma, Mayhabib of Writer, McKess and CIO, Elery Fisher and others focused on practical execution. What's working, what's not, and what it actually takes to move from pilots to real scaled impact across strategy, data readiness, governance, workforce, and value. And of course, it's co-hosted by me, Nathania Whitamore. Go listen and subscribe at www.kpmg.org.us slash AI Podcasts. That's www.kpmg.comg.com slash AI podcasts. Here's why most legacy modernization projects fail. The AI doing the work can't understand codebases at scale.
Starting point is 00:14:03 It sees a small slice of context, examines syntax, and misses years of decisions distributed across the global application ecosystem. Blitzy solves this the way it solves everything. Grounded in your code before any migration begins, Blitzie's agents reverse engineer the entire legacy system into a persistent knowledge graph, every dependency, every constraint, every piece of tribal knowledge that used to live in one engineer's head. From that understanding, Blitzy autonomously executes language migrations, framework upgrades, and monolith to microservices transformations, all validated end-to-end.
Starting point is 00:14:31 One Blitzy customer modernized a $10 million monolithic insurance stack in 16 weeks against a 137-week baseline with coding agents. That's 9x compression. Retire technical debt while accelerating your roadmap. See how at blitzy.com. That's BLITZY.com. Here's a harsh truth. Your company is probably spending thousands or millions of dollars on AI tools that are being massively underutilized.
Starting point is 00:14:56 Half of companies have AI tools, but only 12% use them for business value. Most employees are still using AI to summarize meeting notes. If you're the one responsible for AI adoption at your company, you need Section. Section is a platform that helps you manage AI transformation across your entire organization. It coaches employees on real use cases, tracks who's using AI for business impact, and shows you exactly where AI is and isn't creating value. The result, you go from rolling out tools to driving measurable AI value. Your employees move from meeting summaries to solving actual business problems, and you can prove the ROI.
Starting point is 00:15:29 Stop guessing if your AI investment is working. Check out section at sectionaI.com. That's SECT-I-O-N-AI.com. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together. Forget local agents and chat workflows waiting on your laptop to be prompted. HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing agents turn competitor moves into landing pages. Sales agents enrich leads, draft emails, and updates the CRM.
Starting point is 00:15:59 OpsAgent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at HyperAgent. Get $100 in credits at hyperagent.com slash AI Daily Brief. Welcome back to the AI Daily Brief. I got to say, friends, it is so nice to be fully back in the back to school, back to work, out of summer mode. Relative to other industries, AI certainly has less of a summer slowdown, but you can still feel the difference, man, when September hits.
Starting point is 00:16:35 In the last nine days alone, we have gotten Fable 5.1, GPT6 Astra, and the three models and one agent product that we're going to cover in today's episode. I hope you are as excited as I am because there is a lot of new stuff to check out. First up, last week, right as I was leaving for vacation, of course, we got a new model from Google. It still is not a pro-series model, but it is notable how quickly Google is iterating on their smaller Flash series models. The new Gemini 3.8 Flash comes just a few weeks after 3.7. The central claim from Google around 38 Flash is that it will work harder than It's trained to call tools iteratively and perform more reasoning steps on complex tasks, yielding much better results.
Starting point is 00:17:23 On the benchmarks, the model looks solid if a little spiky. It scored 73.7% on coding benchmark deepswee, just a hair shy of Opus 5 score of 74%, and outperforming GPT-56 sole by 1%. Terminal Bench was another story. The model scored 89.4% on version 2.1, in line with Opus and Soul. However, the scores tanked on version 4.0 falling to 19.1% compared to, for example, Opus 5's 51.8%. On GDPVal, the scores were very middle of the road at 1545 Elo points, around 300 points shy of opus, and closer to Sonnet 5 and GBT 56 Terra.
Starting point is 00:18:02 Artificial analysis found the model was pretty solid on their benchmark run, scoring 59, slotting it in just behind GLM 5.3 in 7th place at the time of release, and only a few points off the frontier. However, this was the old formation of the intelligence index, the one that gave GPT6 Astra a fairly low score, prompting artificial analysis to rush forward their new version of the index, and once AA updated their formula, 3-8 Flash slipped from 7th to 12th place behind GPD 56 Terra and GLM 53 Flash. Now, the idea of this artificial analysis intelligence index update, was to re-weight, reprioritize, and add some new tests that better reflected the computer use and broader agentic paradigm, as opposed to just general knowledge tests, which are now
Starting point is 00:18:48 pretty much table stakes and saturated. Gemini Flash remains the undisputed leader in speed, outputting around 20% more tokens per second than runner-up new Spark 1.3, which if you're wondering what that is, we will get to in just a moment, and almost four times faster than GLM-53 Flash. The question is, of course, which use cases require that much speed at the cost of trade? in performance. Maybe the biggest bright spot from their write-up was cost, with artificial analysis writing that 3-8 Flash was, quote, the cheapest we've measured at this level of intelligence. They continued, this is up 40% from Gemini 37 Flash despite unchanged per token pricing, driven by a 30% increase in output tokens per task and more turns on agendic evaluations.
Starting point is 00:19:29 Unfortunately for Google, as we will see with the release of Muse Spark 1.3 the following day, Google would very quickly lose their place on the Pareto frontier. Certainly cost an event. efficiency is a big part of the pitch from Google. Announcing the new model that Google AI account wrote, while solving ambiguous and high friction tasks is immensely helpful, it can also be expensive. Fortunately, 3-8 Flash features the usual effort controls, ensuring that the amount of thinking required to accomplish the task at hand
Starting point is 00:19:54 is proportional to the token spend. Logan Kilpatrick from Google emphasized the speed at which Google is putting out these new versions, pointing out that it's just their third updated flash model in six weeks. Now, when it came to user testing, people validated that it was really fast, but found a lot of performance lacking. Building the same sticky ball game in Kimmy K3 versus 38 Flash, Aditya from Intelligence AI wrote,
Starting point is 00:20:17 Flash was insanely fast and used way fewer tokens, but the actual game was nowhere close. K3 had much better mechanics, movement, and overall game design. Flash clearly has the speed and efficiency part down, but the gap in what it can actually build is pretty big, wrote Ethan Mollick, It is a very good Flash model, but not equivalent to a frontier model. Others are more optimistic about what that speed could represent.
Starting point is 00:20:40 In another head-to-head, Noclip Pepe wrote, Opus 5 won, but Gemini 3-8 Flash was 39x faster. Opus 5 took 24 minutes, Gemini 38 Flash took 37 seconds. Opus is clearly more detailed and polished, no debate there. But getting a result this good in 37 seconds completely changes the trade-off. In the time Opus finished one run, Flash could theoretically finish around 39. At what point does speed matter more than the last bit of quality? And it will come, of course, as no surprise to anyone here, that as always, I think that the right
Starting point is 00:21:10 way to look at these new models is not whether it's going to replace your daily driver, but instead whether there are specific use cases for which its particular set of tradeoffs are the right fit. Is there something you're doing right now where being able to do it 39 times in a row to iterate is likely to be better than just letting something like Opus do it once? Now, the bigger other model release was Meta's MewSpark 1.3, and Meta chief AI officer Alexander Wang was not shy about promoting the progress that's been made. He wrote, This is our most capable model yet.
Starting point is 00:21:43 Frontier performance almost too cheap to meter, much stronger and agenic encoding with better usability. We think users will really notice the jump. Meta chose to compare their new model to GPT-56 Sol in Opus 5, and in that grouping it was pretty competitive. Generally, it lagged a little behind on agendic benchmarks, but was a little ahead on coding. Spark 1 3 scored a 75.4 on Deep Swee compared to 73 for 56 sole and 74 for Opus 5. On Terminal bench, it scored 88.8%, which tied it with 56 sole and beat Opus 5 by a couple of points.
Starting point is 00:22:15 Wang claimed that the model used 20% fewer tool calls and 25% fewer tokens compared to their previous Spark 1.2, while also producing a very clear jump on the benchmarks. A couple of days after the initial release, Meta also added a max effort setting that boosted performance even more. Now, the part of the release that really made everyone sit up and pay attention came when artificial analysis released their benchmark run. On Mac's setting, Spark 1.3 scored a 68 on the coding agent index, making it tied for first place with Opus 5. Now, worth noting that at the time testing for Fable 51 and GPT6 Astra hadn't been completed,
Starting point is 00:22:49 but still a pretty impressive result. On the overall intelligence index, Spark 1.3 on max setting scored 62, placing it in third place behind Fable 51 and Opus 5, tied with Fable 5 and a point ahead of GPT-56 sole. Now, the assumption for many is that this model was absolutely benchmarked max to achieve the maximum possible score. Certainly that's what semi-analysis argued, writing, Gemini 38 Flash and Mew Spark 1.3 are two of the most clearly bench-maxed models we've seen yet, despite being comparable to both GPT6 and Fable 51 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse. How is this possible? All of the tasks in
Starting point is 00:23:27 Terminal Bench 2.1 are fully public. Though Meta and Google would never train on the tasks directly, they absolutely will buy data from RL environment startups that's designed to mimic TB2.1 tasks as closely as possible. The net effect is the same. You'd typically expect improved TB2.1 performance to generalize to other agenetic tasks, but Gemini and Muse don't even generalize to TB4.0. Ultimately, semi-analysis concludes, this is the fate of all good public benchmarks. TB4.0 is no exception. It's only useful signal now because it was released two weeks ago. Since all the tasks are similarly public, it won't be long until it's hill climbed by all the aspiring, quote-unquote, frontier labs. Alexander Wang actually responded to that one saying,
Starting point is 00:24:06 we don't claim Mew Spark 1.3 is as strong as Astra or Faber 5.1. But it is significantly more cost-effective. Our future models will compete more directly with those models. It is worth also noting that after artificial analysis revised their index, no doubt in part because the original formula had Spark 1.3 outranking GPT-6 Astra, after the revision Spark 1-3 remained in fifth place with a score of 48, which was slightly ahead of GPT-5-6 sole and behind Fable 5, Opus 5, and Astra. And despite the big jump on the benchmarks, the model is still extremely cheap. A.A found that the model spent 55 cents per task, which made it slightly cheaper than Gemini
Starting point is 00:24:44 3-8 flash, 20% cheaper than GLM-53 and about a quarter of the cost of Opus 5. Summing up, Artificial Analysis wrote, MewSpark 1.3 Extra High is the most cost-efficient model at its intelligence level. No model scoring 59 or above costs less per task. Now, the first impressions on this one were pretty positive. Spack 89 wrote, I've been testing Mew Spark 1.3 max and honestly, it's insanely good and surprisingly efficient. Dara does code writes, MewSpark 1.3 is kind of ridiculous for a free small model on open code. It also avoids some of the obvious AI design traps like purple gradients and all that. I want to push it into nastier edge cases next, but for the price, this thing is already very good. Now on the topics that we were discussing in the headlines about trust in the labs,
Starting point is 00:25:28 for some meta is a tough one. After writing about Muse Spark 1.3 being good and basically Opus 5 for cheaper, Z win on Axe ads. There's a catch though. Meta may use your inputs and outputs for training, which is the whole reason the tier on open code is called contributor and the whole reason it's free. Certainly many people were excited to try the discounted open code version. With Dax from open code writing, Meta Muse Spark has dethrone Deepseek as the most used model of the day. First time an American model tops this list. And yet, if Muse Spark 1.3 made some start to wonder, is Meta back? It was the launch of their new personal AI assistant Muse that garnered even more attention. On Tuesday, September 8th, Meta announced their long-promised personal agent called Muse. The product
Starting point is 00:26:11 has been rumored to be in the works for months under the codename hatch, with the basic pitch that we had heard being open-cloth for normal people with a bunch of usability improvements. Presenting the agent Alexander Wang wrote, today we're rolling out Muse, our new personal AI assistant. Muse is always on, wicked fast, can use a browser, connect to your apps, and is designed to be secure. Meta claims that Muse can do everything we've come to expect from personal agents. It can triage your inbox, organize your calendar, make bookings, or shop for you. It also has some of the more impressive features introduced in recent months, such as operating a separate virtual computer, which is the same way that Grockbot works. Meta also made a solid attempt, it seems,
Starting point is 00:26:49 at dealing with the security nightmare associated with the earliest versions of personal agents. Wang again wrote, A big focus for us here was making sure it was safe to give Muse access to your inbox, calendar, and finances. Each muse runs in its own secure VM, an isolated computer dedicated to you. A separate system, the Sentinel, checks every action before anything leaves the VM. your muse never sees your actual passwords or card numbers. Now, certainly reviews from inside meta were glowing, with CTO Andrew Bosworth, aka Baas writing,
Starting point is 00:27:18 very excited for the launch of Muse today. I've been using it internally for months, and I am hard-pressed to think of any product that I've come to rely on more in such a short period of time. I have it linked to my email, calendar, and credit cards. I use it to help me plan travel, pack for trips, research, and make purchases, and sort through all the communication I get from my kid's school.
Starting point is 00:27:36 This is a tool for everyone. You don't need to be an expert or even think about AI. You just talk to it from the app or from WhatsApp like your own personal assistant, except it can do lots of tasks in parallel at the same time. And of course, you'll soon be able to talk to it from your meta glasses too. Jason Toff from Meta said, when I moved to California this summer, I unplugged my Mac Mini and Mac Studio, both running Clause locally, and switched entirely to Muse. They're still unplugged. My favorite thing about Muse is how natural it feels. You talk to it like a person, and it responds like a top-notch personal assistant.
Starting point is 00:28:06 And even from the outside, early reviews are pretty positive. EAC spiritual guru, Beth Jzos, writes, Got to try this product early. It's very solid and quite feature-rich. As model intelligence is no longer the bottleneck for utility, context on your life is, and personal agents running on secure compute is the way. Right, signal,
Starting point is 00:28:25 Muse has been a genuinely impressive product to play with. It has all the functionality of the I-Message agents in flight today, but also has all of the ingredients that will expose it to hundreds of millions including a massive friend graph through Insta, increasingly rich context from email and other services which you connect, and more importantly, stuff like Facebook Marketplace. Marketplace in particular is an incredible distribution wedge. Millions of normal people could encounter Muse simply because an agent helps them try to find something, negotiate the price, and arrange pickup. Pretty good execution here from FB. Olivia Moore from A16Z said that she likes the rich library of connectors that are available in app,
Starting point is 00:28:59 thinking that the native connectors will be more reliable than browser use, and also said that she liked that she can set up goals connected to that data and attach artifacts to them to visualize progress. She worried that the UI was still too cluttered and there was a few too many things to do, but concluded this could be one of the first true mainstream consumer agents to get adoption. Still, Meta has some hills to climb when it comes to consumer trust. That same Olivia Moore wrote, I was more reluctant to press the Connect email button on Muse than on 10-plus startup agent products I've tried.
Starting point is 00:29:28 In my opinion, meta's distribution advantage cuts both ways here. Do I really want to give an agent my personal data and then set it loose on networks where all my friends are? One of the things that makes meta interesting and worth paying attention to in the broader AI race is that they are the only company at their scale that is primarily focused on a consumer rather than a business use case. Now obviously this is all a little blurry, especially when you consider the legions of small businesses that use meta products as their key communication channels, but ultimately I think it's pretty uncontroversial to say that what meta cares about is consumers more than B2B. For a while, OpenAI looked like it was going after both, and nominally they
Starting point is 00:30:05 still are, but of course the pressure from Anthropic has meant that they have really had to focus a lot more resources on the B2B and work use cases of late. Especially as there are more and more questions about whether general consumers will ever really care about AI agents, these sort of experiments from meta have significance that goes beyond just them. Does agentic shopping actually become a thing? Do people really like having a personal agent assistant to help them with daily things like booking travel. We're not really going to know until those things are available and broadly good enough that they actually do what they promise, and it feels like meta is finally playing at the level
Starting point is 00:30:38 where that promise might be real. Wrights, Boxes Aaron Levy, personal assistant agents are going to be a very exciting AI category. It's the first time you can have high token volume agentic use cases that make sense for consumers. Lots of different approaches emerging right now, and it's going to be hypercompetitive because these agents will mediate a lot of consumers spend over time. But this certainly plays directly to meta strengths. Lots of compute required, can monetize with ads in commerce, software-focused experiences so can distribute it at scale, and so on.
Starting point is 00:31:06 Sums up why Combinator President Gary Tan. Harness Wars are full on now, and Muse is very impressive. Now, I don't have a horse in the harness wars or the model wars, but I will certainly be rooting for this as a product category. If for no other reason than people actually getting value out of a personal assistant agent, might make them just a little less hostile to AI. in the first place. Lastly, one more model release to talk about.
Starting point is 00:31:31 Also on Tuesday, OpenAI released ChatGPT Images 2.5. This is the latest in the series of models that power built-in image generation in ChatGPT, a feature which still gets a ton of use. OpenAI says that users are generating more than 3 billion images a week and write that the new model will provide sharper details, more precise editing, and faster generation with a 50% reduction in latency. Alongside the model, OpenAI is releasing a new ChatGPT feature called Sketch, which, as the name suggests, allows you to draw an input to help guide your image generation
Starting point is 00:32:00 directly in the app. Users can add a text prompt to describe a particular style or provide additional details to guide the model output. The model comes into variants, Flair, which is the fast version designed for quick iteration, and sunbursts, which is optimized for professional workflows that require better control across edits. And I think that that word control is really key here. In the same way that the big innovation and update of nanobanana was more fine-grained control over the editing process, that seems to be a big part of what OpenAI is going for with this new model as well. Exulton Alam Kulov, the head of product at Higgsfield, wrote,
Starting point is 00:32:34 What impressed us most about GPT Image 2.5 Flair is how well it understands what not to change. You can make a meaningful edit without losing the character, composition, or visual identity of the original image. That's incredibly important for the way creators and teams actually work across film, UGC, and advertising. And when you combine that level of control with the speed, quality, and cost, image 2.5 flair really stands out. One example that you're seeing a lot of, of what the new better controls and image consistency can lead to is entire new genres like stop motion animation
Starting point is 00:33:04 that become viable for the first time. One of the interesting things that I increasingly feel is I think right now in general, we underappreciate the value of images, not just as a consumer differentiator, but actually as a business use case differentiator for OpenAI. As a for example, while in general I still like the aesthetics of fable-created websites better than GPT-created websites, the fact that I can call upon GPT image to generate aspects of the UI or certain types of aesthetics makes a pretty big difference and leads me to use the integrated GPT models and image generation in codex more often than I otherwise would. Point is, although this update feels routine, don't sleep on how significant it could be.
Starting point is 00:33:44 So that is the new model story for now. Like I said, lots of exciting goodies to try out. And I'm sure there is more on the way. For now, that is going to do it for today's AID Daily Brief. Appreciate you listening or watching, as always. And until next time, peace.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.