The AI Daily Brief: Artificial Intelligence News and Analysis - Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models

Episode Date: October 1, 2026

Google’s Gemini 4 Argon impresses on benchmarks, but does it signal a real comeback? NLW examines Gemini 4, Claude Sonnet 5.5, and what Muse versus Dots reveals about choosing AI tools. In the headl...ines: Trump’s super intelligence accord, the America.gov AI portal, and an FTC investigation into OpenAI and Anthropic’s rogue agents.Next Cohort - Learn How to Build Agents - ⁠⁠https://register.besuper.ai/register?program=ati⁠⁠AIDB Fall Listener Survey - ⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.ai/survey⁠⁠⁠⁠⁠⁠⁠Multiplayer AI Sprint - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://multiplayerai.ai/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://kpmg.com/us/Sophisticated⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Harbor - Invest in the AI ecosystem. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.harborcapital.com/aidaily⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Hyperagent - Hire a team of always-on agents. New users get $100 in free credits. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠hyperagent.com/aidailybrief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.rackspace.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Section - Section turns AI investment into workforce transformation and ROI - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.sectionai.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The AI Daily Brief helps you understand the most important news and discussions in AI. Newsletter: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Interested in sponsoring the show? sponsors@aidailybrief.ai

Transcript
Discussion (0)
Starting point is 00:00:00 Coming into 2026, Google was looking pretty good in the AI race. 2025 had been a good year. A lot of the questions inside DeepMind had been answered. They were putting out competitive Gemini models. They were pushing forward with interesting new products. And all the natural advantages that they had always had in terms of consumer distribution and data and all those things had a lot of people very bullish on them as a contender. But then vibe coding happened.
Starting point is 00:00:23 And with the increase in coding capabilities paired with the power of the new harnesses like Claude Code and Codex, those new capabilities unlocked agents and a way that hadn't been possible before. And all of a sudden, Google found itself very behind. Charitably, you would say that in 2026, the company has been playing catch-up, but even that's not really accurate to what's been happening. For most of this year, Google has been firmly outside of the conversation as a top model lab. And yet I think that those who didn't have a partisan bias towards one of the other labs would never be fully comfortable writing Google off. This week, the company announced Gemini 4, their first new frontier model in more than six
Starting point is 00:00:56 months. By the benchmarks, it looks like Google is so back. But is that the whole story? So let's dig in to Gemini 4 Argon and where it lands in the current AI race. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, robots and pencils, Harbour, and Granola. To get an ad-free version of the show, go to patreon.com.com slash AI Daily Brief, or you can subscribe on Apple Podcasts. To learn more about sponsors, during the show, send us a note at sponsors at AIdailybrief.aI. Also note that we have our next cohorts of superintelligence agent training coming up. These are paid programs, which in very short order
Starting point is 00:01:41 will get you far ahead when it comes to your agentic understanding and your ability to use agents in your daily work. We have both the executive catch-up program and the agent intensive, which we call the executive agent leadership program. The next cohorts for those start next week, and you can find links to all of that at the very top of AIdailydief.aI. President Trump really looked at OpenAI Dev Day and said, nope, absolutely not. I don't want that to be the biggest thing happening in AI this week and invited basically every big AI CEO to the White House
Starting point is 00:02:09 for what David Sachs would later call the Bretton Woods of AI. Now, there was a lot of chatter that came out of this meeting. But one of the first things that people noticed was that Trump pulled Anthropic CEO Dario Amade to be the one to speak to the press following the meeting. Now, Dario insisted that he's just saying what he's always said. that AI has some incredible benefits and some very real risks. However, in an act which some saw as public deference, others saw as Dario finally playing the game, he commented, as the president has said,
Starting point is 00:02:37 whoever wins AI wins, I think that's very important. The mechanism, how we address the risks, is still under discussion. But we all need to work together to make sure that we can win and we can win safely. If we do this right and work with the president and everyone here, we can all win safely. Now, X was absolutely filled with people psychoanalyzing the moment, breaking down the expressions from other tech leaders, Dario's anxious tics, and even dissecting his fashion sense. While some thought that Trump and the other tech leaders might be bullying Dario as he stepped forward to speak, the president followed up with some kind words later in the press conference, saying, Dario has been fantastic. We had dinner the other night and he agrees with everyone.
Starting point is 00:03:14 Whether or not this was Dario truly bending the knee, it was a significant moment for a CEO who previously told staff that Anthropic had been targeted for refusing to give, quote, dictator-style praise to Trump. At least in the immediate wake, the press conference completely overshadowed the subject of the meeting, which was the head of the frontier lab signing an accord on superintelligence. The accord is a one-page commitment to develop internal controls for frontier model testing and deployment, as well as partnering with external auditors for verification of those controls. Trump told the press that the document is, quote, morally binding, suggesting a continuation of the voluntary approach to AI safety testing. Trump added,
Starting point is 00:03:50 The biggest people in the world signed that and I signed it as president, and it really is a form of protection. He also floated the idea of a 10-member industry oversight committee during the press conference, but for now seems satisfied commenting, I'm seeing tremendous self-policing. The general tone, at least among the people at the meeting, was that this was an important first step. Mark Zuckerberg told the press, the idea isn't that this is the only thing we will ever do, it's that this is a start and an accord the whole industry could come to. Mostly the media response to this was a complete Roershack test about how one feels about Trump and how one feels about the tech CEOs, i.e. whether you're inclined to use the word oligarchs to describe them. But I do think that the skeptical take on this one
Starting point is 00:04:28 might be best summed up by media commentator Chuck Todd, who wrote, if you trust the tech companies to police themselves and you want them to decide how AI will impact society without any input from the public, then you'll be fine with this accord. Trying to find the AI moderate or AI realist position in this one, I think a person could agree that it's better that these folks are frequently in rooms together with key leaders in the government rather than exclusively viewing their role through the lens of competition. while also wanting to see more external involvement, not least of which an articulation of what the independent external auditor or evaluator to carry out independent assessments is actually going to look
Starting point is 00:05:03 like in practice. Perhaps now that this accord has been signed, that particular detail is one that can get built out. Following the lunchtime meeting, President Trump signed a pair of executive orders marking the beginning of the superintelligence age. The first made Trump's name change official, stating, as these capabilities continue to improve, they increasingly represent not merely artificial intelligence, but a new era of superintelligence. The terminology used by the federal government should reflect the transformative capabilities of these technologies, and the limitless opportunities they create for the American people. Much more functional was the second order, which established America.gov as an AI-powered single portal for government services. Trump said,
Starting point is 00:05:41 With America.gov, the federal government no longer stands in your way, it only stands at your service. We're simplifying it, we're glamorizing it, we're making it what it should be. Alongside that second In executive order, we got an Apple-style product unveiling keynote, complete with Secretary of State Marco Rubio playing the role of Steve Jobs. Rubio showed demos of planned future functionality, like being able to use the portal to apply for a passport, change your name after getting married, or enroll for Medicare. The features kind of seem like an MCP for government. A user can provide their information once, then the America.gov agent will track down and complete all the necessary forms across numerous government departments. How well this works remains to be seen, but certainly
Starting point is 00:06:18 it is a real problem. The insane difficulty, for example, of changing your name after getting married, depending on which state you live in, is hard to overstate. Rubio said that the functionality will be available from next year, with the executive order giving government departments 90 days to integrate their services into the portal. For now, America.gov only functions as a knowledge base for government services, searching across 29,000 government websites to answer questions based on official sources. Now, a lot of the chatter on this centered around security and privacy concerns, particularly after the president said, it cannot, in theory, be hacked into. And when they figure out a way to do that, we'll end it, because they always figure out a way, right? The National Design Studio, which designed the website,
Starting point is 00:06:56 came under fire in June for including activity tracking tools and other websites they've designed. And while some raised questions about providing the personal information required to apply for government services, AI jailbreaker, Pliny the Liberator, poked and prodded at the chatbot and found its guardrails were pretty airtight. He wrote, interesting, the America.gov chatbot will flag anything that appears to be personal information, like any text vaguely resembling an address or password, and refused to accept the query until the flag text is removed. Never seen that before. Meanwhile, if you think that them showing up together at this meeting means these companies are going to escape all government scrutiny, it might be worth noting that the FTC has opened an
Starting point is 00:07:31 investigation into OpenAI and Anthropics rogue agents. According to agency sources speaking with the New York Post, the FTC opened an investigation into the two leading AI labs a few weeks ago. The scope will cover the hugging face incident and the dozens of so-called rogue agent events that have been disclosed since. The investigation has advanced to the stage that the FTC is drafting civil investigation demands, which are similar to subpoenas. An FTC source said, The agency has plans to compel the executives at these firms to testify about their product and about the dangers they allege their products may have to consumers, to Americans. Reports state that third-party safety research lab meter can expect to demand alongside the
Starting point is 00:08:05 frontier labs. Now, from very early on in the AI narrative, one big threat, of discourse has been focused on the idea that new AI-specific regulation wasn't necessary. That government could simply enforce product safety legislation that's already on the books. Former FTC chair Lena Khan was of this view, writing last month, law enforcers already have authority to charge companies and their CEOs for creating and releasing dangerous, unvetted, or defective products. We shouldn't let discussions about new legal regimes distract from the fact that there's no AI exemption from laws already on the books. This post was co-signed by former AI czar David Sachs, in yet another example of the AI safety debate
Starting point is 00:08:39 creating very strange bedfellows. Still, based on comments from the FCC source, it's not clear that the agency is necessarily pushing for some sort of harsh penalty. They told the New York Post, we need to win this SI race, absolutely, and we are winning, and that's fantastic. Getting a political jab in, the source continued.
Starting point is 00:08:55 And obviously the other side, the Democrat Party wants to destroy this technology, wants to surrender to our enemies and competitors, and that's not something that the chairman's interested in doing, and that's not really something the president's interested in doing. With that being said, the laws have to be followed. People talk about how we need new laws for these companies, and the chairman's been very clear, we have plenty of laws on the books. So an investigation, that's good, right? That'll appease people who think that the labs are getting a free pass.
Starting point is 00:09:18 Eh, that's not exactly what the discussion suggests. Anti-monopoly activist Matt Stoller wrote, It's a protection racket. The idea is Trump investigates and clears them so other investigators have hurdles in their probes. Former Director of Public Affairs at the FTC, Douglas Farrer, wrote, It would be good if the FTC investigated these companies properly, better starting two years ago, but the public credibility of this new investigation is 0.0%, especially after yesterday's AICEO and Trump Lovefest. The optimistic take comes from Joel Thayer, who writes, I kind of love this trust-but-verify strategy. Even though it's mostly self-regulatory,
Starting point is 00:09:52 if any signatories fall short of these commitments, it could trigger FTC enforcement under Section 5 of its UDAP authority, makes FTC Chairman Andrew Ferguson's presence at the meeting very apt. As with anything policy right now, it's worth having a whole bowl of salt when you look at this, but things continue to move in some sort of direction. For now, that's going to do it for the headlines. Next up, the main episode. If you're leading AI inside an enterprise, you already know that the gap right now isn't capability but execution.
Starting point is 00:10:23 That's why KPMGs You Can with AI is back with a new season featuring conversations with leaders like Sorosia Chatterjee of Emma, Mahabeeva of writer, McKesson, CIO, Elery and others focused on practical execution. What's working, what's not? and what it actually takes to move from pilots to real scaled impact, across strategy, data readiness, governance, workforce, and value. And of course, it's co-hosted by me, Nathania Whittimore. Go listen and subscribe at www.kpmg.us slash AI Podcasts.
Starting point is 00:10:51 That's www.kpmg.org slash AI Podcasts. The best teams don't have a single star carrying everyone else. They know their own strengths and each other's weaknesses and play to both. That's the team robots and pencils has built on purpose. Nobody there is grinding through busy work to pat a headcount number. People come for the hard problems and they stay because everyone around them is leveling up at the same time. In a market full of companies that are just trying to hire fast, that's worth a look. Check out robots and pencils.com slash careers.
Starting point is 00:11:20 If you listen to this show, you likely have a thesis. Maybe it's enterprise adoption, maybe it's compute. Maybe it's a specific lab. Harbor Capital's AI Lab ecosystem ETFs let you express it via five actively managed GTFs, each seeking exposure to the ecosystem around one major lab. Anthropic, Open AI, DeepMind, meta, or SpaceX AI. Your view of the AI race in ETF form. Harbor Capital Advisors AI Lab ecosystem ETF suite gives investors a way to invest in the AI
Starting point is 00:11:46 ecosystem they believe is best position for success. Search Harbor AI Lab ecosystems, ETFs wherever you invest or follow at Harbor Capital on X to learn more. Visit Harbor Capital.com for a prospectus containing investment objectives, risks, fees, expenses, and other important information. Read and consider it carefully before investing. risks include principal loss and artificial intelligence-related risks. Harbor UTS are distributed by Foreside Fund Services LLC.
Starting point is 00:12:07 Harbor is not affiliated with AI Daily Brief and the funds are not affiliated with sponsored by or endorsed by any AI lab. This is a paid advertisement and not personalized investment advice. Investing involves risk, including possible loss of principle. When I'm in a meeting, I'm fully in it. I'm thinking about the iteration and creative back and forth it takes to actually push a goal forward. What I'm not thinking about is capturing takeaways, tracking to-dos or any of that.
Starting point is 00:12:31 And that's where Granola comes in. Granola is an AI-powered notepad that captures what happens in your meetings and turns it into clean, structured notes, with the decisions and action items pulled out and easy to find. There's no setup and no configuration. It just fits into how you already work. For me, it means I get to stay in idea mode, and Granola make sure those ideas actually become action. Once you try Granola on a first meeting, it is hard to go without. You can try it totally free at granola.a.i slash brief.
Starting point is 00:12:57 That's granola.a.i slash brief to get your time back. Welcome back to the AI Daily Brief. And friends, it appears that pigs are flying, hell has frozen over. Choose your metaphor for incredulity because after months of waiting, Google has announced Gemini 4. Although maybe those pigs aren't completely off the ground because although they announce the model, we don't actually have access to it yet. On Wednesday, Google unveiled Gemini 4 Argon. It's their first frontier model in over six months. Now, the 2026 struggles of Google have been well documented.
Starting point is 00:13:34 We never got a Gemini 3.5 Pro, with the company basically deciding that it just wasn't good enough to release. And for a few weeks now, speculation has been coming that Gemini 4 was on the horizon, and Google DeepMind CEO, Karee Kovasoglou, says that the model delivers. Quote, frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. Now, despite a long period where Google's future as a frontier lab has seemed somewhat tentative, to judge only by the reported benchmarks, Google looks to be right back in the game. Gemini 4 Argon is the new state-of-the-art model across many benchmarks.
Starting point is 00:14:11 In agentic knowledge work, Gemini 4 scored 68.9% on the VALs index, beating Fable 5-1, GPt6 Astra, and exceeding the score from the current leader Opus 5-5 by a couple of percentage points. The gap was even larger on Automation Bench and VAL's finance agent. And on Harvey's legal agent benchmark, Gemini 4 more than tripled the score of current leader Fable 51, scoring 19.6%. Agentic coding scores were a little more patchy, but Gemini 4 is definitely in the ballpark of
Starting point is 00:14:38 Frontier rivals. It's the new state of the art on Deep Sway scoring 77.9% compared to Opus 55%, 74.2%. On Frontier Swee, Gemini 4 scored 55%, which placed at 10 points behind current leader Astra 6, 7 points behind Opus 55, and 1.3 points behind Fable 51. On Frontier Swee, however, Gemini 4 scored 55%, which placed at 10 points behind current leader Astra 6, 7 points behind leader Astra 6, seven points behind Opus 55 and 1.3 points behind Fable 51. It was also bottom of the pack on Terminal Bench 4.0 with a score of 57.4%, 9 points behind current leader Opus 55. Now, you can look at this in two ways. One, agented coding is one of the most important, if not the most important use case, so failing to achieve a state-of-the-art score is kind
Starting point is 00:15:21 of problematic, or you can look at this as a huge improvement for Google, which it undeniably is. Coding had been Gemini's biggest weak point for the past year, so to have Gemini 4 be in the same ballpark as the frontier models from OpenAI and Anthropic absolutely represents a big catch-up. For computer use, Gemini 4 is nipping on the heels of Astra, which set a new standard for the category. On OS World, it scored 69.2% against Astra's 72.6%. Artificial analysis confirmed that Google is back in the mix as a leading model developer. Gemini 4 Argon scored 53 on the Intelligence Index, putting it tied for third place with GPT6 Astra and Fable 51, one point ahead of GPT-6-1 sole, and three and five points behind Sonnet 55 and Opus 55, respectively.
Starting point is 00:16:02 In terms of cost and efficiency, it's pretty close to the Pareto frontier, costing $199 per task on the AA benchmark run compared to $326 for Astra and $7.63 for Fabo 5.1, however, that is with Google's 50% launch discount, which will be available for an unstated period, meaning that the full price would make it more expensive than Astra. It's also in that uncomfortable middle ground where it's more than twice as expensive as GPT's 1-1 sole, even with the discount, without a super clear improvement on the benchmarks. For the moment, however, none of this really matters, because unfortunately, Google is not making Gemini 4-Argon generally available to the public. Their stated reasoning has to do with cybersecurity. The model
Starting point is 00:16:43 scored 68% on the C-E-Benzhe cybersecurity benchmark, tied with GROC-4-7 and GPT-6 Astra, and slightly ahead of Opus 55-F, Fable-1, and Mythos 5-1. This was reasoned enough, according to Google, to limit access to what they call a set of trusted cyber defenders to ensure the model is not misaligned. They said that they're engaging with the U.S. government's voluntary testing and pre-release access program and intend to gradually expand access over time, beginning with API and Google Ultra subscribers. In their blog post, Google said that the model is already powering internal workflows and boasting strong performance on long horizon coding tasks like large-scale code-based migrations. Now, as to the choice to announce the model without actually releasing it,
Starting point is 00:17:22 I'm not really sure. Honestly, for us, old heads who have been watching this for a while, kind of has echoes of the first time Google fell behind and felt the need back in December of 2023 to announce Gemini, even though the pro version of the model wouldn't be available for a number of months. Yet maybe it was a good call because there were plenty of people who were excited to see it. There was an entire genre of posts on X that basically came down to Google, so back. Dr. Daria Anut Maz wrote, Wow, wow, Google Deep Mind is back with Gemini 4 Argon, insane benchmarks.
Starting point is 00:17:51 In one fell swoop, they've jumped to the very top of the AI frontier. As I've said before, never bet against Google in the age of AI. Although he did add, and this seems to me to be the key detail, Can't wait to try Gemini 4 ASAP. Ethan Mollick writes, And it's a three-way race again. Nathan Lambert, who does open model research, and is about as far away as you could be from a hypester,
Starting point is 00:18:11 writes love to see Google surprising people with Gemini 4. More labs at the frontier is wonderful for consumers because of competition, and wonderful for the world because of reduction in concentration of power. Excited to see how it goes in real-world scenarios. AI leaker, I rule the world, writes, I've been critical of Google models, but Gemini 4 looks strong on benchmarks. We've seen this before with Google, but I'm optimistic they can release a first-class model with a first-class harness like Codex.
Starting point is 00:18:35 Great work to all involved with Gemini 4, rooting for you. But then on the other side of the coin is Angel who writes, I'm sorry, but I can't with all these Google-is-so backposts. We literally can't even use the model, and it's not like benchmark scores for Google's models have ever been reliable. Do we really need to remember what happened with Gemini 3 Pro? No? Seriously? Ishu Agrawal put some numbers around that, sharing an older series of benchmarks and adding,
Starting point is 00:18:58 by the way, these were the official benchmarks for Gemini 3.1 Pro when it released, practically destroying Opus 46 across the board. I'm not saying Gemini 4-Argon will be bad, but it's crazy that no one has the slightest bit of skepticism given Google's track record. To their credit, Logan Kilpatrick, who honestly should get Google's MVP for hanging in there throughout all of these PR challenges this year, actually engaged with this writing, we've gotten much better at testing our models at scale across Google now. So assume most new Gemini revs go through thousands of software engineers for weeks before getting released,
Starting point is 00:19:29 hopefully has helped close the benchmark to reality gap by a real margin. Now, so far, all we have to go on is the reported benchmarks, along with some reports from grumbling inside the company. A Bloomberg piece called Google grapples with employee skepticism about New Gemini 4 says, while Gemini 4 has performed well on benchmarks widely used to gauge model efficiency, it does less well when employees actually put it to work. model struggles to handle certain coding tasks according to people with direct access to the model. Google, for their part, denied the reporting, commenting that it would be inaccurate to claim
Starting point is 00:19:58 that Gemini 4 is underperforming in some areas such as coding. And one of Bloomberg's other sources said that these grippers are a minority, claiming that a, quote, large consensus internally at the company believe that Gemini 4 is the frontier. What makes the position of Google so challenging is that in the time that they've been away from the top, the nature of the race has fundamentally changed. With much optimism, Peter Yang wrote, Google cooked on Gemini 4. Now they just need to be more competitive on coding harness, anti-gravity, and personal agent, Spark. In other words, the competition has become more than just models. And it's not like there already isn't a ton of model competition. One model released in advance of OpenAI Dev Day that we didn't have a chance to talk about
Starting point is 00:20:39 yet is Claude's Sonnet 5.5. And it's certainly if you feel yourself giving a bit of an eye roll or at least a raised eyebrow wondering if you should actually care, the recent history of the Sonnet and Opus series certainly created some reason for skepticism. Opus 5, you remember, was strongly disliked, and Sonnet 5 seemed to deliver even worse results for a higher price because it was so token-hungry. However, if you had to pick just one theme from the last couple of weeks on this show, it's the beloved return to form for Anthropic with Opus 5-5. So does Sonnet 5-5 continue that trajectory? The short answer seems to be in most people's estimation, yes. Anthropics said that Sonnet 55 is 30% faster and 30% cheaper than Sonnet 5.
Starting point is 00:21:18 And the model demonstrates some significant jumps on the benchmarks. For Agentic coding, it scored 70.6% on Terminal Bench 4.0, which is up from just 10.3% for Sonnet 5, and even beats Opus 55 at 66.4%. Scores on Frontier Code and Curser Bench were very slightly behind Opus 55. The same was true for knowledge work benchmarks, GDPVAL, AA, and NAA briefcase, where Sonnet 55 closed the previously massive gap between Sonnet and open-class models. Artificial analysis affirmed the benchmarks, giving Sonnet 5-5 an intelligence index score of 56 that puts it in second place behind Opus 55 and somehow ahead of Fable 51 and GPT-6 Astra.
Starting point is 00:21:59 However, they didn't find that Anthropics claims about cost-saving stood during their testing. Sonnet 55 cost $7.60 per task and used significantly more tokens than Sonnet 5 at the same per token price. This meant that Sonnet 5-5's benchmark run was almost as expensive as Fable 5-1, 27% more expensive than Opus 5-5, and more than twice as expensive as GPD6 Astra. Now, to give Anthropic the benefit of the doubt, they claimed a 30% cost reduction on similar tasks compared to Sonnet. The main AA benchmark is run at max inference settings, but turning down the settings to extra high cut the cost by two-thirds. Still, the model impressed once people got their hands on it. It is a massive improvement over Sonnet 5 in the rendering tests that grab attention
Starting point is 00:22:41 on social media. Matthew Berman made a series of visual tests and game clones writing. Sonad 55 basically Opus 55, but 50% cheaper and much faster. I've been early testing it and it's incredible. If this is pacing the frontier, sign me up. But for real-world use, Builder Kun Chen found it performed best when it's the right tool for the job. He wrote, Sonnet is definitely not the same as Opus just 50% cheaper. That's true only for problems that don't need much wisdom. There's a clear difference when they are asked to propose plans for ambiguous product problems. Opus is able to approach problems with more strategic thinking, i.e. what's the real goal here? While Sonnet is more just looking at tactically, how do I get this done? Now, Kun's current approach is to use anthropic models as a complete
Starting point is 00:23:24 system, with Opus for planning, Sonnet for implementation, and calling on Fable when things go sideways. Pavel Huron tested Sonnet on his bug fix benchmark and found it outperformed all other models, including Fable and Astra. He wrote, The Secret? 5.5 Max might be cheap and fast for most tasks, but it's the least lazy model I tested. It leads in bug hunting by spending turns. Fascinatingly, YouTuber and entrepreneur Theo's review was that Sonnet 5 is an incredible model that you probably shouldn't use. In his view, the model is a big improvement over Sonnet 5, but on Max reasoning, it overthinks and risks getting things wrong because of it. And on lower settings, there's no point where Sonnet is more cost-effective
Starting point is 00:24:05 than Opus because of how token-hungry it is. This changes a little when using Sonnet as a agent, where it can be more efficient on certain tasks, but it's usually not optimal as a standalone model. So will these results stand after a couple weeks of testing? With Sonnet 55 being a technical marvel given that it can produce the results of Fable 5, just four months after the release, but being too token hungry to make sense for most? That's certainly what I'll be watching and I'll report back as people get more reps in. Still, going back to the new world that Gemini 4 has to compete in, the point that Peter Yang again was making is that just being good on the model isn't enough. In fact, one really interesting question brought up by the success of Muse is whether
Starting point is 00:24:43 an AI product with a great user experience but a less than state-of-the-art model beats a product with a less good user experience but a state-of-the-art model. Mews has hit 3 million weekly active users. Daily users, who send at least one prompt per day, have reached 1 million. Now remember, the product has only been available for a little over three weeks, so this is a strong, positive early indication. At the end of the first week, Muse had 500,000 weekly active users and 250,000 daily users. Now, on the one hand, this is a tiny fraction of meta's distribution potential, which reaches half the people on Earth, and is nowhere near as strong as ChatGPT's growth, but a better comparison might be the adoption curve for Codex. Codex took about three months to
Starting point is 00:25:24 reach 3 million weekly active users, so Muse currently looks like a faster-growing AI app. Now, one might expect that, given that it is a personal agent versus Codex, very developer-focused audience, but whatever the case, the early success of Muse continues on. And yet now that OpenAI's dots is here, we're going to start to have some direct answers to this question of good UX plus worse worse Ux plus better model. Urika co-founder Sina writes, After using Muse for a while and loving it, I'm starting to notice the underlying model isn't smart enough. I'm not sure it's context overload, but it gets stuff wrong, a lot recently. Maybe they've had such huge success that they've had to switch to a different cheaper,
Starting point is 00:26:04 worse model. Anywho, Mew, Mews is great, form factor is great, but as a consumer, my loyalty is to none. If OpenAI will drop MUSE with better models, I will use whatever is best. Intelligence is still the moat. And yet on the flip side, a couple days in, I'm definitely seeing some gripes with the way that Dots works. Habibi Slop on X writes, Overall, the vibes are not good.
Starting point is 00:26:24 This feels directionally wrong. I just want to talk to the model. This tool wants you to be a plumber in a house where you're not allowed to touch the pipes. Dots has so much friction built in that it feels like it's not meant to be used. And unless you think that this is just a general griper account, they then go on to include about a dozen examples. The big one that I've seen repeated comes in their final thoughts when they write, they've launched a new product that wants to help you manage your work in chat GPT and codex without having sufficient access to any of the tools and data it needs from those
Starting point is 00:26:49 surfaces. I am certainly at least not even close to being willing to declare, which is the winning strategy between focusing on the model versus focusing on the product user experience. One last note as we look across the state of the competition that Google finds itself in as it gets ready to release Gemini 4. In one surprise launch this week, DoorDash has announced that their AI agent is getting the Mews treatment and can now take orders via text. DoorDash already had an agent on their site, but they're using some of the features of personal agents like Muse to take it a step further. Customers can now text the DoorDash agent with a natural language prompt, and the agent will figure out the rest. DoorDash gave the example of a customer texting,
Starting point is 00:27:24 order my usual protein bowl to the office. The agent then matches the phone number to their user profile, looks up the usual order and delivery address, then handles the payment. Now, in light of Amazon's decision to block Mews agents, there's a huge open question of whether the best agent strategy is to partner with the category leaders or build a platform-specific agent. The logic for an agent from DoorDash seems pretty clear at this point. The company has said that agentic orders have a 50% higher basket value when buying groceries. And by building their own, DoorDash can attempt to hold the user relationship close. There's a risk, of course, that external agents like Muse could order directly from restaurants and cut DoorDash out as a middleman. At the
Starting point is 00:28:00 same time, user preference might reject platform-specific agents because they have much less control. It's unclear what the optimal strategy will be, but for now, DoorDash seems to be doing a bit of both, keeping their platform open to third-party agents while also building an internal agent, and for that reason it's worth paying attention to, even if you don't much care who wins, how you order lunch to the office. Anyways, bringing it back to Gemini for Argonne, I'm with Nathan Lambert, when he argues that Google doing well is both better, for consumers and better for competition. So put me firmly in the camp of people who are hoping that this release is great. Let's just hope we get it soon so we can actually decide for ourselves,
Starting point is 00:28:34 rather than dining from the table scraps of self-reported benchmarks and gripey employees talking to Bloomberg. That's going to do it for today's AI Daily Brief. Appreciate you listening or watching, as always. And until next time, peace.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.