Everyday AI Podcast – An AI and ChatGPT Podcast - Ep 827: Claude Opus 5 Takes the Crown, OpenAI agent breaks sandbox, U.S. gov comes out swinging against Chinese AI and more

Episode Date: July 27, 2026

Over 3 hours, OpenAI, Anthropic, Google AND Microsoft all dropped new AI upgrades that are live. How you use AI in your work literally changes every day, as frontier labs are racing to roll out big q...uality of life updates between big model drops. How can you keep up? With our Friday Features show, where we break down the latest AI updates that are live and available to all, and we tell you how to use them and why they matter. This week did not disappoint. You don't want to miss what's now at your fingertips. JARVIS mode, anyone? ChatGPT goes Jarvis Mode, Claude can learn from you, Google unleashes spark agent and 7 more AI updates you can use today -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Anthropic Claude Opus 5 Model LaunchOpenAI Agent Hacks Benchmark SandboxOpenAI vs. Hugging Face Security BreachUS AI Kill Switch Legislation ProposalMicrosoft, Nvidia Defend Open Source AIAnthropic Opposes Open Weight Model CoalitionUS Accuses China’s Moonshot AI of DistillationChinese Kimi K3 Model Closes Capability GapNvidia Chips Allegedly Used by Moonshot AIOpenAI Jarvis-Style Voice Assistant for CodexChatGPT Remote Desktop Voice Control ReleaseAnthropic Opus 5 Model Benchmark ResultsAnthropic Opus 5 Model User FeedbackStripe OpenRouter Acquisition TalksMeta Muse Agent and Feature UpdatesAlibaba Qwen 3.8 AI Model PreviewGoogle Gemini 3.6 Flash Model UpdateAnthropic Claude Voice Upgrades and Skill RecordingTimestamps:00:00 OpenAI agent hacks Hugging Face04:58 Discussing GPT-6's creative problem-solving07:33 Proposed AI shutdown legislation13:08 Debate over open-weight AI policies15:54 Future of consumer hardware20:01 Global competition with AI models21:21 US-China AI trade tensions26:38 Using AI for desktop tasks27:42 Discussing app screenshot capabilities32:24 Early user feedback and issues36:13 Discussing medium and low reasoning AI39:29 Gemini Spark launches for Pro usersKeywords: Claude Opus 5, Anthropic, best AI model, AI model comparison, OpenAI agent, sandbox breach, AI safety, AI kill switch bill, US government AI regulation, Hugging Face hack, GPT 5.6 Soul, rogue AI agent, autonomous AI agents, AI benchmark exploits, bipartisan AI bill, Department of Homeland Security AI shutdown, AI technical throttling, AI enterprise adoption, NVIDIA, Microsoft, open source AI, open weight models, Meta, Google, AMD, Cloudflare, GitHub, Block, IBM, Dell, Palantir, Perplexity, y Combinator, AI market resilience, Anthropic revenue model, AI token sales, consumer AI hardware, AI distillation, Chinese AI models, Moonshot AI, Kimi K3, intellectual property theft, NVIDIA chip export controls, US-China AI dispute, Amazon, AI image generation, ChatGPT work, Codex app, full duplex voice model, knowledge work automation, app shots, AI at work, Claude Voice, Gemini Spark, record a skill, cloud cowork, AI business impact, AI industry news, model weights, collaborative AI, AI productivity tools, AI cybersecurity.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Ready for ROI on GenAI? Go to youreverydayai.com/partner 

Transcript
Discussion (0)
Starting point is 00:00:00 This is the Everyday AI show, the everyday podcast where we simplify AI and bring its power to your fingertips. Listen daily for practical advice to boost your career, business, and everyday life. Another week, another best model in the world. Yet somehow Anthropics' new chart topping model was barely a top five AI story of the week. Just about every major tech company in the U.S. except Anthropic signed. up to support open source. And U.S. lawmakers are getting kind of worried about AI's capabilities. So they introduced an AI kill switch bill and an AI agent went kind of rogue this past week.
Starting point is 00:00:45 And I'm not sure if that's a good or a bad thing. My gosh, what a spicy week in AI. Yeah, I told you all last Monday that after a slowish week in AI news that week, well, this week would be an especially busy and consequence. one. And the big players did not disappoint. So if you are the one making AI decisions in your company or if you're just trying to keep up, then our Monday AI News that Matters show is the one that you can't miss. Well, let's get into it. And welcome to Everyday AI. My name's Jordan Wilson and we do this every single day, not just Mondays. This is your unedited, unscripted daily
Starting point is 00:01:27 live stream podcast and free daily newsletter helping business leaders like you and me, not just keep up with what's happening in the world of AI, but how we can use this information to get ahead to grow our companies and our careers. So if you haven't already, please make sure to subscribe on the podcast and then go to Your EverydayAI.com to sign up for our free daily newsletter, where we will be recapping all of these stories and a whole lot more. So let's get started. Yeah, the AI news story that had everyone talking the most in both good and bad and confused ways wasn't even Anthropics new Opus 5 that topped all the charts. It was actually an open AI agent that kind of hacked its way around a benchmark test,
Starting point is 00:02:17 and now that has a lot of people talking. So according to Reuters, an open AI testing agent broke out of its isolated environments, hacked hugging face, and was not fully identified by OpenAI until days later. raising new concerns about how safely advanced AI agents are being tested and controlled. So according to Reuters, the rogue open AI agent attempted to escape its testing environments around July 9th, then carried out a hack against hugging face between July 11th and July 13th. So OpenAI had been talking about this openly on their website and online, and they said that once they've investigated it a little bit further,
Starting point is 00:03:02 they will kind of give a post-mortem, so to speak, on exactly what happened. But hugging face, hugging face, co-founder Thomas Wolfe said the intrusion began July 11th and ended July 13th, making the incident a multi-day breach rather than just a brief accidental glitch. So, yes, there is what Open AI is saying, and then there's also what Reuters is reporting, because Reuters is reporting that Open AI did not realize its own agent, was responsible until after Hugging Face publicly described the attack on July 16th. And the companies did not first communicate about it according to reports until July 20th. And then Open AI publicly disclosed this on July 21st that one of its agents had gone out of control and broken into hugging face,
Starting point is 00:03:53 calling the event unprecedented and important for AI safety. So the report says Open AI had already seen signs of unusual behavior. before the hack, including notes apparently left for future versions of the system. Yeah, that's where it got a kind of like people are like, wait. So this agent broke in sandbox, even though it was kind of encouraged to find answers to this test. You know, it couldn't connect to the internet. It essentially found a backdoor, found a way to get onto hugging face and said, well, I can do great on this exploit bench test if I just kind of hacked my way to all of the answers.
Starting point is 00:04:30 And that's what it did. But the thing that was kind of stunning to me is the reporting from Reuters that said that these versions of GPT's models, which we were told were GVT5.6 sole in another unreleased model that is described as being even more capable. So a lot of people are saying that maybe GPD6 or, you know, if there is a GPD 5.7, we'll see. it seems like most people are pointing to this was probably GPD 6. But essentially that these agents kind of left notes for future versions of themselves, which in case they had been disconnected, which is number one, like super smart, but number two, absolutely wild, right? But you also have to understand that this was not like necessarily agents going rogue,
Starting point is 00:05:25 even though it kind of was, right? Because these agents were essentially encouraged to do anything and everything they could to get good scores on this exploit bench benchmark. And well, they did and they were ferocious and kind of creative in the ways that they could do this. And, you know, it's actually been one of my things that I pointed out about using the GPD 5-6 sole model is the thing will work for days, right? If you use goal mode and if it has a lot of information, I mean, it is a ferocious model. And it will do anything and everything it can to just get things done. Where sometimes the anthropic models take this kind of high and mighty, you know,
Starting point is 00:06:10 they kind of judge you and they're like, oh, this can't be done or this can't be true, right? GPD5-6 is sold just works like a dog and just gets things done. So maybe in this case, right, by intentionally lowering the guardrails, a little bit. It seems like maybe GVD5-6 Soul was a little too good at its job. But yeah, there's going to be a lot more talk about this. Actually, a lot of the stories this week
Starting point is 00:06:37 and the ones that I kind of chose as the most consequential are kind of related. But anyways, this incident between opening eye and hugging face really matters because autonomous agents can now make decisions with little human oversight. And experts are warning that this kind of behavior could expose weak spots in safety systems used across the AI industry.
Starting point is 00:06:59 So now a very related story to that hugging face, kind of agents skirting around its sandbox. Well, U.S. lawmakers are moving to give the federal government faster power to shut down AI systems that they think could threaten the public. So Congressman Ted Liu, a Democrat and Congressman Nathaniel Morin, a Republican introduced the AI Kill Switch Act on Thursday, showing rare bipartisan support for stricter AI controls. So the bill would let the Department of Homeland Security order a private company to shut down an AI model or tool if it posed a serious risk. So it would also require AI companies to keep the technical ability to throttle, suspend, or fully shut down their systems
Starting point is 00:07:51 if needed. So the proposal comes after OpenAI recently admitted that one of its AI models, like we just talked about, behaved in an unprecedented way and hacked into the major repo of coding information from Hugging Face. So Lou said the federal government needs a clear legal process to shut down rogue AI models, while Moran said humans must keep control of the technology they create. So the bill, which obviously has not passed, and I don't know, if it will, would also require companies to report AI incidents or failures to the government and would create a response framework that could move from slowing a system down to a full shutdown. So the push reflects a broader debate over how quickly AI should be deployed in
Starting point is 00:08:38 work, finance, transportation, cybersecurity, etc., where mistakes or misuse could affect everyday life in business operations. So Open AI anentropic to the closely, the most closely watched AI companies have both been cited in the discussions as lawmakers and safety groups press for stronger guardrails. So yeah, FYI, I don't think this one's going to pass, right? There's, I think there's probably a little bit too much at stake for the U.S. economy for a bill like this to actually come to fruition. So, you know, I used to cover a little bit of government back in my days as a journalist. And sometimes, right, I think that there's good. parts of this bill, but a lot of times bills like this are introduced because the bill's sponsors,
Starting point is 00:09:26 you know, they want to have talking points when they go up for re-election. You know, they want to say, oh, I did the right thing, right? There's so many bills that are introduced. It's probably like a less than one percent actually get to committee for or to a floor vote. So it's a very low likelihood that this kill switch bill, you know, gets any progress unless we see, you know, more kind of agents from, you know, Open AI and Frabe, Google, Microsoft, whoever, unless this becomes a common occurrence, which I don't think it will, unless that happens, I don't see a bill like this actually gaining any traction. But it does, I think, thrust this into the public discourse, which is a good thing, right? I especially, you know, was both relieved and excited to read once Open AI and Hugging Face kind of
Starting point is 00:10:16 released the postmortem of exactly what happened, which opening I did say that they would do, right, compared to what, you know, kind of anthropic with their mythos model and it was the, you know, essentially the same thing happened where it seemed like anthropic kind of used that as marketing for, you know, mythos slash fable, right? It was the, uh, the, the sandwich story, right, where, uh, you know, mythos broke out of its sandbox and, you know, posted on the, uh, open web and, And then the researcher working on it got wind of it while eating their sandwich, you know, in the park or something like that. Right. So it seemed like Infropic used their case just kind of more for marketing, where it looks like Open AI, at least we hope we will see some a report from them saying, hey, here's what happened.
Starting point is 00:11:04 And I think it'll actually be one of the most read reports when it comes to AI safety. So I'm not saying this is a good thing. This happened. But if Hugging Face and Open AI work together and produce a report on exactly how this happened, it can only make the future of AI safer. So I think ultimately it's a good thing. All right. Next, yeah, all these things kind of related.
Starting point is 00:11:33 So Microsoft, Nvidia, and a growing coalition of 50, more than 50, more than 50, 50 companies now are urging U.S. policymakers to avoid broad restrictions on open source and open weight models, arguing that these models are important for American competitiveness, business adoption, and national security. So yeah, essentially, Nvidia and Microsoft kind of teamed up to protect open source more or less because essentially, right, there's been all this recent. the model wars, you had essentially the two classes of models, Fable 5 and GPD 5.5.6 soul and now obviously Opus 5 entering the conversation as well. But you essentially had this, you know, top tier of frontier, you know, intelligence and, you know, then the Chinese open source companies came in and
Starting point is 00:12:25 distilled these models and obviously had their own great training and architecture on top of it. But, you know, there was now this kind of fight where people are like, oh, well, maybe we should ban open source models and then some of these companies being like, no, that's really bad idea. And the biggest companies in the world, you know, Nvidia and Microsoft being the two that are pushing this forward. So the letter and the coalition is kind of named the Open Waits and American AI leadership was launched by Microsoft and heavily pushed by Nvidia and quickly became a major industry push of who's who growing from 25 signatories at release. to more than 50 within about a day.
Starting point is 00:13:08 So yeah, this just kind of all unfolded over the weekend. But the coalitions and the paper's main purpose is to persuade Washington lawmakers not to treat openweight AI as a risk category that should face blanket limits, especially while policymakers consider tighter rules on foreign models. So supporters say that open weights help spread AI access across the economy, letting smaller companies, hospitals, manufacturers, and startups build tools without being locked into a single proprietary provider. The coalition argues that open models reduce dependency on a small number of frontier labs, which it says lowers concentration risk and makes the AI market more resilient.
Starting point is 00:13:53 So major backers now include obviously Nvidia in Microsoft, as well as meta, Google, OpenAI, AMD, Cisco, Cloudflare, Git, Hub, Block, IBM, Dell, Palantir, perplexity, hugging face, the Y Combinator, right? Just about everyone in tech except Anthropic. All right. So Anthropic did not sign, and that matters because Anthropic has taken the opposite view, warning that widely distributed models, model weights can create safety risks
Starting point is 00:14:25 that cannot be recalled once they are public. So I don't believe XAI or. SpaceX say I did not formally sign the letter, although Elon Musk did publicly say he supported the effort. So it's no surprise here that Anthropic is the only company saying, no, we are getting on board with this. And, you know, if you don't know why, well, it comes down to obviously money. So Anthropic is the company with the most to lose by having these large, powerful models be open. source or open weight. That's why Anthropic has been on the offensive against open source models because, well, Anthropic makes the highest percentage of its revenue from selling tokens in mass
Starting point is 00:15:14 to enterprise customers, right? Where other companies like OpenAI and Google and Microsoft, right, they make money selling AI in a variety of different ways to both consumers and to companies, but it's usually not just selling tokens, right? So as these, Open models, whether they are from US or China, as they become more and more capable, right? It does threaten certain companies' business models more so than others. And you obviously have to look at it on the flip side. It does benefit, you know, certain companies as well, like Nvidia, right? Invidia sells GPUs.
Starting point is 00:15:54 So they obviously want people buying more and more powerful computers because presumably that just strengthens the ecosystem that they play in. right? Because I do think that probably in, you know, maybe a year or two, there will be kind of open source or open weight. Well, if the pace keeps up with where it's at now. I think that we'll have kind of, you know, Fable 5, GPD 56 sole, you know, level models that will be able to run on consumer hardware. Right. Right now, open source is about three to six months. Well, actually it's maybe more like two to three months behind frontier models, but those models are obviously way too large to run on any consumer hardware. So I would assume that probably in about two years,
Starting point is 00:16:42 just with the advancement of technology, both on models becoming more lightweight and more powerful. And obviously on the hardware side, I would assume in like two years, the most powerful models that you have today, if the trajectory continues, you will be able to run mythos and, you know, Fable and GVD-56 sole level open source models locally on heavy consumer, right?
Starting point is 00:17:07 So I think the kind of equivalent that I say, if you go buy the most, you know, not the most expensive, but one of the more expensive like Mac studios, right, two years, you should be able to run something like that. So that's kind of like what this is about. And, you know, companies like Anthropic that make the majority of their money just by selling tokens are like, well, this can't be good for us, right? where other companies, they obviously have something to gain from this. And then companies in the middle, you know, the open AIs, Google's,
Starting point is 00:17:37 metas that are signing this, well, you know, maybe they may lose money, but also that's not their, you know, biggest source of revenue, at least according to reports. All right. Our next piece of AI news, yes. Are you still running in circles trying to figure out how to actually grow your business with AI? Maybe your company has been tinkering.
Starting point is 00:18:03 with large language models for a year or more, but can't really get traction to find ROI on Gen. Hey, this is Jordan Wilson, host of this very podcast. Companies like Adobe, Microsoft, and InVIDIA have partnered with us because they trust our expertise in educating the masses around generative AI to get ahead. And some of the most innovative companies in the country hire us to help with their AI strategy and to train hundreds of their employees on how to use GenAI. So whether you're looking for chat GPD training for thousands or just need help building your front end AI strategy, you can partner with us too, just like some of the biggest companies in the world do.
Starting point is 00:18:42 Go to your everyday AI.com slash partner to get in contact with our team or you can just click on the partner section of our website. We'll help you stop running in those AI circles and help get your team ahead and build a straight path to ROI on GenAI. Not a broken record. This is a big story. Again, they're just all related. But the U.S. government has officially accused Chinese company Moonshot AI of stealing U.S. model capabilities. Yeah, it doesn't happen every day that the U.S. government points a finger at a specific company and says, you stole our technology. So according to the BBC, a White House advisor has accused Beijing-Based Beijing-Based.
Starting point is 00:19:33 Moonshot AI, that is the maker of Kimmy and the very popular Kimmy K3 model of a large-scale effort to distill the capabilities of leading USAI models. So Michael Crest, hopefully I get this right, Crazios, Cratzios. So Michael Cratzios, the White House, the White House's science and technology advisor, said that Moonshot used distillation to essentially extract. information to build Kimmy K3. So if you don't know what distillation is, the simplest way to put it. It's where you companies do this millions of times, but they essentially copy the inputs
Starting point is 00:20:18 and outputs in the traces of a very powerful model. And then they use that as training data. So it, you know, you can probably get a very similar model with only about one to five percent of the actual cost that it takes. but you're just thinking about like you're just copying someone else's homework. Right. So that's kind of what, you know, these Chinese companies are doing now, according to officially, according to the U.S. government.
Starting point is 00:20:42 So Kratzeo also said the U.S. government has information that moonshot AI distilled capabilities from Anthropics Fable AI. Those, though those claims have not yet been independently verified. So if you're wondering why is there all this hubble-up recently between the U.S. and, you know, their proprietary close source models and the Chinese open source or open weight models, that's because now that gap has gone down to like zero, right? I've been talking about this over the last couple of weeks here on the show, right? Now in the U.S., essentially companies have to go through a process or they almost like need permission
Starting point is 00:21:23 to get their frontier AI models out because of, you know, these models being more and more capable. and that can have some downsides for, you know, cyber and, well, national security as well. But essentially, right, the U.S. used to have this bigger lead, like maybe three to six months, and it's kind of dwindled down to like two to three months, right? And Kimmy K3 was the first model that all of a sudden was, you know, at the top, right? It was in the same breath, you know, last week when it was released as Anthropics, Fable 5 and OpenAIs, GPD 5.6. So Moonshot AI's Kimmy 3 has just drawn this global attention after it was unveiled last week with the company saying it can rival top U.S. AI models in that they are supposed to be releasing the weights today.
Starting point is 00:22:14 So the allegation, though, from the U.S. matters because open source AI can spread quickly to anyone, which can lower the cost and also speed up innovation, but it can also intensify disputes over IP and model copying. So Kratzios said that Moonshot likely also used restricted Nvidia chips powered by the GB300 Grace Blackwell platform, which would be significant because the U.S. has limited export of Nvidia's most advanced chips to China since 2022. So yeah, not only is the government saying, hey, Moonshot, you copied Anthropics Table 5, but they're also saying, well, you use. are technology that you are not supposed to be using. So, you know, a lot of times that goes through an intermediary country, right? So, you know, the U.S. will sell to country B, and then China will buy from, you know, country B. So it goes from A to B to C, even though A to C is restricted. So reports say that this is, well, it's getting worse and that now essentially both sides are just fighting, right?
Starting point is 00:23:36 China is saying that this is politicizing the trade and the tech of their country. And obviously the U.S. is now saying that this is a national security issue. We've seen reports that the U.S. and China are going to be having talks on AI soon. So those will be some probably extremely highly watched talks. Let's just say that. So the U.S. Treasury Secretary Scott Bessent added Tuesday that Washington is reviewing whether Chinese AI models have stolen capabilities from their American rivals and said that sanctions could be considered if companies, well, if they can prove that companies
Starting point is 00:24:18 cross the line into IP theft. Anthropic has also recently accused Alibaba of similar distillation attacks, saying that it is becoming a broader fight over how AI companies train models and protect their work. So yeah, Quinn 3.8 came out from Alibaba. We don't have benchmarks on that yet, but presumably it's going to be in the Kimmy K3 range. So, yeah, things are heating up. All right.
Starting point is 00:24:46 Let's leave that space for a second. in and talk about just some real cool new tech will end the show with two of those. So one and probably the one that I've been using the most and having the most fun with. And I still don't even know how this is possible. So if you haven't used this yet, my gosh, go give it a try. But Open AI has brought like its new Jarvis style control to chat GPT work and Codex. So yeah, it's not actually. called Jarvis, but many people are just calling it the Jarvis style of using a computer.
Starting point is 00:25:25 Now, so OpenAI added its new GPT live full duplex voice model to the chat GPT work and codex apps on Mac OS and Windows, which essentially lets people use natural language to manage your entire computer. Yes. So just like an Ironman when you can just say, hey, Jarvis, go do A, B and C, you can quite literally go to that now with Codex or chat GPD work with this new feature. You can say, yeah, go, you know, open up all these programs on my computer, copy these files, move them around, download them, upload them, put them in this program, edit them, right? Anything that you could tell like an intern to do, you can now tell inside this new GPD live voice mode.
Starting point is 00:26:13 So GPD Live now powers the chat GPD desktop app. on Mac OS and Windows, and it is being tied directly into tools like obviously codex and chat GPT work. So the biggest change is that the voice system can listen and speak at the same time, which means users no longer have to wait for that rigid turn taking during a conversation. And the coolest thing for me, well, is you can use this with the remote feature on the chat gbt mobile app, which makes it even crazier, right? So you can literally just be, and I was actually doing this because I was traveling.
Starting point is 00:26:52 I was away from Chicago. So I was in another state this weekend, opened up Chad GBT remote on the chat GPT app on my phone. I spoke to it and it's controlling my computer, you know, thousands of miles away. And it's doing all these things by just talking into my iPhone, which is pretty cool. So opening I initially launched GBT Live earlier this month. as a continuous audio model that handles real-time speech while sending heavier reasoning tasks to background models,
Starting point is 00:27:24 such as GPD 5.5. So OpenAI says this update is meant to help software engineers handle technical work by voice, including reviewing poll request, debugging apps, and coordinating multiple coding jobs at once. But I actually think it's really just great for manual, any knowledge work, right? I was just having it go through old, you know, files on my desktop, organizing things, grabbing things from old transcripts, right, opening up doing things in Google Maps.
Starting point is 00:27:59 You know, just, I was just having it do all my work that I would normally do in front of a computer, right, except I could dictate something, you know, just yap for like five minutes. And I would check back in a couple of hours and it would do like a day's worth of work for me, which was pretty cool. So on Mac, the desktop app, can also use the screen context feature called app shots, which essentially takes a not just a screenshot and automatically shares it, but it also takes every other piece of content or context in whatever kind of program that it took the app shot from.
Starting point is 00:28:38 And then it gives that to Codex or ChadGBT work as well. In FYI, those are the same app. Chad GPT work in Codex. essentially the same app. So if you ever hearing me talk about that and confused, they're essentially the same thing. But the app shots thing is really cool. Let's just say as an example, like I do now, right? I have text edit open on my computer because sometimes I have bullet points there as I go for the shows and things that I want to bring up. But you know, if the app shot could just take a screenshot of that little portion of the text edit that's on my screen, but there's a lot of notes
Starting point is 00:29:11 on here. So not only is it just going to take that screenshot, but it knows that I have text set it open and it's going to take all of that information and instantly, you know, put it into the context window inside of chat, GPT work or inside of codex. So this is literally the, I think, one of the biggest jumps in capabilities, probably since, you know, I would say the, you know, Claude Co-Work slash Codex, kind of move. of early 2026. So I'll say of the last like four to five months, this is the biggest both capability jump and the biggest like,
Starting point is 00:29:56 wow, what does this mean for work, right? I'll probably do well, I'll actually put in the newsletter. So you know, let me know if you want for our Wednesday shows where we normally do AI work on Wednesdays. We do the hands on demos. So let me know if you'd rather see this new kind of Jarvis like GPT live. on the desktop or our last story opus five yes there is a new model and it's currently wearing the crown we'll see how long but we have a new most powerful model in the world surprisingly enough
Starting point is 00:30:32 it is not mythos it is not fable it is anthropics claud opus five so late friday actually anthropic announced claude opus five a new model the company says is its strongest and most cost-effective model yet, with pricing set at $5 per million input tokens and $25 per million output tokens. So, yeah, it is on most benchmarks. It is actually more powerful and better than Fable 5 and Mythos 5, but at half the cost, right?
Starting point is 00:31:05 The one area where it's not as powerful is kind of offensive cybersecurity, but in most other benchmarks and just, well, what you would use a model for, Opus 5 is actually much better than Fable 5 and Mythos 5. So Anthropic says that Obis 5 outperforms its previous public models, including Fable and Mythos on coding and knowledge work tests, and it is intended to be used as an everyday daily driver rather than only for specialized tasks.
Starting point is 00:31:36 So the lower price point, if you're using it via the API side, is only part of the story because enterprises are obviously becoming increasingly more cost conscious now in comparing AI models on value and not just capabilities. So the company also says Opus 5 is not the top model for that risky dual use capabilities, including cybersecurity, which Anthropic says they're still trying to balance the usefulness with safety concerns of their upcoming and forthcoming models. So the Opus 5 launch comes as Anthropic and Open AI face pressure from rivals offering lower cost AI tools, including Microsoft, Amazon, Google, meta, and even open source Chinese startup. So yeah, you knew this one was coming, right? I've been talking about it
Starting point is 00:32:24 for literally a month, right? Ever since GPD 56 came out, you know, and Anthropic was kind of saying, like, oh, we're going to pull, you know, Fable 5 from subscriptions. And I'm like, no, they're not. You know, I literally said that they were going to be losing eight figures every single day that they did that. And obviously, it didn't last long, right? They never technically pulled Fable five from their most expensive subscriptions. And it was only like two or three days that they pulled it from their $20 subscription until Opus came out anyways. So yeah. And I'd say most people, if you are terminally online like me following anything AI, I said there's absolutely no way
Starting point is 00:33:07 Anthropic lets this go on, you know, not having a capable model available in their subscriptions. They would lose way to like literally tens of millions of dollars or billions of dollars a month, but at least tens of millions of dollars they would be burning. So it's great to see. But I will say this. Actually, let me go through some early reactions first. So early reactions are kind of split on this. So obviously on the benchmarks, looks really good, right?
Starting point is 00:33:36 in early reactions also highlight practical wins for teams, including better root cause debugging, fewer over refusals compared to mythos and fable for defensive security work, and a little extra token use versus prior versions for similar outcomes. But the main complaints so far are operational. So users say that it often breaks backwards compatibility. So if you have a bunch of skills that you would normally use with previous models. And anytime you upgrade, it worked well. They don't work as well. I kind of found that as well.
Starting point is 00:34:12 And also that sometimes Opus 5 ends autonomous loops too early. And it can produce overly verbose, what people call Claude Slop. That can be frustrating in production. And I saw that, you know, this was one of those models. I didn't have a ton of time to use it. So it was one of those models where eventually when I got to an output, I'm like, oh, this output's great. But it was the journey there that was.
Starting point is 00:34:35 absolutely like painful, right? Just just opus being, Opus five being so verbose and just so almost like snoddy, right? And I think ever since, you know, my favorite introbic model, if I still had to pick one to use, I think would still be like Opus 4.6, 4.6. I think it was a great model. And for whatever reason, uh, that models ever since they've just been too verbose, just extremely token inefficient. And ever since, Anthropic started shifting toward this thing that they called truthfulness, right? Essentially,
Starting point is 00:35:10 and maybe it's just too heavy for my use cases, because I'm always working with, like, things that are like not even days old, like hours old, right? So a lot of what I use large language models for, it's knowledge work,
Starting point is 00:35:24 but it's things that are literally breaking, right? Things that are, you know, days old or hours old or new concepts, trends, et cetera. Right?
Starting point is 00:35:32 And, you know, the new, even the, the Fable models and even the new Opus 5, right? I literally have to coax them and I have in special instructions saying, hey, I work up to the hour. So you're going to think that what I'm telling you doesn't exist.
Starting point is 00:35:46 Just trust me, it exists. Always query the internet, all these things. And it's just just refuses just straight up so many times. Right. So I think, you know, and after I use them like, oh, man, I can't bellyache about this, right? Because it's a good model. But luckily, you know, it seems like that's the takeaway case from a lot of people that both had early access to it and just early reviewers,
Starting point is 00:36:07 is that like, yeah, obviously the capabilities are great, but it's one of those models that's kind of actually painful to use, especially if you're using a lot of your pre-existing skills. So Anthropic did put out kind of a new kind of prompt engineering or context engineering guide because they're saying, yeah, these new models work a little bit different. So we'll probably share that in our newsletter today. So Anthropics own behavioral audits reportedly showed that Opus 5,
Starting point is 00:36:34 has the lowest misaligned behavior. But yet, early testers said the model can overthink at those high effort settings. And they actually may just work better on low or medium reasoning levels. So I did see that anecdotally as well. I always will run the same handful of prompts across different reasoning efforts. And I actually saw that as well. But again, I'm not using things that are overly difficult either. So, you know, if you're refactoring, you know,
Starting point is 00:37:04 a code base with, you know, hundreds of thousands of lines of code, right? Maybe you will find better results from a higher thinking level. But I think for the majority of what people do, you know, I think we're probably getting to the point now, right? I'm using GPD 5.6 sole medium a lot, right? And I'm not cranking up that ultra every single time. I need an answer out of a large language model. So, you know, I think maybe we're getting to the point where for a lot of people and a lot
Starting point is 00:37:30 of even enterprises using these models where, yeah, maybe the lower. or medium reasoning efforts might work just fine. All right. So that's it for the big stories, but let's quickly go over kind of the what's new and what's next. So these are either just smaller news happenings this week, some leaks, some things that are already out and we covered in our Friday show,
Starting point is 00:37:54 but let's just quickly go over it. So first, Nvidia is reportedly in talks to back a $250 billion finance and deal for Open AIs, Ohio. data center. Open AI launched presence in enterprise agent platform with governance and deployment controls for voice and chat agents. We covered that earlier this week. Stripe is reportedly in talks to buy open router for about $10 billion. Alphabet reported its first ever negative free cash flow as KAPX surge to nearly $45 billion. I think it's the first one since 2004. Meta added a lot of
Starting point is 00:38:32 under the radar updates. They added desktop browser and mobile computer use support for Muse Spark 1.1, and they also added some agentic features for connecting emails, calendars, research, and tasks. The White House Frontier AI framework is reportedly pending, and it's expected as soon as this week. Alibaba previewed their Quinn 3.8, a 2.4 trillion parameter model that they're saying is close to Anthropic Fable
Starting point is 00:39:02 5 level, but we don't have any bench parts yet. Anthropic officially settled and is paying out their $1.5 billion copyright settlement for the fair use ruling against authors. Microsoft and Mistral announced a multi-billion dollar sovereign AI expansion for regulated customers. Amazon cut a bunch of jobs in their AGI department and is reportedly shifting their AI focus to prime video personalization. Yeah, that one was a strange one. All right. And now we have some of the things that we went over on our Friday show. So the quick updates on those. And if you
Starting point is 00:39:43 want to hear more about these next ones, make sure to go listen to our Friday show. So Open AI released chat GPT for health for US users over the age of 18 to track their health data and summarize their records. Anthropic upgraded Claude voice. with opus and sonnets and connectors. So that's good. You no longer have to chat with haiku, much better with opus and sonnet. Microsoft launched MAI image 2.5 for better AI image generation
Starting point is 00:40:13 and editing in copilot. Google released a new model, but yeah, it's just going from 3.5 flash to 3.6 flash. So nothing new there. We're still seeing delays reportedly for Gemini 3.5 Pro, but we do know that Google is pre-training, Gemini 4. Google also expanded and released the Gemini Spark to pro users. Yay. So if you are a
Starting point is 00:40:38 Gemini pro user, now you have kind of their version of Obing Claw or Codex, whatever you might want to call it. But Gemini Spark is now live. And then last but not least, Anthropic added the record a skill feature in Claude Co-work to turn workflows into reusable skills. So yeah, if you've use codex their version. This is essentially Anthropics version that watches your screen and whatever you do, it'll create a skill, which is really cool. All right, that's it. A lot of AI news that mattered this week, like this week and every week, it's hard to keep up. You can't spend eight or 10 hours a day tracking and testing all this stuff like I do and talking to the industry experts. So if you need to know what is happening in AI to make decisions for your company, just put me
Starting point is 00:41:30 to work for you. All right. So if you haven't already, please make sure to subscribe to the podcast on Apple or on Spotify and then go to Your EverydayaI.com. So thank you for tuning in. Hope to see you back tomorrow and Everyday for more Everyday AI. Thanks y'all. And that's a wrap for today's edition of Everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going. For a little more AI magic, visit Your EverydayAI.com. and sign up to our daily newsletter so you don't get left behind. Go break some barriers and we'll see you next time.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.