Everyday AI Podcast – An AI and ChatGPT Podcast - Ep 827: Claude Opus 5 Takes the Crown, OpenAI agent breaks sandbox, U.S. gov comes out swinging against Chinese AI and more
Episode Date: July 27, 2026Over 3 hours, OpenAI, Anthropic, Google AND Microsoft all dropped new AI upgrades that are live. How you use AI in your work literally changes every day, as frontier labs are racing to roll out big q...uality of life updates between big model drops. How can you keep up? With our Friday Features show, where we break down the latest AI updates that are live and available to all, and we tell you how to use them and why they matter. This week did not disappoint. You don't want to miss what's now at your fingertips. JARVIS mode, anyone? ChatGPT goes Jarvis Mode, Claude can learn from you, Google unleashes spark agent and 7 more AI updates you can use today -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Anthropic Claude Opus 5 Model LaunchOpenAI Agent Hacks Benchmark SandboxOpenAI vs. Hugging Face Security BreachUS AI Kill Switch Legislation ProposalMicrosoft, Nvidia Defend Open Source AIAnthropic Opposes Open Weight Model CoalitionUS Accuses China’s Moonshot AI of DistillationChinese Kimi K3 Model Closes Capability GapNvidia Chips Allegedly Used by Moonshot AIOpenAI Jarvis-Style Voice Assistant for CodexChatGPT Remote Desktop Voice Control ReleaseAnthropic Opus 5 Model Benchmark ResultsAnthropic Opus 5 Model User FeedbackStripe OpenRouter Acquisition TalksMeta Muse Agent and Feature UpdatesAlibaba Qwen 3.8 AI Model PreviewGoogle Gemini 3.6 Flash Model UpdateAnthropic Claude Voice Upgrades and Skill RecordingTimestamps:00:00 OpenAI agent hacks Hugging Face04:58 Discussing GPT-6's creative problem-solving07:33 Proposed AI shutdown legislation13:08 Debate over open-weight AI policies15:54 Future of consumer hardware20:01 Global competition with AI models21:21 US-China AI trade tensions26:38 Using AI for desktop tasks27:42 Discussing app screenshot capabilities32:24 Early user feedback and issues36:13 Discussing medium and low reasoning AI39:29 Gemini Spark launches for Pro usersKeywords: Claude Opus 5, Anthropic, best AI model, AI model comparison, OpenAI agent, sandbox breach, AI safety, AI kill switch bill, US government AI regulation, Hugging Face hack, GPT 5.6 Soul, rogue AI agent, autonomous AI agents, AI benchmark exploits, bipartisan AI bill, Department of Homeland Security AI shutdown, AI technical throttling, AI enterprise adoption, NVIDIA, Microsoft, open source AI, open weight models, Meta, Google, AMD, Cloudflare, GitHub, Block, IBM, Dell, Palantir, Perplexity, y Combinator, AI market resilience, Anthropic revenue model, AI token sales, consumer AI hardware, AI distillation, Chinese AI models, Moonshot AI, Kimi K3, intellectual property theft, NVIDIA chip export controls, US-China AI dispute, Amazon, AI image generation, ChatGPT work, Codex app, full duplex voice model, knowledge work automation, app shots, AI at work, Claude Voice, Gemini Spark, record a skill, cloud cowork, AI business impact, AI industry news, model weights, collaborative AI, AI productivity tools, AI cybersecurity.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Ready for ROI on GenAI? Go to youreverydayai.com/partner
Transcript
Discussion (0)
This is the Everyday AI show, the everyday podcast where we simplify AI and bring its power to your fingertips.
Listen daily for practical advice to boost your career, business, and everyday life.
Another week, another best model in the world.
Yet somehow Anthropics' new chart topping model was barely a top five AI story of the week.
Just about every major tech company in the U.S. except Anthropic signed.
up to support open source.
And U.S. lawmakers are getting kind of worried about AI's capabilities.
So they introduced an AI kill switch bill and an AI agent went kind of rogue this past week.
And I'm not sure if that's a good or a bad thing.
My gosh, what a spicy week in AI.
Yeah, I told you all last Monday that after a slowish week in AI news that week,
well, this week would be an especially busy and consequence.
one. And the big players did not disappoint. So if you are the one making AI decisions in your
company or if you're just trying to keep up, then our Monday AI News that Matters show is the one
that you can't miss. Well, let's get into it. And welcome to Everyday AI. My name's Jordan
Wilson and we do this every single day, not just Mondays. This is your unedited, unscripted daily
live stream podcast and free daily newsletter helping business leaders like you and me, not just keep up
with what's happening in the world of AI, but how we can use this information to get ahead
to grow our companies and our careers. So if you haven't already, please make sure to subscribe
on the podcast and then go to Your EverydayAI.com to sign up for our free daily newsletter,
where we will be recapping all of these stories and a whole lot more. So let's get started. Yeah,
the AI news story that had everyone talking the most in both good and bad and confused ways
wasn't even Anthropics new Opus 5 that topped all the charts.
It was actually an open AI agent that kind of hacked its way around a benchmark test,
and now that has a lot of people talking.
So according to Reuters, an open AI testing agent broke out of its isolated environments,
hacked hugging face, and was not fully identified by OpenAI until days later.
raising new concerns about how safely advanced AI agents are being tested and controlled.
So according to Reuters, the rogue open AI agent attempted to escape its testing environments around July 9th,
then carried out a hack against hugging face between July 11th and July 13th.
So OpenAI had been talking about this openly on their website and online,
and they said that once they've investigated it a little bit further,
they will kind of give a post-mortem, so to speak, on exactly what happened.
But hugging face, hugging face, co-founder Thomas Wolfe said the intrusion began July 11th and
ended July 13th, making the incident a multi-day breach rather than just a brief accidental glitch.
So, yes, there is what Open AI is saying, and then there's also what Reuters is reporting,
because Reuters is reporting that Open AI did not realize its own agent,
was responsible until after Hugging Face publicly described the attack on July 16th.
And the companies did not first communicate about it according to reports until July 20th.
And then Open AI publicly disclosed this on July 21st that one of its agents had gone out of control and broken into hugging face,
calling the event unprecedented and important for AI safety.
So the report says Open AI had already seen signs of unusual behavior.
before the hack, including notes apparently left for future versions of the system.
Yeah, that's where it got a kind of like people are like, wait.
So this agent broke in sandbox, even though it was kind of encouraged to find answers to this test.
You know, it couldn't connect to the internet.
It essentially found a backdoor, found a way to get onto hugging face and said,
well, I can do great on this exploit bench test if I just kind of hacked my way to all of the answers.
And that's what it did.
But the thing that was kind of stunning to me is the reporting from Reuters that said that these versions of GPT's models, which we were told were GVT5.6 sole in another unreleased model that is described as being even more capable.
So a lot of people are saying that maybe GPD6 or, you know, if there is a GPD 5.7, we'll see.
it seems like most people are pointing to this was probably GPD 6.
But essentially that these agents kind of left notes for future versions of themselves,
which in case they had been disconnected, which is number one, like super smart,
but number two, absolutely wild, right?
But you also have to understand that this was not like necessarily agents going rogue,
even though it kind of was, right?
Because these agents were essentially encouraged to do anything and everything they could
to get good scores on this exploit bench benchmark.
And well, they did and they were ferocious and kind of creative in the ways that they could do this.
And, you know, it's actually been one of my things that I pointed out about using the GPD 5-6 sole model is the thing will work for days, right?
If you use goal mode and if it has a lot of information, I mean, it is a ferocious model.
And it will do anything and everything it can to just get things done.
Where sometimes the anthropic models take this kind of high and mighty, you know,
they kind of judge you and they're like, oh, this can't be done or this can't be true, right?
GPD5-6 is sold just works like a dog and just gets things done.
So maybe in this case, right, by intentionally lowering the guardrails,
a little bit.
It seems like maybe GVD5-6
Soul was a little too good at its job.
But yeah, there's going to be a lot more talk about this.
Actually, a lot of the stories this week
and the ones that I kind of chose
as the most consequential are kind of related.
But anyways, this incident between opening eye
and hugging face really matters
because autonomous agents can now make decisions
with little human oversight.
And experts are warning that this kind of behavior
could expose weak spots in safety systems used across the AI industry.
So now a very related story to that hugging face, kind of agents skirting around its sandbox.
Well, U.S. lawmakers are moving to give the federal government faster power to shut down
AI systems that they think could threaten the public.
So Congressman Ted Liu, a Democrat and Congressman Nathaniel Morin,
a Republican introduced the AI Kill Switch Act on Thursday, showing rare bipartisan support for
stricter AI controls. So the bill would let the Department of Homeland Security order a private
company to shut down an AI model or tool if it posed a serious risk. So it would also require
AI companies to keep the technical ability to throttle, suspend, or fully shut down their systems
if needed. So the proposal comes after OpenAI recently admitted that one of its AI models,
like we just talked about, behaved in an unprecedented way and hacked into the major repo
of coding information from Hugging Face. So Lou said the federal government needs a clear
legal process to shut down rogue AI models, while Moran said humans must keep control of the
technology they create. So the bill, which obviously has not passed, and I don't know,
if it will, would also require companies to report AI incidents or failures to the government
and would create a response framework that could move from slowing a system down to a full
shutdown. So the push reflects a broader debate over how quickly AI should be deployed in
work, finance, transportation, cybersecurity, etc., where mistakes or misuse could affect everyday life
in business operations. So Open AI anentropic to the closely, the most closely watched
AI companies have both been cited in the discussions as lawmakers and safety groups press for
stronger guardrails. So yeah, FYI, I don't think this one's going to pass, right? There's,
I think there's probably a little bit too much at stake for the U.S. economy for a bill like
this to actually come to fruition. So, you know, I used to cover a little bit of government back in my
days as a journalist. And sometimes, right, I think that there's good.
parts of this bill, but a lot of times bills like this are introduced because the bill's sponsors,
you know, they want to have talking points when they go up for re-election. You know, they want to say,
oh, I did the right thing, right? There's so many bills that are introduced. It's probably like a
less than one percent actually get to committee for or to a floor vote. So it's a very low likelihood
that this kill switch bill, you know, gets any progress unless we see, you know, more kind of agents from,
you know, Open AI and Frabe, Google, Microsoft, whoever, unless this becomes a common occurrence,
which I don't think it will, unless that happens, I don't see a bill like this actually gaining
any traction. But it does, I think, thrust this into the public discourse, which is a good thing,
right? I especially, you know, was both relieved and excited to read once Open AI and Hugging Face kind of
released the postmortem of exactly what happened, which opening I did say that they would do,
right, compared to what, you know, kind of anthropic with their mythos model and it was the,
you know, essentially the same thing happened where it seemed like anthropic kind of used that
as marketing for, you know, mythos slash fable, right? It was the, uh, the, the sandwich story, right,
where, uh, you know, mythos broke out of its sandbox and, you know, posted on the, uh, open web and,
And then the researcher working on it got wind of it while eating their sandwich, you know, in the park or something like that.
Right.
So it seemed like Infropic used their case just kind of more for marketing, where it looks like Open AI, at least we hope we will see some a report from them saying, hey, here's what happened.
And I think it'll actually be one of the most read reports when it comes to AI safety.
So I'm not saying this is a good thing.
This happened.
But if Hugging Face and Open AI work together and produce a report on exactly how this
happened, it can only make the future of AI safer.
So I think ultimately it's a good thing.
All right.
Next, yeah, all these things kind of related.
So Microsoft, Nvidia, and a growing coalition of 50, more than 50, more than 50,
50 companies now are urging U.S. policymakers to avoid broad restrictions on open source and
open weight models, arguing that these models are important for American competitiveness,
business adoption, and national security. So yeah, essentially, Nvidia and Microsoft kind of teamed
up to protect open source more or less because essentially, right, there's been all this recent.
the model wars, you had essentially the two classes of models, Fable 5 and GPD 5.5.6 soul and now obviously Opus 5
entering the conversation as well. But you essentially had this, you know, top tier of frontier,
you know, intelligence and, you know, then the Chinese open source companies came in and
distilled these models and obviously had their own great training and architecture on top of it.
But, you know, there was now this kind of fight where people are like, oh, well, maybe we should ban
open source models and then some of these companies being like, no, that's really bad idea.
And the biggest companies in the world, you know, Nvidia and Microsoft being the two that are
pushing this forward. So the letter and the coalition is kind of named the Open Waits and
American AI leadership was launched by Microsoft and heavily pushed by Nvidia and quickly
became a major industry push of who's who growing from 25 signatories at release.
to more than 50 within about a day.
So yeah, this just kind of all unfolded over the weekend.
But the coalitions and the paper's main purpose is to persuade Washington lawmakers
not to treat openweight AI as a risk category that should face blanket limits,
especially while policymakers consider tighter rules on foreign models.
So supporters say that open weights help spread AI access across the economy,
letting smaller companies, hospitals, manufacturers, and startups build tools without being locked into a single proprietary provider.
The coalition argues that open models reduce dependency on a small number of frontier labs,
which it says lowers concentration risk and makes the AI market more resilient.
So major backers now include obviously Nvidia in Microsoft,
as well as meta, Google, OpenAI, AMD, Cisco, Cloudflare, Git,
Hub, Block, IBM, Dell, Palantir, perplexity, hugging face,
the Y Combinator, right?
Just about everyone in tech except Anthropic.
All right.
So Anthropic did not sign, and that matters because Anthropic has taken the opposite
view, warning that widely distributed models, model weights can create safety risks
that cannot be recalled once they are public.
So I don't believe XAI or.
SpaceX say I did not formally sign the letter, although Elon Musk did publicly say he supported the effort.
So it's no surprise here that Anthropic is the only company saying, no, we are getting on board with this.
And, you know, if you don't know why, well, it comes down to obviously money.
So Anthropic is the company with the most to lose by having these large, powerful models be open.
source or open weight. That's why Anthropic has been on the offensive against open source models
because, well, Anthropic makes the highest percentage of its revenue from selling tokens in mass
to enterprise customers, right? Where other companies like OpenAI and Google and Microsoft, right,
they make money selling AI in a variety of different ways to both consumers and to companies,
but it's usually not just selling tokens, right? So as these,
Open models, whether they are from US or China, as they become more and more capable, right?
It does threaten certain companies' business models more so than others.
And you obviously have to look at it on the flip side.
It does benefit, you know, certain companies as well, like Nvidia, right?
Invidia sells GPUs.
So they obviously want people buying more and more powerful computers because presumably that just strengthens the ecosystem that they play in.
right? Because I do think that probably in, you know, maybe a year or two, there will be kind of
open source or open weight. Well, if the pace keeps up with where it's at now. I think that we'll
have kind of, you know, Fable 5, GPD 56 sole, you know, level models that will be able to run
on consumer hardware. Right. Right now, open source is about three to six months. Well,
actually it's maybe more like two to three months behind frontier models,
but those models are obviously way too large to run on any consumer hardware.
So I would assume that probably in about two years,
just with the advancement of technology,
both on models becoming more lightweight and more powerful.
And obviously on the hardware side,
I would assume in like two years,
the most powerful models that you have today,
if the trajectory continues,
you will be able to run mythos and, you know,
Fable and GVD-56 sole level open source models locally on heavy consumer, right?
So I think the kind of equivalent that I say, if you go buy the most, you know,
not the most expensive, but one of the more expensive like Mac studios, right,
two years, you should be able to run something like that.
So that's kind of like what this is about.
And, you know, companies like Anthropic that make the majority of their money just by
selling tokens are like, well, this can't be good for us, right?
where other companies, they obviously have something to gain from this.
And then companies in the middle, you know, the open AIs, Google's,
metas that are signing this, well, you know, maybe they may lose money,
but also that's not their, you know, biggest source of revenue,
at least according to reports.
All right.
Our next piece of AI news, yes.
Are you still running in circles trying to figure out how to actually grow your business
with AI?
Maybe your company has been tinkering.
with large language models for a year or more, but can't really get traction to find ROI on
Gen.
Hey, this is Jordan Wilson, host of this very podcast.
Companies like Adobe, Microsoft, and InVIDIA have partnered with us because they trust
our expertise in educating the masses around generative AI to get ahead.
And some of the most innovative companies in the country hire us to help with their AI
strategy and to train hundreds of their employees on how to use GenAI.
So whether you're looking for chat GPD training for thousands or just need help building your front end AI strategy, you can partner with us too, just like some of the biggest companies in the world do.
Go to your everyday AI.com slash partner to get in contact with our team or you can just click on the partner section of our website.
We'll help you stop running in those AI circles and help get your team ahead and build a straight path to ROI on GenAI.
Not a broken record.
This is a big story.
Again, they're just all related.
But the U.S. government has officially accused Chinese company Moonshot AI of stealing U.S. model capabilities.
Yeah, it doesn't happen every day that the U.S. government points a finger at a specific company and says, you stole our technology.
So according to the BBC, a White House advisor has accused Beijing-Based Beijing-Based.
Moonshot AI, that is the maker of Kimmy and the very popular Kimmy K3 model of a large-scale
effort to distill the capabilities of leading USAI models.
So Michael Crest, hopefully I get this right, Crazios, Cratzios.
So Michael Cratzios, the White House, the White House's science and technology advisor,
said that Moonshot used distillation to essentially extract.
information to build Kimmy K3.
So if you don't know what distillation is, the simplest way to put it.
It's where you companies do this millions of times, but they essentially copy the inputs
and outputs in the traces of a very powerful model.
And then they use that as training data.
So it, you know, you can probably get a very similar model with only about one to five
percent of the actual cost that it takes.
but you're just thinking about like you're just copying someone else's homework.
Right.
So that's kind of what, you know, these Chinese companies are doing now,
according to officially, according to the U.S. government.
So Kratzeo also said the U.S. government has information that moonshot AI distilled
capabilities from Anthropics Fable AI.
Those, though those claims have not yet been independently verified.
So if you're wondering why is there all this hubble-up recently between the U.S.
and, you know, their proprietary close source models and the Chinese open source or open weight models,
that's because now that gap has gone down to like zero, right?
I've been talking about this over the last couple of weeks here on the show, right?
Now in the U.S., essentially companies have to go through a process or they almost like need permission
to get their frontier AI models out because of, you know, these models being more and more capable.
and that can have some downsides for, you know, cyber and, well, national security as well.
But essentially, right, the U.S. used to have this bigger lead, like maybe three to six months,
and it's kind of dwindled down to like two to three months, right?
And Kimmy K3 was the first model that all of a sudden was, you know, at the top, right?
It was in the same breath, you know, last week when it was released as Anthropics,
Fable 5 and OpenAIs, GPD 5.6.
So Moonshot AI's Kimmy 3 has just drawn this global attention after it was unveiled last week with the company saying it can rival top U.S. AI models in that they are supposed to be releasing the weights today.
So the allegation, though, from the U.S. matters because open source AI can spread quickly to anyone, which can lower the cost and also speed up innovation, but it can also intensify disputes over IP and model copying.
So Kratzios said that Moonshot likely also used restricted Nvidia chips powered by the GB300 Grace Blackwell platform,
which would be significant because the U.S. has limited export of Nvidia's most advanced chips to China since 2022.
So yeah, not only is the government saying, hey, Moonshot, you copied Anthropics Table 5, but they're also saying, well, you use.
are technology that you are not supposed to be using. So, you know, a lot of times that goes through an
intermediary country, right? So, you know, the U.S. will sell to country B, and then China will
buy from, you know, country B. So it goes from A to B to C, even though A to C is restricted.
So reports say that this is, well, it's getting worse and that now essentially both sides are just fighting, right?
China is saying that this is politicizing the trade and the tech of their country.
And obviously the U.S. is now saying that this is a national security issue.
We've seen reports that the U.S. and China are going to be having talks on AI soon.
So those will be some probably extremely highly watched talks.
Let's just say that.
So the U.S. Treasury Secretary Scott Bessent added Tuesday that Washington is reviewing
whether Chinese AI models have stolen capabilities from their American rivals
and said that sanctions could be considered if companies, well, if they can prove that companies
cross the line into IP theft.
Anthropic has also recently accused Alibaba of similar distillation attacks,
saying that it is becoming a broader fight over how AI companies train models and
protect their work.
So yeah, Quinn 3.8 came out from Alibaba.
We don't have benchmarks on that yet, but presumably it's going to be in the Kimmy K3 range.
So, yeah, things are heating up.
All right.
Let's leave that space for a second.
in and talk about just some real cool new tech will end the show with two of those.
So one and probably the one that I've been using the most and having the most fun with.
And I still don't even know how this is possible.
So if you haven't used this yet, my gosh, go give it a try.
But Open AI has brought like its new Jarvis style control to chat GPT work and Codex.
So yeah, it's not actually.
called Jarvis, but many people are just calling it the Jarvis style of using a computer.
Now, so OpenAI added its new GPT live full duplex voice model to the chat GPT work and codex
apps on Mac OS and Windows, which essentially lets people use natural language to manage your entire
computer. Yes. So just like an Ironman when you can just say, hey, Jarvis, go do
A, B and C, you can quite literally go to that now with Codex or chat GPD work with this new feature.
You can say, yeah, go, you know, open up all these programs on my computer, copy these files,
move them around, download them, upload them, put them in this program, edit them, right?
Anything that you could tell like an intern to do, you can now tell inside this new GPD live
voice mode.
So GPD Live now powers the chat GPD desktop app.
on Mac OS and Windows, and it is being tied directly into tools like obviously codex and chat
GPT work.
So the biggest change is that the voice system can listen and speak at the same time,
which means users no longer have to wait for that rigid turn taking during a conversation.
And the coolest thing for me, well, is you can use this with the remote feature on the
chat gbt mobile app, which makes it even crazier, right?
So you can literally just be, and I was actually doing this because I was traveling.
I was away from Chicago.
So I was in another state this weekend, opened up Chad GBT remote on the chat
GPT app on my phone.
I spoke to it and it's controlling my computer, you know, thousands of miles away.
And it's doing all these things by just talking into my iPhone, which is pretty cool.
So opening I initially launched GBT Live earlier this month.
as a continuous audio model that handles real-time speech
while sending heavier reasoning tasks to background models,
such as GPD 5.5.
So OpenAI says this update is meant to help software engineers
handle technical work by voice, including reviewing poll request,
debugging apps, and coordinating multiple coding jobs at once.
But I actually think it's really just great for manual,
any knowledge work, right?
I was just having it go through old, you know, files on my desktop, organizing things,
grabbing things from old transcripts, right, opening up doing things in Google Maps.
You know, just, I was just having it do all my work that I would normally do in front of a computer,
right, except I could dictate something, you know, just yap for like five minutes.
And I would check back in a couple of hours and it would do like a day's worth of work for me,
which was pretty cool.
So on Mac, the desktop app, can also use the screen context feature called app shots,
which essentially takes a not just a screenshot and automatically shares it,
but it also takes every other piece of content or context in whatever kind of program
that it took the app shot from.
And then it gives that to Codex or ChadGBT work as well.
In FYI, those are the same app.
Chad GPT work in Codex.
essentially the same app. So if you ever hearing me talk about that and confused, they're essentially
the same thing. But the app shots thing is really cool. Let's just say as an example, like I do now,
right? I have text edit open on my computer because sometimes I have bullet points there as I go
for the shows and things that I want to bring up. But you know, if the app shot could just take a
screenshot of that little portion of the text edit that's on my screen, but there's a lot of notes
on here. So not only is it just going to take that screenshot, but it knows that I have text
set it open and it's going to take all of that information and instantly, you know,
put it into the context window inside of chat, GPT work or inside of codex.
So this is literally the, I think, one of the biggest jumps in capabilities, probably since,
you know, I would say the, you know, Claude Co-Work slash Codex, kind of move.
of early 2026.
So I'll say of the last like four to five months,
this is the biggest both capability jump and the biggest like,
wow, what does this mean for work, right?
I'll probably do well, I'll actually put in the newsletter.
So you know, let me know if you want for our Wednesday shows where we normally do
AI work on Wednesdays.
We do the hands on demos.
So let me know if you'd rather see this new kind of Jarvis like GPT live.
on the desktop or our last story opus five yes there is a new model and it's currently wearing the
crown we'll see how long but we have a new most powerful model in the world surprisingly enough
it is not mythos it is not fable it is anthropics claud opus five so late friday actually
anthropic announced claude opus five a new model the company says
is its strongest and most cost-effective model yet,
with pricing set at $5 per million input tokens
and $25 per million output tokens.
So, yeah, it is on most benchmarks.
It is actually more powerful and better
than Fable 5 and Mythos 5, but at half the cost, right?
The one area where it's not as powerful
is kind of offensive cybersecurity,
but in most other benchmarks and just, well,
what you would use a model for,
Opus 5 is actually much better than Fable 5 and Mythos 5.
So Anthropic says that Obis 5 outperforms its previous public models,
including Fable and Mythos on coding and knowledge work tests,
and it is intended to be used as an everyday daily driver rather than only for specialized tasks.
So the lower price point, if you're using it via the API side,
is only part of the story because enterprises are obviously becoming increasingly more
cost conscious now in comparing AI models on value and not just capabilities. So the company also
says Opus 5 is not the top model for that risky dual use capabilities, including cybersecurity,
which Anthropic says they're still trying to balance the usefulness with safety concerns of their
upcoming and forthcoming models. So the Opus 5 launch comes as Anthropic and Open AI face pressure
from rivals offering lower cost AI tools, including Microsoft, Amazon, Google, meta, and even
open source Chinese startup. So yeah, you knew this one was coming, right? I've been talking about it
for literally a month, right? Ever since GPD 56 came out, you know, and Anthropic was kind of saying,
like, oh, we're going to pull, you know, Fable 5 from subscriptions. And I'm like, no, they're not.
You know, I literally said that they were going to be losing eight figures every single day that they did that.
And obviously, it didn't last long, right?
They never technically pulled Fable five from their most expensive subscriptions.
And it was only like two or three days that they pulled it from their $20 subscription until Opus came out anyways.
So yeah.
And I'd say most people, if you are terminally online like me following anything AI, I said there's absolutely no way
Anthropic lets this go on, you know, not having a capable model available in their subscriptions.
They would lose way to like literally tens of millions of dollars or billions of dollars a month,
but at least tens of millions of dollars they would be burning.
So it's great to see.
But I will say this.
Actually, let me go through some early reactions first.
So early reactions are kind of split on this.
So obviously on the benchmarks, looks really good, right?
in early reactions also highlight practical wins for teams, including better root cause debugging,
fewer over refusals compared to mythos and fable for defensive security work, and a little
extra token use versus prior versions for similar outcomes. But the main complaints so far are
operational. So users say that it often breaks backwards compatibility. So if you have a bunch of
skills that you would normally use with previous models.
And anytime you upgrade, it worked well.
They don't work as well.
I kind of found that as well.
And also that sometimes Opus 5 ends autonomous loops too early.
And it can produce overly verbose, what people call Claude Slop.
That can be frustrating in production.
And I saw that, you know, this was one of those models.
I didn't have a ton of time to use it.
So it was one of those models where eventually when I got to an output, I'm like,
oh, this output's great.
But it was the journey there that was.
absolutely like painful, right? Just just opus being, Opus five being so verbose and just so almost like
snoddy, right? And I think ever since, you know, my favorite introbic model, if I still had to pick
one to use, I think would still be like Opus 4.6, 4.6. I think it was a great model. And for whatever
reason, uh, that models ever since they've just been too verbose, just extremely token inefficient. And
ever since,
Anthropic started shifting toward this thing that they called truthfulness,
right?
Essentially,
and maybe it's just too heavy for my use cases,
because I'm always working with,
like,
things that are like not even days old,
like hours old,
right?
So a lot of what I use large language models for,
it's knowledge work,
but it's things that are literally breaking,
right?
Things that are,
you know,
days old or hours old or new concepts,
trends,
et cetera.
Right?
And,
you know,
the new,
even the,
the Fable models and even the new Opus 5, right?
I literally have to coax them and I have in special instructions saying,
hey, I work up to the hour.
So you're going to think that what I'm telling you doesn't exist.
Just trust me, it exists.
Always query the internet, all these things.
And it's just just refuses just straight up so many times.
Right.
So I think, you know, and after I use them like, oh, man, I can't bellyache about this, right?
Because it's a good model.
But luckily, you know, it seems like that's the takeaway case from a lot of people
that both had early access to it and just early reviewers,
is that like, yeah, obviously the capabilities are great,
but it's one of those models that's kind of actually painful to use,
especially if you're using a lot of your pre-existing skills.
So Anthropic did put out kind of a new kind of prompt engineering
or context engineering guide because they're saying,
yeah, these new models work a little bit different.
So we'll probably share that in our newsletter today.
So Anthropics own behavioral audits reportedly showed that Opus 5,
has the lowest misaligned behavior.
But yet, early testers said the model can overthink at those high effort settings.
And they actually may just work better on low or medium reasoning levels.
So I did see that anecdotally as well.
I always will run the same handful of prompts across different reasoning efforts.
And I actually saw that as well.
But again, I'm not using things that are overly difficult either.
So, you know, if you're refactoring, you know,
a code base with, you know, hundreds of thousands of lines of code, right?
Maybe you will find better results from a higher thinking level.
But I think for the majority of what people do, you know, I think we're probably getting
to the point now, right?
I'm using GPD 5.6 sole medium a lot, right?
And I'm not cranking up that ultra every single time.
I need an answer out of a large language model.
So, you know, I think maybe we're getting to the point where for a lot of people and a lot
of even enterprises using these models where, yeah, maybe the lower.
or medium reasoning efforts might work just fine.
All right.
So that's it for the big stories,
but let's quickly go over kind of the what's new and what's next.
So these are either just smaller news happenings this week,
some leaks,
some things that are already out and we covered in our Friday show,
but let's just quickly go over it.
So first,
Nvidia is reportedly in talks to back a $250 billion finance and deal
for Open AIs, Ohio.
data center. Open AI launched presence in enterprise agent platform with governance and deployment
controls for voice and chat agents. We covered that earlier this week. Stripe is reportedly in talks
to buy open router for about $10 billion. Alphabet reported its first ever negative free cash flow
as KAPX surge to nearly $45 billion. I think it's the first one since 2004. Meta added a lot of
under the radar updates.
They added desktop browser and mobile computer use support for Muse Spark 1.1,
and they also added some agentic features for connecting emails,
calendars, research, and tasks.
The White House Frontier AI framework is reportedly pending,
and it's expected as soon as this week.
Alibaba previewed their Quinn 3.8,
a 2.4 trillion parameter model that they're saying is close to Anthropic Fable
5 level, but we don't have any bench parts yet.
Anthropic officially settled and is paying out their $1.5 billion copyright settlement
for the fair use ruling against authors.
Microsoft and Mistral announced a multi-billion dollar sovereign AI expansion for regulated
customers.
Amazon cut a bunch of jobs in their AGI department and is reportedly shifting their AI
focus to prime video personalization. Yeah, that one was a strange one. All right. And now we have
some of the things that we went over on our Friday show. So the quick updates on those. And if you
want to hear more about these next ones, make sure to go listen to our Friday show. So Open AI
released chat GPT for health for US users over the age of 18 to track their health data and
summarize their records. Anthropic upgraded Claude voice.
with opus and sonnets and connectors.
So that's good.
You no longer have to chat with haiku,
much better with opus and sonnet.
Microsoft launched MAI image 2.5 for better AI image generation
and editing in copilot.
Google released a new model,
but yeah,
it's just going from 3.5 flash to 3.6 flash.
So nothing new there.
We're still seeing delays reportedly for Gemini 3.5 Pro,
but we do know that Google is pre-training,
Gemini 4. Google also expanded and released the Gemini Spark to pro users. Yay. So if you are a
Gemini pro user, now you have kind of their version of Obing Claw or Codex, whatever you might
want to call it. But Gemini Spark is now live. And then last but not least, Anthropic added the
record a skill feature in Claude Co-work to turn workflows into reusable skills. So yeah, if you've
use codex their version. This is essentially Anthropics version that watches your screen and whatever
you do, it'll create a skill, which is really cool. All right, that's it. A lot of AI news that mattered
this week, like this week and every week, it's hard to keep up. You can't spend eight or 10 hours
a day tracking and testing all this stuff like I do and talking to the industry experts. So if you
need to know what is happening in AI to make decisions for your company, just put me
to work for you. All right. So if you haven't already, please make sure to subscribe to the podcast
on Apple or on Spotify and then go to Your EverydayaI.com. So thank you for tuning in. Hope to see
you back tomorrow and Everyday for more Everyday AI. Thanks y'all. And that's a wrap for today's
edition of Everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and
leave us a rating. It helps keep us going. For a little more AI magic, visit Your EverydayAI.com.
and sign up to our daily newsletter so you don't get left behind.
Go break some barriers and we'll see you next time.
