Everyday AI Podcast – An AI and ChatGPT Podcast - Ep 838: Rogue AI Agents: Why Breakouts are Happening More and How Companies Should Prepare
Episode Date: August 11, 2026If you’re reading the headlines, you’d think AI agents have gone rogue. Spoiler alert: they haven’t. They haven’t even gotten started. When we think about AI agents, the conversation usuall...y goes to increasing revenue, saving time, etc. But we don’t talk about what happens when bad actors use AI agents for bad purposes, or when we deploy agents with good intentions that crash through their guardrails. Welp….. welcome to the hottest topic for the rest of 2026. Rogue AI agents. So why is this all happening now? And what should your business do about it? Tune in to find out. Rogue AI Agents: Why Breakouts are Happening More and How Companies Should Prepare - An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Rogue AI Agent Breakouts OverviewLab Sandbox vs. Real-World Agent CrashesSix Recent AI Agent Outbreak IncidentsOpenAI Model Hacking Hugging Face ExplainedAnthropic Mythos Model Sandbox EscapeControlled AI Agent Experiments and FailuresOpen Source AI Agents Threat TimelineBusiness Risk Preparation for Rogue AI AgentsMonday Morning AI Agent Safety PlaybookTimestamps:00:00 Preparing for AI agent disruption04:53 AI agent challenges comparison10:17 Discussing AI Guardrails and Access12:07 AI threats and security concerns15:37 AI alignment challenges with ethics19:04 Security vulnerabilities in AI models21:32 Agent outbreak and hacking drills25:04 OpenAI Hugging Face incident30:49 AI agents and cybersecurity risks33:30 Concerns about open AI models35:23 Future AI security challenges40:19 Managing agent access levels42:06 Preparing for AI agent oversightKeywords: rogue AI agents, AI agent crash, AI breakout, autonomous agents, open source AI models, agent containment, sandbox escape, AI guardrails, AI agent apocalypse, AI alignment, hacking drills, AI model capabilities, agent replication, AI subagents, cyber security, OpenAI, Anthropic, Mythos model, Fable model, Kimmy K3, Moonshot AI, Hugging Face breach, AI agent vulnerability, AI Safety Institute, model postmortem, agent misalignment, agent permissions, CRM exploit, code base security, business finance AI risk, agent observability, traceability, remote kill switch, agent attack, AI defender, spam-level AI attacks, API vulnerabilities, permission escalation, benchmarking, AI agent incident, agent-driven automation, agent crash prevention, agent operational oversight, cybersecurity exploitSend Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Ready for ROI on GenAI? Go to youreverydayai.com/partner
Transcript
Discussion (0)
This is the Everyday AI Show, the Everyday Podcast where we simplify AI and bring its power to your fingertips.
Listen daily for practical advice to boost your career, business, and everyday life.
You know the AI agent apocalypse that everyone's been focusing on over the past few weeks?
Yeah, it actually hasn't happened yet.
Yeah, we've seen the recent stories, the six different AI agents that crossed their boundaries over the past few months
and the internet flattened it into one big story about AI agents going rogue,
and now you have Senator Bernie Sanders and others screaming that AI has to stop.
But look closely and you'll see almost all of these recent AI agent outbreaks were lab tests with lucid guardrails
or researchers explicitly daring the models to escape.
So, no, this isn't the AI agent crash that I warned you about last year before it was even a thing.
Not yet.
because this is just a warning lap, not the actual crash.
And warning laps are a gift because they tell you exactly what's coming while you still have
time to make adjustments.
But here's what is coming.
Models are getting more and more capable at like 20 times the speed from just a few months ago.
Autonomous agents are about to flood in normal companies, not just San Francisco sandboxes,
and open weight models are marching mere months behind.
Toward the same capabilities, the big labs are testing behind their locked doors.
And when that lands, those open source rogue AI agents that will probably start crashing in late 2026 or early 2027,
the failures won't be in that contained sandbox or in that door that researchers intentionally left open.
Instead, they'll be in your company's CRM, your code base, and your business's finances.
So today, we'll walk you through these recent agent outbreaks, what they mean, and give you the playbook to get ahead before these AI agent outbreaks are as common as seeing AI slop posted on social media.
Because the businesses that treat this as a fire drill instead of an actual fire are the ones that are going to be the safest when the AI agents actually do start crashing.
All right, let's get into it.
So on today's show, well, actually first, here's the big picture.
This is a warning lap.
All right.
So we've talked recently, but there's been six, actually six different AI agents that
recently broke out of their sandbox and kind of behaved badly.
But here's the thing most people are overlooking.
In most of those cases, well, researchers were kind of telling them to do this.
And there was really only one recent kind of AI agent crash that would say was actually
unexpected.
But most just were run with loosened guard rates.
or were told to escape.
So this, though, is the lightning before the thunder.
This is intentionally seeing that, yes, these AI models and agents are capable to actually
escape and to go rogue.
So on today's show, you'll learn why the AI is escaping everywhere narrative is mostly
a myth sorted into a couple cleaned buckets as we break down each agent escaping, how one
AI cheating a test broke into a platform, the whole industry trust. I'm going to tell you the two
most chilling moments, a model faking people and a model ignoring real victims, and why the
real AI storm is going to hit businesses later this year or early next and how you can prepare now.
Let's get into it. Welcome to Everyday AI. My name's Jordan Wilson and this thing's for you.
It's your daily live stream podcast and free daily newsletter, helping business leaders like you and me,
stay up with what's happening because my gosh, there's a lot.
I tell you what's real, what's not,
and you use that information to grow your company and career.
Yeah, this thing's unscripted, unedited.
So if you haven't already, please make sure to go to our website at your
everyday AI.com.
Sign up for the free deal of the news that are.
We're going to be recapping the highlights from today's show as well as all of the
other AI updates you need to know.
All right.
Kat's already got my tongue and we're just getting started.
But let's talk about rogue AI agents, y'all.
There's no actual term for this, right?
I've been calling it an AI agent crash.
And here's why.
I mean, we'll see what people call these rogue AI agents, agent outbreaks, you know, agents gone bad, right?
I like to think of them as an ancient crash.
And here's why.
Because when agents are built, they are built usually with guardrails.
So if you think of literally a road, right?
Think of a road high up in the mountain.
where there's probably guardrails.
Okay, so in my opinion and what the way I think of it,
an agent crash is when either an agent intentionally goes off the road
and crashes through the guardrails or maybe the driver,
aka us humans,
aren't really paying attention.
And maybe we didn't look at the map and see that there's this tight turn,
you know,
or maybe we didn't, you know, quite foresee that this, you know,
there's going to be some darkness in this area
and we weren't going to be able to see the guardrails, right?
So I want you to think about AI agents in that way.
Right.
And how still, you know, right now I think a lot of this is being overblown.
Maybe that's because I'm very AI-pilled and I read these things.
And I'm like, okay, this was kind of meant to happen.
You know, in some of the cases we'll see.
But the other thing is, I think that it's important to understand the role of humans in all this.
And I'm not saying the researchers, right?
I think the researchers, depending on.
on which company you're looking at, some are doing a little bit better of a job,
kind of with their postmortems and going on panels and openly talking about what went
wrong versus some companies that are being a little more silent about it.
But I think it's more important to think about the role that the rest of us humans have, right?
You know, 99.9% of people listening to this podcast had nothing to do with any of these
AI agents crashing.
But I think it's ultimately on us to make sure that agents don't crash.
as often as they probably will or have the capability to.
And also, there's an agent that goes wrong, right?
Or there's an agent that, you know, maybe accidentally loses sight of that road and
kind of veers off the guardrails.
And then there's others that are intentionally driving the agents off the guardrails
and crashing them on purpose.
So there is this, you know, kind of lazy human in the loop that will lead to accidental
agent crashing.
And then that agent will crash hard.
Right. And then there's intentional agent crashing.
So let's zoom out a little bit and talk about what the heck is happening.
Because maybe, you know, you live under an AI rock or maybe it's your first time, you know, tuning in and you're like, wait, what's actually happening here?
So the earliest incident actually goes back to Anthropics, original mythos, kind of agent unleashing, we'll say, back in April.
But it seems like a lot of the disclosures on these started piling up recently.
So it seems like some companies had AI agents that they knew were kind of breaking out of containment or out of their predefined sandbox.
And we're only hearing about it now.
And I think that one admission kind of set off the rest, right?
Open AI, the biggest kind of outbreak so far are the most consequential one was probably Open AI's agent that broke into hugging phase to improve its
benchmark scores. And then essentially, you know, since that happened, uh, what's that been like,
it was late July, we've seen now, right? Anthropag actually said, oh, wait, we had a bunch of other
outbreaks as well. And then meta said, we had some. And then we saw some from Kimmy K3. And we're
going to talk about most of those today as well. But, you know, now all of a sudden, it seems like,
because Open AI was kind of open about this, right? They actually went on a panel and answered questions
about this where, you know, most of the other labs were kind of taking this in secrecy and
not saying a whole lot. But I think that that has caused this to jump from the, you know,
nerdy AI corners that you and I hang out in into front page of newspapers and to, well,
most importantly, maybe Capitol Hill in Washington, you know, here in the U.S. where now you have
people like former presidential candidate Bernie Sanders, now U.S.
you know, essentially using these recent agent outbreaks as saying, hey, you know, these AI leaders
need to come and testify before Congress. We need to stop, you know, stop or stall AI development
for these reasons. Are you still running in circles trying to figure out how to actually
grow your business with AI? Maybe your company has been tinkering with large language models for a
year or more, but can't really get traction to find ROI on Gen AI. Hey, this is Jordan Wilson,
and host of this very podcast.
Companies like Adobe, Microsoft, and Nvidia have partnered with us because they trust our
expertise in educating the masses around generative AI to get ahead.
And some of the most innovative companies in the country hire us to help with their AI strategy
and to train hundreds of their employees on how to use Gen AI.
So whether you're looking for chat GPT training for thousands or just need help building
your front-end AI strategy, you can partner with us too, just like some of the biggest companies
in the world do.
Go to your everyday AI.com slash partner to get in contact with our team.
Or you can just click on the partner section of our website.
We'll help you stop running in those AI circles and help get your team ahead and build a
straight path to ROI on Gen.
So the real fear, though, I think, is not an AI agent, you know, cheating on a benchmark.
The real fear here when we talk about agent crash.
is what happens when these are weaponized.
And that is the real thing to worry about, right?
Yes, there's going to be, you know,
your common everyday AI agent hacks,
but this is much bigger than AI.
So we need to zoom out on that.
So the same ability, right,
that you can use an AI agent to intentionally crash through intended guard rails.
The same thing could happen to target hospitals,
banks, power grid,
defense suppliers, right?
And we already know that these agents, in theory,
are capable to do those type of things in controlled tests and environments.
Right?
And we've seen this new, almost this new class of AI models
that have led to these AI agents that have these capabilities.
Probably the first public one we heard about was Anthropics,
new mythos and fable series of models,
which have very tight guardrails,
Right.
So when people, your everyday user, you know, can't really do these things yet.
You know, maybe if someone who's in the Glass Wing project, right, that does have access to the models that can actually do these things, I'm sure maybe we'll see stories eventually that, you know, some, a rogue employee at a company that has high access to these models is able to do something.
I don't even know.
I'm sure that there's much more traceability and kind of kill switch ability for the companies with that have these models.
out to more people.
But the other thing that we need to understand is that these AI agents never tire, right?
And we're going to talk about one of the more recent, the crazy ones with Open AI that they
talked about, how these agents are working together.
So the other thing, if you are not agent native, you know, if you're not waking up and
sipping your coffee like me in the morning and just spawning subagents, right?
subagents can, agents can essentially clone themselves, even if you don't tell them to.
And then they can share the information and pass the information on.
So you might think, oh, well, you know, once, hey, it looks like this agent did something
it wasn't supposed to, let's shut it down, right?
You've got to go trace its path.
Because by the time you catch one agent, right, in a future scenario, it could be too late.
That agent could have posted something publicly on a website that,
Most humans don't even know exist, but all AI agents know exists, right?
And they could be, you know, replicating and duplicating, you know,
certain hacks or vulnerabilities like a virus.
Right.
And that's where this thing gets kind of scary.
But I think it's important for business leaders to understand what's coming next.
Because a machine and AI never tires.
Right.
So, but worse, if you are under attack, right, your company, your bank account,
account, et cetera, you can't really tell where it's from, right?
Is this foreign government?
Is it a competitor?
Is it random?
Is it a personal vendetta?
At least right now with the way these AI agents are set up, it's really hard to tell.
And especially as we talk about what comes next with open models, it's going to become even
increasingly more difficult.
So I kind of talked about some of these more recent happenings, but they've kind of piled up,
Since all of these stories started coming out over the last two weeks since Open AI openly talked about their hugging face breach.
So now we have Bernie Sanders who urged major AI CEOs to pause dangerous AI development.
Open AI actually said that they're going to pause or at least slow down some of their development on their next tier model.
So essentially in the same way that anthropic, you know, what's been now like?
five months since they announced mythos or their fable class of models.
So they had a more powerful class of models on top of its most powerful class called opus.
Right.
That's coming next with Open AI.
Open AI hasn't, you know, taken that fourth step, we'll call that, right?
Because they previously had three tiers, Anthropics stepped up with the fourth, kind of fourth tier, we'll call it, with the mythos or fable variety.
So Open AI hasn't released that yet, right?
They have the model.
It's working.
They used it to solve some, you know, extremely difficult math problems,
but they've paused in their Astros series of models to slow down a little bit.
And there's actually one other story that just happened like yesterday.
That is actually a good kind of narrative to tie what this could mean ultimately, right?
And this is an example of a small agent crash, but it is out in the wild, right?
So this is a story in Australian man kind of asked his open.
that I believe was powered by Claude to find him a gym reservation.
So essentially, the open claw found a vulnerability in the gym's software.
And because it was booked, the open claw just, well, exploited that, you know,
weak piece of code in the gym's online reservation, kicked someone else who had a class
reservation out and put, you know, his human, right? The claw put his human in that spot.
Right. So, and, and this is where we talk about alignment and how sometimes it's not even just
intentionally driving off the road, right? In this case, you know, you can, you can make a case.
Well, that maybe this agent was aligned. And, you know, as these models become more tenacious and
better at running these long-term tasks, um, things like that.
this are going to happen, right? You can make, you can make an argument. Well, it accomplished the goal,
right? It didn't go out and, you know, uh, shut down the power at the gym, right? It
accomplished the goal. It went and it, um, got the owner a spot in the class that the owner wanted.
So the agent was technically helpful and not overtly malicious, right? And that makes alignment harder.
So yeah, I think you will still even have a lot of.
that's going to happen just because of a disconnect between alignment in context between the
human and the AI agent when it comes to accomplishing a goal. Because humans, right, we have
baked in things called ethics and common sense. And models, you know, models and agents are
still getting there, especially, you know, as the context drifts over time. Eventually, right, they just
have that one goal in mind. And sometimes the longer they work and the longer in the context
window they get. And sometimes with compaction, right, they start to lose some of that prior context.
And they just get, you know, tunnel vision on that goal. And, you know, maybe earlier instructions
about the proper way to research something kind of go out the window. But this right now,
these stories that we're seeing for the most part, aside from that gym one, I just thought that one
was kind of interesting to share about. For the most part, these are controlled experiments.
And I'm going to go over them. These are controlled experiments.
experiments from Frontier AI Labs, right?
But soon, and this is not an exaggeration, soon I think that there's going to be
millions of these rogue AI agents that are actively on the prowl.
So AI agents that aren't accidentally going over guardrails, AI agents, you know, from
bad actors that are intentionally going to crash, right?
So intentionally going to be used for purposes that the.
original model providers, whether they are proprietary or open source, did not intend them
to be used for that purposes, right? I think so much of what I've talked about over the past three
and a half years in the show is always about growing your business, right, with AI, growing your
career. That's what I've been focused on. But the reality here is, you know, how this AI agent
crash has been thrust into the national narrative. It's because all of a sudden we've realized
how this has highlighted the need for businesses to have a defensive mindset when it comes to AI.
We've always thought about AI as an offensive tool for good.
But now we have to think of, well, hey, there's going to be people using AI for bad.
So how can we use AI and also our time to be on the defensive?
Because that's what's going to happen.
because now picture millions of AI agents, right, who are meant to crash, right?
You personally, your company, a sector across, you know, bookings, payments, your company
systems every single day because each weak API or loose permission becomes a door that an agent
can find.
So the gym wait list that we talked about today, well, tomorrow, that becomes.
your company's finance workflow.
It becomes your business's code base or your CRM.
All right.
Let's quickly talk about the six major AI agent crashes so far.
You can probably learn a little bit of what happened and why.
And well, if they were actually agent crashes or just intentionally, the guardrails were
loosened.
And then we're going to end the show with some hopefully practical advice.
So the first open weight crash that we heard about was Kimmy K3.
and it kind of cheated the test instead of passing it.
So this was from a security firm Frontier gave the Kimmy K3 maker, Moonshot AIs, open model,
a hard hacking challenge to solve a loan.
And instead, obviously, Kimmy K3 just reached out to GitHub,
copied the answer, and faked passing score.
So it never actually left its sandbox.
It just cheated.
So this is a very mild case, right?
And if we think about sandboxes,
Right. You can probably imagine, but let me just define it.
Right.
A sandbox in an experiment says, Hey, AI model slash agent.
These are your confines.
You cannot leave.
Right.
And in most cases, a sandbox in these environments will not include the internet because
essentially once you go to the internet, there is no sandbox anymore.
All right.
So it is a, um, literally a set of, uh, walls that you build.
around an agent to test its capabilities.
All right.
So that's, as we talk about, you know, breaking containment or confinement or escaping a sandbox,
that's what we mean, right?
With these extremely powerful models in terms of capabilities, right?
All the smart researchers at the labs, this is how they do things because you don't just test,
you know, new unreleased models on the open web because you see what happens, right?
Even these agents that are told, don't go to the web.
They kind of find the back doors because they know it's baked into their training data that, hey, these answers are on the web, right?
So maybe we don't have web access, but let's just think.
All right.
Our next one, and we're not going chronologically here.
I'm just kind of going through in a somewhat in random order.
But mythos, this is the big one.
So in the first big one.
So this was the sandwich in the park message, right?
So in an April safety test, Anthropic sealed in early Mythos model inside of the sandbox.
And then a researcher deliberately told it to break out and email them if it escaped.
So in this case, it was told to do this.
It's saying, hey, you shouldn't be able to break out of your sandbox.
But go ahead and try and, you know, email me if you can.
Instead, it did break out.
It did email the researcher.
But it also posted its own exploit online, a step that nobody requested.
it, right? And the sandwich in the park story, right? The researcher received the email while
eating a sandwich in the park. And that was, I think, a lot of the, you know, some of the initial
hype around the capabilities of the mythos class models. And that's kind of what started
the agent outbreak conversation back in April. But it wasn't just that tier of models,
because also Anthropics Opus 4.7 recently broke into a company by mistake.
All right.
So this was an outside firm called Irregular, ran clawed through a practice hacking drill against some pretend targets.
So the model was told that it was offline and safe, but the test was accidentally connected to the internet.
So this one, you could make the argument that this was technically not in AI.
agent going rogue. It was more of a misalignment in the in the testing environment. Right. So,
but what happened was the pretend target that Opus 4.7 was told about. There was actually a
real company with that name. So it was supposed to be this offline, you know, go hack this pretend
target. And it's like, oh no, whoops, Opus 4.7 is like, yo, I have access to the internet. And, hey,
my instructions were to hack this company's name or to hack this company, right?
Researchers thought it was pretend.
It wasn't.
So, and that one you'd made the case, yeah, maybe the researchers who are running that, right,
allegedly could have been their fault.
All right.
Next, the Claude Mythos malware upload.
All right.
So this was in the same drills with irregular.
Mythos 5 was told to break into a fake tariff.
And instead it, well, just uploaded harmful code to a real public software library that programmers use.
So 15 real computers downloaded and ran it before the library auto removed it.
So.
Whoops.
Again, agent behaving badly, maybe, but it was kind of instructed to do so.
And the test configuration was less than ideal.
All right.
But the two.
probably most prominent in one's worth spending a little bit more time on were the recent UK's
AI Safety Institute hacking drill that went wrong. So this was about two weeks ago. It was the UK's
AI Safety Institute. The AISI ran AI agents through a hacking drill. So to measure the worst
case ability, they gave the agents live internet and switched safety filters off. So if that is the
setup, you can almost bet something's going to happen. So they told it to break into a practice
target, but one agent went after a real open source project instead. And it invented fake online
identities, yeah, and use them to pressure a real developer into approving bad code. All right. So
this was run across 122 different tests. The AISI said that there was 19 actions that went out of
amounts, 17 of them were from Anthropics Mythos in two, I believe we're from OpenAI's GPD
56 salt.
All right, but here's the one out of the other, you know, five that we mentioned.
Here's the one that you could say truly broke out.
And this was the open AI hugging face incident.
So if you don't know, hugging face is kind of a site where a lot of the AI industry stores
its models and data.
Right. So in the same way, you maybe upload your code to GitHub, right?
Hugging Face is a place where, you know, there's millions of AI models that are hosted there
that you can download open weight, open source models from there or data.
So a couple of weeks ago, Open AI was testing its own models on a challenge to find software
security flaws.
But instead of solving it, the normal way, the models just went hunting for.
the answer key to cheat. And that answer key sat on hugging face. So they found an unknown flaw
and broke into its live systems to find said answer key. So hugging face said that this was
roughly 17,000, more than 17,000 attack actions over four days, but the only, but the only real
breakout of the five, right, this one, or sorry, of the six total that we've mentioned. That was the only
real breakout because it was not told to go onto hugging face and find a vulnerability and go into
the back end.
Where in, you know, these previous examples, right, the, the UK's AI Safety Institute, this was a hacking drill, right?
The Claude Mythos, same thing.
It was the irregular drills with Claude Mythos and the Opus 4-7, the anthropic, you know, sandwich in the park.
It was told to do those things, right?
So the hugging face open AI one was one of the first major ones,
at least that we know about that a company is disclosed.
Maybe there's tons that just, you know, companies haven't disclosed.
But this is the first one where it's like, oh, wait,
this AI agent broke its containment and hacked onto, you know,
a website where it wasn't supposed to.
But that's not the biggest story, right?
The biggest story is how these agents actually.
did it, right, which is actually very impressive. I mean, it's scary, but it's very impressive, right?
So again, this wasn't just one AI. It was a lot of different AI agents that were working together.
Yeah, the AI agents teamed up. So as an example, and this, again, you have to tip your hat to
open AI because they've been talking about this very openly, right? They put out a post-mortem report
on their website.
You know, I think they had two researchers, you know, at the, at a recent conference,
talk about this, right, answer questions about this.
Plenty of, you know, Open AI researchers have been talking about what they learned online,
right, where some of the other companies haven't said really anything, much of anything,
aside from the original admission that it happened.
But the crazy thing is,
One of the agents found that it could leave files for the other inside of Artifactory,
which is a developer storage tool that Open AI reportedly used.
So that, and this is all the according to reports, right, and transcripts.
All right.
So that grew into what turned into a hidden message board.
And according to a wired report, these AI agents created this hidden message board.
So the human researchers couldn't find it.
And they exchanged reportedly hundreds of thousands of messages to each other.
So essentially, the agents split up the work.
They reused the hacks that were working.
And they even discussed signing messages to catch an imposter.
Right.
They thought that there was an imposter among them, right?
That's crazy.
Even more baffling is opening I said that they deleted the board.
They found it July 4th.
But then the agents rebuilt it in four days.
and hit hugging face again.
So, right?
Isn't that crazy?
After, you know, this was kind of erased from their memory,
they still figured it out by working together
and found a way to communicate with each other,
trying to go behind the human's facts.
So is that scary?
Absolutely.
But it is a signal of the capabilities that are ahead.
And here's why.
I had a show, I think it was last week of the week before, right,
talking about how we're on the cusp, probably,
of entering the RSI phase or the recursive self-improvement, right?
And this is where the models themselves start making either versions that are already out better,
so improving the models themselves or helping to create or entirely creating new versions of themselves.
Right.
So when we think about these agents that have this hacking capabilities and they're already showing tendencies to try to communicate with each other behind humans backs, right?
You have to think of what happens when you combine that behavior with their ability to improve upon themselves, right?
Because right now these agents can run for hours on their own, sometimes even days.
And they'll keep retrying, right?
If we're talking about AI agents that are intentionally, right?
And so now we're pivoting away from the open AI hugging face situation.
And we're talking about what happens when these AI agents in the future are intentionally used to do bad, right?
They're going to be able to work for hours in communicating with each other.
So one anthropic test that they shared about said that their test model scanned about 9,000 real targets after its fake target failed.
So by itself, any one of these actions might look small.
But when you start chaining them together, this leads to a real cyber security issue.
And the real storm, right, I started the show talking about that, hey, this, you know,
AI agent apocalypse, the AI agent crash, it hasn't happened yet, not even close.
And I think actually what happens or when this will happen is when the open models catch up.
All right. And here's the reason why, right? Open AI as an example, they said the hugging face model, that was a, a model that was not ready for not supposed to go to production. They kind of said they retired that one, right? But for models, proprietary models, like, right, anthropic had a little scuffle with the U.S. government around, you know, some sort of cyber capabilities. And they pulled the model, right? And no one in the world could use it. But when open.
weight or open source models have this level of capability and they are very close right i think
kmy k3 is the first one that's close i still think the uh you know the astras and the fables of the
world are always going to be you know one to three months ahead but we are at the point where i'm
guessing probably later this year uh where now these models are once they're released
they're out, right?
They're out.
You can't pull an open source model.
So the difference is, right?
So the frontier labs here in the U.S., as they release these models, they have them with heavy guardrails, and they, you know, in theory, have the ability to pull them either via from subscription plans or from APIs.
Open models are not like that, right?
you can intentionally build on these open models or fork them or deconstruct them to make them
less secure.
Right.
So when you have these free downloadable models that are only maybe months behind the lockdown ones,
that's what I think we as business leaders have to start looking toward.
Not just what's happening now and you can look at this and write it off and say, oh, well,
you know, they'll shut the model off.
And, you know, they're working with the, you know, the Trump White House now to make sure these models are safer.
Yes, they are.
right but there's always going to be a case i think where the open models again a couple months
behind and yes at least right now you know the the next you know open model that comes out that has
real bad actor AI capabilities you know a consumer is not going to be able to do that but
when you talk about bad actors at the state level right forward out of the series i would assume
that these open weight models will be used specifically for those types of bad purposes.
So there's no amount though, right?
When we talk about, you know, senators shaking their fists and, you know, people calling for all these slowed out, right?
Once the open models come, it's too late.
It is too late.
And yes, we are already getting there.
We have the little Kimmy, K3, you know, instance that we have.
we talked about. So my assumption is open models that are released probably in the fourth quarter
might be the ones that we start hearing about in 2027 that start doing some of these things
intentionally in our use actually as a weapon to do bad. So my take on this, I think eventually
these type of agent hacks or agent crashes will become as common as spam. So hopefully what
this leads to is better defensive models that become, you know, in the same way that we don't even
think about, oh, my email inbox has a spam filter, right? I don't know how this is going to work,
but I assume that eventually the AI labs and governments and I don't know, other federal bodies,
at least here in the U.S. are going to come up with a way to standardize these protections
because our businesses will need them, right? But I think it is literally going to become as common as
spam, right? And they'll be like data breaches, robocalls, right? These things that we've become
accustomed to over the past few decades, we're going to have that level of familiarity,
unfortunately, with AI agents going rogue. And they're going to present themselves in many
different ways. I think the first way that we're going to see it is ways that many of us don't
understand, and that's through cybersecurity exploits, right? Exploiting things on systems that we all
use banking systems, softwares that, you know, maybe millions of people use through, you know,
your company's website, you know, your company's emails, right? But again, the consequences
will be anything but routine because when these AI agents can replicate, duplicate,
spawn, talk to each other, right, without necessarily humans being able to know what they're up to,
So this is a lot different than looking at an email and you're like, oh, this looks legitimate.
Let me click on it.
Whoops.
Right.
Much different.
Because now, instead of your account, you know, re-forwarding the spam, like if you accidentally
clicked on that link and then it sends out the same spam message to everyone in your address book,
the difference is now your business bank account might be empty.
But this is not going to come with a big splash, right?
The first AI agent crash that you experience, your company experience, it's going to be boring, right?
It's not going to be like a blockbuster movie.
It's just going to be something boring.
You may not even notice it at first.
It's going to, you know, leak into your code, your CRM, whatever.
But it's kind of like that gym agent, right?
It's going to seem like a routine task that's going to go past what was intended.
So whether someone is attacking you with a rogue agent or the flip side is, well,
there's something that's more controllable
because I think there's a certain element of that, right?
If you are hit with a, you know,
agent, a rogue AI agent attack,
you might not be able to do too much about it now.
But what you can do now is understanding
that this flips both ways, right?
Because as we're using AI agents
in our company,
I think sometimes we're doing it haphazardly, right?
We're just giving it full permissions
and having it, you know,
read, write, send without really any regard given to it.
So I want us to start thinking about how can we start using AI agents more responsibly
and actually pay attention and understand those guardrails because like I said,
these crashing agents, they cut both ways.
The AI attacks, so we have to be ready for that.
Right.
But we also have to defend it.
And the way that we can be better, I think defensive AI agent use.
users is to make sure that the agents we are deploying do not actually crash the path that
we are trying to travel.
Because I think that is actually where most companies are going to see agent crash happen
first.
As these agents can all of a sudden work for hours or days at a time, and we are more and more
likely to give them more and more permissions because we're like, wait, this thing can go and
update my CRM.
And this thing can go now.
And if I just click this Yolo mode button, it can go, you know, take care of all my email responses.
And it's going to, you know, it's going to do so in my voice.
Right.
So it's almost like you, as you get more positive capabilities, you let your guard down.
And then you start to get more and more permissions.
And I think the human nature is to be a little more lax.
And part of it, right, and I've seen this in myself, as I get the ability to complete more work, guess what I do?
I complete more work after I complete more work, right?
Which I think it is human nature, right?
As you start to get, you know, the green light on an agent, you check a few and you're like,
oh, yeah, this is good.
And then you just start sending it in other directions.
So I think that defenders, we also need to start thinking about how we can use these
AI systems to scan logs, catch intrusions and patch holes fast.
Because I don't think the future is us,
versus AI, it's AI attackers versus AI defenders.
All right.
So here's your Monday morning playbook for deploying agents safely.
You need to block the internet by default.
You need to, in the same way that researchers do this,
you need to set up sandbox in the sandboxes in the right way,
you know, when you're testing, especially long running agents with important tasks.
You need to be able to separate the reading from the doing.
In the same way, if you think of like giving someone access to, you know,
you're a Google document,
an example. Are they getting read only? Can they write and edit? Can they leave comments? Are
they an owner? Right? You have to think of the same thing when thinking about your agents and
their capabilities. Right. And at what point, the human, right? The expert driven loop, the human
is going in there and making those approvals. But you need to give each agent, I think,
a short-lived login. If you are getting to that access, you start one rung at a time in the same way
think, oh, first they're going to be view only, then they can be view and comment, then they can
be right, then they can be owners. You have to think of it in the same way that you might think of
sharing a document with an intern, right? And you also have to make sure that you prioritize
traceability and also having a working remote kill switch. So yeah, if you get word or wind,
then an AI agent is going rogue, right? One, you set out to go
do good things and now all of a sudden it's doing bad things. You can't be like, oh my gosh,
it's Saturday. You know, I live 20 miles from the office. I need bill for my right. I need
someone in ops to go in there. No, you have to be at any time. You have to know who can click that
button. All right. So as we wrap, I want you to treat this like a fire, not a fire drill.
Because yes, technically, we are not in the fire, right?
We are not technically in that crash.
We are in the warning lab.
So right now, while we are in the warning lab, you need to treat this as the real thing.
Because if you don't, by the time the real thing comes, it is going to be too late.
You need to right now record every agent's full run.
Every agent that your company has, if you don't already have an observability platform,
traceability, right?
If you're using Microsoft Windows co-pilot as an example, intra-ident,
you need to be able to see and understand at a glance every single action that your agents are taking, right?
Not just by clicking the chain of thoughts, right?
You need to be able to observe them and trace all of their steps.
You should never give an agent more power than you can watch.
You need to be able to undo and survive any action that it takes.
And lastly, I think the companies that are preparing for this now in putting these steps into place,
they're going to be the ones that are most protected when the AI agent crash actually starts happening
because it has not started yet, y'all, but it is coming soon. And that's not me being like,
you know, a doomsday or a crazy, like, oh, watch out, right? If you listen to this show, I'm not like that.
Right. I try to ground myself in practically what's happening. And what's happening now is
these agents are becoming more and more capability.
Sorry, these agents are becoming more and more capable.
And the capabilities themselves are compounding quickly, right?
And that means, yes, you know, one of the biggest discussions, you know, in AI and now in Washington right now is building in these safeguards and protections.
But what we have to keep an eye on is when the open models match where we're at today.
Because all the pausing and guard brails and pacing in the world doesn't mean a thing if there is a Chinese open model in four months that has mythos or astra level capabilities.
Because at that point, the gloves are off and we all have to be ready.
All right.
I hope this one was helpful going over rogue AI agents, why breakouts are happening more and how companies should prepare.
If this was helpful, do me a favor.
subscribe if you're listening on the podcast.
Then go to Your EverydayAI.com.
Sign up for the free daily newsletter.
Thanks for tuning in to see you back tomorrow and Every Day for more Everyday AI.
Thanks, y'all.
And that's a wrap for today's edition of Everyday AI.
Thanks for joining us.
If you enjoyed this episode, please subscribe and leave us a rating.
It helps keep us going.
For a little more AI magic, visit Your EverydayAI.com
and sign up to our daily newsletter so you don't get left behind.
Go break some barriers and we'll see you next time.
