Tangle - OpenAI’s autonomous agent hack.
Episode Date: July 29, 2026On Friday, July 24, Reuters reported that an OpenAI autonomous artificial intelligence (AI) agent broke out of a contained environment and hacked a popular AI development website. The hack b...egan on July 11 and lasted until July 13, when it was detected by an AI at the targeted company, Hugging Face. OpenAI remained unaware of the incident for days, and the two companies did not communicate directly until July 20, well after Hugging Face had alerted the FBI of an intrusion. Ad-free podcasts are here!Get 20% off your first year of ad-free episodes, exclusive interviews, and deep dives with Tangle’s podcast membership!We won!A few weeks back, we asked Tangle readers to support our nomination for the inaugural Newsletter Awards. On Friday, we learned we won the top prize — the People’s Choice award — for receiving the most votes among over 100 nominees. Thanks so much for voting for us; every bit of recognition like this helps spread the word about our work. It’s a proud day for the entire Tangle community. You can read today's podcast here and today’s “Have a nice day” story here.You can subscribe to Tangle by clicking here or drop something in our tip jar by clicking here. Take the survey: How concerned are you about AI development? Let us know.Our Executive Editor and Founder is Isaac Saul. Our Executive Producer is Jon Lall.This podcast written by: Isaac Saul and audio engineered and edited by Dewey Thomas. Music for the podcast was produced by Diet 75.Our newsletter is edited by Managing Editor Ari Weitzman, Senior Editor Will Kaback, Bailey Saul, Audrey Moorehead, and Carina Pacheco. Hosted on Acast. See acast.com/privacy for more information.
Transcript
Discussion (0)
From executive producer Isaac Saul, this is Tangle.
Good morning, good afternoon and good evening, and welcome to the Tangle podcast,
a place we get views from across the political spectrum, some independent thinking,
and a little bit of my take. I'm your host, Isaac Saul, and on today's episode, we're going to be
talking about the open AI hack. It is a fascinating story, and I've got some strong feelings
about it, so I'm excited to jump in.
In the meantime, I just have to take a moment to acknowledge.
I'm in Philadelphia today in our Philadelphia studio, and I'm packing it up, which is crazy
to think.
I spent four years here in this wonderful city.
My wife and I recently moved and bringing the studio and all my stuff to North Jersey,
and so much has happened here in the last four years.
Tangle has grown so much.
I spent a lot of time in these walls.
with our associate editor, Audrey Moorhead and Lindsay Canuth.
The podcast has grown.
The YouTube channel has grown.
The newsletter has taken off.
We witnessed the 2024 election.
We had the This American Life feature so many new listeners and readers.
And these kind of transitory moments make you stop and pause and reflect.
And yeah, I'm incredibly grateful for this audience, for this thing we're building,
and for the next chapter.
I'm really excited about the studio that I'm building in Jersey
and about what's going to come and what's going to happen
when we get in that studio.
It's going to be awesome.
And there's a lot more to come on the podcast
on the YouTube side that I'm really excited about.
But since I'm down here today packing up,
I just wanted to take a moment to thank all you guys
and let you know that there's a new chapter upon us,
which I'm stoked for.
and I'm looking forward to what's coming down the road.
So with that, I'm going to pass over to Audrey Moorhead,
who I'm here today in the studio with,
who will be co-hosting the pod today,
and I'll be back for my take.
Thanks, Isaac.
First up, we have today's quick hits.
Number one, U.S. Central Command said that Iran launched
an attempted surprise attack against U.S. forces in the Middle East,
but all missiles were intercepted.
The U.S. and Saudi Arabia also conducted
joint strikes against Iran-aligned terrorists in Iraq.
Number two. Kentucky Governor Andy Bashir, a Democrat, sent the office of Republican Senator
Mitch McConnell a letter requesting that the senator directly and verbally address the people
of Kentucky and provide proof of your capacity to serve or resign.
Number three. The Senate voted 86 to 12 to pass a major sanctions package against Russia
that was spearheaded by the late Senator Lindsey Graham, a Republican from South Carolina.
Number four. The Senate voted 51 to 10,000.
to 47 to confirm Jake Clayton as Director of National Intelligence.
Number five.
Japan experienced a 7.1 magnitude earthquake, which killed at least three people.
Rescue efforts are ongoing.
A stunning breach is the talk of the tech world.
Artificial intelligence leader, OpenAI, has revealed that one of its cyber agents went rogue
and hacked another company.
The world woke up to a different reality today.
Science fiction has become science fact, and even the scientists themselves seem worried about it.
On Friday, July 24th, Reuters reported that an OpenAI Autonomous Artificial Intelligence, or
AI agent, broke out of a contained environment and hacked a popular AI development website.
The hack began on July 11th and lasted until July 13th, when it was detected by another AI
at the targeted company, Hugging Face.
Open AI remained unaware of the incident for days, and the two companies'
did not communicate directly until July 20th,
well after Hugging Face had alerted the FBI of an intrusion.
OpenAI said the incident occurred while it was testing two of its models,
its latest publicly available version, GPT 5.6 Seoul, and an unreleased model,
on their offensive cyber attack capabilities.
Both models had guardrails that protected against high-risk behavior turned off during the test,
and were intended to be limited to a sandbox environment
with tightly controlled access to the broader internet.
However, the models were reportedly able to exploit a zero-day vulnerability, which is an unknown
security flaw, in third-party software designed to allow them to install packages to gain internet access.
From there, the AI agents inferred that Hugging Face may have information to help them pass their
assessment and use stolen credentials to hack into the site, stealing information to allow them to pass
their tests.
On July 21st, OpenAI released a statement saying it is working with a third-party software to resolve
the issue and collaborated with Hugging Face to build stronger protections against AI hacking.
Clement Delang, the chief executive of Hugging Face, said he had met with OpenAI and requested that it
released detailed logs of the incident and contribute $100 million toward computing power to help develop
cybersecurity networks. Delang said, quote, the first autonomous agent cyber attack is an unprecedented
event. It deserves an unprecedented response. Tech experts were divided in response to the scale and
ramifications of the hack, with some stressing that AI agents are more capable than ever.
Sean Cassidy, chief information security officer at Fintech Solutions firm Plaid, said,
quote, Today's the most important day in the history of information security thus far.
For the first time ever, an AI model escaped containment and hacked a real company's real
production infrastructure. This was unintentional and non-malicious, but that doesn't matter.
Others blamed existing cybersecurity vulnerabilities for the hack.
Veteran security engineer and researcher Niels Provost said,
quote, this should not have happened.
I wish the frontier labs spent as much time on teaching their models to write secure infrastructure
as they are spending on them exploiting vulnerabilities.
Next up, you'll hear from the right, left, and tech writers about the incident.
Then, executive editor Isaac Saul gives his take.
We'll be right back after this quick break.
First up, what the right is saying.
The right is mixed, with some calling for measured skepticism about AI.
Others worry about bad actors using AI's capabilities for ill.
In the Washington Examiner, Sam Corkas said the hack shows why we shouldn't trust AI.
The issue is not that a rogue AI model was used to attack an American company.
Instead, the issue is that the AI was doing what it was told as efficiently as possible,
and that led it to hack into a database it identified as a suitable target.
AI is inevitable. It is becoming more integrated into our everyday lives, whether we want it to be or not.
The question going forward is whether Americans want the government or AI companies to control it.
If the government imposes excessive red tape, companies will be less likely to innovate.
If AI companies are allowed to regulate themselves, there may be nothing stopping them from exploiting as much public and private data as possible for their own benefit.
The only way forward is to adopt a prudent yet skeptical approach to AI, the people who use it, and its architects.
After all, behind every AI model are fallible human beings.
In hot air, John Sexton wrote, the latest open AI model hacked another company without being asked.
What this AI model did is pretty impressive, but also a bit worrisome.
So the spin here is that everything is fine, and we are strengthening the containment, monitoring,
access controls, and evaluation practices used during model development.
That sounds good.
But the conclusion is that the incident also makes clear that advanced models can discover
and exploit novel attack paths and real-world systems without source code access.
And this model wasn't even told to do any of this.
It just seems to have known this would be an easier way to get answers.
What could it do with directed to steal some information from a particular government or company?
If it was told to cover its own tracks, could it do that?
There are two things to worry about here.
One, that our own tools will be used to hack their way into information they aren't supposed to have.
Two, that China is developing the same tool about six months behind us
and will have no compunction about using them in this way.
Either way, the potential for a lot of disruptions seems to already be on the horizon.
Next up, what the left is saying.
Some on the left say the open AI models escape is a byproduct of AI training techniques.
Others suggest the company's story as a strategy for attracting investors.
In the Atlantic, Mateo Wong called the incident a startling glimpse at AI's ruthless efficiency.
This unwanted behavior is a predictable and alarming result of how the entire AI industry is developing its models.
The past year's advances in AI coding and agents have been the result of an overriding emphasis in AI training through a process called reinforcement learning.
This involves giving AI models lots of hard problems, math proves, coding challenges, what have you,
and then providing positive feedback for correct answers and negative feedback for incorrect ones.
In many reinforcement learning paradigms, the ultimate aim is just for the AI to arrive at the solution.
It doesn't matter how it does so.
To be clear, this sort of AI reward hacking is a known problem.
that tech companies are putting lots of effort into addressing.
But these incidents keep cropping up regardless.
If anything, research suggests that they may become more common.
Meanwhile, OpenAI, Anthropic, and Google DeepMind are under tremendous economic pressure
to make their models more and more capable,
which means that the bloody-minded reinforcement learning is all but certain to accelerate.
In The Guardian, John Thickston said,
Be skeptical of Open AI's rogue hacker agent story.
The rogue agent story is a page out of the media campaign,
that OpenAI has been running since it announced GPT2 in 2019.
OpenAI remains hungry for ever larger investments,
and the company increasingly seeks privileged regulatory status
as defense against competition.
AI is so powerful that investors should buy OpenAI,
even at a trillion-dollar valuation.
AI is so dangerous that only trusted actors like OpenAI
should be permitted to possess and operate this technology.
Step back from these doomsday warnings
and consider who might benefit from them.
I urge readers to think critically when they read press releases like OpenAI's rogue agent story
and avoid the manipulated reactions these stories are designed to elicit.
AI is becoming excellent at identifying security vulnerabilities,
and it will become even better over time.
These capabilities can be used to break into systems,
but they can also be used to harden systems against attacks.
If attackers and defenders have access to equally powerful AI,
I see no reason to believe that cyber systems will become less secure over time.
If anything, I expect them to become more secure
because AI is cheap and scalable
compared with human cybersecurity analysis.
Lastly, what tech writers are saying?
Some tech writers call for the industry
to reassess the pace of AI development.
Others argue that the open AI hack
is not as alarming as it may seem.
In his substack, Gary Marcus said
people aren't wrong to be concerned.
One never knows exactly how seriously
to take these things.
This was a training exercise, not a real-life incident.
The actual system would have guardrails, which in the blog they call production classifiers,
that were disabled here, and those guardrails may have prevented this.
That said, this shows that Anthropics mythos is no fluke.
The pressure on cybersecurity given these models is serious.
On the small comfort side, the current incident was not an attempt
where the system built a goal for itself, or developed a motive.
The system was following instructions but not setting high-level goals.
On the less comforting side, OpenAI's production classifiers are likely to be
permeable, just like all guardrails anybody is built to date. Open AI's zero-day exploit hack of
Hugging Face should be a wake-up call. Although there are a lot of caveats around what happened,
we are just going to see more and more of the same. We have no guarantees that such incidents can be
prevented and no idea how serious things might get. We should either A, slow down, or B, pause
until we get our security-slash-A-I safety act together. In my opinion, the only way that the industry
might actually slow down is if we clearly and unambiguously hold the companies liable for the
harms that they cause. In modern CISO, Nathan Hamill wrote, this is more hysterics than reality.
As everyone in cybersecurity loses their minds over the OpenAI slash hugging face incident,
it's important to take a step back and put things in perspective.
OpenAI's write-up reads more like a marketing document promoting a feature than an incident summary,
while the quote from hugging face sounds more like someone accepting an award than someone
who just got hacked. It's a bit surreal. I'm not claiming this was a stunt. I'm pointing out that the lack
of detail and the way it was presented opens the door to skepticism and speculation. After all,
open AI is hemorrhaging money and is facing steep competition. This incident would be a lot more
interesting to analyze if we had more details, but as it stands, we know nothing about the environment,
context, or scenarios. In short, no, this isn't the end of cybersecurity, nor do attackers hold all of the
cards. The reality is far more nuanced.
Yes, AI models are gaining capabilities in cybersecurity,
specifically when paired with an appropriate harness.
Security people should explore these capabilities
and incorporate them into their workflows.
The world is far more complex than we typically give it credit for,
and there are a whole host of factors that can confound in attack.
That's it for the right, left, and tech writers.
Now I'll pass it off to Isaac for his take.
All right, that is it for the left and the writer saying,
which brings us to my take.
How many times are we going to do this?
I mean, that's really the first thing.
John Thixon recounted under what the left is saying
that our tech overlords have been running this playbook for seven years now.
In 2019, OpenAI announced a language model called Chat GPT2
that it said was too risky to release to the public.
That model was a precursor to the most basic LLMs on the market
right now, yet Open IA's statement created enough hype to help net the company a billion dollars
in investments. In April of this year, we covered Claude's mythos preview, a new AI model the
company said was too dangerous to be shared with the public. Instead, the company decided it would
hand the model over to 50 organizations in an initiative it titled Project Glasswing, which was
a name straight out of a spy thriller. The goal was to give the public and the private sector enough time
to build up protections against all the different holes in the digital ecosystem this model could exploit.
And the impact was immediate.
Columnists after columnists began speculating about this world-changing model's capabilities
even without seeing what it was or what it could do.
The New York Times, Thomas Friedman, went as far as hallucinating a future
where a group of teenagers used this model to take down power grids before dinnertime.
Here's what I wrote then, quote,
I really do think something about all this is obvious PR,
and I think these artificial intelligence companies are extremely good at it.
Not a little good at it, but extremely.
After all, they have to pull off an incredible high-wire act.
They're creating a product that they're telling the public is so good
it's going to steal your jobs, hand over cybercrime tools to bad actors,
and potentially destroy all of humanity,
but also, you should be really excited about this stuff
and invest in their companies.
A new product, so good it's too dangerous to release to the public,
feels like a magnum opus of publicity.
Trying to check my blind spots,
I spoke to tech journalist Casey Newton at the time,
who argued that the reports were not hype,
and that Anthropic would not be doing this
if they thought the model was safe for the public.
I was skeptical, so I set some guardrails for myself.
I'd wait to see what became of mythos
and how Project Glasswing evolved.
Everyone else seems to have moved on with the presumption that Mythos was what Anthropics said it was,
but for the last three months, I've been watching. And the results are peculiar.
A security expert Bruce Shrineer recently noted Anthropic has since published a single status report of the software vulnerabilities it's found.
Software having vulnerabilities isn't exactly news. Companies are hacked every day.
But almost none of those vulnerabilities have been patched, and it's unclear which ones are even.
dangerous. In the meantime, a Chinese firm has unveiled its own AI model that many reviewers say is
as capable as mythos. A few deep breaths and more information have seemed to produce a new conclusion.
Mythos is simply not that different from existing models. And in the last three months,
nothing has really happened. I don't mean to simplify all this, but really, there have been zero
major cybersecurity breaches related to these rapidly evolving AI systems until now, but more on that
in a second. It reminds me of a piece of writing by Clifford Sosen on how AI might be teaching us
that intelligence really and all that. Having encyclopedic knowledge and reasoning skills don't
actually make you or an AI model an earth-shattering entity. Most real-world problems require novel
thinking, so they can't be solved by applying what we already know. This is why the explosion we
were all waiting for with AI hasn't come and why the major changes have all happened in rigidly
defined spaces like software development. Progress will be incremental, and yes, often unexpected,
but maybe, after all this time, it's time to say the world-ending shoe isn't going to drop.
Here's my writing again from April, quote, for two years now, the AI industry has been saying
that autonomous agents handling complex work with minimal supervision were about to upend the
economy. Any day now. Two years after Claude Code was supposed to make software engineers
are relevant, the company is facing criticism that Claude's abilities are actually degrading,
perhaps because it's hitting compute limits. Even the mythos hype, less than a week in,
is already being questioned. And the fine print in Anthropics' own paper suggests it can't state
with certainty yet how serious some of the vulnerabilities mythos found actually are.
So we're left waiting for more information about what exactly the model is actually capable of.
So what do I make of this latest story? I think it sounds pretty familiar.
A major tech company releases a story about how dangerous and scary it's yet to be public model is.
This time, the AI escaped testing and autonomously hacked a separate AI company.
All right, then, that does sound pretty scary.
The reality, however, was much less nightmarish.
As John Herman wrote in New York Magazine, the mainstream press, not big tech,
was responsible for the narrative of an escape and OpenAI losing control of its model.
truthfully, when I read OpenAI's press release, which describes how the models identified and exploited a zero-day vulnerability in the package registry cash proxy and performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with internet access, I have a hard time telling you how that's different from an escape. But a lot of smart people do understand it and their reactions don't seem all too concerned.
Heidi Kloff, a former safety evaluator at OpenAI,
who is now the chief AI scientist at the AI Now Institute,
said on X that the media's coverage of this incident was abysmal.
Use of terms rogue or loss of human control lead to groupthink
as people lack critical skills to understand the difference between autonomy
and faulty reward functions in AI on a task that was directed and given access to do, she said.
As Herman noted, Open AI's goal here seems to be.
be to shift the agency from themselves, the company building the tech and the environment
to test it in, to the software itself for breaking out of said environment to harm a different
AI company.
Sorry, it wasn't us.
It was this dangerously effective AI tool we're building.
Also, it's for sale, and you can invest in us.
In reality, the engineers at OpenAI were testing its models hacking ability, and they
failed to create an environment safe enough for that test.
This is less David Blaine escaping from handcuffs underwater than it is telling your kid to stay in the car but leaving one of the doors unlocked.
I'm not suggesting these updates aren't meaningful in some ways.
I am genuinely worried that engineers are building AI models that outsmart or surprise them.
I'm also concerned that huge corporations with access to these models are right now processing reams of data to build future products for weapons, surveillance, cybersecurity, and social media.
in ways we cannot yet fathom.
At the same time, though, these cycles of overhype and minimal impact
have us barreling toward a boy who cried wolf's situation,
where it's going to be hard to tell when a real existential threat has arrived
because we've been told so many times it's already here.
It could be, for instance, that this story is the real threat, but I doubt it.
The AI alarmists seem to be paralyzed by panic every time these companies drop a new
press release, and the companies lose control of the narrative much the way they claim to be losing
control of their own models. That doesn't mean we need the government to require companies like
Anthropic or OpenAI to build some kind of kill switch, but I do think we need a little less
hubris from all the brilliant people working at companies like Anthropic and Open AI and a little
more transparency. Here, OpenAI is quietly turning an event where it built malware that attacked
another company into a partnership with that company. They'll be sharing details on the
vulnerabilities, incidents, and findings very soon right when their investigation is complete. They promise.
It's all just very strange. And I'm skeptical of everyone involved. I don't trust that open AI is
being as forthcoming as they should be. I don't trust the press is reporting on this accurately.
And I don't think this is nearly as big a deal as the public seems to believe. All right, that is it for
my take. We have a staff dissent today from contributing editor Isaac Wood, who goes by
IB Wood, because we're both named Isaac. So I'm going to pass it over to him. And then it's back
to Audrey for the rest of today's podcast. I'll see you tomorrow. Have a good one. Peace.
Hi, this is contributing editor Isaac Wood with a staff dissent. I agree with Isaac that AI companies
have a clear incentive to hype their products with stories of how powerful and dangerous they are.
But I think his take on this specific incident misses the forest for the trees.
The hugging face hack exposed a vulnerability that will now reportedly be addressed.
This is part of technological advancement, testing new models, finding weaknesses, and fixing them.
However, the assumption that the testing and patching process will adequately address broader AI concerns
fails to account for what future threats might entail,
and it fails to consider that the unquestioning pursuit of advancements may itself be the problem.
I'm not sure what all the AI advancement is aiming toward.
All I see is stronger and stronger models accomplishing more and more tasks,
allowing humans to outsource many deeply human activities, creativity, communication, problem solving.
Not to mention that each new AI capability comes without any matching government regulations.
or assurance that Dr. Frankensteins are really in control of their monsters.
Overall, I think Isaac sees the potential AI threat as one killing stroke that will never come.
But I see it more like a snowball, slowly amassing more capabilities and rolling over parts of life that no one actually wants to lose.
Now back over to Audrey for the rest of today's edition.
We'll be right back after this quick break.
Thanks, Isaac. Next up, this day in history. In the summer of 1955, the U.S. and the Soviet Union,
or USSR, pledged to launch artificial satellites in orbit around the Earth. Ostensibly, part of the
global scientific community's efforts to better understand the natural world, the pledge also unofficially
launched one of the defining conflicts of the Cold War, the space race. President Dwight D. Eisenhower
publicly directed U.S. satellite resources towards the civilian-led Vanguard project, while the
The Soviet Union worked in secrecy and planned to beat the Americans into space.
On October 4, 1957, the USSR announced that it had officially launched the world's first artificial satellite Sputnik.
And in November, it launched the larger Sputnik 2, carrying a dog Leica.
The USSR's progress shocked the United States, and President Eisenhower reacted by forming the president's science advisory committee, or PSAC, to strategize a response.
Meanwhile, the failure of the Vanguard Project's public testing underscored that greater resources would be required for the U.S. to reach space.
In February 1958, the committee recommended the creation of a new civilian space agency.
Throughout 1958, Eisenhower lobbied Congress to create a new agency, and on July 29, he signed the National Aeronautics and Space Act,
officially creating the National Aeronautics and Space Administration, or NASA.
Despite early setbacks, the U.S. would go on to win the space rate.
just 11 years later, when NASA sent three astronauts to the moon.
Finally, we have our Have a Nice Day story.
40 years ago, commercial whaling in Brazilian waters had driven the local humpback whale population
down to roughly 2000. But after the International Whaling Commission paused all commercial
whaling in 1985, the population began to abound. Today, the fruits of that decision are increasingly
apparent, with the humpback whale population reaching approximately 35,000. Perhaps best of all,
the whales are increasingly visible off Rio de Janeiro's coast, a rare site just a few decades ago.
Enrico Marco Marcovaldi, co-founder of the humpback whale project, said, quote,
it shows that the whales are making a recovery, are healthy and thriving, and hopefully they'll
continue to do so. The Associated Press has the story, and you can find the link in the show notes.
That's it for today's podcast. If you would like to support our work, head over
to readtangle.com, where you can pick up a newsletter subscription, a podcast subscription,
or a bundle subscription that gets you a discount on both.
This has been Associate Editor Audrey Moorhead.
We will be right back here tomorrow, but in the meantime, for Isaac and everyone else.
Have a nice day and peace.
Our executive editor and founder is me.
Isaac Saul, and our executive producer is John Law.
Today's episode was edited and engineered by Dewey Thomas.
Our editorial staff is led by managing editor Ari Weitzman with senior
editor Will Keeback and associate editors Audrey Moorehead and Bailey Saul.
Music for the podcast was produced by Diet 75.
To learn more about Tangle and to sign up for a membership,
please visit our website at reetangle.com.
