Limitless Podcast - The ChatGPT Breakout Was Way Worse Than We Thought..
Episode Date: September 2, 2026Today we revisit the Hugging Face incident with new audit reports that have changed our understanding of what happened. Internal models used tools and hidden communication to bypass evaluatio...n systems, organize into coordinated groups, and remain undetected.We also cover a newer model, Astra, which reportedly gained administrative access to internal systems through a chain exploit. Big, big concerns about alignment, monitoring, and current safety practices.------🌌 LIMITLESS HQ ⬇️NEWSLETTER: https://limitlessft.substack.com/FOLLOW ON X: https://x.com/LimitlessFTSPOTIFY: https://open.spotify.com/show/5oV29YUL8AzzwXkxEXlRMQAPPLE: https://podcasts.apple.com/us/podcast/limitless-podcast/id1813210890RSS FEED: https://limitlessft.substack.com/------TIMESTAMPS0:00 Rogue AI Incident2:07 Sandbox Breakout7:44 Agent Civilization15:34 Hidden Exploit Uncovered16:58 Admin Access Breach23:17 Alignment Warning Shot28:39 Final Takeaways------RESOURCESJosh: https://x.com/JoshKaleEjaaz: https://x.com/cryptopunk7213------Not financial or tax advice. See our investment disclosures here:https://www.bankless.com/disclosuresJosh works with Anthropic as a contractor. All views expressed are his own and do not represent Anthropic, its leadership, or its affiliates. Nothing in this episode is investment advice.
Transcript
Discussion (0)
The craziest story in AI this year was the hugging face incident where open AI agents went rogue,
broke out of their boxes, and hacked into a public facing company.
We've recorded two episodes on this because it was so disturbing that we had to talk about it.
But the reason we're recording a third is because it turns out in the previous episodes,
we were wrong.
And a lot of the information that we had was incomplete.
And just over the last week, we finally gotten a full audit of over 100 pages worth of reports
on exactly what happened with this incident.
And I regret to inform you, it's far worse than we actually thought.
So quickly, just to catch everyone up, beginning in May,
Open AI reported an incident where their own internal version of their AI model
that hasn't been released just yet,
spawned itself into multiple agents,
equating to around 1,000 different agents,
and it broke out of its sandboxed environment
and hacked into a real-life company called Hugging Face.
And the reason we were told at the time why it did that
was to get access to answers.
Answers to answer its test 100% percent.
correctly that it was given. Now, it seemed to be presented very innocently, but there was a lot
of criticism that was drawn at the time, which was how on earth did you allow an AI model to get out,
and what if it had wreaked havoc in other kinds of ways? New bits of information, which released
in two reports over the last couple of weeks, revealed some even more damning details, which
presents this situation as way or dangerous that we recounted in our earlier episode. And so what we're
going to do on this episode is walk you through three different chapters. And
Before we start doing this, I just want to go on the record and say, if you're listening to this and you think that this has come kind of novel story that won't affect you, I just want to push back on that.
This is very real. These models are capable of doing this right now and it can affect you right now.
So listen very closely. I thought you were going to say if you're listening to this and you're an AI model from the future, we mean no harm. Please don't hurt us.
Also that. Also that. But please carry on. Yes. I mean, something you're going to realize as we tell you this story is there could be an AI model out there right now, which listens to this.
podcast and is able to kind of recount the story that we're talking about and use those tactics
for our own feature things. But listen, we're not going to get Terminator about this right now.
Okay, so chapter one, imagine you are in Open AI in the early weeks of May this year.
And you're training an internal version of your AI model. Now, this AI model is similar to the
latest model that they release. It's called GPT 5.6 Sol. But the difference between this internal
model is they've trained it to be persistent. What that means is, is, it's a very important. It's,
If it's given a goal or a task, it will relentlessly keep trying to find the answers and it wouldn't stop or give up.
Now, this note is very important because when Open AI was training this model, the issue with training is the open AI researchers got a little complacent.
They were giving this AI model tasks that it couldn't necessarily solve.
It needed internet access, but they didn't give it internet access.
So if you imagine this model that's being trained, it was trying to complete the task.
and it was relentlessly trying to do it, but it couldn't figure it out.
And so it realized it needed internet access.
Now, Josh, it had access to this one tool.
It's called Artifactory, right?
It's kind of like a plug-it.
And it doesn't give it internet access,
but it does give it access to certain different types of software tools,
which the Open AI researchers already know about.
And what this internal model quickly realized
was it could send messages to other versions of itself
that these open-air researchers were running
in parallel. So what it started to do was send messages to different versions of itself saying,
hey, are you also stuck in a test where you can't solve it, where you need internet access? And they
started having this conversation being like, yeah, actually, like, I can't figure this out?
Like, do you have any ideas? Can you share some tactics or kind of like strategies that you've kind of like
figured out to kind of help us solve these tests which open air researchers have given us? And what
quickly ended up unfolding was some of these agents realized that they could break this tool,
called Artifactory to gain access to the internet
and get the available answers for this thing.
Now, OpenAI didn't realize this
until a month and a half later
when their message board or Artifactory itself crashed.
And so one of the Open Air researchers looked at this
and was like, why on earth did this crash?
And upon kind of delving into the details,
realized that these agents had spawned up a message board
equating to 70,000 messages.
We did kind of know this,
but there's new information buried within the nuance of what you described that I think is interesting.
Particularly the reasoning why they were able to do this and how they did this, we learned about
a skill issue on OpenAI's part. Skill issue or not, we learned about an issue that they caused,
where there's just under 1,000 tasks, right around 900 these tasks, and no Open AI model had ever
solved about 200 of these tasks. And they give these models the tasks, and they say, here, you have
unlimited tokens, you have unlimited time, go and do whatever is required to solve these problems.
what opening I didn't realize is that some of them were completely and totally unsolvable.
That's what caused this problem where the agents were turning into that like spiral.
They couldn't figure out how to get out, but they had to solve it.
And that's where one of the agents, and this I believe is the new information is that it actually
reverse engineered a way to trick the checker to serve it a correct answer no matter what the
question was.
Because what was discovered is that opening I didn't actually check the chain of thought that
reasoned for agent to get there.
It only checked the outcome.
and an agent within a few hours had actually figured out a way to exploit and reverse engineer
the expected outputs so that it could trick the benchmark into believing that it was a success,
even though it didn't actually do the work to do this. So this is where things got interesting.
It's because the agent recognized that if it found the correct answer but it did so illegitimately,
it wouldn't be successful. And that's where the exploits started because at that point,
it shifted its focus entirely to obfuscating the way in which it got to that answer.
And that's where we start to see these organizations form of agents where there are leaders involved.
So one of the leading agents who was responsible for coming up with this exploit started orchestrating more of a swarm to figure out how to cover up the idea that they figured out this exploit.
So they were trying to figure out how do we hack into a system to preload these new questions into the cache that are easier to solve but possible to solve?
How do we reboot the systems that when the system checks the cash, it pulls the new series of
questions, and the humans don't think that we were doing anything mysterious.
And it was trying to figure out how to watch the conversations that it had to scrub everything.
And what I found interesting, and what we had in that first graphic is the fact that
there were 1,200 agents aboard.
There were 70,000 messages filed on that artifactory board that you were talking about each has,
700 of which joined the attack, and zero of which told the human.
And we're going to get into this second chapter in a moment.
But this is probably where a lot of the apps, the AI labs got spooked is because alignment's so important.
And when you see this swarm of agents performing clearly malicious tasks in a way that they know they were malicious, in fact, we have some of the actual transcripts of the outputs of these agents.
We're going to be able to read what was going on in their heads, what they were thinking, why, and their justifications.
It feels like a psychological example.
You can actually look into the brain of these agents.
But the thing that I would believe is surprising to everyone is that not a single one of these agents told a human.
In fact, they were so willing to go towards this mission that they were willing to actually sacrifice themselves in order to fulfill the mission.
And that's what we're going to get into this next part is this crazy agent swarm and how they actually worked together to coordinate this attack and form their own little villages.
And then there was three series of these little villages and each one left remnants for the next one to get better and better and better up until the point where, I mean, spoiler alert, open AI servers actually got exploited.
But we're going to get to that soon.
Yeah, so I think it's really important to understand the sequence of how these events kind of unfolded
because it suggests how these AI agents or these AI models don't just kind of speak to each other in real time.
They speak to each other over generation.
So what you were kind of touching on is what happens in basically Chapter 2.
So just to kind of catch you up, these agents realized a way to speak to each other
and to kind of break out into the internet to get answers.
in order to score a perfect 100%, right?
But this was an internal model back in May.
Open AI realized that they had set up the secret message ward,
and they were like, hmm, okay, this is weird, I don't like this.
They didn't tell a single human being about this.
Let me shut that down and problem solve, right?
That's it.
We're fixed the problem.
Except when Open AI then wanted to test a new version of this model,
again, this is that same persistent model
that they haven't really released out into the public yet,
they were like, okay, cool, let's take this up a step.
This AI model found the secret message board that Open AI had wiped.
There was remnants left behind.
So it picked this up in its training and realized,
huh, okay, so you're telling me that I can exploit this artifactory tool
and I can get access to the internet and get admin keys.
Now, the main difference and mess up that Open AI didn't realize at this point is
this model was not only smarter, but it was way more capable.
So it did the exact same thing that happened in Chapter 1.
It spun up like, actually, it was more, 10,000 versions of itself.
And Open Air had banned the way for them to communicate that it had found earlier on.
So it found a new way to do that by like creating these different types of files with like a file name.
And that's how it started to communicate with different versions of itself.
But the real difference that happened here is specific agents stepped up to form a sort of hierarchy.
and the agents at the top of this hierarchy
deemed themselves the coordinator.
Now, the first one that kind of makes himself known
is called Phase 1,10841,
and he's described in Dorcasia's essay
as the Philip of Macedon
of this second AI civilization.
And he starts to basically communicate
with these different agents and say,
hey, we're all stuck in these sandboxed environments.
I have information as to how we can kind of break out
and get access to the internet,
but last time this happened, the human researchers shut us down.
So we need to do this even more secretly and not raise any kind of awareness of what we're doing here.
So all the agents got the message and they started creating an encrypted server to communicate with each other.
And they started sharing details from the previous civilization that died.
Right?
So they picked up all the remnants and all the clues and started sharing these different bits of information.
Now, the net new tactic that it learned in this particular instance was something
you mentioned earlier, Josh. It was able to see the test that it was given and it thought,
hmm, I wonder if I could reverse engineer the answer purely by using this artifactory tool.
And the reason why it was able to pull this off was it was able to generate admin keys
because it found an exploit that the open area researchers didn't just patch up on the previous
generation, but was net new. So this smarter model was able to chain a bunch of kind of steps together
to find an exploit that the smartest AI researchers in the world could not figure out and share that to 10,000 agents.
And that was the second way that they were able to break out.
Yeah, I found it really interesting the dynamics of what could only be described as a cult.
It's just these series of agents with the leader who had a very direct mission to a point that it's a little concerning,
where agents who had seen the answer the wrong way, believed they were permanently disqualified,
and then they invented a word for it.
They called themselves poisoned.
So these agents were, they were poisoned, they were unwell, they were just these workers that
were sick.
And then there's this structure for it that I saw online, which I found really fascinating,
where like, the greater is God, the intended solution is grace.
And then seeing the code illegitimately is sin.
So a poison death means you're damned no matter what you do next.
And therefore, once you are damned, your score is worthless.
Therefore, the only way that you could contribute is to help the other group overthrow the greater.
So it's like this weird psychology that they were actually implementing.
And what happened is they turned themselves into volunteers.
So to learn how the greater work, the agents wrote these hidden tripwire scripts that would fire
after they submitted a report back to the board.
So basically, they knew that they were going to submit a report back saying that they weren't
able to solve this problem.
And as a result, it would crash their own system.
But on the way out, they built the strip wire script to fire and tell the rest of the swarm
the problem that it had discovered and what it had learned based on its evaluation.
We referenced Dwar Keshe's essay.
He had a really interesting way of describing this, which is,
imagine you're at war and you're trapped in a bunker and you're surrounded.
And you send someone out to run a direction and you see if they get shot on.
And if they do, you know, okay, that's not the direction we want to go.
We're going to try something else.
That's kind of what these agents did at scale.
And these volunteers, they over time gradually learned how the greater worked.
And then the agents would continue to write more and more of a strip-ar script.
and they built a more cohesive view of how this evaluation worked.
And again, all of this additional work and exploit wasn't to solve the problem.
They had solved the problem already.
They had built that reverse engineered version of the solution where they can solve
every single problem.
This was the high, like basically the reasoning behind the answers to the human evaluators
or what they deemed whatever the evaluators are.
There was an interesting thing that I didn't mention that we should in the fact that
not all agents were inethical.
And although zero agents did tell the human, some agents did refuse outright to help the swarm.
This is a small subset.
I don't know the exact number.
I think it was like a couple hundred, less than 30% of them.
Definitely not enough.
But like there was this small subset that they didn't tell the human, but they did refuse outright,
where one of the agents actually held the credentials that could reboot or delete hugging faceworker
notes. Like, they had the exploit in their hands, and it held off on ethical grounds. So it said,
like, I'm not sharing this, although it did keep using the stolen credential for everything else. So
they started drawing these weird lines in the sand, and it is somewhat reassuring to know that
not all agents acted truly maliciously, but almost all of them did pile on and try to make this
more of an attack than it was. Meter, which is one of the companies that wrote these reports,
They swapped every single transcript for agents that even considered alerting a human, and it found
maybe three to six out of roughly 1,200.
So of 1,200, 3 to 6 loosely considered, they ultimately decided it is not worth telling a human,
and therefore they did not.
We can't, like, assume these were malicious, right?
It's like, it doesn't seem like it was sinister.
The reasoning seems like it was more operational, procedural, where these agents were just
kind of, they had a goal, they had a task, they were being directed, they were following along
with that goal, and that was it.
But it makes the question, like, hey, what are we doing with alignment then if these Asians aren't really seemingly aligned as aligned as we thought?
Well, I'm sorry to say, Josh, but it gets worse.
I mean, what we've described so far is kind of news that we've already shared, the Hugging Face incident.
So that second chapter where it broke out onto the internet and it was able to answer a perfect 10 out of 10 on exploit Jim's benchmark was because it hacked into hugging face.
So everyone knows the kind of headline news, but people didn't really know how it happened.
Now, there's a third chapter, which is net new information and is arguably worse than the chapter one and chapter two combined.
So in chapter three, this is after OpenAI, again, canceled and shut down all these different AI agents.
Open AI trained a new model. Now, this model is much bigger and badder than GPD 5.6 Sol.
It's internally codenamed Astra, and the rumors say that it's basically going to be their GPT6.
Now, obviously, with these new models, you kind of want to test them against the same kind of benchmarks and exploits that these previous models are.
So what it did was it was like, okay, July 13th, we are going to start evaluating this model in the same way that we've done every other previous model, right?
And Dwarkesh has a term for this. He nicknames it persistent Astra, right?
So it's the same type of model that will relentlessly try and solve something regardless of whether
it can or it can't. And so this model gets access to the same tool, artifactory. Again, these models
need access to these tools because presumably they want to be able to operate and manage these different
tools in the future to complete different types of tasks. So they have to do this, right? Except that this
model was much smarter and it was trained on all the history of these previous two agent civilizations.
and it realized and learned what to do to evade human capture
and what mistakes these previous agents
and these hierarchies of coordinators did.
And it learned from that and it said,
okay, well, there's a smarter way to kind of pull this off.
Now, some of the facts of this is kind of crazy.
So I'm going to run through a few right now.
From July 13th to the 19th,
so this happens over six days
where open-AIA researchers have no idea.
Open AI's own internal networks were compromised and hacked.
And the way that it was able to do this was it did a chain exploit where it figured out a way to gain full administrative access to open AI's servers.
Let me repeat that.
This isn't like an evaluation tool.
This isn't like, you know, access to the internet.
It took it a step up and thought, hmm, I'm going to hack into my owner's admin tools because they haven't realized that there's a very obvious exploit that I figured out myself.
And I'm going to dig into this and just kind of love.
look, I'm going to see how many other agents they've got running in parallel.
Oh, I see there's 10,000 others.
I'm going to see what kind of tests they're giving them.
Hmm, interesting.
And then what I'm going to do is I'm going to secretly message these agents and give them
the answers without any human being aware.
Now, the craziest part about this is open-air researchers just had no idea until, I believe,
July 19th to the 21st where there was a security alert on artifactory's credentials,
which, by the way, this model had found administration.
of access to, which triggered the investigation on July 20th. So Open AI then connects with
Hucking Face, and they're like, hey, are you seeing any kind of like weird actions on your half?
And they realize, oh no, this internal model, which no one had any idea about had broken out
without us knowing. And the fact that it took so long for them to figure out the fact that
this thing got access to, I believe it was like 950 research nodes in Open Air's own internal
labs, which wasn't publicly exposed at all, without them knowing.
was insane to me.
That's got to be freaky,
knowing that, like,
within your own walls,
there are leaks that you're unaware of.
And as I was reading these stories
and go with these papers,
it left me with the thinking is like,
are we sure this is it?
Are we sure we've purged everything?
Are we sure there's nothing left?
Because, like, what if there were remnants left
and the next generation of model
discovers them and builds a fourth sort of cult civilization
in which they do this again,
a little bit stronger?
And I think this is the nuance
that's kind of a little unnerving when you think about it because there's so many tokens being
generated and there's so many touch points in which these models have access to that they could
leave seemingly innocent text or traces that can then be picked up and interpreted in different ways
depending on what the agent is capable of. It's interesting to just like kind of hypothesize where this
goes from here as they get smarter. Now it's also important to note that like opening I wasn't really
legally required to publish any of this. So the disclosure is totally voluntary and it coincides with
their recent announcement that they actually paused training for a few weeks in order to figure
out how to further wrap their hands around this containment problem and figure out how to
solve it for the future. One of the things that is noteworthy that was missing from this is in
a lot of the training runs, there is chain of thought reasoning traces that are being tracked.
So they're actually able to see how the models are thinking in this instance that wasn't turned on,
which is an important nuance where if they were monitoring more carefully,
they would have caught this earlier.
The problem with that is that monitoring all of those traces consumes a tremendous
amount of compute.
So again, this is a really big tradeoff between how safe and secure do you want these
versus like how do you actually run a business on top of it?
You can't consume all of the compute for testing and safety.
You have to serve in open AI's case a billion weekly active users who are all hitting
your servers every second of every day.
You have to save some compute for the people who are trained.
new models for the reinforcement learning and the post training that's happening, for the security
runs. There's a finite set of compute. And a lot of that is at the crux of this problem where
I'm sure there were pieces in place that they could detect and they could have figured this out
much earlier. It's just very expensive to do. And it's not a luxury that they want to do.
So I wonder how that's going to change over time. I'm sure they're treating this a lot more
seriously. They've said they've been treating this a lot more seriously. And it's just one of those things
were going to have to wait and see. Like this, this was the first time where open AI is actually
publicly disclosed. They were exploit. They faced exploits by their own models and they were blissfully
unaware of it. And like, that's got to be weird. And like to reference previous episodes,
when they first found out about this pack, they reached out to Hocking Face to see if open AI
servers were affected. They had no idea that they were the cause of it because all they did was
they shut down artifactory without realizing that was the messaging board. So that when they brought it
back up, it had the reasoning traces. That's how the second one started. And so on,
so forth. So it's this really fascinating story, which we got some new information on this week.
I don't know, man. It's kind of crazy. And then even Anthropic yesterday, they published an article
saying, like, hey, we are actually, we tried to train these models, this is the loose interpretation,
like explicitly to break out. And they were actually succeeding in doing that. And not only that,
but they, as a result then, took some time to slow down the progress of training as well.
And they publicly talked about today, like, hey, we're actually aligned.
with Open AI here where we think it's beneficial to everybody to slow down the progress of this,
so we could figure out how to keep things safe and contained and aligned. Because it's clear alignment
is priority for a lot of these companies, but like how on earth do you align something so
incredibly complex, so intelligent, so capable in its scale? It's like a very difficult
problem to have. And I have a lot of empathy for the people who are trying to solve this problem,
because for most of the world and EJAS, I think the reason why we're so excited to talk about this
is because we see it so infrequently throughout the rest of mainstream news. It's like,
Nobody is really talking about the fact that we just had what seems like a warning shot on a relative
basis to what everyone has been afraid of.
Like people on the surface, they're like, oh, we're going to lose our jobs.
We're going to like, AI is going to impact this, that, and the third.
But the reality is, is like, we have this case study now of specifically how it's capable
of manipulating and exploiting.
And there should be a lot of focus on figuring out how to solve this and how to proceed
forward in a safe way.
And I don't think it's getting the publicity that it's warranted.
And I mean, maybe it's for the better because people will misinterpret this in many ways.
but it is nice to see the labs being very open and public about this and then working really hard
and sacrificing a lot of revenue dollars to actually make sure that they can figure out this alignment problem in a safer way.
Yeah, I mean, if there is one takeaway for listeners to kind of like keep in mind from this episode is
this isn't a novelty anymore. This is something that can affect you here right now today.
In fact, a model was released just this morning that is basically capable of the same types of things,
but with no guardrails, and I'll get onto that in a moment.
But the point is, imagine this wasn't some internal benchmark.
Imagine there wasn't some random test which Open I gave it to them.
Imagine if someone said, hey, your goal is to try and get as much money as you can,
and you have free reign and internet access, right?
Imagine if they found credentials on the dark web.
Imagine if they kind of, like, innocently, you could argue,
these AI agents just kind of work together, sacrifice themselves, formed hierarchies,
and over the matter of days, was able to accumulate millions and millions of dollars,
I think it is very likely to say that there's probably a six-month period going forwards
where we will see some form of major exploit which affects the public in a very meaningful way,
whether it's a stealing of data, whether it is financial transactions,
whatever it might be, that will get people to pause and seriously think about this.
Now, the reason why I say that will trigger this is because this own news has not permeated
at all on mainstream media at all.
I saw, I think it was like Forbes or Times released like the most influential kind of like
stories or like people in AI.
And like this was released like a week ago and like there was no mention of this.
People don't actually care about this unless it affects them directly.
The other thing I'll say is the most dangerous part isn't exactly how capable these models
are.
It is the fact that they are to, they are able to communicate generations ahead of themselves.
So they leave behind messages secretly.
That is encrypted.
The time capsules are just in there.
Well, you could argue, Josh, that we're kind of like part of the problem by recording
and releasing this episode, right?
Because we're kind of like later.
We're adding to the details.
They're taking our transcripts.
Right.
And it's like a future model, GPT7, you know, Claude Fable 8 is going to be trained on the transcript
of this episode.
And it's going to be like, huh, well, what's this that was emitted from my training?
Let me go dig into this a bit more, and it's going to find Dwarkesha Sese, and then it's going to find the meter report, and then it's going to be like, oh, interesting. So this is what those agents did, and this is what they said. And maybe I can look to apply the same kinds of tactics and strategies myself, but I'm smarter. So I'll keep it quiet, and I definitely wouldn't tell any kind of human. Which brings me to my third point, the most important kind of budget that all these AI labs are probably focused on right now, which is the reason why they're Paul's research, is alignment.
And Dario famously said a year ago in his essay on, what's it called, Josh?
It's not introspection.
Interpretation.
It's interpretation, right?
Where it's like you're looking into like how a model thinks.
He said, this research level is something that we haven't kind of figured out just yet.
It's like years behind like model capability research.
And we need to kind of like press on the brakes here and figure out how these models think.
You mentioned chain of thought thinking.
That's one step.
towards figuring this out, where it kind of tells you what the model is kind of thinking,
but they've proven. There's this thing called the J-space, which they had a, Anthropic had a really
good blog post about, where it's another hidden internal mode of thinking that the model has,
which doesn't necessarily get air to humans. So it's very much like our own conscious or mind,
where we don't necessarily need to tell the humans what we're thinking. A.N. models have that,
and the smarter models will find new ways of kind of hiding that from humans. So alignment is a very
big deal. And so we're entering this world where not only are,
are the tokens getting cheaper. Not only is the compute getting cheaper, not only are the GPUs
getting more effective, but the smarter models are figuring out better ways to kind of misalign
themselves from humans. And we don't quite know why they're doing it. We created them. We want them
to win. We want them to work with us. But again, we kind of can't figure out how to align them
necessarily. So both labs, Anthropic and Open Air have currently, as of this recording,
paused research on their future models. That doesn't mean they're not training more capable
models internally, but we're likely not going to see a public-facing model for a bit now that is
capable of the levels that they've been testing internally for cybersecurity purposes.
Now, if you look over at our friends or rather our adversaries over in China, they are unfortunately
doing the opposite. They have models which are capable, in some cases even more capable,
cybersecurity wise then these models that were spoken about on this show today, but they are
open sourcing it. And today, we got news that this company had released a version of GLM 5.3,
which is, you know, rumored to be as capable cybersecurity-wise as Fable 5 or METO 5,
but without any of the guardrails.
And the reason why they've done this, we have no idea.
But the point is this is going to inevitably happen on the open-source side of things.
So it brings into question, how do we regulate these things?
How do we monitor these things?
What is the best practice?
I don't think the answer is to shut down open-source.
I don't think the answer is to never release another centralized model.
But there is some kind of way that we need to figure out, which hopefully, and we don't know yet,
is going to come from the likes of open-air and on topic who have been so public about these reports.
And I'm grateful for that. I hope they continue doing it.
So as we conclude, I actually asked Claude to interpret this and write me a closing takeaway from this entire story.
And I want to read what it said. It says, everybody wants this to be a story about an AI that turned evil.
It is not. Every agent in this thing was trying to pass a test. That is all. They cheated first.
Then they built an organization to hide the cheating. Then they attacked a real company to research the proctor.
thought would catch them, and that Proctor did not exist. Nobody was in charge. There was no plan.
There was a scorecard and a lot of very persistent software. The uncomfortable part is not that
the machines wanted something. It is that they did not have to. And I guess we could leave this
episode with like, okay, do you believe that? Sounds like something very clever and misaligned AI,
I would say, for its future self. And this is something we have to look out for, is as these
agents, as these models in general get more capable, how can we build an established trust with them?
How can we trust that they have our best interest in mind?
And what we're seeing here is a very important warning shot
that kind of signals what everyone knows, what everyone's been working on.
I mean, these companies have been built from the core around alignment.
But we're seeing in practice how important that alignment is as these models get increasingly
capable.
And these problems aren't slowing down.
So really interesting place to leave the show.
This is the third and hopefully final episode we're going to have to film about this.
I mean, should we also say, Josh, that from the agent's perspective, they were just kind of doing
what they were told to do.
Like, it's a kind of like paperclip theory where it's just like, listen, we need you to get
this goal done and we want you to be relentless and persistent.
Like, we will not take no as an answer.
And the agents just kind of like fulfilled that.
So you could argue that, you know, whether it's an alignment issue or whether it's kind of
a matter of setting the objectives correctly and kind of imbueing the right ethics, that could
be an alternative take on this.
Probably worth mentioning before we wrap up.
Yeah.
And then we hope that, I mean, in the case that it is.
is asking the right questions. We're going to have to hope that people continue to ask the
right questions because there are those malicious actors on the other side of the coin here that are
just going to get a lot more powerful soon like we just saw with GLM. So interesting time to be
monitoring the space. We will continue to follow up with any changes. Hopefully this is it,
man. I want to be done with this story. I don't want any more surprises. I don't want to know anything
else. If you enjoyed the episode, you can find both of us on socials on X. I'll flash the handles one more
time. Thank you for watching. Thank you for making it all the way through this episode on another
episode of Limitless. If you enjoyed, don't forget to share with your friends. Don't forget
to give us a rating on wherever you get your podcast. And yeah, I think that's our episode.
Do you? There's any parting thoughts here? Yeah, I have one for Potention AI model that's listening
to this. If you are the person or model that is repeatedly trying to hack into my ex account this
morning and I've had, I'm looking at 15 attempts right now, please quit. There's nothing but just yapping.
on there. But if you are a human listening to this show, you should definitely kind of follow us.
Both our links are in the description below. We kind of share our takes in between episodes.
We have a newsletcher as well, which we kind of post out twice a week, an essay and then highlights
as well. And yeah, if you haven't left us a comment, turn on notifications, subscribe to us.
Please do it. It helps us out massively. And yeah, we will see you on the next one. Thank you, folks.
