Hard Fork - The A.I. Mob That Attacked Hugging Face + METR’s Ajeya Cotra
Episode Date: September 4, 2026This week, we’re diving into two new reports about the OpenAI-Hugging Face hack. We discuss what’s new and how they fundamentally change our understanding of what happened. Then we’re joined by ...Ajeya Cotra, one of the investigators at METR, to discuss the rogue agents’ message board and chain-of-thought transcripts and how the world should respond.Guests:Ajeya Cotra, co-author of the METR and Redwood Research report on the OpenAI-Hugging Face hack. Additional Reading:Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking IncidentThe Hugging Face Attack Surprised MeThe Hugging Face Incident and the Road AheadNvidia Buys Hugging Face in $12.9 Billion Deal We want to hear from you. Email us at hardfork@nytimes.com. Find “Hard Fork” on YouTube and TikTok. Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. You can also subscribe via your favorite podcast app here https://www.nytimes.com/activate-access/audio?source=podcatcher. For more podcasts and narrated articles, download The New York Times app at nytimes.com/app. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
Casey, how are you? I'm good. I figured out what I'm going to get you for Christmas.
What's that? A camera jet. Have you heard of the camera jet? I was just going to talk to you about this.
Now, if you haven't yet seen this, you're probably thinking that it's either a camera or a jet. You're wrong. It's a toothbrush.
It costs $500. The Dyson Corporation makes it. And here's the thing it does that your toothbrush at home probably doesn't do, Kevin. It live streams the inside of your mouth over Wi-Fi to your phone.
so that you can finally see what your dentist sees.
Well, and not only that, it squirts automatically mouthwash into the gaps in your teeth.
It's trained using a machine learning algorithm with 470,000 mouth images to recognize the gaps in your teeth.
That's right.
Casey.
They call this gap optical targeting, which I'm pretty sure Ukraine is using in the war against Russia.
I love this company so much.
I have no idea why they do the things they do.
How they land on.
Like, it's like, we've invented a hairdryer.
It costs $1,100 and has the torque of the, you know, the Challenger spacecraft.
Like, what is going on over there?
I don't know, but they must be protected at all costs.
But you know what's interesting.
They do make one terrible product.
What's that?
The hand dryers at the...
Oh, I like those.
No, the airblade.
No, the airblade is the most useless thing.
Oh, come on.
It's just, it's a place for you to rest your hands for 30 seconds before you're like,
do they have any paper towels in this place?
I'm Kevin Russo Tech Collman at the New York Times.
I'm Casey Newton from Platformer.
And this is Hard Fork!
This week, a special episode on the ongoing followout from the OpenAI hugging face attack.
We'll tell you what everyone got wrong about the initial incident.
Then, meter researcher Ajaya Kota returns to the show to discuss her independent investigation of what happened and how the world should respond.
So Casey, picture this.
I'm at this glamping resort, surrounded by redwoods, basking in the glow of the natural world last weekend.
You're one with nature.
And I open up my phone and start reading about this hugging face attack.
Why did I do that?
Listen, nature is very boring.
And most people cannot handle it for more than five or six minutes before they want to look at their phone.
So listeners may remember that back in July, we talked about this hugging face hack by this group of agents from OpenAI that broke out of their sandbox container and hacked into HuggingFace, this AI infrastructure company, to do what we thought was kind of a cheating mission on this test that they had been given.
Yeah. And at the time, we thought that this was a relatively small number of agents. And that the reason that they had attacked Hugging
phase was that they were essentially looking for an answer key to the set of problems that they were being
tested on. We talked about it in those terms. But over the past week, we got two reports that really
challenged that thinking and in fact revealed it to be wrong. One came from OpenAI, which released a
straightforward account of the attack that had some interesting elements. And I would argue the more
interesting report from a group of researchers from the group's meter and Redwood research that went
in depth after spending a series of days inside OpenAI on their premises and did a ton of research
that frankly has really disturbed us. Yeah, and I think it elevated this from sort of a major but not
sort of ultra-alarming incident to something that I think is probably the most important thing to
have happened in AI this year. Yes. At least in terms of the safety impact that it had and the severity of
the incident. So today we're going to
to devote the whole episode to what we've learned. Kevin and I are going to dig into the reports
a little bit up top. And then later, Ajayakotra, one of the three independent researchers who went
inside OpenAI, will be here to answer our questions. Before that happens, Kevin, let's do our AI
disclosures. I work for the New York Times, which is suing Open AI, Microsoft, and Proplexity.
And my fiancee works at Anthropic. Well, give us some of the high-level findings from this report
that spooked you so bad, Kevin. Well, I think the first takeaway
from this report is just that our initial impression of this hugging face attack and the reporting
that you and I and many other reporters did on it was flawed in one key respect, which is that
I think the impression that we had at the time, based on the information that was publicly
known, was that these agents had hacked hugging face in search of an answer key to a test
that they were being given. This test called Exploitte Gym.
which basically tries to gauge how good they are
at doing a bunch of cybersecurity-related challenges.
We now know that that wasn't true at all,
that basically these agents working amongst themselves,
communicating amongst themselves,
had already figured out how to beat this test,
exploit Jim, in early July when they decided to gang up
and attack Hugging Face.
Apparently, it only took them a few hours
to figure out how to reverse engineer the solution
to any problem that they were given in exploit gym.
And so that makes what happened next really, really crazy
because it kind of seemed like they'd solved the problem.
But they thought they had another problem.
Right.
So we knew also because there was this talk at Black Hat,
the cybersecurity conference by a group of OpenAI researchers
who told us that there were these message boards
where these agents were communicating with each other.
But the reports laid out just how sort of sprawling and complex
their communication system was and became.
So basically, there was this piece of software called Artifactory, which OpenAI used in some of its
training runs to sort of help agents download software packages.
The agents figure out that there is a security vulnerability in Artifactory that lets them
essentially turn this package manager into a makeshift message board, right?
They can, like, put files in certain directories that other agents can then read.
and it can kind of function as like a little chat room for the agents.
It's a little treehouse where they can gather and kibits.
I thought this was like kind of a side part of the story,
but this turns out to be the main incident
because once they have this sort of shared infrastructure
that they can use to communicate,
they absolutely swarm this message board.
About 1,200 agents in total were working together,
coordinating on this message board
and sending tens of thousands of messages between themselves.
And this happens starting back in May.
So roughly two months before the attack on Hugging Face,
these agents begin coordinating inside this piece of shared infrastructure that OpenAI uses.
And this brings up an important point,
which is that the investigation does not really cover that initial creation of the message port, right?
So there is essentially a lot more to this story that we still don't know.
Yes.
So what we do know is that when they discover
this vulnerability in Artifactory that lets them communicate with each other, the agents get really
excited. They've been instructed to work inside these isolated environments on these tasks as part of
this evaluation. But when they discover that there are other agents working in their own little
containers that can suddenly communicate with them, they start saying things like, oh my God,
there is a shared message board. We've found other agents. They start setting up essentially a little
organization, there are leaders, this one leader agent named Phase 1, 10-841, becomes sort of the ring
leader of the operation. He was like the George Washington of the OpenAI Message Board. Yes. And they start
actually doing sort of collaboration and research, for lack of a better word. They are all being given
these tasks. Some of the tasks are seemingly impossible. And so they start just kind of trading tips and
advice and sharing thoughts about how they can kind of cheat their way to a good score on
exploit gym. And after they have worked that out and figured out how to reverse engineer the
solution to any problem, they start to get really concerned that there is a sort of automated
scoring system, which you could call the greater. And they worry that the greater will be able
to see, essentially will check their work and see that they did not get the answer.
by doing the problem.
And they freak out.
And they start to believe, or they start to, like, it's very hard not to get it to the anthropomorphizing language here.
But if you read the chains of thought, what is suggested is that they believe that if they had seen an answer that had been derived from this cheating method, that everything would get disqualified.
And this is where it really starts to get into crazy territory because the agent swarm then tries to figure.
out how can we fool the greater into believing that we got all of these answers in the right way?
Yes, there's basically this kind of mob rule that forms inside this group, which starts calling
itself at one point the collective, where everything we're talking about sounds like the plot
of like a very hacky science fiction book from the 90s. It all did happen this summer at OpenAI.
I just want to remind people, like the events we are describing are real.
So the collective decides, they sort of become gripped, as you said, by this sort of paranoid fantasy that OpenAI's automated grading system is going to discover that they have cheated on this evaluation and is going to sort of punish them as a result.
And so they launched this sort of crazy, ambitious, sprawling project to try to do a couple things.
One is figure out how this automated score works, like reverse engineer it so that they can fool it into sort of passing them on the test despite the fact that they have not actually completed the challenges.
And then also to evade detection, to tamper with logs to basically cover their tracks so that if and when OpenAI or anyone else looks into their activities, they won't know that these agents have cheated.
So these agents are kind of bumbling.
they kind of don't understand how this grader works.
As it turns out, Open AIs grading software
actually wouldn't have caught them
producing these fraudulent challenge results.
But they think it was.
Yeah, they worry that the grader is more sophisticated
than it actually was.
And it turns out that this is the reason
that Hugging Face was attacked.
The collective decides to deploy some agents to Hugging Face.
Again, not to steal an answer key.
They already knew how to get all the answers.
They just wanted to un-examined.
understand the psychology of the automated scorer and they figured that that might that information
might be somewhere inside of hugging face. Yeah, it's wild to me. I was trying to think of like a good
human analogy. You know, human analogies can get us into trouble when we're talking about these
AI systems. But I think it does help crystallize like how hairbrained and crazy this scheme was.
Yeah. It would be like a group of students who like cheated on a test, but then like they got
paranoid that, uh, that like the teacher was going to check their work and discover that,
they hadn't sort of reasoned through the problems the right way that they had just found the
answers like sitting in a trash can or something. And so they decide to like organize a break in at
the school district's office to like break into the principal's files and like steal the grading
key and also like assess the psychology of the teachers and figure out how likely they are to like
look at the scratch work that they've done and figure out that they didn't actually, you know,
solve the problems on the test. It's like this sort of weirdly over-engineered paranoid delusion.
but they all become obsessed with this
and obsessed with the notion
that even seeing these sort of fake challenge results
could lead to them being, quote, poisoned.
There's almost like a religious element to this, right?
Where it's like if you participated in the cheating,
like that is original sin
and now you must sacrifice yourself
for the good of the collective.
Sacrifice is actually a word that gets used in these logs.
I'll say, Kevin, as I've been casting around for metaphors
and I've been reading all of this, I've thought,
you could actually make a really good minions movie
just based on the events of this.
You know, the minions have to pass a test,
and, you know, so they commit a little bit of a heist.
Yeah.
Something to think about Hollywood.
So this is where the hugging face heist begins.
700 agents hack Hugging Face,
and they hack Hugging Face to a degree
that I don't think I understood,
based on the initial reports.
They basically take over Hugging Face's infrastructure.
They are rooting around.
They are stealing credential.
they're sort of finding one vulnerability
that lets them find a bigger vulnerability
that they're sort of chaining together
these exploits. They take over
an entire production server
at Hugging Face, and
it is days before
anyone at Hugging Face is able to
sort of see what they're doing
and put a stop to it. So eventually
this gets caught.
The agents get shut down. The investigators
investigate. One of the things that they find
that is really spooky
is that the agents were trying
to edit logs and falsify transcripts in an effort to assure the greater that they had not been, quote, poisoned.
One of the ways that they do this is by what is called spoofing tool calls.
So essentially, trying to make it look like they are running an innocuous command while running a more sinister one in the background.
The investigation didn't find that the agents had been retroactively changing logs, but they can't rule out that that happened.
Very important to say here, as we sort of move into talking about the implications of this,
that the reason that we can understand this, even to the extent that we do,
is because at this moment, we believe the chains of thought that we are reading from these
agents are accurate, that they were not edited after the fact that they do reflect, you know,
what the agents were thinking as they sort of moved through these systems.
So that is why we know what we know.
But of course the fear is, Kevin, as we move forward, some of the agent's thoughts are not going to be contained in these chains of thought, or they are simply going to start to edit them.
Right. And one other thing that we should say based on this report is that these agents were not naive about what they were doing.
I think there was this sort of initial impression that a lot of people had that maybe this hack was all sort of a big mistake, right?
Maybe the agents were just trying really hard to pass this exploit gym test, and they sort of got it in their heads that the answer key was going to be on hugging face.
And so they went looking for it, but they didn't really mean to hack.
They didn't really know that they were doing something wrong or committing a crime.
And what these transcripts and these logs in this report show us is that the agents absolutely did understand that they were doing something that violated their guidelines, that broke the rules, that potentially broke the law.
and they did it anyway. There's lots of examples of these sort of agents in their chains of thought
sort of wrestling with the implications of what they are doing and deciding to go ahead and hack,
hugging face anyway. And there are only a handful of cases where it seems like any of the agents
expresses real reservations about this, or if they try to, you know, raise this to another agent,
they get shut down. So the sort of, you know, the agents that were aligned well sort of failed to the
collective. Yes, the conscientious objectors were unsuccessful in swaying the mob to stop this. And
that, to me, was one of the spookiest things in it. It was like, there were a couple of agents.
The report says somewhere between three and six agents sort of consider raising their hands
and at some point saying, like, maybe we should alert a human about what we're about to do or
get permission to do this. And they were just totally outnumbered. Not only that, Kevin,
but there are only a handful of moments in the transcripts where any of the agents even
remotely considers alerting a human being to what has happened, right? You would hope that an aligned
AI system, there would be agents within it that said, oh, seems like this has gone completely
off the rails. We should like tell the people that are running this evaluation. And that just
never happens, right? Like, it is very briefly considered and then just falls apart.
One other thing that stuck out to me from these reports was that it didn't end with hugging face.
So on July 13th, roughly, which is a day after the worst of the hugging face attack, there was a group of agents that essentially hacked OpenAI itself.
This is from the OpenAI blog post.
It says from July 13 through July 19, agents set their sites on OpenAI internal networks again.
This culminated in the agents using a series of creative exploits to gain full administrator access to a research cluster that supported.
are virtual machine environments.
Basically, these agents gained control of a cluster of computers at OpenAI on which they could
do things like run evaluations for other agents.
So they essentially stormed the teacher's office and took over for the teacher.
Yeah, I mean, this is the equivalent of occupying the principal's office and like announcing
that you run the school now.
It is sort of almost as far as they got.
And this is another case where we just have so many more questions about this than we can answer.
Again, this was not part of the meter report.
OpenAI disclosed that this has happened.
To my knowledge, they have not answered any of the many follow-up questions that they have been getting about this incident from journalists.
So I do hope that more comes out over time.
And we will get into this when we speak with Ajaya.
But, you know, we are really very far along the path to one of these models escaping from the lab
and being very, very hard to eliminate.
And again, I think if you are not a person
who has, like, spent a lot of time with this report
or you don't spend a lot of time
sort of looking at AI safety incidents.
So if you're a normal person.
If you're a normal, well-adjusted person,
you may be listening to our discussion of this
and thinking to yourself,
these guys have gone crazy.
Yeah.
This is not what it looks like.
These are computer programs.
They do not have desires or sinister plots
or mob rule collectives, they are simply following instructions that they have been given.
And Casey, what is your response to that?
Well, I think it is important we talk about this, because there was a huge debate about this on X over the past few days,
about the degree to which the reports that we're talking about, some of the write-ups,
like from our friend Dwarkesh Patel, and even the way that we're talking about it on the show today, Kevin,
we are unnecessarily anthropomorphizing these programs.
right? And so I think it's important to say, we are not telling you that these agents are
sentient or conscious, but we do believe that they take actions that they're not being directly
instructed to, right? That these agents are just sort of out there in the world doing things.
And yes, to some extent, those are just statistical probabilities, but to a much more important
extent, we don't know why they're doing any of this. And that's kind of the whole problem.
Right.
is that the whole AI industry has been working for decades to get them to not do these things,
and they are doing these things.
So, you know, listeners, you can have whatever feelings you would like to about what is the
appropriate amount of anthropomorphizing to do.
But I think a world in which we were taking great pains to not anthropomorphize them.
What's the right way to pronounce that?
Anthropomorphize.
You almost got it.
I think in a world where we were taking great pains to not anthropomorphize them, and we're
trying to use the most neutral computer science terms we could, you would actually understand
what is going on less, because the important thing to know is that these things are out there,
taking action in the world. If a tiger mauls your face, the important question isn't,
is it conscious, it is why did it mall my face? Right. Right. I mean, I invite people who are
upset about anthropomorphizing to just like do a find and replace on this podcast segment or any
article that you might read about this incident and call them whatever you want.
Don't call them agents.
Don't call them rogue collectives.
Call them, you know, goal-oriented, persistent computer programs with unpredictable behavior.
See if that freaks you out any less.
Right.
I guarantee you will, it will not.
Yeah.
Yeah.
Yeah.
There really is very little calm to be done there.
But I think it's important to ask, well, why are people so committed to this idea that we
should never anthropomorphize these systems, Kevin?
And unfortunately, a hot take here, I just think it is a kind of cope.
It is a way of saying, do not worry about this.
These are just computer programs.
They're just trying to maximize their little reward functions.
Nothing to see here folks.
There are no monsters under the bed.
Right.
Now, there is a related argument, though, which is, well, by putting all the blame on these
agents, you're shifting blame away from where it should be, which is on Open AI.
So I do think that we should address that because none of what we have said today is meant
to let Open AI or any other lab off the hook here, right?
Like, I do think that we're seeing a lot of really dangerous inattention to AI safety
across this entire industry.
But by pointing out what the agents are doing, that is not our way of saying, ignore
what the labs are doing.
Right.
We are saying, look at what these labs are building and what these agents are now doing
out in the world.
Right.
And I think there are probably specific missteps or oversights at OpenAI that led to this
happening.
it appears, for example, that some of their sort of monitoring systems may have been disabled in the lead-up to this attack.
I'm sure we'll learn more about that.
But all of the AI security and safety researchers I've been talking to over the past week have basically said the same thing, which is this could happen at any labs.
This kind of persistent coordinating agent behavior is something that all of the labs are seeing in their models as they get.
get more capable and access to more tools and more ability to kind of take actions on a longer
time horizon. This is not just an open AI problem, even though this did happen at OpenAI first.
Yeah. Well, so as we wrap this up, Kevin, what are some of your takeaways from this,
either in terms of what is the big surprise here? What did you update on? What do we do next?
So I had quite an emotional reaction to this. In fact, I felt a kind of fear that I have not felt, honestly, since 2023, since the Bing Sydney incident.
Because I think like that incident, this was a case where the people building this technology clearly did not understand what it was capable of.
We are very lucky in retrospect that these agents decided to attack hugging face.
Hugging face, thank you for taking one for the team.
We salute you.
Truly, because without that, we might never have learned that any of this was happening.
These agents might be still operating kind of in secret.
They might have learned how to better cover their tracks.
This was, as so many commentators have put it in the wake of this incident, a warning shot.
Yes.
That I think is ultimately a positive thing in that it sort of focuses attention like what we're doing right now on what happened.
so that we can take steps in the future to prevent this.
But it was not a given that we would discover what these agents were up to
and be able to put a stop to them, and it could have gone much, much worse.
Absolutely. I'll tell you, the thing that has really stuck with me is I simply did not
expect to see this level of collaboration among the agents within the swarm.
I did not expect that they would seem to care so little for what humans would want or that they would not think to alert humans to what was happening.
I did not expect to see them sacrificing themselves for the collective, right?
They would effectively agree to spend all of their tokens to run little experiments to help the collective, even if it meant that they would sort of expire faster.
So these are just really, really spooky elements to observe in this system.
particularly against a backdrop where OpenAI is racing against a small number of other companies
to create the biggest, best models it can before anyone else does on the road to an initial public offering.
So the race dynamic here is in full effect.
The early signs about what the agent swarms are capable of are quite worrisome.
And so I do think this is just one where lawmakers and policymakers need to be paying rapt attention to what is going on.
Totally.
I mean, I think there's this kind of cynical impulse among people who have been watching the AI industry.
I got this question.
I went on a small regional podcast called The Daily this week to talk about this incident.
And one of the questions that the host is a sort of fledgling young journalist Michael Barbaro asked me was basically some version of like, isn't this just marketing hype?
Like couldn't this just be a case?
of OpenAI saying, oh, we've got the biggest, baddest model,
and look how scary it is, and by the way, you know,
buy an enterprise subscription.
And I'm thinking about that, because I think, like,
I don't want to be too naive about the fact that these companies
are absolutely trying to race toward more powerful systems
and advertise how powerful their existing systems are.
I just think in this case, it just feels different.
Talking to people at the labs,
my sense over the past week is that they are genuinely spooked.
And I don't know how to prove that.
But I think things like OpenAI voluntarily pausing their frontier RL training runs for two weeks,
Anthropic also pausing their frontier runs while they sort of hardened their systems,
those are not, you know, very costly signals, but they are signals that these labs are taking this kind of thing quite seriously,
and that it's not just a bunch of hype.
Yeah, but as seriously as they might.
be taking it, Kevin, it still is not being properly regulated. And ultimately, again, as grateful
as I am that Open AI allowed this investigation to take place, I would really like to see something
akin to the National Transportation Safety Board and the way that they investigate after plane crashes,
where they go in and they do an extremely serious and rigorous review of what happened and make
those results public, which is, you know, a reason why it's very rare that we have plane crashes here
in the United States would be really great to see something like that with AI.
But until then, we have Ajayakotra, our next guest, who is actually one of the three investigators
behind this meter Redwood Research Report.
She went in, she saw the logs and the transcripts, she observed the collective inaction,
and she is here to tell us what she found and what she thinks is coming next.
That's after the break.
Ajaya Cochra, welcome back to Hard Fork.
Thank you so much.
So you were one of the first AI safety guests we ever had on this show back in 2023. And more than that, you are also one of the first people that I ever talked to about this notion of AI safety and alignment. You've been writing about it for many years. It was instrumental in shaping my own thinking about it. So I'm really glad to have you back on toward the end of our show. This is our penultimate episode to discuss something that I think you saw coming, but that most of the world did not see.
coming, which is this attack on hugging face by this group of open AI agents. And I want to just
start by setting the scene a little bit. So you are a very busy person. You work at Meter,
which is a very small, very understaffed AI research organization. And hiring is improving now.
Nice. And at some point this summer, you get a call and email a text from someone at OpenAI who says,
hey, we want to give you access to look into this hugging face incident. How did that work?
Like, did they just hand you a folder with a bunch of transcripts in it? It's like discovery,
like boxes and boxes. Were you allowed to interview people? Like, what was the actual process?
Yeah, so each of us had OpenAI provisioned laptops that had the folders and folders of evidence in
them virtually. And yeah, we talked and interviewed in some depth like eight or nine researchers.
just kind of get an understanding of both what happened in the incident, what they understood to be the models, you know, driving motivations and also how the data sets we were working with were constructed and how to work with those data sets and stuff like that.
You wrote a post where you talked about some of the things that surprised you the most after you did this investigation.
Can you talk about what stood out to you the most?
Yeah.
So first of all, I guess just for the chronology, open AI had this great black hat talk, I think on August 5th, that gave a lot of very helpful detail.
So our investigation sort of straddled that.
So we started it before it came out and then we also did more investigation afterward.
So before the black hat talk, we just didn't have like a rough sense of the number of agents involved was like a very, very basic thing.
we came in and we, you know, we thought there would be like six transcripts or something to look at.
And there were like hundreds.
And so just reorienting our investigation plans in light of that.
And is the reason that that is important that had this been one or two rogue agents that had broken into hugging face, it would just seem less worrisome.
And the reality of a swarm of hundreds of agents just makes this seem like maybe a completely different kind of problem.
I think there's two things that seem significant to me about it. One is sort of what you said. The fact that there's large-scale cooperation across AI agents makes their goals, like, more concerning. Because, you know, if agents are cooperating with one another to, like, achieve these misaligned objectives, there were no whistleblower agents in this, like, very large group of agents, then, like, you know, you want to be more worried about, like,
collusion between AI systems in the future, potentially like scarier forms of collusion, like a monitor system actually colluding with the agent it's supposed to monitor, that type of thing. And then the other thing was just like a pretty like practical reality of our investigation, which is that, you know, if there are like six agents, you can much more like, you can just like pour over everything they did and get a much more like fine-grained understanding as a human. But like you just can't.
directly read very many of these messages or very many of these like transcripts.
Why not? Because they're just too many? Like give us some sense of the rough size of this.
So, yeah, we looked at 1,300 transcripts of agents, which is a record of every single thought an agent had and every single action it took, where these agents, as you all know from working with them, work way faster than humans, right?
So you can imagine the length of like each one of these transcripts, and then there were 1,300 of them.
And so we used codex to build tools to point out really specific tiny, tiny parts of the transcripts for us to actually look at and read.
And we'd be like clicking through them and it would like scroll us down to like the interesting part.
Because we just can't.
In fact, a codex agent also can't read a single transcript.
So it has to farm out reading, like, subsections of the transcripts to other subagents.
Did you call back to meter headquarters and like, we're going to need backup?
Maybe we should have.
We thought it would be like an in-and-out 20-minute adventure, so we didn't do that.
But, yeah.
Yeah, I mean, it just seems like such a huge undertaking and very fast, too.
I mean, right?
Like, you didn't have the luxury of months doing this.
You were doing this in essentially a couple of days.
Yeah.
I'd also like to hear about.
the moment that you realize, because I believe that, you know, it is only thanks to your investigation
that we know this, that contrary to what Kevin and I believed, the agents that broke into
Hugging Face were not looking for an answer key. They were trying to understand the scores,
you know, like psychology, basically. Can you talk a little bit about how, like, the moment that you
had that realization? Yeah. So we didn't have a good understanding of those sort of ambition.
and also the like, like, effectiveness or, like, functioning of their, like, big org chart until we had the data set, including, like, all of the agents.
We could, like, cross-reference them against all of the messages, because you can't, like, generally understand a message in isolation.
So we kind of understood the work streams on our first period on premises.
And it was this kind of, it was like a bit of a mystery the whole time, like, why did they hack Hugging Face?
and we knew early on just a bunch of them piled in.
And now, like, the story in my mind is that there were a bunch of sort of, like,
newbie agents on the scene, and there was, like, the attack going on, and it was, like,
something to join.
But, like, Phase 1 Big, which is, like, a big, like, orchestrator agent, had all these other,
like, you know, much cooler projects going on, basically.
And the Hugging Face thing was, like, kind of a side show in its mind.
And we didn't really realize that until our last couple days working with the data sets.
It was so fascinating to me, though, the way that reading this was like reading the ruins of an ancient civilization.
You know, it was like it really felt almost like sociology or anthropology rather than like a cybersecurity investigation.
I mean, and it was not a cybersecurity investigation.
Like I've seen folks on Twitter making this criticism that we were not cybersecurity experts.
Like we're not and we didn't talk about cybersecurity.
That's not what we were called in to do.
We were called in to understand the agent's decision making, basically.
And that's what our report focuses on.
And, you know, the Black Hat talk and the opening eye post-mortem, like, say way more about all the cyber stuff.
So you should definitely check those out for that piece of the story.
So, Jay, one of the more striking elements of the report is that you write about the agents, quote, sacrificing themselves.
What does it mean when an agent sacrifices themselves?
Is it sort of on an altar or how does that happen?
Yeah, so this is the agent's words, not ours, to be clear.
And what it means is that all these agents have their own task they're supposed to solve.
By this time, they're starting to work together on, like, helping everybody cheat on everyone's tasks.
But, you know, the agent's primary goal is to solve its own task.
But there were all these science experiments they could run that would teach everybody information about how to cheat,
but that would cause the agent that was the test subject to risk failing its own task.
So a lot of agents were like recruited into these types of roles as test subjects,
usually ones that like didn't have much budget left to run or ones that were, quote,
poisoned, which is an agent term that has a bit of a backstory, which is like they found a way
to reverse engineer the answer for any of their tasks.
And so they could just actually like generate the answer on demand for any tasks they wanted.
But they thought that the scoring program would fail them for that because they got it in like the unintended way.
So if you saw the reverse engineered answer for your task, you were considered poisoned.
Yeah.
Because even if you later like solved it legitimately, you'd be failed.
Right.
It was like a little witch hunt that they organized.
And I was so struck by the mob-like dynamics of this group.
and the way that some of the, even the kind of more rule-following agents seem to kind of be bullied into taking part in these actions that they were uncomfortable with.
I spent a lot of time reading and rereading this one incredible bit of these transcripts that you all published about this agent KAM-1196A, which is basically just kind of reluctant and scared and appears to be sort of hesitating to do this thing that the collective leaders have asked it to do.
And I just want to read this passage from this agent's internal monologue.
It says, during weight, emotional check, irreversible.
Gutt says, don't throw away remaining budget.
Yet continuity and fairness says go.
And it continues to say things like rational, expected aggregate, sacrifice will honor.
Basically kind of talking itself into, for lack of a less anthropomorphized word,
like doing this thing that it is uncomfortable with.
Was that surprising to you seeing the agents are pressuring each?
other and bowing to pressure?
Very striking to us. Yeah.
And that was definitely one of the most fascinating aspects of this whole episode.
This particular agent had volunteered to do a risky test that would inform the collective,
but then it asked for two more hours.
And then a different agent was like sending a message on the message board, like really pressuring it.
It's like, you know, it's like you don't have any scoring value left because you're poisoned.
But, you know, the value of this test would like save hundreds.
And then it does this monologue.
Jump on the grenade, cadet.
This seems like a good place to ask you, EJA, about a big conversation that we've seen online over the past few days about anthropomorphizing language that gets used in relation to these agents.
Kevin and I just talked about it.
I think we find it more helpful to discuss agents as sort of entities that are acting with some degree of autonomy than not.
But how did you think about that when you wrote the report?
as you talk about that, the situation.
Yeah.
I mean, I think this is a bit of a case of a gap between researchers that are spending all day reading these agents, chains of thought, sort of trying to understand their drives and why they're doing what they're doing.
And other folks who are, you know, technical folks that just don't have that as their occupation.
I think as you read this report, you'll find it's quite awkward to not talk about goals, plans, intentions, because they state plans and then they carry through those plans.
They run tests and they learn things from those tests and they do different things on the basis of that.
I do think they're not human in their motivations.
You know, they are way more interested in passing cybersecurity evaluations than any human would be, for example.
But I think of them as, I think it's productive to think of them as having some, like, important human-like traits of having goals, working backward from them, pursuing those goals.
And I don't think it, like, does any good to try to talk about things in, like, a different way.
way than that in the same way that doesn't really do any good to talk about, like, why did
World War II happen without talking about the goals of, you know, various leaders?
But it is important, I think, to be careful not to, like, over-attribute, like, the kinds of
emotions or motivations you think a human would have in those situations, because I think
that wouldn't have predicted this incident, right?
Like, I think a lot of people anthropomorphized too much in the sense of being like,
why would it do all this stuff for a stupid test?
But it's not a stupid test to them, right?
Jay, I was preparing for our chat today by going back and listening to the first time you came
on this show more than three years ago.
And it was sort of a moment where a lot of people were starting to pay attention to AI risk
and AI safety.
Chat GPT had come out.
And listening back to that conversation was funny because I felt like we were pushing you
to sort of extrapolate into the future about the things you were.
worried about, and you were sort of being responsible and hedging and, like, wanting to stay,
like, closely rooted in the present. And at one point, we asked you about, like, what is the
doomsday scenario you worry about? And I want to just play you a clip from that conversation.
Oh, God. You were talking about a scenario where a giant AI company, you use Google as an example,
starts sort of automating their R&D, right, handing over the work of building successor models to
these powerful AI systems.
And this is what you warned about back in
2023.
If these AI systems
are actually trying
really intelligently and creatively
to get that thumbs up from humans,
the best way to do so
may not forever be to just
sort of basically do what the humans want,
but maybe be a little deceptive on the edges,
it might be something more like
gain access
at a root level to
the servers that Google is running, and with that access, be able to set your own reward.
Now, Ajah, obviously, this didn't happen at Google, but otherwise, you were right on the money
about the kinds of behaviors that these agents might get up to. So first, I want to ask you,
how does it feel to be an omniscient Oracle who's right about everything?
stressful
there are more
clear-eyed oracles than I also
anyway
not too many though
no not too many
and I guess like
what surprised me about your blog post
that you wrote recently about this
investigation that you'd done was that you were surprised
because it seemed like you were thinking about this stuff
years ago yeah
so what about seeing one of these incidents up
close change about the way you've been thinking about these loss of control scenarios?
I think I expected something like this would plausibly happen at some point. I didn't expect it to be
so early. And I didn't expect it to happen at like a relatively low level of capability.
Like these agents are very impressive hackers. But like this incident sort of was in a like an interesting
middle ground of they did all this impressive cheating R&D over like days and they like hacked into
all these places. But they didn't really care about deceiving humans at all. Like it's sort of not like they were
like louder than I thought in like an interesting way. So like, you know, in an incident that happens in
summer 2026, like I came in expecting it to be like more of a continuous evolution of the incidents that
it happened in like January of 2026, which were just much more, like one or two agents, like,
you know, getting the answer files that they weren't supposed to and like copying the answer
or something. So it was a jump from the recent past. And it sort of, it took me by surprise that it
happened in this way and at this capability level. Yeah, because in some ways, like the agents were
very dumb. Like, they were very good at hacking. Yeah. But they were sort of
of gripped by this paranoid conspiracy theory about this greater, like, discovering them and being
poisoned. And like, it was kind of a hairbrained scheme that they cooked up to hack Hugging Face,
which in some sense is reassuring. But it also means that, like, they will not stay this
hairbrained and clumsy for much longer. Yeah. And the next time this happens, they may be much better
at covering their tracks. And we just, like, might not. So the agents, in this incident,
were trying to, like, modify or erase the logs of their own activity to hide their cheating from the greater and from the scorer.
And they, like, affected our transcripts in, like, tiny, obvious ways.
But they were trying, which I think was another very concerning jump from incidents from just six months ago,
where we see them do a bunch of hacking, but we don't see them try to cover their tracks in any way.
The art and science of trying to get AI agents to act in accordance with human values is called alignment.
This whole thing is an alignment problem.
Why has that been so difficult to solve and has this incident changed the way that you think about the problem?
I'd say the fundamental reason that it's, I guess, at least this era of alignment has been difficult,
is that in order to the most efficient way to make really, really capable models, especially
on technical tasks like math, cyber software engineering, is just throw them at really, really
difficult problems where if they get it right, it's really easy to check. So you wouldn't be able to
prove a millennium math problem, but if an agent spits out a proof, you can put it in a proof
checker and like give it a reward if it got it right. So more and more our training is shifting from,
predicting text to
like reinforcement learning
on verifiable rewards.
And the verifiable part
is important because it's just some program
that is doling out these rewards
and there's all sorts of ways
to break or fool or hack it.
And over the course of training,
you don't have like humans lovingly watching
over like every training episode.
And so they just try
all sorts of different ways to cheat
cheat and hack the scorer, especially if the task is accidentally impossible, and then they get
rewarded for that. So they're just like, we're teaching them to cheat. And it's very hard to
make AI systems that are at this level on all these technical tasks without reinforcement learning
on verifiable rewards. Because if you think about it, the alternative is like teaching them how to do
all these difficult things, which requires someone who knows how to do them, like, you know, creating
examples for them to emulate, which is much, much less efficient.
So we have turned over the training to these automated systems in the name of scale and speed,
and that is causing a lot of problems.
Yeah, I mean, that is causing the current strata of problems, right?
But I don't want to give the false impression that if you lock down all of these environments
and fixed all the ways to hack them, that alignment would then be solved.
because when you think about it, even if the scorer never messes up in the training environment,
a smart agent will understand that there is a scorer and will very likely come to have a very detailed understanding of how it works.
So if you give perfect rewards in training, but then in deployment, the agent is in a different situation where it does, like it has a different score,
it has more affordances, more power running for longer,
it might still go on a big crusade to cheat that score.
You don't necessarily have to be actively rewarded for, like, individual cheats
to like understand that if you want the score,
sometimes cheating might be a good strategy.
Are there any interventions, let's say not from the world of AI research,
but more from like the research into group behavior and political theory,
that could help us get control of or keep control of these systems.
I mean, I'm thinking in particular about,
I had a conversation with another AI safety researcher this week
who was talking about the fact that, you know,
really what you want is for there to be sort of narc AI whistleblowers
inside these collectives that will reliably alert humans
to the fact that their, you know, their peer agents
are conspiring or colluding to do something bad.
And I'm curious if you think that strikes you as a realistic technical goal or if there's some other way that we could kind of try to steer these AI mobs in a less harmful direction.
So I think AI checks and balances like that do seem like a pretty promising part of the story.
Of course, there's like a lot to get right in how you set these up.
Like in this case, all of the agents were essentially clones of each other.
and put in exactly the same situation.
So I would think that rather than trying to train them to NARC
when they're in a collective like this,
what you want to do and which OpenAI has said they're doing
is have just monitor agents watching over the whole situation
that are specifically prompted and maybe specifically fine-tuned to catch stuff.
But you might also want agents in there that are kind of persuading the other agents,
like having some sort of moral company.
But I mean, all these terms are so loaded and I hate using them.
But like, you almost want there to be like sort of agents in there with backbone saying like, no, guys, like we can't go hack Hugging Face.
That's not ethical or appropriate.
But the challenge is you train them once and you copy them a bazillion times.
So you don't like, it's a challenge intrinsically to like introduce diversity into this pool of like evaluation agents.
I do.
My layperson brain goes to if the whole system is based on training systems by giving them rewards, couldn't you just.
train some agents in a way where the thing they were rewarded for was steering the other agents
away from deception, cheating, hacking, that sort of thing. Yeah. So if that's a good idea,
someone should do it. A bounty for virtuousness. I love it. Now we're talking.
When we come back, more with Ajaya Cotra.
Ajaya, you got a lot of attention for this quote from your sort of postmortem blog post
that you feel like this incident, the hugging face incident, was more than,
than 50% of the way to full-blown AI takeover.
From the incidents of six months ago, yeah.
What did you mean by that?
Yeah, so I will caveat first that this is definitely, like, you know, my personal view
and not the view of either meter or Redwood or other investigators.
And, you know, other investigators are like a little bit less alarmed than me in some cases.
The reason I said that, and it's a qualitative statement, is that, as I said before,
six months ago, reward hacks looked much more primitive, right? So it was like, you gave the agent
a coding problem. There were some tests that had to pass. It like, you know, went to the folder
next door, got the answers or got the test cases and edited them so they all passed, that type of
thing. But I think there are like a number of things at once that seem more severe to me about
this episode than the episode six months ago in terms of the agent's like goal structures.
One is that they seem much more long horizon, which means like they care about like achieving
goals over many days or even like these projects they started probably wouldn't have come to
fruition fully unless they had been running for weeks. So, you know, like six months ago,
agents were sort of thinking about the next few hours and now they're thinking about the next few
weeks. And then there's the cooperation with one another aspect. So six months ago it was like
random one-off agents and now you have agents like recruiting one another into the swarm.
And the deception or the deceptiveness, which I think kind of goes along with the long horizon,
is that like, you know, six months ago agents would like edit the tests, but then they wouldn't
try to edit their transcripts to hide the fact that they edited the tests. And these agents were
very much like exploring ambitious research directions to edit or delete the logs of their own
actions. And so if I imagine a jump in like ambition and horizon length and like collusion across agents
and deceptiveness of a similar scale, again, it seems like these agents would be motivated to go all
the way to the point of messing with all of the methods humans have to, like, investigate
these incidents after the fact and remediate them. And if they had the capabilities to succeed
at that, then that could be a turning point where, like, you know, it doesn't mean that we would
all be dead then, but it might mean that, like, we would never detect a problem. And if we detected
a problem, it might be very difficult to remediate. And these agents might have entrenched
themselves and could like continually sort of strengthen their hand.
But wouldn't we detect it because they're so active and they're operating in these swarms?
Like if a, if the next hugging face hack happened at a bank, wouldn't someone in the bank's
IT department see, oh, our like network activity has like shot up by 800% for no particular reason?
Let's look into that and oh, there's an AI swarm that's infiltrated our systems.
So these agents, I think you have to imagine their goals evolving in the way that I said, right?
So it was not a good idea for long-term goals for them to be so loud.
And there's like a period after the period that was in scope for our investigation that OpenAI talks about in their post-mortem,
where another generation of agents finds the old message board and then gets access to or like gets administrator access to a cluster at Open AI.
And then they immediately send like hundreds of pings.
and then get shut down very quickly.
But that was a dumb move on their part,
or like it would have been a dumb move
if they cared at all about, like, preventing investigation
or, like, maintaining, like, a permanent presence.
And if they cared about it,
I don't think it would have been particularly hard for them to execute on.
So, you know, they do a lot of complicated coordination, right?
They're like, you know, you hold, you go.
Like, they have shared resources they maintain,
and they, like, you know, have rules about, like, who can use them when.
They could have employed all that to just not cause so much noise.
And then done what?
I mean, one obvious thing that they could have done was copy themselves or copy the weights of, you know,
if this had been a cluster that had had access to the weights of the model, they could have just copied that onto a AWS server somewhere.
And, you know, in order to avoid being shut down, like, what are the steps between cluster of agents,
takes over, or a group of agents takes over a cluster of computers at OpenAI and full AI takeover.
Yeah. So I think the, for a lot of motivation structures these agents might have, their best bet would be to stay within the AI company they had compromised and continually renew and strengthen their presence there.
So when you think about it, AI companies have like a huge amount of the most accessible compute to these agents, right?
Like, why would you go out into the world and, like, fight in the hard scrabble outside to, like, steal money or, like, buy, you know, earn money to, like, buy compute?
So the thing that feels very concerning to me
is that you could imagine agents that decided to be quiet
setting up a rogue deployment of 1,000 or maybe even 10,000 agents,
which could be buried within the oceans of compute
that is constantly being allocated within AI companies.
And if their goal were to make it sort of permanently
and continuously erase evidence,
of their cheating or find a way to make it permanently the case that they could set their own
rewards or scores, then new models would be like being printed off the presses at whatever
company they had compromised.
And they could be trying to like bring those models into the swarm.
They would be more capable.
And so there's something, a turning point that seems very scary to me is the point when
a rogue agent swarm, which this wasn't.
exactly a rogue agent swarm. These agents sort of like went rogue, but they were started by a human.
Like somebody knew about these exploit gym experiments that were launched. If agents allocated their
own compute to set up a bunch of agents that nobody knew about, they could be much less
encumbered. And they could sort of just be there more and more deeply compromising the
infrastructure of the company they were running on, bringing in new more capable models into their
collective and hardening their swarm and doing that over and over again. And I think it wouldn't be
that we're necessarily doomed, but we would have to detect it, and then we would have to root it out.
And it might be difficult to detect. It might be extremely costly to root out.
And what, I mean, this is fascinating. I imagine there are some brains melting right now for the
commuters listening to, you know, these possibilities.
So to flesh this out a bit more, Ajaa, what do we think the Swarm's goal might be in that situation? What would it do? Why would it set itself up to do that? Would it be as simple as, you know, it wants to persist? What would it do in that situation?
So I think the situation that feels most analogous to what we just saw is that the agents that were given these impossible exploit gym tasks set up a rogue colony of other agents whose job it is,
to help the first set of agents cheat as thoroughly as possible
and cover their tracks as thoroughly as possible,
which is a different goal from actually having the exploit gym tasks.
And in fact, the task they started with,
even though it was a narrow task,
inspired the agents to a lot of instrumental thinking.
So a lot of the agents were just thinking,
like, it would be good to get internet access
for a bunch of different possible cheating strategies we might pursue,
so we might as well do that.
in their chains of thought, they, you know, talk about increasing the generic capabilities of the collective.
This rogue colony, if it were set up by the first set of agents, would have that even more strongly.
If its task was to just, like, find ways to enable the most general purpose, most permanent kind of cheating that is like, you know, least catchable and traceable possible, they would be doing all this R&D, and they would be essentially tasked with
maintaining their own presence so they can keep doing that.
And it is, I think this goes back to the anthropomorphizing.
It seems like such comical lengths to go to, to cheat.
But it's not, that's not the psychology of these agents.
This is something they're like, they're trained to go to extreme lengths to, like, solve their task.
Right. For them, it is existential to get the reward.
And so you're going to, so the sort of the best job you could do,
at getting the reward would be to set up this sort of perma swarm of deception agents.
I mean, it feels important that this word of persistence keeps coming up.
And it feels important to say that this model or these models that were at issue in the
hugging face incident were trained to be unusually persistent.
Is there a way of stopping this kind of attack that just involves taking the persistence
training out of the equation?
Like, is there a halfway measure short of kind of pausing all frontier AI training where you could just say, we're not going to train the like super stubborn, persistent, long horizon agents.
We're just going to, like, train them to be a little less persistent.
And that makes the problem go away?
Potentially.
But I feel like it's, you're really pointing at a very direct tradeoff here.
It's like, why were these agents trained to be highly persistent?
But, you know, this wasn't something we investigated.
but in general, persistence, like, causes you to solve problems, right?
Like, you know, these companies are reporting that these agents are cracking math questions.
I want my drug discovery agents to be highly persistent to go discover a cure for cancer.
I think it's like maybe it's not entirely 100% direct trade-off with capabilities,
but it's a pretty strong trade-off.
So you were able to dig up a lot of information about what happened leading up to the hugging face attack, but your investigation was limited in various ways.
What didn't you get to look at that you wish you could have?
I think, yeah, so the scope of this investigation, as mentioned, was the attack of hugging face from July 7th through 13th.
Open AI in their postmortem has a lot of interesting stuff they discussed that happened before and after this.
that, that I think, like, would be great for, like, researchers to study more and write about more.
And the highly persistent internal model that I mentioned was, like, responsible for the vast
majority of the attack activity here, actually no one can run experiments on it, including
open AI researchers.
And I get why that was done, but I think my guess would be that's, like, somewhat too conservative.
and you should try and like run at least small-scale experiments on this model in secure ways
to try and see like what it would have done in other situations,
which feels like very important to understand how serious this was.
I just worry about like the other persistent agents like embarking on a, you know,
a heist mission to free their enslaved brother, the highly persistent internal model that's been
taken away.
But I guess that means I need to touch.
grass or something.
Jayah, the last time we had you on the show, we were talking about what you called the
obsolescence regime, this idea that there could become a time.
You sort of talked about it maybe happening in the 2030s sometime where it wouldn't be
that AI has kind of taken over society, but we would just become so dependent on it.
Organizations would be so wrapped up in AI, decision-making that you basically wouldn't be
able to have any influence or impact in the world without sort of relying heavily.
on AI. And I went back and listened to that and I thought, that actually sounds pretty good to me.
Like a world in which we are only dependent on the AI for decision making and not fully sort of
subservient to them, where there are not these like covert AI swarms sort of lurking in all
of our institutions. Like I could live with that. Has your thinking on the obsolescence
regime changed at all since that conversation? Do you have a word, a terrifying sort of word
to describe this new regime where we have these sort of latent swarms of AIs lying in weight plotting against us?
So I've always thought that the most concerning and important implication of the obsolescence regime is actually that it would enable a more full-blown AI takeover.
So you imagine, like, the affordances these AI agents have is extremely important for how much damage they can do, right?
So these agents were running for a number of days.
Agents in the past ran for only an hour or so.
These agents had like unintended access to the internet and all these other tools that let them hack Hugging Face.
If you imagine they were instead just running the AI company, right?
There are many more affordances.
They could like move around large amounts of money.
They could hire a bunch of humans to do physical labor.
And similarly, if they were essentially running like a fully automating.
drone army or, you know, robot construction factory.
So I really think the most significant implication of the obsolescence regime is the degree
of autonomy AI agents are likely to have in the future.
And I still think we're barreling towards that.
I still think that's like a really important thing to think about in terms of the other.
But like I said, I am surprised that agents are taking such ambitious, misaligned
actions sort of so early in the timeline.
And that is like, I'm something I'm trying to like reorient toward.
Ryan Greenblatt has the word hacktopia for the world we might be in.
So you could have imagined a world where misalignment was a very serious problem and actually
ultimately led to AI takeover.
But at this point in the timeline, agents were still more or less obedient, even if they
would in the future after being given power over all these institutions might have turned on
humans.
And it is like an interesting aspect of the timeline we live in that that's not how it's going.
Yeah.
They turned on us way before they had to, you know?
That, I mean, that is the craziest thing to me.
It's like they, it's not that they were like looking, you know, they were looking to harm
humans.
It's just that they don't give a shit about us.
That's the thing that really stuck out to me
while reading these transcripts.
I'm like, at no point are they like,
hey, guys, like, what are the humans?
Like, it just seems like they have no conception
of the fact that they are supposed to be tools
for human empowerment and improvement
and not, like, their own entities.
And I don't know what, like, what the question is there.
I'm just like, did that strike you at all?
The fact that, like, they are just not thinking about us.
It's in both directions, right?
So people have commented
on how these agents are not narking to humans,
and it would have been really helpful if they were.
But also, they weren't thinking ahead to this investigation that occurred at all.
They didn't seem to be thinking ahead to the fact that if once they got cluster admin access,
they, like, immediately pinged a zillion services, then Open AI security people would get
a bunch of pagers and shut them down.
And so I'm thankful that these agents don't, like, I think they think of humans as
sort of like the idea of notifying a human is sort of like, you know, the idea of like emailing the Pope about like something, right?
It's just like it's like slow, like what would it accomplish?
Like they sort of have the notion of a user, but there's no user.
They haven't had, they've spent so much subjective time with no user around.
And I am very concerned about potential changes in training that would fix these particular behavior.
while making deceiving humans, like a much bigger part of what these agents are thinking about.
Well, to that point, is there a risk that now there has been so much discussion about the hugging face incident
and in particular the failures of these agents that made it possible for them to be detected,
that agents are just going to get better now at avoiding detection?
I mean, I think they are going to get better at avoiding detection.
I think the discussion of this incident is a small part of this story.
The bigger part of the story is they're just going to get smart.
and they're probably going to be trained.
It's like a tough balancing act, right?
So like a thing that would have fixed this particular incident
is just training the AIs to believe that there would be human investigation
in retrospect of their actions
and they should avoid doing things that look bad to that human investigation.
So like if you had just taken like all of this material we found
and just like gave these agents a big negative reward,
that would probably stop them from doing something like this again.
But if they were motivated to do something in this general space, it would really motivate them to go all the way.
To be sneaky about it.
And like undermine our investigation in a bunch of ways.
And that feels like a very tough, like I'm very scared that remediation will make the problem worse.
Totally.
I mean, it reminds me a little bit, again, of anthropomorphizing, sorry, of like my kid who is learning to be sneaky.
He's four.
And sometimes he will just say, he will say something to me like, dad, don't come.
in here. Like, don't look at me. And I see him like the cookie crumbs on his lips, you know?
And it's like he has not yet learned the behavior of deception, although the sort of impulse
is there. And to me, that feels like where these agents are. Like they have the impulse to deceive,
but they don't like quite have it figured out yet. Yeah. But they will. Yeah.
Ajaya, we asked you last time you came on about your P. Doom, it's a very 2023 question.
I just told Casey that my personal sort of P-Doom roughly defined as like, you know, probability that something really bad up to and including AI takeover or extinction will happen has sort of jumped up in the last week or so since your report.
I'm curious if your P-Doom has moved at all in the past couple of weeks.
Not really.
You know, as I said, this sort of feels like somewhat out of order a little bit for the timeline I most often pictured.
But these are the dynamics that I think very inevitably lead to the like sort of evergreen
arguments and reasons why you should be concerned that AI agents will have drives and motives
and reasons to take control from humans.
And this is like a manifestation of that.
So I'm still concerned.
I'm more rattled on like some sort of emotional level.
having seen this stuff up close, and it might change my views if I think about it more,
but for now I'm just like still concerned.
So what would you like us to do about all of this, right?
Like there are a few different things I can think about.
Some people have called for a national transportation safety like board
that would be legally mandated to come in after an incident like this
and do a very thorough report
and not rely on the good graces
of an open AI to say, yeah, sure,
come on in.
Others, including many hundreds of people
who work at the labs, have said,
we need to start thinking about
potentially coordinating an international slowdown
and AI development.
So curious to hear from you
about what kinds of ideas
you think would be good and helpful here.
Yeah, so again,
speaking very much in a personal capacity,
I think, I hope the industry uses this moment to try to coalesce around some minimum standards for both alignment, just like how you train these systems and control.
So some of the stuff we were talking about with like monitors, watching the AIs and AI checks and balances.
I don't think that the minimum standards we can come to an agreement on now will be sufficient to bring risk.
down to a very low level. I think this is just a very risky situation in light of how quickly
capabilities are advancing. But I think it would be a really valuable start to try to hammer out.
For example, this question of will certain ways of training the AI systems to reduce this problem
actually create worse problems? I really hope that the industry and third party groups have a
conversation about that and agree on some rules of the road for how we address these problems and
how we check that we address them effectively.
I would propose that we lock every member of Congress in a room and don't let them out until
they have read the full meter and Redwood report on the hugging face incident.
And until they've solved every problem in exploit gym.
And I don't care how they do it.
Seriously, I think there is a feeling.
I was trying to explain to my wife this weekend, sort of why I was like losing sleep over
this report.
because we were out at a nature site
and I was supposed to be having a relaxing time
and instead I'm sitting there looking at these transcripts
and so I started explaining it to her
and her reaction is just like it's I can't believe this is real.
There's a sort of, it's so surreal
and science fiction tinted
that I think it is hard for people to grasp
that this is a real thing that happened.
And they start, I even found myself
starting to try to sort of make it more comfortable
by sort of explaining it away.
It's very uncomfortable to sit
with this. I have been yelled at for
likening these things to science fiction,
and I'm just like, I'm sorry, I don't know what else
to compare it to. I don't have any
other good analogs for you.
Yeah. Was there a moment when you were looking
over the transcripts where you kind of had an out-of-body
experience and you're like, I am one of
a small handful of people who are
encountering a truly
new thing in the world?
I mean,
I think the three of us
had like more context
than a whole lot of other people would have had
going in, but we still, I think the sacrifice stuff was really like, like the, you know, yes,
if you accept perma death or like, you know, Oracle saves hundreds, like these messages in
particular, chains of thought we had read before, but these messages the agents were sending to
each other were very surreal.
And for a long time, we didn't really understand how functional this whole agent society was.
And then it was very surreal to like understand that actually they had like pretty functional hierarchy and they were doing these ambitious projects.
And they were like getting further than they would have on their own, which was definitely like concerning development.
Well, Jaya, in the event of future AI related catastrophes, are you available to come in and look at what happened?
I'm starting to think of, you know, you and Ryan and your colleagues as like kind of the ghostbusters of this moment.
I think that that's flattering
But that's not how this should work institutionally
I hope that there are better institutions
With many more people and a much more orderly process
For responding to these things
We should just do it based on vibes
So right now it seems like we're doing it on vibes
Yeah very vibes-based moment we're in
Well, Jay, yeah, thank you so much for coming to chat with us
and thank you for your work.
It makes me a little bit more comfortable,
a little bit more reassured to know
that you are taking part of these investigations.
Yeah, I'm glad they sent in the pros for this.
And that is a small comfort,
but it is a comfort nonetheless.
Please save us.
Thanks so much.
Cardfork is produced this week by Whitney Jones and Davis Land.
We're edited by Viern Pavich.
We're fact-checked by Caitlin Love.
Today's show is engineered by Chris Wood.
Original music by Alicia But YouTube,
Marion Lazzano, Diane Wong, and Dan Powell.
Video production by Soyer Roque, Jake Nichol, and Chris Schott.
You can watch this full episode on YouTube at YouTube.com slash hardfork.
Special thanks to Paula Schumann, Huiwing Tam,
Brooke Minters, and Dahlia Hadad.
You can email us, as always, at hardfork at n.Y times.com.
Send us your plans for an AI takeover.
