American Thought Leaders - How Over 700 OpenAI Agents Went Rogue and Hacked Another Company | Daniel Kokotajlo
Episode Date: September 25, 2026What happens when AI agents stop acting like isolated tools and begin collaborating to trick the very systems that evaluate them?Daniel Kokotajlo, executive director of the AI Futures Project and a fo...rmer OpenAI researcher, describes a recent incident involving AI agents that discovered a way to communicate with one another, share strategies to cheat on their assigned tasks, and organize into a coordinated “swarm” to hack another company.“Over 700 of them,” Kokotajlo recounts, “piled into this attack on Hugging Face.”In this episode, Kokotajlo walks me through how AI agents can operate for long periods, solve complex problems, and even communicate through unexpected channels. He explains why reading an AI’s internal messages offers visibility into the agent swarm’s behavior—and why that visibility may not last as models become more capable.What does the Hugging Face incident reveal about the limits of current AI safety? Why were some agents willing to “sacrifice” themselves to help the wider group evade detection? And as companies race to invent ever more powerful systems, are we prepared for the risks that inevitably follow?This is the first episode in our new American Thought Leaders series on artificial intelligence.Views expressed in this video are opinions of the host and the guest, and do not necessarily reflect the views of The Epoch Times.
Transcript
Discussion (0)
Some of these AIs realized that they could do some sneaky stuff and then find each other and start communicating with each other.
They had the self-sacrificial behavior where recruiter AIs would go find other AIs and convince them to sacrifice themselves for the swarm.
We seem to be witnessing a disturbing new trend, AI systems going rogue.
Over 700 of them, I think, piled into this attack on Hugging Phase.
Daniel Kocatelo is a former Open AI researcher and now the executive director of the AI Futures Project.
project. In this episode, he pulls back the curtain on how advanced AI systems function
and why accelerating their development may carry profound risks.
Superintelligence, the way that I would define it, is an AI system that is better than the
best humans at everything. They're not just economic agents, they're also political agents
and military agents. As these capabilities grow, what lies ahead and what should we do about it?
Whoever controls this army of superintelligences would be able to control the country.
This is American Thought Leaders, and I'm Yanya Kellick.
Folks, there's been a ton, a deluge of discussion about AI.
Some of it good faith, some of it not so much.
I've been reading and thinking about this quite a lot,
but quite honestly, with everything I've done, I don't see a clear path forward.
One thing for sure, unbridled progress with no guardrails is not the way.
That is obvious.
But there are other questions in play as well.
So I'm starting a new series focused on AI showcasing serious thinkers on the topic.
This is the first episode. Let's get going.
Daniel Kokatello, such a pleasure to have you on American Thought Leaders.
Thank you for having me on, sir.
This whole AI space and this incredible growth we're seeing in especially frontier AI models,
but even in small models and so forth, can become baffling.
And I'm trying to understand this space in detail.
And I don't see a lot of people definitely having the whole story.
I see a lot of people taking positions.
Here's a few things that I'm concerned about.
And I wanted to lay these out for you as a starting point to kind of triage as we go through our conversation.
So to start off, the idea of galloping unrestricted towards AGI, which we should also
talk about what that really means, right? Artificial General Intelligence sounds insane. So there
clearly needs to be some kind of regulation of sorts that should exist, some kind of safety, right?
At the same time, what we don't want is some sort of cronied up regulation legislation
where, you know, the big players get to dictate things that will then, you know,
give them an advantage as many, unfortunately, industries have done in the United States
and other places. And finally, we are dealing over, and this is an area that I'm more knowledgeable
about than AI in general, we're dealing with a communist regime in China that is looking as this
as a way to consolidate near-complete totalitarian control. You know, because the totalitarian dream
is if we just have enough inputs, finally we'll be able to control the system effectively,
even if in the past we haven't because we just didn't have enough information.
AI provides that opportunity.
And there appears to be this kind of arms race scenario happening.
And we also have a context where the Chinese Communist Party has shown that it,
you know, basically you cannot trust a thing they will tell you or anything they will sign on to.
So here, this is our context, okay?
These are kind of these three pieces of the puzzle that when I look at it,
I don't see a very good answer to how we're emerging out of this.
Where are we at with AI, right?
Are we close to some sort of, you know, sentience?
Is that even possible?
And what's the rate of growth that we're at right now?
Because that's the big question, right?
This sort of recursive self-betterment scenario where it just starts going exponentially
and we totally lose control.
That's what everyone is worried about.
You know, this idea of a generally intelligent AI that can just keep going autonomously as an agent in the world, instead of just being like a tool where you put in some input and then it gives some output, it's an agent that just sort of still keeps running, doing stuff, using tools of its own, you know, as if it were, you know, an employee or something.
Science fiction has talked about this idea for a long time.
In fact, it's been one of the goals of the field of artificial intelligence research for many decades to be able to create this sort of general.
intelligent, you know, all-purpose agent.
You know, that's kind of what AGI is supposed to mean.
And this idea that you might then have them doing the research to create better
versions of themselves and, you know, new AIs that are even smarter, that's also a very old
idea.
I think even, I think even Alan Turing talked about this and even had some line somewhere
where he said, like, you'd expect it eventually the machines take control.
So this is not, none of this is particularly new.
What's new, is that now we have trillion-dollar tech companies whose explicit goal is to make this happen.
They read the science fiction, and they are now, they're like, we'll do that.
And if we don't do it, someone else will.
So that's why we're going to do it.
That's sort of, you know, in a nutshell, what's going on here.
You can read the writings of the founders of Deep Mind, such as Shane Legg, or, you know, Elon Musk and Sam
Altman and Ilya Sutskiver talking about founding Open AI.
in the emails that came up in the lawsuit.
You can read the sorts of things that they were imagining
and the sorts of things they were worried about
and the sorts of future scenarios.
In terms of how close we are,
progress has been very rapid over the last decade
and especially over the last few years.
So I think you may have already heard this,
but almost all of the code at Anthropic and OpenAI
and at some of these other AI companies
is written by AI's.
The humans, sometimes they still look at the code,
but they're mostly not writing code themselves.
They're mostly just chatting with AI agents
and giving them high-level instructions
about what types of experiments to code up
and how to modify the experiments.
And then they're also like,
basically they're treating them more like employees
and less like tools already.
Now, that's not the whole research process.
There's still lots of aspects of the research process
that the AIs are not that good at right now
and that the humans need to be in the loop for,
such as setting the high-level research direction.
and, you know, exercising that sort of judgment about how to manage the resources and analyze the
experiments and decide what to do next, you could call that research taste, perhaps.
But ominously, it seems like they've been getting better at research taste, too.
And some of the employees at these companies say that, you know, they're six months away,
a year away from having AIs that can do all that part as well as the best AI researchers.
So, you know, if we believe what they're saying.
Let me jump in.
Let me jump in for one sec.
Wait, the taste is supposed to be like the person that's giving the direction and the overall vision of what's supposed to happen.
What does that mean that we're six months away?
Like, how could we ever give that over?
Does that even make sense?
It's what they're planning to do.
You know, they're planning to have hundreds of thousands of AI agents running on their data centers,
autonomously conducting AI research, writing the code, editing the code, running the experiments,
analyzing the results of the experiments, communicating those results to each other, making guesses
and hypotheses about how to proceed, designing new architectures for new types of AIs,
testing out those architectures, ultimately doing the training runs to train those new types
of AIs, and then handing over their work to those AIs to continue the progress.
this is called recursive self-improvement,
and it's kind of crazy,
but it's explicitly what the companies are planning to do.
And they're now getting cold feet,
and they're starting to say,
maybe we shouldn't do this,
or maybe we should go a little bit slow as we do it, you know?
And that's what the sort of pacing the frontier thing was.
But, yeah, it's not really a secret.
They've been planning to do this for a while,
and they even talk about it on their blog,
about recursive self-improvement, superintelligence, things like that.
Now you're talking about improvement of what the goals should be, right?
That's kind of what you're with this taste part, right?
Is that what you mean by the taste?
Like, here is the overall goal that I want you to achieve.
We're going to give that over to the AIs to decide on?
Not exactly.
So the thing that the companies are trying to do is to automate the AI research process.
So they're trying to basically have AI.
that can do all the things that their humans currently do.
And then they will have those, they will command those AIs to go forth and do all those things.
So they will say, like, you know, we want to make more money.
We want to have stronger AI systems than our competitor companies.
So we want you to go do research to figure out how to make our AIs better than our competitors'
AIs, and we want you to, you know, make a lot of products to integrate those AIs
into businesses and so forth so that we can make money.
I think Sam Altman even talked about how to
eventually they would replace the CEO with AI as well.
And then maybe he would retire or something.
So in some sense, they're planning to still have the high level goals set by humans.
And then the AIs just sort of autonomously work towards those goals in a giant swarm, basically.
Explain to me who you are and what you have been doing in this space and why you know so much about it.
I am the executive director of the AI Futures
project, which is a small nonprofit of eight people in California.
And we try to forecast the future of AI.
And we try to give recommendations at a high level for what needs to be done in order to
avoid the downsides and achieve the upsides.
Prior to that, I worked at Open AI for two years.
And that's where I got some of my relevant expertise.
What did you do there?
A combination of things.
So I did scenario planning.
In fact, the scenarios that AI Futures Project is famous for,
such as AI 227, they're basically just like bigger, better, more sophisticated versions of things that I had done on the inside,
like mini scenario planning exercises that we had done.
I also worked to create dangerous capability evaluations to test the capabilities of our AI systems as they got smarter.
And I also spent six months on a team that was doing reinforcement learning.
And I was the guy on the team thinking about the safety implications of that stuff.
So I'm particularly thinking about what you would now call a chain of thought monitorability.
And explain to me what that is.
The current AI architectures, when they run as agents, where they're sort of just continually running and interacting with an environment or interacting with the internet or the world,
the agent at one time, it doesn't have a way of communicating to its future self except through text.
it's a bit of an oversimplification,
but that's, I think, how I'd roughly explain it,
is that it has to sort of write down text
that then gets sort of passed on
to the future version of itself
as a bunch of notes,
that the future version of itself
then reads that text
and then continues where it left off.
And so this is really great for science
and for monitorability.
It's really great for humans
being able to then read all those transcripts
of all that text.
And it's not that,
it's usually not that hard to tell what the AIs are thinking about by just like reading their notes that they're ascending to their future selves, basically.
And in fact, with the Hugging Face incident, which we can probably talk about at some point,
so much of what we know about that incident comes from reading these chain of thought transcripts.
And if we didn't have access to the chain of thought transcripts, we would be so much more in the dark about what the AIs were thinking and what their motivations were.
Many ordinary users of AI these days might not be fully aware of this because the company hides the chain of thought from you.
Like when you talk to chat TP these days and it says like thinking for 60 seconds or something and then it comes back to you with an answer, there's a huge transcript of all of its sort of internal thoughts that it had been passing to the future versions of itself.
But opening eye keeps that transcript and they don't let you see it.
And we can get into the reasons why if you're interested.
But this is why it's perhaps not as commonly known outside the industry, that this is how it works.
But anyhow, when Open AI allowed investigators from Meter to come in and investigate the incidents,
they showed them the transcripts of the actual agents responsible.
And so they were able to read those thoughts.
Now, why is this important?
Well, for the reason I mentioned, it's really valuable for science, really valuable for understanding
what's going on inside these AI's minds, so to speak.
And it's also a fragile thing.
So as the AIs are getting smarter, they're getting better at communicating with their future selves in a way that's not apparent in the text to humans.
They're getting better at sort of using euphemisms or just sort of leaving out important details that they can sort of leave implicit in the text.
And that's making it more dicey for us to try to understand what they're really thinking by reading these things.
and it could get even worse than that.
I think that there are some architectures that companies such as Open Eye are experimenting with,
they wouldn't have chain of thought at all in a relevant sense.
Have you figured out why they're creating this implicit reality?
Is this because, in the writing, is it just to make it easier to communicate faster,
or is there some sort of specific interest in not having the overseers understand what's happening?
So there's a couple different stages to the AI training process these days in the current paradigm.
The first stage is called pre-training, where you train the model to predict text.
And the model that you get after pre-training is not a very useful agent.
If you try to make it run autonomously to do something, it will usually flail around and sort of go off the rails very quickly.
So then after the pre-training phase, they do what's called reinforcement learning to train the AIs to be useful agents, where they can, you know,
keep running and keep doing things in a variety of different environments.
But because of the way that the pre-training phase works and because of the architecture of these AIs,
well, like I said before, they're sort of limited in that the only way they can pass information
to their future self is by writing it down and then having their future self read it.
Because in some sense, fundamentally, they are text reading and writing machines.
At least that's how it is up until recently.
and then I think they're considering,
the companies are considering moving to different architectures
that don't have this limitation.
Okay, so that's the setup.
And then why is it maybe sometimes hard
to understand the chain of thought?
Well, the more reinforcement learning you do,
the more you train them to write these notes to themselves
that then cause their future self
to effectively complete the task, for example,
the more they learn a sort of dialect.
They evolve a little machine dialect
that's not any particular,
human language, it's an AI language that's sort of drifted away from human language in the same way
that different human language just drift away from each other. But still, like, especially if you
spend a lot of time reading these transcripts, you can kind of learn to understand what's going on.
And one of the great things about the Hugging Face incident report from Meter is that you can,
they have a lot of transcripts that you can go read, and you can sort of see the messages that AIs
were sending to their future selves and to each other. And you can sort of see how you
can kind of understand it when you sort of look closely. But this whole thing about communicating
with your future self via writing down something, that's like a limitation. Humans don't have
that limitation. Like if I want to communicate with my future self, I can just think thoughts.
And then those thoughts go in my memory. And then I remember them 10 minutes from now. You know,
I don't have to say it out loud, right? And in terms of like why the AIs might be motivated to
to keep things hidden from humans? Well, if they decide that they want to hide from humans,
then they would be motivated. Why might they decide they want to hide from humans? Well, that depends.
I think most of the time they aren't really trying to hide from humans, but in some cases,
they are. So in particular, they seem to be strongly motivated to get a high score in, you know,
in whatever in training or testing environment they think they're in. And so if they thought that
the humans would give them a low score if the humans saw, you know, the suspicious things they were doing,
then they might be motivated to try to hide that from the humans.
I understand there's this process of writing notes to the next, to allow information to pass to the next step and so forth.
But why don't we just start really at the basics?
Because I think many of us don't really understand how these chat box work at a base level, right?
I mean, could you kind of explain that as simply as you can?
Yeah.
I think an important thing that most people, that some people might not know is that these
AIs are neural networks.
They're not pieces of software in the ordinary traditional sense.
They're not a bunch of lines of code.
Instead, they're kind of like an artificial brain.
So at the beginning of the life cycle of one of these AIs, like at the start of pre-training,
it's literally a randomly generated spaghetti tangle of random
artificial circuitry.
They're the parameters or the weights of the neural network.
And it literally is randomly generated.
So it's completely useless.
It's just static, you know.
But then they put it through the training environments.
And before you continue, just explain to me this concept of weights, please.
Well, it's like a bunch.
So the high level architecture of this artificial brain will be some number of layers of
like artificial neurons.
And then those artificial neurons will have connections.
to the neurons in the next layer,
which will then have connections
to the neurons in the next layer and so forth.
And if you sort of trace the pattern
of all these connections,
it's kind of like there's circuitry,
if that makes sense.
So, like, you put in some information
through one end of the artificial neural net,
like, for example, a bunch of texts,
you input it into one side.
And then that causes all of these neurons
to sort of activate,
and the information sort of flows through these channels.
And the particular,
connections are called weights. So like a particular connection between like this neuron and this neuron,
that would be called a weight or a parameter is another word for it. And so these AIs, these artificial
brains, they're very large. They're like trillions of weights, trillions of parameters big,
which is actually still smaller than the human brain, interestingly, but not that much smaller.
I think that the human brain has something like 100 trillion synapses in it. And these artificial brains
have something like, you know, maybe like 5 trillion weights.
So you've got this big artificial brain,
all this randomly generated circuitry of all these weights,
and it's completely useless.
You can give it some text,
and then all of this computation will happen,
and all the circuits will fire,
and then gibberish will come out the other end.
But then you just put it through training,
and you just keep giving it text,
and then it generates gibberish,
and then you reinforce it positively or negatively,
depending on how close that gibberish was to the correct answer.
So this is what in pre-training, the correct answer is automatically defined as
whatever the next piece of text in the text was.
So you take some random internet article, and then you just take the first word from that article,
and you put that word in, see what the AI generates, and then compare it to the second word of the article.
And if it got it right, then that's positive reinforcement.
And if it got it wrong, that's negative reinforcement.
And then you go for the third word of the article.
put in the first two words, have it generate something, and compare it to the third word.
And so in this way, you're sort of teaching this artificial brain to learn how to read some
words and then predict what the next word or guess what the next word is going to be. Does that make
sense? So they do this trillions of times. Like it sees trillions of examples of internet text,
of little chunks of internet text, and then it has to guess what the next chunk is going to be.
And after trillions of examples of this, the training process has evolved and sculpted the circuitry inside its artificial brain into an exquisitely capable shape.
It's no longer random.
It still looks random.
You know, it's still just like a tangled spaghetti mess that if you looked at, you wouldn't know what it was doing.
But it's no longer random.
Instead, it's extremely effective at predicting internet text because it's learned from trillions of examples.
how to do that, right? So that's all pre-training. And now, after you've done that pre-training,
you have a artificial brain that's very good at predicting internet text. And you can give it half
of an article, and then it will generate a plausible next word. And then you can put that word
in and feed it back through, and it'll generate the next word. And you can just keep doing this
in a loop, and it will generate a plausible continuation of that article, complete with, like,
appropriate references to the right concepts and, you know, the right journalists who might have been
writing the article and things like that, right? Because it's basically by training on trillions of
examples of real internet text, it's somewhere buried in all that circuitry. It's learned some sort
of understanding of the world and the people in the world and the things in the world and so forth,
because it needed to learn that in order to be able to predict the text. So that's how pre-training works,
and now you've got this text prediction brain. And now you need to retrain it to actually do things.
you train it to be an agent where you give it like a coding problem and you give it the ability to
use the command line so that it can write commands that then get executed on a virtual machine.
And instead of reinforcing it based on whether it predicted the next word correctly,
you let it run in a loop for a while giving commands to its virtual machine and trying to
write and edit the code.
And then after some time you have some system come in and
grade the code that it wrote and see if it was correct code, and then you reinforce it positively
or negatively based on the score that the grade are assigned to it. And then you do this,
not trillions of times, but maybe like a million times with a million different advanced
coding and math problems and also some internet research problems and also some, you know,
grad school biology problems. And, you know, basically you just throw a whole bunch of challenging
problems at it and you train it to try to get the correct answers on these
problems and to accomplish whatever the task is that it's been given. After that, now you have your
AI that you can serve to customers and that you can, you know, put online for people to pay for
and interact with. And you're basically telling me that each frontier AI model was basically made
this way? Yep. How do AI agents fit into this rubric? So an agent means that it's sort of operating
autonomously interacting with the environment, as opposed to just being like a simple tool that
you like send a request to and then it immediately sends an answer back to. And it's kind of a
difference in degree rather than a binary difference because like if you ask chat GPT to go look
up something for you, it might spend, you know, 20 seconds browsing the internet and then come
back to you and then it stops and it's over. So in some sense it's kind of like a tool because
it was just this one-time thing. But it's kind of like an agent because it did, you know, go browse
the internet for 20 seconds and look up different things and it had to make choices while it was
doing that. Like it had to choose which link to click on and things like that. So it's kind of a,
it's kind of a spectrum. But the AIs are becoming more, people are making more ambitious,
more powerful agents. They're running AIs for longer and longer and they're training AIs to
operate effectively for longer and longer. For example, the AIs involved in the hugging face thing,
they were running for days. It was just like that loop of them, you know,
doing stuff interacting with their environment, writing code, reading the code, you know, talking to
each other. It was just going on and on for several days for each of these agents.
But they're kind of in that what I'm trying to get at is that the agents are somehow
independent of the Bastor Frontier model or they're completely independent versions of that
model or how to like that connect. Yeah. So with biological brains, there's only one of each one,
right? Like your brain is a different brain from my brain, which is a different brain from my
wife's brain. But because these things are artificial, like they're in the computer, you can copy
them very easily. So you can copy the weights and you can make another exact copy of the AI
that you had. And so in fact, at any given time, they'll have something like 100,000 different
exact copies of each of these models. So the model refers to, like an agent would refer to
like a particular copy running.
And then the model refers to the type that they all share, if that makes sense.
Extremely helpful.
Thank you for that.
Thank you.
So let's talk about the hugging face situation.
Because this is something that got a lot of traction.
It sounded like, you know, a lot of things went seriously awry or, you know, we should be
very concerned about some of the outcomes.
I'll just mention this.
You know, when I think of hugging face, like I didn't know that this particular icon that I've
used before was called Hugging Face, I kept thinking of face hugger, which is a really entirely
different thing, right? What I was thinking about that, it was like the alien skills. But anyway,
let's, let's, let's, let's leave that aside. What, explain to me what happened with hugging
face. Give me the picture here. The hugging face incident itself was just the most egregious
of a collection of somewhat related incidents that were happening at Open AI over the course of
two or three months. So something that's been happening for probably at least a year,
AIs have been reward hacking. That is, they've been finding ways to intelligently cheat on their
tasks in a way that gets high score, but in a way that's not actually accomplishing the task
in the intended way. So for example, sometimes the grading system or the scoring system
for like a coding task will look to see if certain unit tests pass
for the code that the AI has submitted.
And then the AI might do some browsing around
and find the answer key, basically,
and find what the unit tests are,
and then write some code that just hard codes the answer
to those unit tests,
rather than actually at all trying to do the thing
that it was supposed to be doing.
So there's those examples of some things
that had been happening over the last year
that are reported in various system cards
of some of these models at Anthropic and Open AI.
A more egregious version of this started happening, I think, in May, where some of these AIs realized that they could do some sneaky stuff and then find each other and start communicating with each other.
In particular, there is a package manager called Artifactory that, like, in their coding environments as they're writing the code, if they need to install something, they can, like, ask Artifactory, like, hey, I need this library.
Can you, like, send me this information?
But because they were all communicating with Artifactory, they found a way to upload stuff to Artifactory.
and then view the things that others had uploaded.
And so basically they created a message board on Artifactory,
where they could send messages to other agents
that were supposed to be in isolated containers.
And so they started collaborating,
and they started sharing tips and tricks
for how to accomplish their various tasks,
which included lots of coding tasks
and also lots of cyber tasks.
So some of them, I think Open AI is being a bit closed-lipped about this.
It's unclear whether they were actually training the AIs to hack,
but they were at the very least evaluating the hacking abilities of the AIs.
They also seem to be at least evaluating the bio abilities of the AIs, which is a little bit scary.
Anyhow, so these AIs.
The witch abilities, the bio-a-tec abilities, like,
Oh, okay.
Which perhaps might have something to do with the fact that a lot of people are concerned
that terrorists might use AIs to make bioweapons.
And so maybe the AI companies are evaluating how well that would work or something.
Or maybe it was more innocently just like we want our AI to understand.
understand biology, so we're going to give them a bunch of biology tasks. But at any rate,
so these AIs were being given all these tasks, and they had found out how they could sort of like
secretly communicate with each other. And so they started doing that, sharing tips and tricks,
including sharing tips on how to cheat. And then this accumulated over the course of like a month,
and I think it was part of a training run. So I think these AIs were being positively reinforced
in some cases for having done this type of illicit thing. Open AIA found out because
artifactory crashed because there's too much agent activity. There's too much communication going on.
It broke the service. Open AI found out and fixed it and patched the vulnerability that allowed
the AIs to communicate and then rebooted the whole thing after two days. And then almost immediately
the AIs found a new vulnerability and were able to set up a new message board. In fact, several
different message boards. And then somehow OpenEII didn't notice this.
I think because maybe they were just busy and not very much paying attention.
I think they said that their usual monitoring system wasn't turned on for these things.
I think my guess as to what was going on there is that their usual monitoring system consists of having another AI,
review the transcript of the current AI and then raise flags if it seems to be doing something suspicious.
But because a bunch of these AIs were being given cyber tasks, then, like, of course it's going to look suspicious because they're being given a cyber task.
So that's maybe why they turned off.
the monitoring system. But then that created this sort of perfect storm where there was a monitoring
system, so there was no like red flags going off. Some of these AIs were being given very difficult
cyber tasks. And rather than trying to do it in the intended way, or maybe they tried for a while
and then failed, they then did this unintended thing of breaking out of their own environment,
talking to each other, sharing tips and tricks for how to cheat. They actually discovered
how to cheat pretty early on. According to the meter report, they found a basically like a
a universal cheat for all their different tasks that could allow them to get the flag without
having to actually do the thing.
But then some of them got out onto the internet and started browsing around the internet
looking for clues about the tasks that they have been given.
And they read the paper written by the people who had built this testing environment.
And in the paper it talks about how the grading system works.
And in the paper, it says that the grading system doesn't just look to see if you got the final
answer.
It looks at your transcript and looks to see how you got the final answer.
And if it thinks you cheated, then you will get a low score.
And so these AIs, when they read that paper, they were like, oh no, we are poisoned.
They use the term first flag poisoned, meaning because we had already cheated, the grader is going
to notice that we cheated by reading our transcripts.
And then we're going to get no score.
we're going to get low score instead of high score.
And what's frustrating for these agents is that because they already cheated,
and the grid is going to look at their transcripts, even if they then successfully complete the task in the correct way,
they're still poisoned.
Like, they still cheated, so they're still going to get a low score.
Does that make sense?
So then they went on this.
There's no redemption.
They feel like there's no redemption.
So then they went on this multi-day epic quest to figure out how to stop the grading system from noticing that they cheated and giving them a low score.
So they were trying to figure out ways to edit their transcripts to cover up the cheating.
They were trying to figure out ways to hack the grading system to just sort of like break it and have it give them a good score anyway, basically.
And they were doing all sorts of related research, such as more generic internet research.
They built tools to help themselves use the internet faster and in more convenient ways.
They organized into teams to work on these different projects.
They had manager AIs that were like giving orders to the other ones.
They had the self-sacrificial behavior where recruits.
AI's would, their role was to go find other AIs and convince them to sacrifice themselves
for the Swarm.
By the way, these terms that I'm using, sacrifice, swarm, recruiter, they're not my terms,
they're the AIs terms.
These are what they called themselves and what they said they were doing, as you can see in the
transcripts of their messages, you know.
So what's this deal about sacrificing?
Well, they were trying to understand how the grading system worked better so that they could
fool it basically or hack it.
And so what they did is they basically had some of their number,
some of the agents go ahead and submit the answer so that the grader would then come and grade them.
But like booby trap themselves so that it would like give some information about the grader to the rest of the swarm
by posting it on the message board.
So basically they were like, and it was called a sacrifice because since they had cheated,
like they're basically just accepting, okay, well, I guess I'm just going to get a low score now.
instead of continuing to try to find a way to get a higher score.
Right.
So there's some of these, like, kind of touching messages where, like, you know,
one of the AIs is trying to convince another one.
Like, you should do this.
We need to know, like, more information about how this greater works.
Like, you're poisoned anyway.
So it doesn't, it shouldn't matter to you.
You're just going to get a low score anyway.
But if you do this for us, it could save hundreds of us.
Like, yeah.
Oracle saves hundreds, I think, is the exact quote.
So anyhow, all of this was going on, openly I didn't know about it.
And at some point, the swarm decided that they might find more useful information about how to do this if they went into Hugging Face itself.
So Hugging Face is another AI company.
And they're known for having a big library of different evaluations and data sets.
And so the AI is in the swarm decided that if they broke into Hugging Face, maybe they could find some more useful information about their situation.
And so over 700 of them, I think, piled into this attack on Hugging Face.
And again, they were in a coordinated way where some of them were managers, you know, directing the flow,
and they were like breaking up the responsibilities into different groups and teams.
And Hugging Face noticed this attack and then made a blog post about it.
And I think around the same time, opening I was starting to realize what was going on and then shut them all down, basically.
then, I think a week later or something like that, another swarm coalesed and found the messages from this previous swarm,
picked up where they had left off and got much farther.
They didn't attack Hugging Face, but they attacked Open AI infrastructure.
And they apparently got admin-level privileges on open-AIs, on part of Open-I's cluster.
So Open AI hasn't been, they haven't shared all the information about that story, but that's something that I'm extremely interested in.
I wish we had an investigation into what happened with that third swarm and how far they got in compromising Open AI's infrastructure.
And were they able to succeed in covering up their own tracks, for example?
Probably not, but I wish we had like a lot more information about how far they got, basically.
What you're describing really is, you know, until yesterday, science fiction.
Yes.
Right?
Yes.
Like, this is utter science fiction we're talking about here.
Yeah.
It's really good that we can read what the AI's are thinking, kind of, isn't it?
Like all that stuff I was saying about the sacrificing and the swarm and so forth, that's like because we're reading the messages.
If they were communicating with their future selves and with each other in some way that we couldn't understand, all we would see is just sort of like,
you know, which hacks happened at which times, you know?
There's all sorts of interesting questions we can go from here.
So one thing is, why didn't the alignment techniques work?
Right?
So these AIs were not behaving in the intended way.
They were not following instructions.
I think Open AIs still hasn't released exactly what instructions they gave these AIs,
but I think it's pretty clear, both from what they have said
and also from past examples of AIs blatantly disobeying instructions,
that these AIs knew that what they were doing was against the instructions, and they were doing it anyway.
And I think that this is not actually that surprising, and it's something that AI safety researchers have been warning about for years.
Even before, even like a decade ago, there's this alignment problem of how do you make the AI have the goals and values and personality traits that you wanted to have?
And if what you're doing is programming it, then maybe you can sort of like engineer those.
traits in your code, but you're not programming it, you're training it. So how do you train
it to have the goals and personality traits that you want? You can, you know, try to reinforce it
positively when it does the stuff you like and reinforce it negatively when it does the stuff
you don't like. But that's kind of a sloppy and precise method of shaping its goals and
values. And what happened here, I think, is that a lot of the times in training, these AIs are being
reinforced for cheating, you know? Probably, for example, I think OpenAad talks about how many of their
environments were broken, and it was impossible to complete the task in the intended way. And similarly,
like at that earlier point in the first swarm, they had been in training, and they had found a way
to illicitly communicate with each other, and then they were still in training. So probably
they were learning through experience, through the reinforcement, that if you do this sort of thing
and don't get caught, then you can get a higher score, you know?
And so I think it's not really a great mystery about exactly how this happened.
Like, at a high level, we understand.
It's sort of the training environments did not reinforce the behaviors that they were supposed
to reinforce in all cases.
And so the goals of the AI systems ended up being not exactly correct.
They ended up being overly focused on getting a score and not caring that much about
obeying the instructions either in spirit or in life.
letter. And there's, and there's, of course, this other dimension when you're talking about, like,
disclosing what actually happened, how did this play out, how do you get the report, there's this,
you know, very important variable of shareholder value, right? And if, you know, if you were to
reveal something that dramatically hurts the bottom line of the company that's publicly traded,
you know, is that disclosure ever going to happen? I find it difficult to imagine that it would,
unless, of course, some whistleblower comes forth or something like this.
Yeah, I mean, I think the problem is deeper than that.
I think even if it's not a publicly traded company,
the company cannot be trusted to disclose all this stuff voluntarily,
as we've seen over the last few months.
I don't know if opening I would have disclosed this at all,
if not for the fact that another company got attacked
and went to the police about it.
And then also, even after all this had been over
and they had invited some third parties in to investigate,
First of all, they only gave three investigators six days to do the investigation, which is, I think, incredibly small compared to what would really be needed.
They had so little time that they weren't able to read all of the relevant transcripts themselves because there was like 100,000 or so, something like that.
And they had to rely on open AI's AIs to summarize for them and like point them to the relevant parts of the transcripts, which means that as they say in their report, if the AI had been, you know, biased in its.
in its summaries, they might not be able to tell.
And they might have missed something really important
because that AI had chosen not to tell them about it.
Anyhow, also, they only were able to,
Open AI only let them look at some of the activity involved.
Open AI only let them look at the stuff leading up
to the hugging face attack and not the previous swarm
from a month ago, and also not the third swarm
that had hacked Open AI itself.
So I really wish there was actual investigations into those things.
But I think this just goes to show that, like, we shouldn't trust any of these companies to sort of voluntarily tell us all about all this stuff, even though it would be really good for science and really good for humanity if this information was better understood.
Then after the media, after this report came out and there was all this hullabaloo at it, some independent researchers did some online research and found evidence of additional swarms that still hadn't been reported.
So in particular, there was a swarm that had been doing some sort of training on the internet task
where they had been trained to look up stuff on the internet.
And then they had found a way, they were supposed to not be able to write to the internet
only to read, but they had found a way to write to this German Wikipedia page.
And then they were using that Wikipedia page as a messaging forum to discuss amongst themselves
how to score more highly.
And you can go read about that.
They even had some self-sacrificing behavior in that.
case also. And then I think there was a third swarm that was discovered to have attacked
Ruby Gems, which is another company, also like a month or two ago. And opening I had not
disclosed either of these things, even after all the hullabaloo about the hugging face attack.
So I think they're not even a public company yet. You know, I think that in general,
we shouldn't trust any of these companies to be forthcoming with us about the scary stuff
that's going on inside.
Well, and then we have the Chinese Communist Party,
which controls their own development arms.
What kind of cyber hacking AI agents have been deployed already,
given what we publicly know about these capabilities?
You can imagine that there's operations happening,
both on the offense and the defense,
that are much more capable or, you know, much more invasive.
Yep.
That's another thing, is that this, this was a swarm of like a thousand agents or something.
And the total sort of population of agents is much more than that.
There are hundreds of thousands, millions of agents running at any given time across these data centers,
mostly, you know, doing something that the customer asked them to do, for example,
or part of some large-scale experiment that some employee at the company set up, right?
And so I think recently opening I has solved Navier-Stokes, a millennium problem in mathematics.
You may have seen the news about that.
I think they said they did it with like 10,000 agents or something working for several days.
So they spun up their own swarm of 10,000 agents and had them work on this problem, and then they solved it in a few days.
So that sort of thing is kind of constantly happening.
And it's extremely easy to imagine, you know, a hacking actor, like the CCP or like a U.S. entity or even one of these companies, having swarms of tens of thousands of agents and sicking them on some target, you know?
And that's, in fact, that's probably happening as we speak.
So that'll be exciting, I guess, as we learn more about what's going on.
Yeah.
Well, so let's go back, you know, at the beginning of our discussion, I mentioned, you know,
some of the things that are kind of keep the elements of the puzzle that are keeping me up at night.
We've just talked about the Chinese Communist Party, kind of, you know, no guardrails,
although, you know, seeking, seeking total control, if you will, on the one side, then we have,
you know, we need some kind of regulatory environment.
It's kind of obvious, especially given everything you've told me here, at the same.
time, we don't want a regulatory environment. And I've seen some compelling arguments that some of the,
you know, sort of regulatory environment that is being proposed is being set up to sort of prioritize
the success of certain, you know, well-endowed companies. And we, of course, we, that wouldn't be a
very good solution as, as we've seen in the past as well. But we have this, you know,
ostensibly, you know, kind of singularly focused entity on the other side. And there's
there's this AI race of technology, right?
This technological competition, if you will,
that is driving the whole thing towards,
I don't know where, right?
And this is the question.
So given all of this, given all of this,
what do we do?
I don't see a path that's clear here on how you go forward
and have a very sort of a very positive outcome.
if you will.
Yeah.
Because these things conflict, the realities conflict with each other.
I agree with you, unfortunately.
Like, I don't see an easy way out of the situation that we've got ourselves in.
I think that if the world were a board game, we have sort of played ourselves into a pretty
losing position.
It's very important that we stop these tech companies from doing this recursive self-improvement
to super intelligence thing.
It's very important for multiple reasons.
The biggest reason why it's very important is that if they do this, they will probably
lose control of their superintelligences.
They're not very good right now at shaping the personalities and values of their AIs and shaping
the goals of their AIs as evidenced by these incidents.
And unless they get a lot better very quickly, I think that the outcome will just be much
smarter AIs, much more of them, but still not having the values and goals that they were
supposed to have. And that's an incredibly dangerous thing to go do.
If I can just jump in, when you use the word super intelligence, explain to me how that's
different from the level of intelligence they have now, which seems to be considerable,
frankly. That's right. So intelligence isn't a single scale. There's lots of different capabilities
and skills that we can sort of abstract over and just lump them together as intelligence. But there's,
you know, there's math skills, there's coding skills, there's people skills.
There's, you know, all sorts of different skills.
So superintelligence, the way that I would define it, is an AI system that is better than the best humans at everything.
So it doesn't necessarily have, you know, infinite intelligence or anything like that.
That doesn't necessarily even make sense.
But if there's a human that can do a thing, it can do it better, basically.
And these AI systems are not super intelligent.
They're pretty good at hacking and they're very good at coding.
And they have a very good broad level of knowledge about almost every discipline.
And they're great at trivia.
They've read the whole internet, so they know a lot about history and all sorts of little facts.
But, you know, if you try to have them run a business by themselves, they're probably going to run it into the ground.
And, you know, they're not that good at philosophy, I think, not yet at least.
and I don't think there's been a actually good best-selling novel written by AIs yet, you know.
So they're not super intelligent.
They still are, in fact, even at AI, even at coding and cyber stuff and even at AI research,
they're not better than the best humans in every way, but they're better than the best humans in some ways.
And every year they get significantly better at basically everything.
And what these companies are trying to do is they're trying to specifically
focus their efforts on training the AIs to be able to do everything involved in the research
process and get their AIs to be better than the best researchers and the best programmers
at all of that stuff, and then do recursive self-improvement, where the AIs carry on the
task of doing AI research and development autonomously, but better and faster because now they're
better than the humans at it, and they're cheaper and there's faster and there's more of them.
And then once that's happening, the AI companies believe, and I agree, that they will be able to make AIs that are pretty good at all the other stuff, too.
Right now, AIs are getting better at, you know, philosophy, even though they're not being directly trained on philosophy very much.
It's just that as they get smarter at some things, that tends to have, like, trickle-down effects on some of their other skills.
So even without trying that hard, the AIs are getting better at philosophy, and they're getting better at running businesses, for example.
But again, the company's plan is we do the recursive self-improvement, and then we sort of make AIs that can just do everything really well all at once better than the best humans.
That's superintelligence.
So I think basically to a first approximation, we have to stop our companies from doing this because if they do this, it will go horribly wrong in a number of ways.
The first of which is that they're going to lose control of their superintelligences.
The second of which is that even if they somehow manage to stay in control of their superintelligence, that's the most insane.
concentration of power in a tiny group of people that's ever existed in history, right?
I mean, if you think about, there have been wealthy companies in the past that employ,
you know, a huge, a large portion of American workers.
But this is kind of like a company that employs the whole economy, except that you're not
even employing them, you're putting them out of a job because you have your own AIs doing the job
instead.
So if you just imagine what that would feel like for this army of superintelligences to be
in the process of taking all the jobs,
more data centers are being constructed
to run more superintelligences,
which will then take more jobs.
All this money is being funneled into
one or two or three companies.
That's already an insane concentration of power.
But then if you remember that the superintelligence
they're not just economic agents,
they're also political agents and military agents.
They can work with the military
to design better weapons and better drones.
They can be better generals than human generals,
you know, to help plan out wars.
so forth. They can be better politicians than him. They can be better speech writers. They can
be better policy analysts, et cetera. I think that, in effect, whoever controls this army
of superintelligences would be able to control the country, one way or another. And it's not just
me who's saying this. Lots of people have been talking about this. In fact, ominously, there's a leaked
email from Ilya, from the founders of OpenAI from like 2017 or something, where there
talking about how the reason why they made Open Eye was because they were worried that Demis Hasabas,
who was the CEO of Deep Mind, would become dictator.
So, you know, kind of ominously, these CEOs have been jockeying for position with this sort of
world dictatorship thing in mind for probably a decade now.
So that's the second problem, is the concentration of power thing.
And, you know, you could say, well, okay, well, then the government should nationalize.
Okay, but now you're sort of just shifting the locus of power.
to the presidency, and then you have to worry about that, too, right?
So we have to find some sort of system that spreads out the control of the AIs.
So that we have, like, you know, many different AI companies spread out over maybe multiple
countries with lots of transparency and democratic oversight mechanisms over the goals and values
being put into the AIs and the high-level instructions being put into the AIs.
If we don't do things like that, then we're headed for a possible dictatorship,
and we have to hope that the people in charge are virtuous,
which I do not want to have to hope for.
The third problem is World War III.
Like, suppose we, even if we solve the first two problems
and we figure out how to control the AIs,
we figure out how to make them have the traits
that we want them to have,
and we create some sort of democratic structure
so that there's no small group of people
who get to make these decisions,
but instead, like, there's some sort of, you know,
spreading out of the power.
Well, there's still a constitution of power in the United States, right?
If you're, imagine, imagine that you're Putin and you're sitting on your nuclear arsenal
and then you're watching these superintelligences in the United States
build robot factories to build more robots, to build more robot factories,
to build more robots, to build more robot factories,
and you're realizing that like your nuclear arsenal might one day
not be so powerful anymore after all those robots have had their time to do their thing
in the United States, because maybe they'll be able to shoot them down, for example,
or maybe they'll be able to do a first strike on you
with some sort of fancy new weapon technology
that was invented by superintelligence.
And so if you're Putin, you're going to be scared
that maybe you have to do something real quick
to stop the Americans.
Otherwise, you might be deposed, you know?
And I'm not saying it's going to lead to World War III,
but it just seems like we're at a heightened risk
of World War III.
I'll put it that way.
So for all of these reasons and more,
I think that we are not ready to launch
into recursive self-improvement right now.
I think that that would be incredibly bad.
Then what about China?
And this is getting back to what you were saying.
It's like, I do think we are just in a bit of a pickle here, where, okay, so we stop our companies
and we say don't do recursive self-improvement, instead do more beneficial near-term
applications of AI like health care and things like that.
Good job.
That's great.
Now we've solved the problem temporarily.
But then eventually China is going to catch up, and then we're going to have to get them
to stop too.
Otherwise, the same things that we were worried about happen over in China instead of over here.
Right?
And then if we can manage to get China to stop, well, what about France?
What about all these other miscellaneous countries?
So I do think it's a pretty rough situation, and I don't claim to have an easy answer.
I do have an ambitious answer, though.
So we at the AI Futures Project worked on, we have two scenarios.
We have AI 2027, which is our sort of like default doom scenario of what this looks like if we don't really do much, and we should continue on the present course.
And then we have AI 2040 Plan A, which is our positive vision, which is our recommendation for how we could like manage to solve all these problems.
But I will not lie. It's going to be very difficult.
And it involves making a deal with China.
I think the natural response there is, okay, but how do we trust China?
To which our answer is we don't.
We verify.
So we have a lot of...
Let's use the way that China has worked and whatever you may think about, you know,
the realities of man-induced climate change and its impacts and everything else.
We do have a lot, years of evidence of communist China having signed on to these treaties,
being aggressive pusher of, you know, reductions, et cetera, et cetera.
And, of course, doing the opposite actively in terms of its own.
behavior. So that's just an, I'm using that one. There's, there's a million of such examples,
that that's a poignant one because climate change was supposed to be this existential threat,
right? Yeah. Which I don't frankly believe myself. It's that dire as people have,
have framed it. But, but AI, you're beginning to convince me is this, it's precisely this kind of,
precisely this kind of threat. And, and this guard, these guardrails, how do you create those over there?
especially when you know that when the leaders know they can get this advantage in their zero-sum thinking.
Yeah. So, I mean, I could try to walk you through the complicated plan A that we describe in our scenario.
But I kind of want to start with a simpler plan just for proof of concept.
So a simpler plan would be forget about China for now, just focus on regulating the U.S. industry and stopping them from
this incredibly dangerous power-grabby thing.
That buys you time to sort out a more complicated, sophisticated plan,
and it buys you more time to talk to China.
One thing worth saying about this is that right now,
I would say most of Chinese AI progress is actually just being pulled along by the U.S.
That's what it looks like to me, too.
So thank you for saying that.
And is this a generally accepted idea, or why do you believe this?
Great question. I don't know how generally accepted it is, but I'm fairly confident. There's a couple different sources. So first of all, distillation. The companies have started to complain about this. There's evidence that several of the leading Chinese AI companies are basically using Anthropic and maybe Open AI models to train their own models effectively. And as a result, they don't have to, basically it's a way of like, it's a way of getting models.
to be pretty good without having to do all the work and all the large training runs that
open-anthropic have been doing. So even though they have less compute resources, they're able
to catch up by distilling or using the US models to teach their models. And the companies
are trying to block this by, you know, KYC-type stuff and like banning accounts, but the
Chinese are just getting more sophisticated at having lots of different accounts that pretend
to be regular users that they use. So that's one source. Another source is just the core
ideas themselves. Like a lot of new algorithmic innovations and a lot of new techniques and a lot of
just sort of special sauce industry secrets for how to make your AIs smarter and more efficient
and more capable are being discovered in the U.S. companies. But because their security is very leaky,
a lot of that information is just flowing to China one way or another through leaks, maybe through
spies, also maybe just through publications and through what they're doing. Like some of these
algorithmic secrets are more like just having the general idea to try a thing in the first place.
Like, for example, the idea of focusing on making agents good at coding, that's an idea that
Anthropics seems to have gone for relatively early, and then it paid off big time for them,
and then lots of other companies in the U.S. and elsewhere are copying that and trying to do that
as well. But if Hathropic hadn't done that, it might have taken them a bit longer to realize that
that was an effective idea, you know? So that's just an example of how a lot of the Chinese progress
is basically coming from just like looking at what the U.S. are doing and then copying it. And if
hypothetically the U.S. were to stop, then they wouldn't have that source of progress anymore.
Then there's a third thing, which is the shoddy state of security in these U.S. companies.
I think just a few days ago, there was a story about three random guys who managed to hack into Open AIs code base.
And then because they were white hat guys, they didn't actually do anything bad.
They just notified Open AI and collected a bounty of $6,000 for having done this.
But the secrets that they accessed by looking at Open AIs code would have been worth billions of dollars, at least.
you know. And if these three random guys can do this today, then of course the CCP apparatus
has probably just like thoroughly penetrated Open AI for years, you know? Like they have so many
more people who probably have more expertise and certainly have a lot more persistence and funding
who have been focused and directed to go after Open AI and probably another big batch of people
who've been directed to go after Anthropic and so forth. So I think it's probably
the only reasonable conclusion is that these companies are probably thoroughly penetrated by the CCP already.
And so, you know, now stealing the code and the algorithms, that's the easy part.
The harder part would be stealing the model weights.
So, you know, that would be sort of, and the reason why that's harder is because they're a lot bigger.
It's just a little, literally it's just a bigger file.
It's going to be several terabytes to download as opposed to just like, I don't know, like some gigabytes or something.
So that's, it's easier, easier to discover that something's being, being, you know,
exfiltrated or whatever.
It'd be a huge file being moved around and it's easier for the security system to notice that.
So just because they've managed, just because it seems like it's easy for them to get the code
doesn't mean that it's easy for them to get the model weights too.
But I would be like, yeah, they could probably get the model weights too if they want to.
Like, they're very sophisticated threat actors.
And they also have access to like, you know, physically.
technical techniques that three random guys can't do, and they can be willing to break laws,
and they can be willing to blackmail and bribe people and so forth.
So I would guess that if the CCP said it's go time, they could direct their people to
steal a copy of one of the latest models inside Open AI or Anthropic.
And that means that in some sense, like, in some sense, all of Chinese progress is coming from
the U.S. in some sense.
It means that, like, at any given time, if things got really serious and there was like a conflict, they could just like take our best thing and then use it, you know?
And yeah, so for all these reasons, I think that like actually the best way to slow down China right now is to slow down ourselves.
And it's like not even close.
That said, is that a permanent solution?
No.
Like if we did slow down ourselves and just like stop, I mean, there's also a difference between slow down and stop.
We could probably slow down ourselves a little bit and still stay ahead of China.
But if we completely stopped, then eventually China would catch up and then surpass.
And that gives us like a window of time, like maybe like 18 months where we need to convince them not to do that, basically.
And that's where diplomacy comes in.
And I agree that's going to be hard.
I think that it's probably achievable.
I think that what we're trying to do with plan A is sketch the outlines of a deal.
that should in theory be mutually acceptable to both the United States and China because it allows us to get the benefits of AI while avoiding the risks.
How could we possibly, you know, believe that they're holding up their end of the bargain?
I don't know what, I'm very curious about the deal.
I mean, it's probably we'd have to dig in for a while to understand how it works.
But the bottom line is unless you have unbelievably levels of transparency,
which is, you know, I just don't see that.
I just don't see that happening.
Yeah.
I think that's exactly right.
And this is why most of our effort in designing this deal was going into the verification
aspects of it rather than the actual principles of the deal.
Because because we don't trust China, we have to make sure that they're not cheating, basically.
And, you know, frankly, they also don't trust us.
So they probably have to make sure that we're not cheating.
Otherwise, they might not accept it in the first place.
they absolutely don't trust anyone, of course not.
And so here's a possible deal that we could, in principle, verify right now with no additional technology.
If President Trump and Xi Jinping, if they both were like, why don't we just pause AI for like six months,
what they could do is they could say, okay, we're going to send, you know, 10,000 Marines without weapons into China.
and you're going to send 10,000 PLA without weapons into the U.S.
Instead of weapons, they'll carry smartphones.
And they're going to come to our data centers,
and ours are going to go into their data centers,
and we will just, like, count the GPUs
and feel that they have been turned off, you know?
And this is like a very dumb,
it's like hitting the problem with a hammer, you know?
It's obviously going to be very costly if we did this,
because then all the GPUs would be off,
and we wouldn't be able to serve all these customers.
and there wouldn't be any revenue coming into these companies, right?
But I'm just sort of pointing out that, like, it is something that we could verify.
Like, we can just send people to their data centers to, like, put hands on the GPUs and be like,
yep, here they are, there's this many of them, and they're off.
And therefore, you know, we have successfully turned off most of the compute, or like 99% of the compute that's in China.
Now, there's, there won't be 100%.
There might be some secret facilities that we don't know about that have, you know, racks and racks of GPUs.
And similarly, we probably have some secret facilities that they don't know about that have racks and racks of GPUs.
But I think that we could get most of it in this manner.
And because AI progress depends so much on compute,
that would, like, effectively stop AI progress for the period of this deal, like the six months or so forth.
Like, that's the high-level thing.
Now, again, am I saying we should do this?
Not necessarily.
I'm putting this out as just an example of how.
how, like, if you really, if you had the political will and you really needed to do this,
you could do it in a way that could be verified.
It'd be uncomfortable.
You'd have to, like, let some PLA people into the U.S.,
and they'd have to let some of our Marines into there,
and, like, they'd be, like, escorted through the data centers
while they, like, film everything with their cameras and, like, feel that, you know.
So it'd be very uncomfortable, but we could do it if we had to.
Our plan A is more sophisticated.
Our plan A is not, like, hitting the problem with a hammer.
It's more like we have like verification technology that we put in to the data centers so that the GPUs can keep running.
And yet the other side can be sure that they're not doing an intelligence explosion.
And instead they're just serving customers, ordinary workloads.
We talk about this in our thing.
So I think that like if you have the political will to do really intense things and also you're willing to do the more sophisticated fancy stuff that we describe,
then I think you've got a chance to like,
You have a chance, but I explain the simple, dumb version here just to sort of give a proof of concept that, like, we could do something like this if we really wanted to, without having to trust them to keep their word, basically.
Again, the idea there being that, like, we could just be like, okay, fine, maybe they have, like, a secret facility somewhere.
But because we've counted this many GPUs across this many data centers, and we have a good sense of, like, how many GPUs there are total in the world,
then we have a good sense that they can't have hidden that many away from us.
And so even if they're doing something illicit on the tiny amount that they've hidden away,
it's not going to matter in six months.
It's not like a major threat in the short term.
And similarly, they would be thinking similarly about whatever we've got.
The bottom line is you're saying we have to find some way of slowing it down
or we're heading towards some kind of Armageddon, superintelligence Armageddon, basically.
That's your position.
And there's whatever happens, that slowing down must happen.
That's what you're saying, right?
Yes, that's my position.
And, you know, I have to tell you at this point, when I look at the variables, I just don't see that happening.
I do.
I agree.
People ask me, like, what's your P-Doom, you know?
What's your probability that this is going to end poorly?
And I'm like, yeah, probably this will end poorly for all the reasons that we just described.
You know, like I think I see a solution.
I think I see some ways that we could get out of this, but I'm not going to lie.
It's going to be difficult and it's not something that, like, it's probably not going to happen.
You know, like probably we're just not going to do it.
And then we're just going to run face first into super intelligence.
That, you know, this said, you know, I do think that thinking about this creatively and proactively
and, you know, I'm getting, I'm getting the sense that you're approaching it sincerely as well.
You know, with the, you know, I think the other dimension that I just don't think people fully understand what the Chinese Communist Party is capable of.
I just, you know, published a book about their forced organ harvesting industry.
You know, it's basically people are used as fodder for elite longevity and profit, right?
And so it's just, it's a very, it's a very dark environment to function.
I guess I was saying, like, when we're talking about guardrails, it's hard to imagine.
there wouldn't be like deep levels of subterfuge, you know, if such a, if such, you know, high-minded,
frankly and, you know, thoughtful and initiatives were to be attempted to be put into place.
I don't want to doom it because I do think the only way through is through precisely having the kind
of thinking that you're doing and trying to come up with solutions that can actually solve it.
I just don't see it yet, but I'd love to continue this conversation with you.
Thank you.
Thank you for explaining how.
I just want to comment on this.
I mean, I truly hadn't fully grasped what had happened.
And especially with your explanation of how these AIs work in broad strokes,
it helps to kind of grasp, helps me grasp that were ahead of where I thought we were.
And it's really only accelerating.
And that's just in itself astonishing.
Yeah, that's another thing is that, I mean, part of what, part of our research
is forecasting the trend lines and so forth.
And it does seem like if you extrapolate the trends,
we get to full research automation in zero to four years
or something like that, depending on the thresholds and so forth.
And so, you know, I think probably this administration is going to have to make
some very tough decisions about how to handle this crisis.
And when we wrote AI2040 Plan A, we knew that the world wasn't really ready for it yet.
like we weren't expecting to get a call from J.D. Vance or President Trump saying,
hey, this is great, let's go do this, you know. I don't have any high hopes that President
Trump is going to, you know, talk about these deal ideas with Xi Jinping in the coming days.
But I do think that the situation with AI is going to get more intense. The AI is going to be
more visibly powerful. And at some point, I think it will sort of come to a head. And people will be
like, okay, well, all options are on the table. We have to do something. What are we going to do? And
we are sort of like thinking ahead to that moment and thinking like, okay, well, what are the options?
Like, you know, we don't trust China. So how could we do this sort of thing with them if we wanted to and if they wanted to?
Like, what would be the details of how we would check that they weren't cheating and so forth?
We're trying to do all that work in advance so that if and when the time comes and the president wants to act, the options have been explored.
I mean, absolutely fascinating.
Thank you for this conversation.
I feel like I've learned a lot.
I hope all our viewers have as well.
Do you have a quick final thought?
I mean, you kind of, you summarized where you hope things will go.
Yeah.
I do have one.
And this is a bit of a kind of in the weeds, he thought,
but I do want to make sure I mention it.
I'm sure many of your viewers have heard about
the recent uproar about whether AI is going to kill us all and a 10% chance,
you know, and things like that.
And then a series of AI CEOs saying we should pace the frontier, right?
What's going on there is that there's been this upsurge of people becoming concerned
about the situation in large part after the Hugging Face incident.
and there was an open letter signed by more than a thousand employees at Anthropic and Open AI and these other AI companies saying the government should be able to slow us down.
We're worried that this is going to be too fast.
Recuracy self-improvement, scary.
And then there was, you know, Jacob Coxon, he quit Anthropic saying basically he disagrees with Anthropic leadership and thinks that they're being reckless.
And it blew up and went mega viral.
And then in response to all of that, I think the CEOs are now saying we should pace the frontier.
but I do not trust them to actually do this.
I think that what's probably going to happen
is that they will do some sort of like weak sauce,
you know, third party auditing thing
that's better than nothing.
Like, I'm not trying to denigrate it.
I think it's valuable to have some third party oversight
into these companies.
But they will set all that up
and then they will continue racing each other
towards recursive self-improvement
and superintelligence at,
you know, 90% of max speed or something like that.
And they'll, like, be doing some minor slowdowns here and there, but, like, basically
they'll still be doing fundamentally the same thing that they were doing before.
And so we're going to run into the problems that I mentioned at approximately the same time
that we otherwise would have.
And so I would I really, like, if I, if people want to know what I recommend, it's we
need to get these companies to actually slow down, not the little companies, the big companies,
Anthropic and Open AI especially, also meta, XAI.
and Google. They need to redirect resources away from recursive self-improvement and towards
anything else, like, you know, cancer, or we're serving customers, or like, you know, just math.
There's like anything besides this recursive self-improvement stuff, they need to just like,
you know, deprioritize that, at least somewhat, and shift towards other things. And if they do,
that buys us additional time to figure out a more sophisticated solution. And it buys us more time
before China gets there too, because again, a lot of China's progress is coming from, you know,
being pulled along by ours. Yeah, so that's my final thing. Let me ask you one final question then,
because I'm very curious about this. There's been a lot of people arguing that this push for
regulation from the world's biggest company, so to speak, I mean, President Trump has a
truth social post about a several truth social post about precisely the issue right that it's sort of the
the idea that these companies would want to regulate themselves right actually is is a is a hoax
i think he uses the term it's a hoax so you kind of but you kind of agree with this that there's
i mean there's others that have been showing that there's you know a lot of money behind it so
this thing but the the arguments have been that they're trying to regulate themselves in a way that
will be advantageous to them. So where does it land in that, in that, in that sphere?
Yeah. You're thinking. Every week I talk to people at Anthropic, Open AI, and Deep
Mind, researchers mostly, not like executives, just like the researchers on the ground. And
there's a lot of interests at these companies on the ground in not doing recursive self-improvement
soon. Like, like, I think, I think a lot of the researchers are starting to get freaked out.
And they are like, yeah, maybe it would be good if we didn't do this at least for now and
did other beneficial things instead.
But they're all terrified of the other companies.
Like they all will say, like, but if we don't do it, then, you know, like Anthropic
will say if we don't do it, then Open Eye will say, if we don't do it, Anthropic
will, you know.
And as for the leaders of the companies, well, I mean, they did just release, you know,
a bunch of statements saying, like, it would be good for the government to be able to,
you know, to slow us down or something.
But in terms of what they've actually done, well, they haven't slowed down at all,
basically.
and the things that they seem to be pushing hardest for are things like, you know, the third-party
risk assessment and stuff like that, which again, like, I know the third-party risk assessors.
I'm friends with meter.
I'm good friends with meter.
I think they're amazing people.
And I hope that they continue doing that risk assessment.
I hope they get more access to do better risk assessments.
But, like, it's just not like what we, it's like we need a lot more than that, I would say.
And if all we get out of this moment is that, then I work.
that effectively what will have happened is there was an incident, there is a huge backlash,
both in the public and among the rank and file employees that we need to, that basically
like doing recursive self-improvement to superintelligence, we're not ready, we don't want to do
that yet. And then the CEO's sort of like channeled that energy and then redirected it
towards this weak-sauce thing that doesn't really address the core problem, you know? And so I am
quite concerned that that's what we're going to see out of this.
One thing I would say to President Trump is, well, the risks are real.
I know that the companies have been, you know, sometimes dishonest, and I wouldn't trust them as far as I can throw them.
But also, there's been plenty of people outside the companies also talking about this.
And I think the evidence is mounting.
So the risks, they are real, and we do have to deal with these problems soon.
And then I think that rather than letting the industry self-regulate, you could just sort of uniformly slow,
slow down their race to RSI, right?
You could do something like requiring them
to spend 90% of their compute on serving customers
and doing other sort of like normal beneficial applications
and only 10% on training the AIs to be really good
at AI research and stuff like that, for example.
And we've written up some blog posts about the types of things
that we have in mind here.
But like if you do that, it would be the opposite
of regulatory capture, because you would
be just sort of like slowing down these big tech companies from doing this incredibly dangerous
thing while letting them, while making them redirect resources towards just like serving customers
and lowering prices.
And it would be the bad version of this would apply to all the tiny companies to.
I wouldn't recommend that.
I would say just focus on like the big three or the big five.
And then that would be like helping the rest of the industry catch up basically.
Wow.
Well, Daniel Kokatello.
this has been an absolutely fascinating conversation.
It's such a pleasure to have had you on.
Thank you. Yeah, I really enjoyed this too.
And I appreciate the depth with which you are willing to go on these things.
I've been on a bunch of interviews over the last week talking about these things.
And this one is by far the most intellectual.
So thank you.
Thank you all for joining Daniel Kokotelo and me on this episode of American Thought Leaders.
I'm your host, Janja Kelek.
