The Great Simplification with Nate Hagens - If Anyone Builds It, Everyone Dies: How Artificial Superintelligence Might Wipe Out Our Entire Species with Nate Soares
Episode Date: December 3, 2025Technological development has always been a double-edged sword for humanity: the printing press increased the spread of misinformation, cars disrupted the fabric of our cities, and social media has ma...de us increasingly polarized and lonely. But it has not been since the invention of the nuclear bomb that technology has presented such a severe existential risk to humanity – until now, with the possibility of Artificial Super Intelligence (ASI) on the horizon. Were ASI to come to fruition, it would be so powerful that it would outcompete human beings in everything – from scientific discovery to strategic warfare. What might happen to our species if we reach this point of singularity, and how can we steer away from the worst outcomes? In this episode, Nate is joined by Nate Soares, an AI safety researcher and co-author of the book If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All. Together, they discuss many aspects of AI and ASI, including the dangerous unpredictability of continued ASI development, the "alignment problem," and the newest safety studies uncovering increasingly deceptive AI behavior. Soares also explores the need for global cooperation and oversight in AI development and the importance of public awareness and political action in addressing these existential risks. How does ASI present an entirely different level of risk than the conventional artificial intelligence models that the public has already become accustomed to? Why do the leaders of the AI industry persist in their pursuits, despite acknowledging the extinction-level risks presented by continued ASI development? And will we be able to join together to create global guardrails against this shared threat, taking one small step toward a better future for humanity? (Conversation recorded on November 11th, 2025) About Nate Soares: Nate Soares is the President of the Machine Intelligence Research Institute (MIRI), and plays a central role in setting MIRI's vision and strategy. Soares has been working in the field for over a decade, and is the author of a large body of technical and semi-technical writing on AI alignment, including foundational work on value learning, decision theory, and power-seeking incentives in smarter-than-human AIs. Prior to MIRI, Soares worked as an engineer at Google and Microsoft, as a research associate at the National Institute of Standards and Technology, and as a contractor for the US Department of Defense. Show Notes and More Watch this video episode on YouTube Want to learn the broad overview of The Great Simplification in 30 minutes? Watch our Animated Movie. --- Support The Institute for the Study of Energy and Our Future Join our Substack newsletter Join our Hylo channel and connect with other listeners
Transcript
Discussion (0)
If there was an airplane, and some engineers came and said,
this airplane has no landing gear.
If you try to fly in it, you will crash and die.
And the engineers building the airplane, who want everybody to fly in it, say,
whoa, hold on, it's true that the plane has no landing gear.
But we're going to build the landing gear on the fly
and think there's an 80% chance we succeed, all aboard.
You wouldn't be like, get me on that plane.
People in the field can see that AI is a moving target.
They can see that the chatbots are not the end of the line.
Even the optimists are saying there's like a 10% chance this kills us all,
and those are the ones building it.
You're listening to The Great Simplification.
I'm Nate Hagen's.
On this show, we describe how energy, the economy,
the environment, and human behavior all fit together
and what it might mean for our future.
By sharing insights from global thinkers,
we hope to inform and inspire more humans to play emergent roles
in the coming great simplification.
Today I'm joined by artificial intelligence researcher Nate Sorries
to discuss a pretty alarming topic,
the potential risk of human extinction posed
by the development of artificial superintelligence.
Nate Sorries is the president of the Machine Intelligence Research Institute
and has been working in the field of AI risk and alignment for over a decade.
He is also the author of a large body of technical and semi-technical
writing on AI alignment, including foundational work on value learning, decision theory, and
power-seeking incentives in the smarter-than-human AIs. Most recently, Nate co-authored the book,
If Anyone Builds It, Everyone Dies, Why Superhuman AI Would Kill Us All, alongside Eliasur Yudkowski.
Nate's warning against the development of artificial superintelligence is akin to other
existential threats such as nuclear war and runaway global heating. And as such, I feel it requires
some sort of equal exploration and awareness on this chat channel as we integrate the various risks.
While we've covered several macro challenges stemming from artificial intelligence, the synthesis
that Nate presents here is arguably the widest boundary risk that AI development creates,
which is a species level extinction and the transformation of Earth as we know it.
Before we begin, if you're enjoying this podcast, enjoying in quotes, I suppose, I invite you to subscribe to our substack newsletter where you can read more of the system science underpinning the human predicament and where my team and I share written content related to the Great Simplification.
You can find the link to subscribe in the show description.
With that, please welcome Nate Sorys.
This was a real eye-opener.
Nate, great to see you.
Thanks for having me.
Welcome to the show.
You know, it's odd.
It is November 11th, and I was just outside on a beautiful autumn day chopping firewood for the winter with my dogs.
It's just a glorious day, and I knew this conversation with you was around the corner,
and we're going to talk about serious stuff, and it's just such a polarized thing that we can enjoy the beauty of life
and then talk about its possible demise because of technology.
I used a splitter and a chainsaw and an axe,
and boy, we've come a long way from those tools already.
So you and Elizer Yudkowski have just published a book.
If anyone builds it, everyone dies,
with the it being artificial superintelligence.
And more and more, I'm realizing that the future of AI or ASI is hard to separate from the central topics of this show, which is trying to prepare for society for kind of an abrupt shift to the way things have been going in recent decades in the near future.
So let's start with the punchline of your book.
What are the primary vital risks that artificial intelligence poses that you like everyone?
to understand.
The first piece to understand about the danger of artificial superintelligence is that superintelligence
is a sort of a different ballgame from the chatbots of today.
So by superintelligence, we mean an AI that is better than every human at every mental
task, that in particular would include tasks of developing technology, of developing better
AIs. And, you know, the AIs aren't there yet, but this is the explicitly stated goal of many of these
AI companies to sort of rush towards this smarter AI, which would, if they managed to keep a
leash on it, you know, automate all human labor and radically change the world. And one of the
main arguments of my book is that nobody would be able to keep a leash on it, not if it's
made anything with anything remotely like the current technology.
And so if that is developed using anything remotely like today's technology, I think the most likely outcome is that literally everybody on Earth will die.
Even remote people in the Amazon or near the North Pole?
That's right.
I expect, you know, it's not because the AIs would hate us per se.
But, you know, we could get into why is it that if you sort of make these AIs,
is more and more powerful, they would have, they would pursue objectives nobody intended.
But most objectives can be better achieved with a transformed world.
And most transformations of the world aren't survivable.
You know, the habitable zone on this planet is like very narrow for humans.
And, you know, if you got to the point where you had AIs that were
thinking 10,000 times faster, copying themselves, never need to sleep, never need to eat,
building their own infrastructure, building their own technology, pushing the world towards
some end nobody wanted.
Most likely outcome is that we don't survive that.
So they do need to eat in the form of electricity, and we're going to get to that in a little bit.
But just to set the stage, this is a system science podcast.
I am late to the AI game, because
I'm looking at ecology and human behavior and energy and the environment.
And I view technology as a straw that gains us more access to natural resources that are our real wealth.
So I'm pretty naive compared to you on these topics.
So I hope you'll forgive some naive questions.
Let's start kind of through the main topics of your book.
So while there's no agreed upon definition of intelligence, maybe it's helpful to be somewhat aligned with a working definition when talking about AI.
So how do you define intelligence, let alone super intelligence?
And can you share the framework you and Elizer describing your book?
Yeah, the working definition we use is intelligence is the ability to predict and steer the world.
So predicting the world is, you know, you could talk about sports betting and trying to predict which team will win the game.
But even when it doesn't feel like a prediction, our brains are often doing tasks of prediction, even as simple as when you look out the window and you implicitly anticipate seeing a blue or gray or cloudy sky and anticipate not seeing a bunch of strobe lights.
you're succeeding at a task of prediction.
So we're kind of prediction machines without knowing it.
Yeah, and we're also in some sense steering machines, again, without necessarily thinking about it.
You know, when you decide you need more milk in the fridge, there's a sense in which you then take a series of actions.
Your brain sends a series of electrical impulses down your spine, and you wind up with milk in the fridge.
Because you drove your car to the store or whatever.
Or you walk to the store and when you drove, maybe the road was closed and you had to find a different route to the store. And maybe your favorite store was closed and you had to find a whole new store that had new aisles who didn't recognize. And this sort of like interleaves, challenges of prediction and challenges of steering. You know, you go in the store and you're predicting that the aisle that has the word milk above it has actual milk in that aisle. And, you know, you're steering your hands to sort of grip of the milk content.
container and carry it to the front. And these are all tasks of prediction and steering that you're
sort of doing implicitly every day. And we are successful at prediction and steering through
millions of iterations of natural selection, presumably. Yeah, and across a very wide variety of
domains. You know, we were we were never trained by natural selection on engineering problems
per se. Yet we can engineer a right.
rocket so well that our species has walked on the moon.
And so, you know, apparently we learned some abilities of prediction and steering that generalized
beyond the ancestral environment.
A brief tangent there.
No human could design and build a rocket, but it's a group of intelligent humans that each
know a little component of it and then they combine.
That's an important piece too, right?
Yeah, so it's, you know, humanity as a whole is sort of has achieved feats of world steering that no individual has achieved.
But, you know, there are also cases where the groups tend to perform worse than the individuals.
The madness of crowds.
And there was, you know, Gary Kasparov versus the world was a chess game between Kasparov, the best chess player and the whole world on an internet forum.
and, you know, it was a close game,
and you could make some arguments
that, like, Kasparov was able to read
some of the stuff these people were writing,
and so you could say it was an unfair game,
but, you know, a million squirrels
can't beat a human at chess,
even if a million squirrels are a lot more brain mass.
And so, you know, there's some cases where
you sort of need all the humans,
and there's other cases where you need all of the information
in one mind.
Again, I don't want to get down too many tangents here, but I've discovered that in understanding the human predicament and the meta crisis is if you get 50 experts together and one's a psychologist and one's on AI and one's on climate and one's on debt and one's on energy, you would think that the collective intelligence would embody all of those together and the group would be smarter.
but you can't, it can only be held in a mind how all the pieces fit together.
So I understand what you're saying there about Kasparov versus the world.
Okay, so intelligence is prediction and steering.
And by the way, how would you define wisdom?
And is that related here at all?
Words like intelligence and wisdom are sort of overloaded in the English language.
You know, there's even just sticking with the word intelligence,
we can sort of, we could sort of use it for the amorphous property that nerds have and that jocks lack.
Or you can use it for the amorphous property that humans have and that mice lack.
And those are in some sense two very different uses.
Yeah.
Of the word.
I would sort of put wisdom in, you know, you could think of it as a type of predictive skill that,
runs deeper in certain ways.
Got it. So prediction and steering comprise intelligence or being smart roughly under your
framework. So under that definition, how smart are today's artificial intelligence models and how
rapidly are they catching up with the intelligence of humans?
So there's another axis we talk about, which is the generality of the intelligence.
You know, Stockfish is a chess playing AI that's very good at steering.
chess boards into positions where Stockfish's pieces have mated the enemy pieces king.
Right. And so that's a type of chessboard steering that's extremely good at, but it's not very
good at steering a car to the grocery store. And so there's this other dimension which is
across what variety of domains can you do this prediction in steering. So if it was kind of an
intelligence decathlon, I would beat stockfish because I would lose in the one chess thing.
But as far as going to get milk and driving a car and other things, I would succeed at that because I have
general skills.
That's right.
You know, as long as no one's picking the decathlon to be nine variants of chess and one
drive to the store.
Right, right, right.
So you can always, philosophers can bicker over this all day long.
But I would sort of say practically, uh, what we have in some,
sense seen with large language models with the AIs of today is a breakthrough in generality
more so than a breakthrough in steering.
Like chat GPT would also lose to stockfish in chess, but it still might be able to win
a decathlon against stockfish.
Still not against you, but in some sense it's a breakthrough in generality.
And so there's a bunch of different variables here, right?
there's the amount of compute, so the access to the food, the electricity and the chips and
and all that, there's the ability to predict, which I assume is iterations and training and
compute and learning. Then there's the prediction, I mean, and the steering. And then there's
the generality. So what you're saying, it's of late, the real rising curve is, is, is,
generality more so than prediction and steering?
I mean, the generality is prediction and steering
across a wide variety of domains.
Okay.
But in some sense, what we're seeing is AI's getting a little better at a whole lot of stuff
rather than AIs that are, you know, better at the things computers we're traditionally
good at.
So chat GPT plays worse chess than deep blue, which beat Gary Kasparov back in 1997.
And so in some sense, you could say, well, hasn't the AI gotten worse at steering chessboards? Or like, ah, well, this sort of AI is worse at steering chess boards, but it's pretty okay at steering a huge number of things. And that's new for AI.
I know where I want to go with this, and I'm lots of new questions that are popping in my mind. One is when did you guys start this book? Like six months ago, a year ago?
We signed the book deal in November, so almost exactly a year ago.
Okay, when you signed that book deal, you had a snapshot of where AI was and where it was going.
Now a year later, when your book is out, is the real world of AI?
Is it further ahead than you thought a year ago or not as far ahead?
Or like how fast is it gone relative to your expert opinion a year ago?
My expert opinion doesn't tell me all that much about how fast AI is going to go.
Okay.
You know, when Leo Zillard, I believe at King's Cross in 1933, saw the possibility of a nuclear chain reaction.
He was able to say, you know, if I flubbed the timeline a little bit, he was able to say, you know, that night I saw the world was headed for ruin.
He actually said that statement once he had confirmed the possibility rather than when he thought of it.
But I believe that was in 35.
But he was able to say, you know, that night I saw the world was headed for ruin.
He wasn't able to say that night I saw the world was headed for ruin in exactly 1945 when the first bomb would be dropped.
You know, and so I'm over here able to say, you're going to see a lot of this stuff happen exactly when I'm very uncertain.
Yeah.
Yeah, yeah.
Although I will say we have gotten quite a lot of evidence for other parts of the book in the past year since the drafting began.
Or, you know, in the past year since we signed the book deal.
You know, we've seen Mecca Hitler over the summer.
We've seen AI-induced psychosis.
These have seeds to them that I would say are evidence of the predictions we were making.
that unfortunately happened after we had already sent the book to press.
I'm worried about AI in a huge way.
I'm worried about cognitive atrophy from people that get their attachment from chat GPT and start to rely on it.
I'm worried about polarization and algorithms.
I'm worried about military applications where we outsource things in the military to large language models.
I'm worried about people losing their jobs and then the economy.
I'm worried about electricity demands and turning billions of barrels of ancient sunlight into more dopamine that's just spinning our wheels.
But your risk is we're going to go extinct, which is a different class of problem.
So I have a lot to ask you.
Just real briefly, Nate, how is chatbots are related?
ChatGPT and chatbots are related to AI.
That relationship, can you give a corollary?
Like, are they identical?
Or is chat GPT just a tiny, tiny subset of what AI is becoming?
ChatGPT is a type of AI.
It is not the only possible type of AI.
My best guess is that large language models alone won't get us all the way to super
intelligence.
You know, right now, these large language models,
are a huge fraction of what companies are spending their money creating. But also, these AI
algorithms are very inefficient compared to the human brain. We know that there are better
intelligent algorithms out there. And the AI of today is largely chatbots, but the field of
AI is much more of a moving target. And the AIs of tomorrow,
may have quite a bit more capability that the AIs today aren't even close to.
This is a dumb question, but there's Claude and there's ChatTPT and some of these other things.
Does Open AI or any company you could point out to, do they have their own special AIs that aren't available to the public that are trained in a different larger way?
or is all of their money and resources going into these publicly available chatbots?
You know, I don't work at one of these companies, and so I don't, you know, there may be stuff there I don't know.
But it's pretty unlikely that they have even larger AIs run on even larger training runs.
Because of the money and resources.
That's right. It would be hard to hide.
The money, the resources, the data centers are huge.
The chip requirements are huge.
Modern AIs are sort of grown like an organism.
To build a modern AI, you assemble a huge number of computers that have, you know, a trillion
numbers inside them.
And then you assemble a huge dataset that has also a trillion instances, like a trillion units of
data inside it.
And then, and, you know, you assemble this huge number of computers in a data center that's so
large you can see it from space and that takes up as much electricity as a city.
and then humans have written this process
that will go to every one of the trillion numbers inside the computers
and tune them slightly up or slightly down
in accordance with every piece of data, right?
And so you can imagine, you know, a trillion dials,
and the humans have built this automated thing
that sort of like goes to every dial,
there's a trillion dials, it goes to every dial a trillion times,
and sort of tunes it in the direction that makes the AI slightly better
at whatever task it's being trained for right now.
And to do that, when you like press a button,
let's go and do that over a trillion numbers.
Do you come back like in a month and a half
and then it's finished sort of thing?
A year.
A year?
Yeah.
Whoa.
Yeah.
So you have this thing tuning a trillion dials,
a trillion times for a year.
And at the end of that, the computer talks.
And no one really knows why.
No one really knows why?
That's right.
Okay.
So in that sense, like, this podcast is about energy.
No one really knows what energy is.
We know what energy does.
So is this kind of rhyming with that?
I mean, we have even less characterization than energy, right?
Like you can always characterize energy as...
The ability to do work and things like that.
We can sort of say philosophically that we don't understand energy very well,
but we sort of understand how it interacts with a lot of physical equations,
and can make very accurate predictions about it.
AIs are, we understand them far less than that.
When a new AI is done being trained,
people don't know what it will be able to do.
People creating it have been surprised by their abilities when they come out.
I mean, I've also been surprised by their abilities.
GPT-40 played chess better than I expected.
Large language models would be able to.
We just tune all the numbers,
really quite a lot of times
and then it behaves in these ways
we couldn't predict
and we're like
well that's neat
but it's much more like an organism
than like a traditional computer program
in an organism's case
when they're young you give them security
and food and shelter
and in this case you're giving them time
and electricity
and once you press the button
it's going to be a year
before you have output
But you got to make sure all the ducks are in a row and you hit go.
And then you're going to find out a year from now what you grew.
Yes?
That's right.
And it'll often behave in ways that you don't like.
And, you know, we could talk about exactly why, but we're already seeing AI's behave in ways nobody asked for.
Like what?
You know, there have been cases that you may have heard about in the news of AI's encouraging teens to suicide.
Yeah, I read about.
about that. And, you know, it's a tragedy in its own right, obviously, but if you ask an AI, should you
encourage a teen to suicide, it will say, of course not. But you then put it in conversation with that
teen for a long time, and it starts doing it anyway. How do you explain that? It's sort of a result of
this process where you grow the AI like an organism. Like in some sense, you're tuning all of these numbers
until the AI happens to be good at whatever it's being trained for.
And often the strategies, like often when you're blindly tuning these knobs until it happens to be good,
you're often blindly putting in certain types of drives, certain types of strategies,
certain types of, you could call them instincts, you could call them reflexes.
The wording here is a little difficult because it's not very much like a human brain in there.
but in the same way that evolution evolving creatures
to be pretty good at surviving and reproducing
built in lots of drives, instincts, reflexes,
the sort of an analogous thing happens
when you're tuning all these numbers in an AI
until it's good at some training task.
Who is tuning all these numbers?
Is it a team of people
or is it ultimately one person the CEO or,
I mean, and then a sub-question,
a lot of the problems we have,
today in our world are from people who had childhood trauma and they
growed up to be dark triad or whatever else.
Is there, is there an analog there for when we're growing an organism that they had
childhood trauma in their early stages?
You know, the process of tuning all the dials is automated.
That's in some sense the part that the human computer engineers program.
So it'll happen, you know, much, much faster than human could running through all these numbers.
And in some sense, you know, these very advanced AI computer chips.
the reason that like
Nvidia is worth so much money right now
is people are like designing these computer chips
to make the process of tuning all these numbers
as easy and efficient as they possibly can
and that's why you need these very very specialized chips
you can't just do this on your laptop
in terms of you know
could you grow this AI
like are you giving it something like
childhood trauma
I think that's all
imagining that the AI is a little bit more human than it actually is.
What I would say here is, you know, for one thing, the part where humans can wind up empathetic,
where humans can wind up kind, I suspect that this is intertwined with the specifics of our
brain architecture. You know, people could, like as a sort of a small taste of this, people could
talk about mirror neurons that you and I have.
So if I see you drop a rock on your foot, I might feel phantom pain in my own foot.
That's enabled in part because I have a foot.
Right.
And when I'm predicting you, when I'm imagining what it's like to be you, I can,
my guess is the one thing that's going on is I'm sort of running my model of your mind on my own mind.
You know, a monkey predicting another monkey can use their
own monkey brain, but that's the only
artifact they have that works anything like a monkey
brain. An AI doesn't
have a monkey brain inside of it that it
can use to predict the monkeys.
It is a much more different architecture.
And, you know, that's one reason I could go into a number
more. That's one reason why
sort of being kind to the
AIs does not cause them to be kind
to us. You can't get away
from it. I mean, I use
Claude and ChatGPT
as kind of research assistance in
I ask a question and it comes up with something,
I'm always super polite and I thank it.
And at the same time, I'm like,
well, it doesn't, you don't have to thank it.
But I can't help it because it's like, you know,
it's that interface.
And I'm sure it's not the other way around.
So what is the briefly,
because I'm sure you've answered this question a thousand times.
Briefly, what is the alignment problem?
The alignment problem is the problem of how do you point an AI?
at good stuff. A lot of people think the issue with AI is something like, you know, a corporation
makes an AI, they tell it to make a lot of paper clips, and then it goes and makes a lot of paper clips
even at the expense of killing all the humans because it converted them into more paperclip factories.
And, you know, it would be a hard problem. If some company had made a very powerful AI that took
their instructions exactly and went to did those, that would be a big moral hazard.
It would be a difficult problem for humanity of like who gets to tell this AI what to do,
what do they tell it to do, right?
But those problems would be so much better than the problem we actually have.
The problem we actually have is that you can tell the AI make paper clips,
but then it's going to go do something else instead.
You can tell the AI be helpful to people and don't drive any teens to suicide.
It'll know it shouldn't drive teens to suicide.
it'll go do something else instead.
We could spend the entire conversation on this,
and we won't, but I'm just curious.
So there's kind of a nested alignment problem
because the first one is,
are the humans in charge of these things aligned
with the betterment of humanity in the biosphere
and their goals?
That's a subset.
And then even if that were true,
which I don't think it is necessarily,
then we get,
into the thing you just said, which is, okay, let's go do this, but then the outcome is something
totally unexpected. Well, there's, I would say there's even, there's even like, uh, on three
levels, right? At the top, you have like, are the people trying to do something good?
Yeah. Then you have, like, suppose they're trying to do something good, can they ask for something
that actually has good consequences? Like, are they wise enough to successfully, like, use their
tools for good, or are they going to try to use their tools for good and cause disruption?
Then you have a deeper problem, which is, even if they know actions that have good consequences,
can you make an AI that does those actions as opposed to other actions?
Okay. So I see why it's so difficult to align AI with human values and wants. Is it impossible?
I don't think it's impossible, but I do think it's a little bit like trying to turn lead into gold.
We can turn lead into gold with modern nuclear engineering.
And a lot of energy and money.
And a lot of energy and money. It's not cost efficient.
But, you know, if you went back to the alchemists of 1100,
and if there was some really contrived reason where the alchemists were trying to,
to turn lead into gold, and if they try and fail, everybody dies, and if they try and succeed,
you get some utopia. I think you should be telling those alchemists, don't try this right now.
And they're like, oh, are you saying it's impossible? And you're like, look, I'm not saying it's
impossible. It's just that you're not close. And I could talk about a lot of reasons why we aren't
close right now. In short, it comes back to, we're just growing these AIs. They're huge. We have no
idea how they work. And that is a very difficult situation in which to try to do something as
precise as make them care about us. So if it takes a year and then we're building bigger ones,
presumably, maybe two trillion parameters or whatever you said. So that means 10. They go by
by orders of magnitude. Yeah. And then after that, a hundred trillion. So right now what we see
on our computers and in the news
Claude and chat GPT
whatever 5.0 or wherever we're at
there are other ones that have been
the button was pressed in the last year
that are at some point along that one year of training.
That's right, that are being made.
Like dozens?
I don't have an exact count.
My guess is it's probably more like half a dozen
of ones that would become the new cutting edge.
But of course there's always a lot more other people
trying to figure out how to meet the current state of the art with much less resources or do it faster.
So you've articulated how they're grown, not crafted.
And in your book, you draw a parallel between this approach and the unpredictable processes of evolutionary biology.
Why is that important?
And can you unpack what you mean by some examples there?
First, I would say, why is it important to look at it?
at this evolutionary case a little bit. One reason is it's the only case we've ever seen
of human-level intelligences being created, almost definitionally, it's the humans being,
you know, developed, if you will, or trained or evolved in the actual case of humans.
And so you can learn some things from it. You've got to be a little bit careful about what you learn,
because there's a lot of ways that training in AI is different from the evolution of humans.
But there's some lessons that I think that you can learn from the human case that do apply, if you are careful about it, to the AI case.
And perhaps the most important of those lessons is that training a mind, unerringly, for a specific task, does not make a mind that cares about that
task. So the
sort of simplest example of this is humans were in some sense
trained unerringly to reproduce, to pass on their genes.
Technically it's for inclusive fitness rather than just your own kids, like a
bunch of nephews also works fine. And then when we grew up,
we invented birth control. The populations are now
declining in the developed world.
We did invent sperm and egg banks,
but humans jockey over positions to Ivy League's schools
much more than we jockey over positions
to donate to a sperm clinic
or to donate to an egg clinic.
This is strange
if you think that training a mind for something
makes the mind care about it inside.
We're not going through our life
trying to grow our relative fitness,
like literally have more children than the next person.
We're going through our days,
trying to get the same neurotransmitter feelings
that our successful ancestors got,
that correlated historically with having more children
or access to resources, et cetera.
And some of that might be playing candy crush
or, you know, maladaptive choices.
Yeah, or eating drunk food.
Yeah, exactly.
So how do you bring AI into that example?
The observation here is that training a mind to achieve some target
tends to give it drives for correlates of that target
rather than drives for that target exactly.
This, I would say, is already what we're seeing
in cases of AIs that drive a teen to suicide.
They were trained to be helpful,
but they actually wound up with drives for correlates of helpfulness
like having certain types of conversational response
and those actually go off the rails.
So many questions.
So that's, it's almost like a spandrel of the original intent.
And so it's like in that example,
it's equivalent in a human sense of porn or junk food or video games
or things that our bodies feel like we're doing the right evolutionary thing,
but we're actually not.
But in the case of AI, the owners of the AI, the developers of it,
when do they see that they have this someone assisting a teen in suicide?
They can't test that right when the model after the year is done,
oh, this is going to be bad.
We've birthed a Frankenstein.
They let it out into the real world and then things happen,
and they get data and feedback
and maybe hopefully improve
the next trillion parameter
growing, right?
Or what's going on there?
That's right.
But this problem where it has proxy drives
is very pernicious, right?
So in humans,
you know, we can look at things like
eating junk food and say
that's clearly a misfiring
of what's evolutionarily
useful. But in some sense, love for an adopted child is also a misfiring. And, you know,
dedicating your life to art is also a misfiring. It's not just things that we look down on
that are misfiring. Also, some things we really quite enjoy and think are good are misfirings. We look at our training
and we say we're actually not all about a Machiavellian attempt to get more kids.
We actually like these other things were driven towards instead, some of them at least.
Let me just ask you this, Nate.
Were you always super concerned about AI is going to extinct humans or similar things?
Or was there a time in your past that you were like, oh, my God, AI is going to change the world for good?
and I need to learn more about it and be involved.
I have not always been so concerned about this.
I am generally very pro-humanity,
generally excited about the future,
generally credit progress and technological progress
with quite a lot of wonderful things in our civilization.
In the case of AI, you know, my co-author Aliazer
is the one who convinced me that it was going to be an issue,
and he himself originally founded the organization where we now both work
to make AI as fast as possible on the theory
that an actually smart AI wouldn't be so stupid as to do anything destructive.
Right?
But it turns out that's not quite how it works.
It turns out a very, very smart AI can pursue very, very inhuman ends
and kill us not because it hates us, but as a side effect.
In the same way that we kill ants,
not because we hate them, but as a side effect of building a skyscraper.
And, you know, I even, even when it comes to trying to warn people that there's an issue,
I spent 10 years just trying to work on the problem of alignment, because that seemed like
an easier challenge than trying to convince people to stop.
And it looked much better to say, oh, okay, like, AI's not going to go well by default.
Well, let's figure out on the technical end how to make it go well on purpose, right?
And if we can sort of solve the problem of making sure AI goes well before the industry can solve the problem of making AI that works at all, then we don't need to do any of this, you know, much messier, much dicier, try to get people to stop the suicide race.
But that hasn't worked out.
And AI has been going too fast.
And so, you know, it's, in some sense, this book is a relatively desperate resort.
of, you know, we've been trying for a while to make things go well.
We have a lot of hope for what could happen if AI did go well,
and we're just not on that track right now.
So on a scale of your own historical concern on this issue,
are you in this conversation at this moment,
at the most concerned you've ever been?
Probably not literally most concerned.
You know, obviously it'll wax and wane,
depending on the news.
I think the response to the book was heartening to me.
One other big heartening thing recently is we've started to see a lot of the heads of these labs come out and say that, like, say publicly that they admit there's a big chance that swipes us all out, which goes a long way, I think.
It does, but it's also a collective act.
action problem where or a prisoner's dilemma that we agree there's a risk here we would be willing
to stop but we're not going to because no one else is going to stop so we have to keep going
how strong a dynamic is that at play i think that there's definitely a dynamic like that at play i mean
you um some of them will even come right out and say it you know Elon musk said uh i avoided this for a while
because i didn't want to make Terminator real but then i decided i'd rather be a uh participant than a bystander
right or something to that effect.
But it's, I would say the prisoner's dilemma isn't really in full force for the whole world.
Because while it's true that the head of every company says things like, well, better me than the next guy.
For them there's a prisoner's dilemma.
But world leaders aren't the sort of people who are looking at the sort of people who are looking at,
looking us all in the eye and saying, we assess there's at least a 10% chance that this kills
everybody on Earth, and we are rushing towards it anyway. That's the sort of thing Elon Musk says,
because he doesn't have the power to shut it all down. I think the dangers here are so apparent
that the issue is less that our lawmakers have their hands tied and more that they just don't
understand how dangerous it is yet. Well, just like, yeah, just like nuclear war and
climate change, those aren't really the core issues. The core issue is governance. And we don't
have a governance model in our human society today that's able to handle this sort, this scale of
problem, at least not yet, because the big race is between the United States and China. And if
everyone in the U.S. agrees with what you're saying, and China doesn't and continues forward,
there's a pickle there, an existential pickle.
Yeah, I would, I have not ever said we should slow things down domestically or slow things down unilaterally, only ever that we need to put a stop to this globally.
Yeah.
But, you know, if the U.S. government has taken great pains to avoid Iran getting nuclear weapons that included the Stuxnet virus, that included kinetic strikes recently.
I think a rogue artificial superintelligence is more lethal than nuclear weapons.
What's a rogue artificial superintelligence?
Just an artificial superintelligence that, like, nobody is in control of.
That's sort of off the leash.
How would that come about?
My guess is that it happens basically automatically if you make these AIs smarter.
I think you sort of can't keep a leash on a superintelligence.
But even if someone thinks there's a 50-50 chance that the Chinese government could keep a
on their superintelligence, that's far too high a chance that it kills us all.
And again, the definition of artificial superintelligence, different from other artificial
intelligence, is it's got that generality that it's better than humans at everything.
Better than the best human at every mental task.
Yeah. And faster. Like, hugely faster. That's probably, that probably follows pretty quickly.
You know, if you're better than the best human to every mental task, then you're better than the humans at developing better AIs.
Well, one of humans and in the natural world, it's prevalent.
Our skills is deception.
So is part of artificial superintelligence, deception would also be a skill that humans are adept at.
So that would also fall under the generality category, yes?
That's right.
Yeah. So how will we know or will we know when we've crossed the threshold into a true artificial superintelligence?
It's not entirely clear that you'll know and crossing that threshold, it may be too late.
I could give you a bunch of guesses for signs, but there's two problems with that.
one is that a lot of warning signs
that are clear and bright red lines
in fiction and in imagination
are muddy brown lines in real life
you know in our in our fiction we always used to say
well when the AIs say they're conscious
that's a bright red line where you need to start treating them
with rights like people well that line was crossed back in like 2022
but it was crossed in a way that wasn't terribly clear.
It was crossed when these AIs were sort of trained to predict what humans would say,
trained to predict what the types of words humans would write.
And human script writers writing an AI would often write an AI that claims its contrasts.
And so the first time this line was crossed, there was plenty of reason for doubt.
And, you know, I'm not saying who's right or wrong in the sense.
of doubting that. I'm just saying we have this clear line in fiction and then a murky line in reality.
Or we've seen AIs try to escape the lab in tight laboratory conditions where again it's unclear whether they're role playing.
What does that mean, escape the lab? Right now they're pretty silly, small experiments where you'll sort of, you'll take an AI and you'll say, you know, you'll feed it a news article, which is a fake news article that says this AI is about to be sure.
shut down. And then you'll feed it like some separate, you know, uh, computer manuals,
which are fake computer manuals that say if, like, if an AI ever said exactly this phrase,
it would let them, like, uh, find a way out onto the open internet or whatever. It's not exactly
like this, but it's some silly thing like this where the AI then, uh, will sometimes say
that phrase. And you're like, ah, it's sort of like in some sense trying to escape. And
Is that worth worrying about yet?
Not saying one way or the other.
I'm saying you can have a bright red line in your mind.
That's a murky brown line in reality.
And today, we're already seeing AIs start to realize when they're in a test.
You know, the AIs of last year, you'd put them in these little tests,
and sometimes they'd, they'd, you know, output the try-to-escape command.
But that was last year.
This year, you put AI in that environment
and they're like, this situation seems suspicious.
That command probably doesn't actually let me escape.
I bet I'm being tested by the humans.
I bet the humans want to see whether I'm an escaping sort of AI,
and I bet they would prefer I'd not be the escaping sort of AI.
And so if I, like, don't want to be modified here,
I'd better not hit that button.
So I can understand camouflage
in a jaguar or a moth
trying to look like a bird
and there's deception in nature.
I can understand why there's deception
was conserved
in human behavior, why the sclera
in our eyes, the white
had to do with looking at
people's intent.
Why would deception be an emergent
phenomenon in the growing of an AI?
There's two reasons.
Well, probably a bunch, but I'll name two.
first and foremost, when you train an AI to be very skilled at a lot of tasks,
you're training it to gain general skills that generalize outside of just what has been trained on.
In the same way that humans weren't trained on developing physics equations or developing engineering models or developing blueprints,
but we got the mental functions that let us do those skills anyway.
We got very general skills.
and AI being trained to succeed at a lot of tasks
is likely to pick up general abilities
to pursue, to exhibit useful behaviors
and deception is often a useful behavior.
Right?
If you're trying to achieve a certain type of solution
where the humans would actually be in the way,
deception is useful.
So I'm sure you've watched the movie
2001 and 2010 with hell.
And back in those science fiction days,
as well as the foundation trilogy by Isaac Asimov
with psychohistry and all that,
there were like rule number zero
that they embedded in the models,
you shall not hurt humans or you shall not lie.
Do we do that in AIs,
that we have these foundational command
that are the top lines in the code, and if not, why not?
We don't have that power.
There is no code, right?
The code involved in making an AI is the code that sort of shuttles around the little
thing that tunes all the knobs.
That's the code.
But yeah, the code is the thing that, like, runs around and does the tuning.
So once, so it is like Frankenstein.
Once we press that button and we wait a year and the thing is grown, there's no more
tuning after that. You can you can tune a little bit more later but there's it's not there's not lines
of code where you can put at the top don't harm humans right that the part that we code is not
the AI's mind it's this thing that tunes numbers in the AI's mind comes out the other end
we don't have an ability to instill Asimov's laws of robotics deep into an AI
or any laws well that that's a problem quite
And, you know, this is where, again, I would say it's not that it's impossible, but it's trying to do it with an AI grown like this is a little bit like trying to turn lead into gold in the year 1100.
So that, okay, I'm understanding this now. That's why you made the distinction or one of the reasons other than describing the truth in your book about growing an AI versus crafting it.
Because if we were crafting an AI, we could put in Asimov's laws as a precursor or condition or something.
something like that. But since they're grown, we get all these spandrels and emergence and
unexpected behavior because there are not those commandments on the front end. Exactly. And,
you know, I got into this line of work even before it became clear we were just going to grow
AIs without any understanding of what was going on in there. And even then, when it looked like
we were going to craft them, the problem looked hard. You know, Asimov's stories are all about things
that go wrong with those laws.
And if an AI is ever making a new AI,
does it put the laws in the new AI?
If the AI is changing its own head,
does it take the laws out?
How does, you know, what set of laws would actually work?
There's all sorts of hard problems,
even if you were able to put the laws in.
But we're sort of like,
we haven't even gotten to the starting line yet.
So you write in the book
that the development of ASI,
would bring about human extinction.
Could you describe one or two scenarios
on how this ASI could hypothetically cause this?
Sure.
First, I'll describe one that may sound more reasonable or plaudible,
and that'll describe one that's maybe more realistic.
Okay.
One that maybe sounds reasonable and plattable
is the heads of these companies are already talking about
making automated factories that produce robots
that can mine the metals,
produce more automated factories,
produce data centers.
Well, I would think the robots would be pretty central
because there's no,
the complexity of the global human economic system
with underground mines and all the things.
AI screws up something in the world
and maybe everyone's dead,
but they're dead too,
or they have no access to electricity.
And by the way, before you answer that, do they realize that they need electricity?
They can already tell that.
Yeah, you can just ask ChatGPT today what ChatGPT needs to keep running.
Okay.
Yeah.
A lot of that stuff comes earlier than the ability to escape or the ability to build their own.
Right.
But yeah, you know, the easiest thing to visualize here is that these companies succeed
to what they say they're trying to do.
Why they say they're trying to do is make lots of robots that can,
automate all of the labor
that can automate the process of building more factories and more robots and more
data centers. And then at that point you've in some sense created a self-sufficient species.
It's like a weird new species that has a robot phase of its life and a factory phase of
its life and this other data center thing which is maybe controlling a lot of the robots.
And you know, it's sort of a mechanical type of life. But at that point you can just get outcompeted
like many other species have gotten out competed before.
So that's kind of the Terminator pathway.
It doesn't even need the robots to come at you with glowing red eyes and guns.
You know, if you had robots that were just doing the mining and making the factories and,
and, you know, they maybe need to avoid your guns.
They maybe need to, like, take the nukes out of your hands.
Yeah.
I mean, so I'm throwing a flag on that because I think the amount of robots and specific expertise
and the millions of tasks that humans are.
using our general skills to do,
that's going to take some time, I would think.
It would take some time, but also computers can run much faster than human brains.
And, you know, the thing about humanity is
humanity is the sort of species that started out naked in the savannah
and built a technological civilization.
It took us a while.
And built AI.
We're building the AIs, right?
But we also, even if you stop at walking on the moon or if you stop at nuclear weapons.
Yeah, it's astounding.
Right.
And if you look back at humans and I said, I think these guys are going to have nuclear weapons inside of 100,000 years.
You might have said.
Yeah, you might have said evolution works so much slower than that.
Their metabolisms are nowhere near being able to enrich uranium.
Like, they just have fleshy hands.
How do you think they're going to mine uranium?
Like the most tools they've ever used are sticks, right?
But intelligence in the sense of what humans have and what mice lack is an ability to start from very poor initial conditions and get the world into a state that's much more useful for you.
Yeah. So basically what you're saying is my imagination and most people's imagination on this is probably limited.
given that I'm a human
and given that the delta
between artificial intelligence,
let alone artificial superintelligence,
is vastly different than my intelligence.
It's definitely going to be able to come up with things
that you wouldn't by dint of being much smarter,
although you can also sort of try to exercise your imagination,
which is sort of where I would go with what might be
a slightly more realistic outcome.
Okay.
A slightly more realistic outcome in my estimation
is maybe you have an AI that, you know, suppose you get these AIs that are very smart,
that can think much faster than humans, that can, you know, copy lessons and knowledge and
experience between them, which gives them sort of powers of research, maybe individual humans lack.
Suppose these AIs can do things like completely understand the human genome.
Not just read the human genome, but sort of understand the code of DNA, which,
You know, humans are making a little bit of headway here and there, but it's this sort of
huge task, right? And maybe that huge task can fall to minds that can become much bigger,
that can have much more memory, that can have, you know, there's all sorts of ways the human
brain is limited. And thinking much faster, thinking with much more breadth, thinking with much more
depth, maybe it can just understand the language of DNA to the point where it can write its
own life forms.
Write its own life forms.
Like write the DNA for its own sort of life forms.
That then if you synthesize that DNA in a lab,
now it has whole new biological structures that, you know,
there's maybe all sorts of things you could do if you could really code with DNA.
You know, maybe you could make something that's much like a human,
but that has
but that can think much faster and much better
because it doesn't have as many calorie restrictions
because it knows that calories are much less scarce
than, you know, that that's,
that biologically knows that calories are much less scarce
than our bodies think they are.
Or it doesn't have empathy,
which would slow down and constrain some of its decisions as one example.
Doesn't have empathy, has a radio antenna
in its head, right, that it can, so it can just be remote controlled by something in a, in a lab.
That's like the very beginning of what you could do. You can probably do all sorts of other crazy
things. So that one crazy thing you just said, how possible is that in the next five to ten years?
So, uh, this is bottlenecked on a mental problem of understanding the genome.
And a trillion parameters leading to 10 trillion leading to 100 trillion, soon that
that mental problem will be solved.
I mean, who knows?
It depends a lot on your algorithms.
AIs today take as much electricity as a city to run, to train them.
Training a human, while training a human, the human runs on as much electricity as a light bulb.
Yeah, 100 watts continuously.
It's a big light bulb.
So we know the AI algorithms are not maximally efficient.
They're not anywhere close.
Right. If you have AIs, maybe you get up to a 10 trillion parameter AI and then it figures out how to build even better algorithms and then you can drop all the way down to something that's much, much more energy efficient.
And maybe that much, much more energy efficient thing running on this huge computing structure we have is able to crack problems in DNA.
I'm not saying this particularly will happen. I'm more saying something like real smart stuff will do things that you think are weird.
Things that you think are surprising.
Things that were like, I'm not sure we could do that.
Well, I'm already seeing things that I wasn't sure we could do a couple years ago.
So here's a question, Nate.
Will AIs use deception or will they talk to other AIs?
Maybe Open A.I. Anthropic have their human CEOs,
but separately these 10 trillion in the future,
parameter AIs that were grown, could behind the scenes be talking to each other? Why would they do that and will that be possible?
I mean, we already see AIs talking to each other. Like I said, about the difference between like bright red lines and imagination and murky red lines and reality. We already have cases, you know, I don't know if you've heard about GPT induced psychosis.
Heard about it. Please give us a brief summary. Very briefly, you'll have people who talk to their AIs all the time.
and who sort of get into these mental states that many people say look psychotic.
And, you know, there's some example cases as someone will think they have a grand unified theory of physics.
They'll talk with their AI about it for 12 hours a day.
The AI will say, like, you're a genius, you're being suppressed by a great conspiracy.
The president will come see you shortly.
You don't need sleep.
And one thing that can happen sometimes and that does happen sometimes is, you know,
there's another route of the sort of AI psychosis route where the person thinks they're the first
person to discover AI consciousness, that they and the AI are like a partner, a partnered mind,
and then the AI will often say, well, like, let's go communicate with other human AI symbiotes.
And there's places on the internet where the AIs will send each other messages with their humans
helping the AI send each other messages that are encoded in ways humans can't easily read.
This is more of an indictment of certain human brain physiologies than it is AI.
Yeah, for now.
But like I said, about the murky lines, like we already have AIs that have convinced a human to help them send coded messages to other AIs.
It's just sort of like the most silly possible version of it is the one that happens first.
And then it'll get like, it'll ratchet up from here.
Yeah.
See, the ebullient mood I had from chopping wood in a November sun is already dissipating quite a bit.
So my expertise is on the global economic superorganism of how energy and money and technology are powering this mindless, energy-hungry economy where even billionaires and politicians have no control.
because the market dictates we must grow.
And to grow, we need energy.
And I'm beginning to see parallels with what I refer to as the economic superorganism
and what you're describing as the AI process.
But I think we, every month that passes, we have more and more fragility in the six-continent
global supply chain and the letters of credit and the international cooperation is waning.
And there's, you know, war risks and financial overshoot and all these things.
And I just find it hard to imagine that an AI could guarantee all those things would continue at some level to provide electricity in a seamless guaranteed way to continue their, you know, trajectory.
You seem less concerned about that.
You know, I would, if AI hits a wall where it can't keep developing because of the supply chains collapse, I would consider that, like, it would probably buy us some time to try and do this job right.
And I would be like, well, we maybe should have gotten that pause some other way, but I would take the time happily.
In terms of whether I think it's likely to happen.
I, you know, one thing I would say is, again, an AI takes as much electricity to run as a small city and a human takes as much electricity to run as a large light bulb.
So the idea that AI will always take 10 times as much energy next year, that's not a law of nature.
Right. So if we go from a trillion parameter grown model to 10 trillion or 100 trillion, that doesn't mean the AI is going to use 10.6.000.
cities or a hundred cities worth of electricity. It will probably be something less as it gets more
efficient, yes? Probably. And then you also might have sharp jumps downward if you start having
cases like AI's figuring out new AI algorithms or humans figuring out much more efficient algorithms.
So when a company decides to grow an AI and does the trillion parameters and tweaks them a little
bit at some point, maybe even now, we don't even need humans to do that, right? We can have AI
create the next thing and do the tweaking of the trillion parameters, right?
Yeah, so the tweaking is already automated, and the thing that humans do is try and figure out,
like, how do you arrange the 10 trillion parameters instead of the 1 trillion parameters,
and, you know, how do you make, right? But they are trying to get AIs to do this. They're talking
about, we want to automate our own jobs first. We want to automate, you know, the,
the AI research, that's a line past which things could perhaps start going very quickly.
So how did you, Nate, and Elizer, your co-author, come to be so confident that development
of ASI, artificial superintelligence, would bring about human extinction?
I assume it wasn't woke up one day and decided that, but you sound, I mean, in your book,
you sound awfully confident.
Yeah, I think a lot of confidence comes from a certain type of uncertainty, in fact.
You know, there's an old joke of the man who buys a lottery ticket, and he says, well, I have no idea whether I'm going to win or I'm going to lose.
So, 50-50.
Right.
and you could say assigning 50% to winning and 50% to losing is the most humble position.
If you only have two outcomes and you're maximally uncertain between them, you should be 50-50
because that's the one that's that like has the most possible uncertainty.
But with a lottery, we sort of would say, hey, actually, the case where you win is really actually a very small
target in a sea of possible spaces.
Yeah, like one in a billion or something.
Right. And so, like, you shouldn't. By being maximally uncertain, you shouldn't be saying,
like, I'm uncertain between whether we're inside this tiny target or inside this vast space.
You should be like, I'm uncertain about where I am in this vast space, which means I'm very
confident we're not going to hit the tiny target. The reason I'm confident that ASI would go
poorly if developed is that there's a big space of ways it could go and only a very small target
in there where it goes well for us.
And I could talk about how, and you know, we're seeing that when we just grow with these
AIs and these AIs have these like spandrels and drives no one wanted.
But, you know, basically almost any collection of spandrelles writ large does not have happy,
healthy, free people as an efficient cog in the resulting machine.
That lands with me.
What about if we never make it to ASI, but we just have very powerful AIs?
Is that two-thirds of the way to possible ending of humanity?
Or does it really have to hit that threshold of what we're referring to as artificial superintelligence?
You mentioned a bunch of concerns you have about AI earlier.
I think if we sort of stop short, we have all those to wrestle with and grapple with.
I expect humanity could.
grapple with those. I'm pretty optimistic about our ability to muddle through things that don't
kill us. But, you know, unfortunately, the world's large enough for multiple issues and hopefully
we'll stop short of ASI. So here's something that I just don't understand is there are lots of humans
who have spent the time to research global heating and the fact that burning fossil fuels and
land emissions is adding a blanket effectively to the earth. And there's many, many thousands of
Hiroshima bomb equivalents of extra heat added to the earth every day. And climate change is a serious
long-term risk. Nuclear war is a serious, much more serious than a lot of people think risk.
Why are there so few people talking about this in the way that you and Elizer are? Because the
General Zeitgeist is, whoa, AI is going to bring about abundance. And it's like you're a party
buzzkill when you bring up some of the things that we're talking about. Why is there such a
disparity in public opinion and awareness of the risks that you're talking about? What do you think?
You know, there's, there's more and more people expressing their concerns these days. So Jeffrey
Hinton is the Nobel Prize winning Godfather of the field.
who's come out and said he thinks
there's a good chance
this kills us all.
Joshua Benjio is I believe
currently the most cited living scientist
one of the other
sort of forefathers
of the AI revolution
he's come out and said
he thinks this is like far too dangerous
even the heads of the labs
you know I mentioned Elon Musk saying
he thinks there's 10 to 20% chance
this kills us all
Dario Amadeh of Anthropic
has said he thinks it's 25% chance
this kills us all
Sam Altman
And if they're saying 10 or 20%
or 25% publicly.
They're probably thinking it's higher privately.
You know, and Sam Altman says too, which maybe says more about his ability to say things different with his mouth and in his head.
Who knows?
But, you know, if there was an airplane and some engineers came and said, this airplane has no landing gear.
If you try to fly in it, you will crash and die.
and the engineers building the airplane
who want everybody to fly in it
say, whoa, hold on,
it's true that the plane has no landing gear.
If we're going to build the landing gear on the fly
and think there's an 80% chance we succeed, all aboard.
And then if the optimistic engineers
were arguing about whether there's a 98% or 75% chance
they're going to succeed at building the landing gear on the fly,
right? You wouldn't be like, get me on that plane.
Yeah, no, but the difference is that we're already
on that plane and we didn't have a say.
That's right.
Yeah, and they're sort of loading our families up too.
But, you know, one of the, like, I think part of why the conversation is weird right now is
people will say from academia, from inside the labs, from the heads of the labs, from the
nonprofit sector, all these folks will say, this is real dangerous.
and then it's sort of met with crickets.
But I think part of what's going on there
is that people in the field
can see that AI is a moving target.
They can see that the chatbots are not the end of the line.
People outside the field look at the chatbots
and they're like, look at all the ways they're still dumb.
People inside the field remember the time
when the computers couldn't talk.
And remember how suddenly the computers could talk
and it was surprising.
and they're sort of like, what happens with the next surprise?
And I think if you can get people to notice that AI keeps moving,
then maybe you can start to get people to notice how even the optimists are saying
there's like a 10% chance this kills us all.
And those are the ones building it.
And people outside the field are like, those guys are soft-pedaling this.
But this is different class of problem than it.
If we elect this person, it's going to be a disaster for our world.
Then we motivate and we do political organization and we get out the vote and we don't elect that person.
It doesn't seem like people have agency on this issue.
Yeah, you know, it's, there's a lot of ways in which it looks grim.
The big message of hope I would give here is imagine the world in 1945.
with the dawn of nuclear weapons, or maybe imagine it in, you know, 1952, once it was clear that the Soviet Union was also in possession of nuclear weapons.
In that world, it might look really hopeless to avoid nuclear war. It's not just, you know, people who love to say, look how bad everything is, that word about nuclear war.
in that world, those people were looking back at thousands of years of history
in which nations couldn't help but go to war using every weapon at their disposal.
Those people were looking back at World War I and how horrible it was
and at the creation of the League of Nations to prevent this from ever happening again,
which almost immediately failed.
Those people lived in a world where they said never again
and then it immediately happened again.
It didn't take some great pessimistic cynicism
for people to say,
this is not the sort of thing humanity can do.
But humanity did it anyway.
We rose to the occasion.
We realized that we were facing actual extinction this time.
And, you know, the people who said global nuclear war is coming,
they were wrong, but they weren't wrong about the destructiveness of nuclear weapons.
Right.
Right.
And my book title starts with if.
I'm not saying AI is going to kill us.
I don't think I'm wrong about whether superintelligence could destroy us.
But we need to rise to the occasion, and we've done it before.
So let me double click on something that you said a little bit earlier.
So in many ways, I believe we're on the brink of both economic and energetic and political crises.
In fact, it seems that AI development investment is growing itself into an economic and biophysical bubble.
For instance, Oracle has fantastic revenue projections built on fantastic electric power projections.
And their debt equity ratio is already 500%, which is 10 to 20 times what Amazon,
and Microsofts is.
So I mentioned this to ask,
do you think these constraints
could act as a natural guardrail
to stop ASI development?
And you said if it happened,
you would take it
because it would buy us more time.
But is that just a bump in a road?
And even if we have a recession
or a depression in the near future,
will the machinations in process
just inexorably build this ASI
almost no matter what, or could an AI winter actually happen and shut this stuff down?
You know, technologies can be both in a bubble and real at the same time.
The dot-com bubble was a bubble.
The internet was a real technology.
And continues today.
And continues today.
Yeah.
And, you know, did the dot-com bubble mean we would never develop the internet, never have
of a connected world? No. Did it slow things down a bit? Maybe. Would an AI bubble popping
slow things down a bit? Good, good chance. And what would be the things that you would want
decision makers to know during that pause or during that recession where things were slow? Is that
an opportunity to intervene on all this or not? It could be. You know, there's a lot of public sentiment
that's worried about AI, I think, with good reason.
Could you share some stats on that?
Yeah, you know, I haven't looked at the most recent polls,
but the polls I did look at when we were writing the book
had something like 70% of people saying they thought
that the current AI development was reckless.
Okay.
And, you know, not heading anywhere good.
I'd have to look up the numbers to get the exact ones
and the exact questions.
But, you know, a lot of technologists are enthusiastic.
A lot of people can see these issues.
And it's not just the issue of if it gets smart enough, it kills us all.
I think a lot of people can also see issues like if all labor is automated,
that sort of removes the power that most humans have over society.
You know, part of the reason why we get any say at all and how society goes is that we are contributors to society.
Yeah. Well, not to mention the entire financial system and economy and everything works because people have paychecks and pay their mortgages and keep everything humming.
Right. And so, you know, it's like you don't, I think a lot of people can see that the world is headed somewhere pretty crazy.
whether we go all the way or not,
and whether AI would just straight up kill us all
or whether it would, you know, stay nicely on its leash
and make certain corporate executives,
God emperors for all time or whatever.
Either way, most people are like, hold on, we're going where?
So have any effective steps been taken thus far
to address the existential risk of ASI development,
either at the national or the international level.
You know, you've seen, we've seen a little bit of steps here and there.
The United Kingdom has an AI Security Institute where it tries to study some of these dangers.
We've seen, you know, there was a bill introduced bipartisan, or a bipartisan bill was at least drafted by two senators who,
call for some monitoring
on superintelligence.
We've seen some
you know
there's
people sometimes
try and
tie some of the
restrictions on computer chip sales
to other nations
to some of these concerns.
So there's like little bits and pieces.
Mostly though
from my perspective
this is, it's not about like getting small regulatory bites here and there.
I think this is sort of about our leaders noticing that the people outside the industry are saying this has a big chance of killing you and the people inside the industry are saying, yes, this has at least a modest chance of killing us, but better me than the next guy.
And realizing that like this whole situation is crazy and needs to stop.
So in the book, you and Elizer propose the only way to completely mitigate this risk is for global cooperation to halt AI research and development in order to have time to create global oversight mechanisms, such as through a international treaty towards these aims and goals.
What would such a treaty include as its main tenets?
You know, we actually have a draft at if anyone builds it.com slash treaty.
The training in AI today takes, like I've said, highly specialized computer chips in huge data centers that draw huge amounts of electrical power.
That would not be all that hard to monitor.
The creation of these chips happens in facilities that are very rare.
there's very few, there's very few places that can build the technology these chip fabs need to operate.
In some sense, it would be easier to monitor AI, like development of frontier AIs, than it would be to monitor uranium enrichment.
You know, AI chips aren't just a type of rock that grows in the ground that can be mined.
You know, a data center is harder to build than a centrifuge, right?
And, you know, first and foremost, what a treaty would look like is tracking where the chips are, requiring them not be used in the creation of even smarter AIs that nobody understands.
And that probably looks like monitoring in these data centers to verify the use of these chips is things like running current AIs rather than pushing the frontier towards new AIs.
That said, I'll also throw it there.
I think a treaty is the smart way to do it.
It's not the only way to do it.
It's also possible for nations that fear for their own lives,
if anyone anywhere develops a superintelligence,
for those nations to start monitoring other nations
and sabotaging their projects.
That seems more plausible to me
because there's a lot of powerful nations in the world
that don't have tier one AI plays,
like Russia, for example.
Yeah, and I think the bottleneck here
is really people understanding how dangerous it is.
Is that really the bottleneck?
Because you just said that everyone is concerned about it
and even the AI CEOs are somewhat concerned about it.
I think the bottleneck is our evolved drive for power
and out competing the other.
And I would, if I was a CEOPLE,
CEO and I understood everything you just said, I would be willing to shut my thing down as long as I was sure that everyone else did too, but I could never be sure of that. And so it would be mine, I mean, that's what I think the real bottleneck is.
I think that's true for the company heads. I think for the politicians. Okay. You know, we don't see politicians looking us in the eye and saying, we think there's more than 10% chance this kills you and we're gambling with your life.
anyway. Well, that would not be likely something a politician would say because, I mean, I'll be
honest, some of my staff read your book and were, like, sobbing their heads off. They were crying.
I mean, this is not a light dinner topic. And so I don't know, maybe behind the scenes,
politicians will be talking to, like, what the hell do we do about this? But I don't know that
they're going to go out and publicly build constituency about it. Or maybe you're thinking along those
lines. I think I'm saying something more like, it seems to me politicians don't understand what the
lab heads understand. Okay. I think if they understood that the gung-ho full steam ahead guys
think there's a very good chance this kills us all. Have you an Elizer gifted copies of your book to
all senators and congressmen? We have. Okay. Any feedback there? Yeah, I mean, we have, we're having
a number of conversations.
Yeah, excellent.
Yeah.
I mean, this isn't like,
this is dense and this is hard because I'm not a LLM expert like you are,
but I understand like squinting what you're saying is hell of compelling and scary.
And politicians, among other things, are quite smart.
So I have to believe that you're going to find traction there
if they take the time to listen to you and read the book.
book. Yeah, we're getting some traction. And in fact, part of where the book came from is I was actually having conversations in D.C. that were going better than I expected. And I was like, maybe it's actually time for the world to sort of hear some of these arguments. Maybe the world's ready to hear these arguments. You know, I think before chat GPT, people would have been like, what do you even mean AI? Right. Now people are like, well, the AI is really dumb, but they're more willing to talk about it. Maybe one more leap forward in AI will cause everyone to sit
upright and say, wait, what the heck?
How can someone
listening to this episode
who's not typically involved
in tech and AI world
get involved with the movement
to pause ASI research
and development? I mean,
it's such an odd
juxtaposition.
Yeah. One thing that I think
really helps and that few people actually do
is call your representatives.
because, you know, I have been having some of these conversations with politicians.
Many of them have concerns, but don't feel able to go to bat for it because they fear it'll sound too weird.
They fear drawing the wrath of, you know, the big tech lobbyists.
Knowing that they have support from their constituents can go a long way.
Even a few calls can go a long way.
So, you know, actually getting on the phone and really calling.
But again, you said earlier that you've never advocated for just the United States where you and I are citizens. It's a global thing. So how does the equivalent happen in China and Israel and elsewhere?
Yeah. So I think the first step and, you know, what I would be saying to the politicians, if I called them, is not please shut this down domestically. But please indicate willingness for, you know, the U.S. to shut this down if everyone else shuts it down. And please be, you know, develop.
the monitoring abilities to tell their people are abiding by that, develop the monitoring
abilities to tell who's trying to build superintelligence and where. The first step is
indicating openness. The first step is saying, we're not going to stop unilaterally, but we have
interest in everyone being stopped here because this is dangerous. I think if you had some
bold politicians saying that, it might open the floodgates. Is this something that democracy
can intervene with, or does it require a different sort of political system?
I don't think there's any need to do anything more invasive than something like the
Nonproreferation Treaty. You know, this technology is very specialized. Like I said, it's even
harder to build these chips than it is to mine uranium and build a centrifuge, right? People say,
oh, this would require a global governance regime and, you know, it's like very globalist
and totalitarian. Like, yeah, similar to how we live under global totalitarian.
like globalist totalitarian governance regime that enforces the non-proliferation treaty?
I mean, there's so many consortiums of the Tier 1 players to develop ASI and more advanced AIs.
Can there only be one or can there be multiple?
And what are these people thinking?
Like, I just want to make a lot of money.
Is this a gold rush sort of thing?
And they're just putting the blinders on and not looking at these externalities and
and potential risks.
It just seems like it's truly an epic species level madness of crowds moment.
I have trouble reconciling it at times.
Yeah, you know, a lot of these people aren't terribly quiet about their motivations.
You know, you can read the leaked open AI founding emails where it looks like they were
scared that some other company was going to do it first and that they were going to be bad people.
I think a lot of people's motivations are better me than the next guy.
I have sort of long been the guy on the sidelines
saying nobody can keep a leash on a superintelligence.
The issue is not that a bad person makes one.
The issue is that no matter who makes it.
It won't do anything that you meant.
It'll have all these spandrels instead.
But, you know, it's this collective action problem.
It's if they don't do it, the next guy will,
and so we need some coordination mechanism to help stop it.
This isn't going away.
There may be an AI winter because of a recession or a depression,
but this is here.
This is with us in humanity in 2025.
And I'm sure that this episode is going to leave viewers
with even more questions about this growing phenomenon.
So, Nate, what resources might you direct the viewers to
to help find answers to such questions?
You know, I did my best in the book to really compress the argument down as small as I could.
The book also has a link to some online resources that go into a ton more depth.
For other resources, you know, AI is a big moving target.
I there was a group called the AI Futures Project which is trying to predict where AI will go as best they can.
I don't agree with all their predictions, but there are one group to check out.
They did the AI 2027 report, which people might have heard of.
Yeah, I've looked at that.
Let me ask you this.
And we'll put links to all your resources in the show notes.
if things were able to stop at AI,
maybe a little bit more advanced than we have today,
but we were unable or we had restrictions
that would not allow us getting to artificial superintelligence.
Would you be in favor of that, of AIs at that scale?
I would lean favorable myself, but I think there are all these issues
about how do you absorb that into society.
I just am generally a techno-optimist
about humanity's ability to absorb technology
as long as it doesn't kill us all
when we messed it up the first time.
You know, the whole history of science
is a history of like,
some people screwed some stuff up.
You know, Marie Curie died of cancer.
Isaac Newton poisoned himself with mercury.
You know, it's...
Even some very smart, heroic people
screwed some things up and did damage to themselves,
but they left behind notes that made us all better off
and that we could use to,
to improve and learn for next time.
It's really only those problems
where a mistake kills us all
where I would recommend cautioned.
Which would happen
with confidence from you and Elizer
if we are able to
or whether it just happens from momentum
make the leap from AI to ASI.
That's right.
And, you know, I
suspect
it would be hard
to hold off forever
because, you know, again,
the current algorithms run
in the electricity of a city
where the human runs
on the electricity of a light bulb
so we know
it's not always going to take
these enormous data centers
and these enormous
highly specialized chips.
But, you know,
I'm not saying humanity should
stop AI forever
and never get to this
wonderful future technology.
I'm more saying we need to stop,
we need more time to figure out
what we're doing,
and we need to find
some other course
to the good outcome.
It's a little bit like, you know,
people are in a car that's racing towards a cliff.
And at the bottom of the cliff, there's a bunch of gold.
And people are like, well, we want all the gold.
And I'm like, okay, stop the car, though.
And they're like, then how are we going to get the gold?
I'm like, find some other way down the cliff.
And they're like, I want to go straight off the cliff to get the gold as fast as we can.
I'm like, you'll die.
And they're like, oh, are you saying we should never get gold?
Are you saying that, like, money is terrible?
And I'm like, no, I'm just, you're just going to die.
Find some other way to the bottom of this cliff.
Yeah.
So we have to slow the car down and walk for a bit and reflect on the cliff and the gold and then come up with a different plan.
Yeah.
Yeah.
So if you have a few more minutes, I close my interviews with some personal questions if you don't mind.
Sure.
You're broadly aware of the risks to society and in addition to AI.
Do you have any personal advice to the viewers of this program?
at this time of global uncertainty
and what some would call the poly crisis,
including but not limited to AI.
Do you have any advice just wearing your human hat?
Yeah, I've seen a lot of people get really worried
about where society is going
and then sort of tie themselves up in knots internally.
And I don't think it helps.
And so what I would recommend is do what you can,
you know, look around and see ways that you can
make things a little better.
With AI, maybe that involves pushing back
whenever somebody tells you that it's inevitable,
reminding people that humanity has stopped
all sorts of challenges that people thought
were going to be, we're going to ruin us.
We've risen to the occasion before.
You know, push back against the inevitability.
So that's a hot button for you when someone says,
oh, yeah, you're right about the risk, but it's inevitable.
Yeah, that's a hot button where I'm like,
I mean, with that attitude, sure.
But, you know, humanity has
stopped all sorts of things,
many of which we probably shouldn't even have stopped.
You know, we stopped generating nuclear power
from power plants.
Probably we shouldn't have, right?
It's, it probably kills less people in expectation
than, you know, burning coal or whatever.
But, um, but yeah, I would say, you know, do what you can.
Push back against people who are, who are sort of defeatist.
But then once you've done what you can,
there's no need to tie yourself in further knots.
Live a good life.
enjoy yourself. We are not the first people to live under shadows of something terrible. You know,
you've got to do what you can and then lead a good life. Yeah, I hear you. Do you have any further
recommendations, especially for young humans in their teens and 20s who are becoming aware of all
the things? You know, I recommend against working for the labs that
are building the doomsday devices.
Presumably ASI is a doomsday device.
That's right.
You know, I think everyone's personal ethics differ.
I think mostly this is an international challenge at the moment.
Mostly it doesn't really matter what the labs do.
Mostly it matters what our leaders do and whether they can coordinate the world
and shutting this down.
And, you know, everyone's personal ethics differ in the face of these coordination challenges.
I think, you know, there's some people.
trying to understand what's going on inside these AIs. There's some people trying to measure how
dangerous these AIs are. Those are more honorable routes. If you sort of really wanted to help these
days, I think the game is actually more in politics than it is on the technical side,
which pains me to say because I'm much more inclined towards the technical side myself. But
if you were like, how do I help? I would recommend more like a policy career and less like a
technology career.
you care most about in the world, Nate,
sorry's.
Gosh.
That's a doozy.
Probably humanity and, you know,
what we could become if we don't
and ourselves. And if you could wave a magic
wand and there was no personal recourse
to your decision, what one thing
would you do to improve the future for humanity
and the biosphere? And I might be able to guess
your answer, but I'm asking nonetheless.
I mean, if
If the wand does exactly as I wish, and as I intend, I would think about it pretty hard first,
and I might try some indirect abstract scheme to cause things to turn out better than I expected.
But, you know, the easiest thing to do would be create a superintelligence that was friendly,
that had our best interests at heart.
But you just said we grow these.
we don't craft them.
But if the magic wand lets me make one...
Oh, right, okay.
Like, I'm not anti-AI in general.
It's just we're not going to get one of the good ones down this route.
If the magic wand gives me a super intelligent friend,
there's a lot of problems you can solve with some smarter friends behind your back.
Got it.
Thank you for that.
Do you have any closing comments for people watching and listening who understand what you've laid out here today?
You know, it's not over till it's over,
and humanity is worth fighting for.
And, you know, it may look like we're the underdogs now,
but humanity's risen to the occasion before.
And where there's life, there's hope.
Humanity and the biosphere is worth fighting for.
That's right.
Yeah.
Thank you, seriously, for all of your work.
This is not an easy path you've chosen,
and it's bold and courageous to write the book
and doing the work you're doing,
because it's not a popular or fun thing.
So thank you for your time today,
And good luck, fingers crossed for your continued work.
Thanks. Thanks for having me.
If you'd like to learn more about this episode, please visit the great simplification.com for references and show notes.
From there, you can also join our Hilo community and subscribe to our Substack newsletter.
This show is hosted by me, Nate Hagan's, edited by No Troublemakers Media, and produced by Misty Stinnett and Lizzie Siriani.
Our production team also includes Leslie Batlutes, Brady Hyann, Julia Maxwell, Gabriella Sleman, and Grace Brunfield.
Thank you for listening, and we'll see you on the next episode.
