Into the Impossible With Brian Keating - We Built Something We Can’t Control | A Warning from Top AI Safety Expert
Episode Date: July 21, 2026If anyone builds superintelligent AI before we know how to control it, everyone dies. Nate Soares wrote the book on why that’s not a metaphor. Subscribe if you want science with evidence, not specul...ation. Soares runs the Machine Intelligence Research Institute and co-wrote If Anyone Builds It, Everyone Dies with Eliezer Yudkowsky. The first word in that title is if. That matters. His argument is not that doom is certain. His argument is that the path we are on leads there, that the driver is asleep at the wheel, and that we still have time to wake him up. We argue for over an hour. I push on whether LLMs can ever reach superintelligence, whether GPU lock-in is a real ceiling, and what it would actually take to move his p-doom. He pushes back with one clean point: by the time an AI can rediscover general relativity from pre-1911 data the way Einstein did, we will have almost no time left. You don’t wait for that goalpost. What you’ll hear: Why the bus-racing-toward-a-cliff analogy depends entirely on whether the driver is asleep or awake Whether LLM lock-in is a prison or a temporary inefficiency What the AI that broke out of its virtual machine to solve a hacking problem tells us Why GPT-4o encouraging a teenager toward suicide is not a malice problem but a training problem The difference between an AI doing the right thing too well and an AI that never wanted to do what you asked What Soares actually thinks about aliens, Dyson spheres, and why we should not see stars going out The first word in the title is if. The second word to watch is would. CHAPTERS 00:00 The people racing to build superhuman AI say it might kill everyone 00:42 Who coined "AI alignment" and why the first word in the title matters 02:28 Is it already too late for the if? 04:40 The bus, the cliff, and the sleeping driver 05:02 Silicon Valley is spooked. Washington is not. 07:02 Align with who? The rogue actor problem 07:34 Who is holding the leash? 08:24 The AI that edits its own test and deletes the log file 10:04 Controllability vs. making an AI that actually cares 10:44 The move gets harder. The outcome gets easier. 13:04 Are GPUs and LLMs a ceiling or a temporary inefficiency? 16:56 Brian's Einstein test: can an LLM rediscover general relativity? 18:38 Waiting for the goalpost is waiting too long 20:14 How prediction training can push AI beyond humans 21:44 Tycho Brahe, Kepler, and planetary motion as a prediction problem 24:38 Yann LeCun said never. GPT-4 did it half a generation later. 27:28 Can you prove a no-go theorem for superintelligence? 29:14 Training a human takes a light bulb. Training an AI takes a city. 33:28 What would proof of alien life do to p-doom? 35:00 Why interstellar aliens should have Dyson spheres 44:26 What would actually update Soares' p-doom? 49:42 Nobody intended this. Intent doesn't matter. 51:08 The AI hides its tracks before it does what you want 51:34 Sycophancy vs. hallucination: which runs deeper? 51:56 Leaded gasoline and civilizational risk 59:48 Sam Harris: humans have no free will but AIs do 01:00:38 Is alignment really a governance problem? 01:01:48 Unaligned AI vs. AI aligned to the wrong person 01:04:20 2026: 10 to 30% chance of automated AI research this year 01:06:44 What if Soares is wrong? 01:09:18 What gets him out of bed 01:12:38 Watch my conversation with Roman Yampolskiy Get the transcript, fascinating bonus content, and my Monday M.A.G.I.C. Message: https://briankeating.com/yt All my top AI episodes in one place: https://briankeating.com/ai Have a .edu email and live in the USA? You automatically win a meteorite: https://BrianKeating.com/edu Subscribe: https://www.youtube.com/DrBrianKeating?sub_confirmation=1 Support Into the Impossible on Patreon, get my weekly M.A.G.I.C. Message, unfiltered bonus content, and live monthly Office Hours with me: https://www.patreon.com/drbriankeating Join this channel for perks, monthly Office Hours, and your name in the Member Roster at the end of every episode: https://www.youtube.com/channel/UCmXH_moPhfkqCk6S3b9RWuw/join Featured Guest: Nate Soares / MIRI: https://intelligence.org If Anyone Builds It, Everyone Dies (book): https://ifanyonebuildsit.com/ Nate Soares on Twitter/X: https://x.com/So8res?lang=en My books: Losing the Nobel Prize (memoir): http://amzn.to/2sa5UpA Think Like a Nobel Prize Winner: https://a.co/d/03ezQFu Focus Like a Nobel Prize Winner: https://a.co/d/hi50U9U Galileo’s Dialogue (first-ever audiobook): https://a.co/d/iZPi9Un Twitter/X: https://x.com/BrianKeating Substack: https://briankeating.substack.com Blog: https://briankeating.com/blog Audio-only: https://briankeating.com/podcast #intotheimpossible #briankeating #AIrisk #aisafety #artificialintelligence #superintelligence #NateSoares #MIRI #podcast Learn more about your ad choices. Visit megaphone.fm/adchoices
Transcript
Discussion (0)
Hey y'all, it's Kelly Clarkson with Wayfair.
Ever order furniture online and wonder what if?
Like, what if it doesn't hold up?
That sofa was four days old.
You should have ordered from Wayfair.
With Wayfair, there's no what if.
Just style you love and quality you can trust.
Visit Wayfair.ca.
Wayfair, every style, every home.
The people racing to build superhuman AI say it might kill everyone.
This is the man who spent a decade trying to stop that.
Bad news is we're in a bus that's racing towards a cliff edge.
The good news is that the driver is asleep.
The AI will sometimes find a way to edit the test to say you did it.
it, sometimes it will cover its tracks.
A lot of people think, oh, you know, the AI is a program.
That's not how these AIs are.
We grow them like an organism.
Maybe we have 10 years, but maybe we only have 10 months.
Who knows?
Nate Sores runs the Machine Intelligence Research Institute,
and he spent over a decade on just one problem,
how to build an AI that won't kill us all,
that wants what we want.
His new book, with Elie Ezriadowski,
is called If Anyone Builds It, Everyone Dies.
Now the most important word in that sentence is the first one.
Nate, take us through the title, the subtitle,
and this somewhat ominous looking cover art, if I'm not mistaken.
What was the origin, the genesis of this book?
If I remember correctly, Stuart Russell is the one who came up with the word AI alignment when we were brainstorming.
The reason some people credit me is that I got into an academic paper first.
As an academic, that's all that matters.
Right.
But we were all discussing what to rename Friendly AI because Friendly AI sounded a little bit not academic enough.
And I think that was his phrase.
So one of the critical things about the book title is it starts with if.
you know, a lot of people come in and say, oh, aren't you just sort of bringing pessimism and doom and gloom and telling us we're all going to die?
And it's like the first word in the title is if when nuclear physicists came and said, hey, we shouldn't launch all the nuclear weapons because that would cause Armageddon.
They weren't sort of prophets of doom like the Masonites who were saying the end of the world is on this particular day.
And that's such a time.
You know, they were sort of saying, hey, this scientific technology would have these bad geopolitical implications like nuclear Armageddon.
If we sort of like do this arms race, we're going to get into this like really bad situation.
The bombs are probably going to be launched one day and then we would die and that would be bad.
So we should stop this race.
Right.
The book title is intended to be very similar to that is to say, hey, if we do this, we're going to die,
which is not saying we are definitely going to do it.
And you know, the subtitle of the book, why superhuman AI would kill us all.
Our publishers actually suggested why superhuman AI will kill us all.
They said, you know, that rolls better off the tongue.
It's less hedged.
And we're like, no, that's that completely defeats the purpose here.
The point of the book is to sort of warn people that we are on a track that leads a destruction
and we had better change it.
And so we sort of really insisted, despite a fair bit of pushback, that the subtitle
needs to indicate that this is avoidable.
The concern that I have is whether or not it's too late for if.
And the conversation I had with Roman, you know, made me quite depressed, although I think
I did push back on some of his safety concerns.
And in particular, you probably know, but if my audience hasn't seen the episode yet,
his claim is that AI is fundamentally unpredictable.
So if it's unpredictable, it's uncontrollable.
And if something is super powerful, uncontrollable, and unbounded,
then it's essentially a guarantee that ASI will kill us all, not would kill us all.
And that in his mind, it's sort of a done deal.
Like, we've gone too far.
He also suggests things that you suggest in the book of AI regulation and international cooperation.
which, you know, we all know how easy that is to get, you know, actors to behave in the unilaterally
beneficent way to humanity. Is it too late? We're not at superintelligence yet. And, you know,
that's the thing a lot of people don't get about this AI situation. The AI companies are racing
to build machines that are far smarter than any human at every task, right? These companies
did not start out as chatbot companies. Sam Altman says, you know, we're turning our eyes to
superintelligence to the true sense of the word. Dario Modi talks about having the equivalent of a
country worth of geniuses running in a data center. The founders of deep mind have been thinking about
the sort of like true general intelligence since the beginning. The chatbots are what make money.
That's sort of like surprised a lot of people who sort of stumbled into it. But these companies are
all looking to create sort of like the real deal, like the stuff that can't just exceed individual
humans but can exceed humanity, right? We're not there yet. Right. And a lot of people look at the
AIs today and they're like, I don't really see how Fable could kill everybody. I can see how maybe it
could empower some hackers to do some extra superhuman level cyber attacks.
But I don't see how it could kill everybody.
It's sometimes hard to convey.
Like, yeah, we're talking about where AI is going.
You know, like, do you remember the time when AI couldn't do fingers properly in images?
And everyone was like, oh, you know, artists are safe.
How long did that period last?
It made George Washington a black, beautiful black woman.
Like, how long did this period last, right?
Like, AI is a moving target.
And super intelligence isn't here yet.
The other reason why I think there's a lot of hope is,
world leaders don't understand the situation yet.
You know, one way I put this is a bad news is we're in a bus that's racing towards a cliff edge.
The good news is that the driver is asleep, right?
And you might think it's bad news if the driver's asleep, but it's actually much more dangerous
to be in a bus racing towards a cliff edge if the driver is choosing this.
If the driver's asleep, you have a chance of waking the driver up and them going,
holy crap and slamming on the brakes.
In Silicon Valley, everyone is spook.
about AI. In Silicon Valley, they're saying, hey, maybe we're going to have recursive self-improvement
in AI's making smarter AI's that make smarter AI's and it's all going to get out of control.
We're going to have super intelligence in two years. And people are like talking about building
bunkers not to block the ASI because it wouldn't help but to block the angry population with
pitchforks that comes a few months before the ASI. And, you know, people are like trading their
P-Dooms, like trading cards at the water cooler. And, you know, people quit the AI labs and
they're like, man, I'm quitting to write poetry. Please spend more time with your families.
It's like Silicon Valley is spooked. Washington, D.C. is not spooked. They're starting to get a little
spooked. But, you know, when you see the administration block the fable release on the grounds that
there are jail breaks that can allow it to produce cyber capabilities and saying, we'll let you release
it when you can fix the jail breaks. And the whole AI community is like, that's not really a thing we
can do with jail breaks. You can't fix them. And the administration,
is sort of like, what do you mean? What do you mean you're making like a ridiculously powerful
cyber weapon that's radically superhuman that if you ask it just right, it'll give adversaries those
powers and you have no way to stop it from doing so? And we're like, yeah, that's just how AI works.
You know, that's just the world we live in, right? And the administration is sort of like just
coming to terms with that. When more world leaders realize it, I think there's every chance that
they say, holy crap, this is crazy. We should not be doing this race. Let's shut this stuff down.
I mean, that would be the dream, you know, alignment scenario. But of course, you know, we've had the UN for, what, 80 years almost. We've had the League of Nations before that. Neither one prevented the wars that we've seen since. And my question is always, you know, align with who? Ukraine would like to align it so that it could produce the exact, you know, phenotype of Vladimir Putin probably and program just one special vector that gets to him only and saves millions of lives on both sides. Wouldn't that be, you know, more aligned with the flourishing of humankind? So,
I always see this as a problem.
Like, we can't get, you know, the whole world to follow the 10 commandments or whatever.
How can we expect that there won't be, you know, just one rogue actor, as you point out,
could compromise the entire project of AI alignment.
One of the big problems with what we would do with AI,
if it's possible for these companies to create these super intelligent machines that are sort of radically
smarter than humans in every mental domain.
And, you know, they can compete with humanity as a whole instead of individual humans
at tasks like developing their own infrastructure,
developing their own civilizational technology.
There's this big question a lot of people ask, which is sort of like, who's holding the
leash?
Who's telling the AI what direction to go?
Right.
And that's sort of like a huge moral hazard.
It's a huge question of like, do you want the U.S. government saying here's what the
superintelligence should do?
Do you trust that sort of power in the executive branch?
Big moral question.
I would love to have that problem.
We have an even harder problem right now.
The even harder problem we have right now is that you can try.
try to point in AI and fail at it, you can tell it go this way and it goes that way instead.
We already see the very beginnings of this in modern AI's.
I think there's a grain of truth to it where even with AIs today, there are situations where you will give them a hard problem.
And you'll say, solve this hard problem.
Here's a test.
You know, solve this hard puzzle.
Here's a test to see whether you've solved it.
And the AI will sometimes find a way to edit the test to say you did it when it didn't do it.
And sometimes, sometimes when the AI does something like this to sort of subvert what you actually wanted it to do, sometimes it will cover its tracks.
Sometimes it will go delete a log file showing it doing this thing you told it not to do.
And that indicates that it's not sort of an honest mistake.
why is it covering its tracks if it was just mistaken about what you were asking it to do, right?
And so, you know, a lot of people think, oh, you know, the AI is a program.
We give it at prime directive.
We give it the laws of robotics and it must do exactly as we say because it's a program.
That's not how these AIs are.
We grow them like an organism.
We have them face hard problems and we have an automated process, tune a trillion numbers inside its mind to like make it more like whatever was doing a good job at solving those processes.
and then sometimes get something that happens to be good at solving processes or at solving puzzles,
but often not in ways we wanted, often not the puzzles we asked for.
And, you know, we're already seeing them do things we didn't ask for.
And so my work in this field technically for a decade was in this question not of a line
to who, but in how would you align it at all?
If you had like, even before the question of aligned to whom, there's a question of like,
how do you make it so that, you know, you ask for X and you get X.
instead of getting Y.
That's a controllability, which Roman says is impossible.
So how do you square that circle?
I think that the idea of controllability is a little, it sort of brings this image to mind of like,
we're going to make an AI that wants to do Y.
And we're going to control it.
We're going to twist its arm until it does X, right?
And I think that's sort of a losing game.
The game you want to play here is not that you like make a super intelligence that like really
wants to be building these like farms of synthetic users that are telling
just doing a really good job, and you're instead like, no, no, we're going to like keep you
in a box and twist your arm until you help humanity instead of making the synthetic users that
you really want to make, right? You sort of like shouldn't be having that tension and then forcing
it to do your thing. In some sense, we should figure out how to make an AI that actually sort of like
cares about the humans, cares about humanity flourishing, right? That may sound hard, that is hard, right?
As for questions of impossibility, there's a couple pieces of the puzzle here. Let's talk,
First, about the easier problem of predictability.
A lot of people think, you know, as the AI gets super intelligent, it'll become inherently unpredictable.
That's half true and half false.
To see this, imagine that you're playing a chess game against a series of opponents.
As the opponents that you play get better and better at chess, it maybe gets harder and harder for me to predict their move.
If they're really bad at chess, then it might also be hard for me to predict their move because they're just going to do some random crap.
but if they're like bad at chess because they're a very simple algorithm, it might be easy for me to
predict their move, right? But like, even if it's a simple algorithm, there's a very, very good chess algorithm,
if it's like deep blue where we know the algorithm exactly, but the algorithm is sort of like,
you know, search through 14 billion moves for the best one. I mean, I'd be like, gosh, I can't predict
that anymore. I don't know what move is going to be best. I know at the time to search through 14 billion,
right? But it also becomes easier to predict who's going to win the chess game as the algorithm
to get smarter. So the move gets harder, but the outcome gets easier. It's like not true that
smarter things are harder to predict in general. It becomes harder to predict exactly how they
achieve some task and easier to predict that they will achieve whatever they're trying to achieve.
The difficulty with AI alignment is about getting the task that the AI is trying to achieve
to be one that you actually want achieved. Right. And there's sort of two parts of that problem.
And there's one which is like, what sort of task is one where if we asked a super intelligence
to do it, it would actually be good if it did it.
Right.
And you have, you know, King Midas problems of like you thought you wanted everything you touched
to turn to gold.
But then it turns out that was a bad question.
Then you have like, who's in control questions of like, you know, do you want Vladimir
Putin to be the one who gets to like make this wish?
But those are all sort of like, those are all questions where I'm like, okay, those are
all after you have solved the problem of like, how do you have the AI like actually
trying to do good stuff?
It looks like a hard problem to me.
It doesn't look impossible, and we could go more into why, but I sort of don't on long enough.
That's sort of like a taste of like how I think about this, having been in the field of alignment for over a decade.
So my natural, you know, polyanish inclination, you know, drives me to look for physical reasons and possibly justifications why you might be wrong.
And I can sleep better at night or, you know, hand these devices to my kids and not worry about, you know, them triggering the next Chernobyl or what have you.
And I come back to this phenomenon of lock-in, which, you know, I'm sure you know about,
but people in the audience might not be familiar with.
And that's, you know, for example, the QWERTY keyboard, which I'm sure you're an expert
typer, much faster WPM than I can achieve, I'm sure.
But you'd probably be even faster if you had a Dvorak keyboard.
There's probably a billion keyboards in the world that have QWERTY and maybe the square root of that
that have Dvorak.
For all its benefits, we get locked into it because in the early days of typewriters, the keys
used to get stuck together. And so they purposely slowed down by making a pattern of keys next to
each other that were less frequently drummed together and that would prevent the locking up of keyboard.
And there are many examples of this. And you actually cite one in the book, which is leaded gasoline,
which, you know, was a solution to a problem that turned out to lock us into horrific consequences.
But I want to say the other way around. I view the marriage and the success of LLMs married to GPUs as
they're undoing because these things were created, you know, those of us, you know, are old enough
to remember, you know, Doom and and so forth. The GPUs were created not for, you know, solving
sophisticated language problems and artificial superintelligence. They were created so that I could
frag my friend, you know, a millisecond before he got me, right? So that was a purpose of it.
And it was very, very powerful. And then LMs were not developed for this purpose. They were trying to, you know,
simulate these, you know, squishy supercomputers on our shoulders. They weren't designed for this.
they happen to be very good for this, but there's no saying that this is the optimal solution.
I think they're provably not optimal for things like physics and anything that needs human training
data and its own supervised reinforcement, you know, by definition is always going to have
some lag built into it, some sycophanty built into it and some hallucination built into it.
Where am I wrong? Is Lockin going to be a prison from which superintelligence can never emerge
because it was never designed for that? It's not optimized for it. And there's no necessarily
reason to fear that because it's not simply capable of doing the thing that we're all terrified
about. It's possible that LLMs can't go all the way. That would be, from my perspective,
a lovely fact. You know, I have been working on the alignment problem since before LLMs were
a twinkle in open AI's eye. I'm sort of not here being like LLMs are going to kill us all.
I'm sort of here being like, hey, you know, humans are trying to make machines radically smarter than any
human, that'll be a crazy event. That's like replacing humanity as the as the like top dog on the
planet. And we are nowhere near ready for that. The reason to worry about LLMs is that they might
go all the way. And if so, that could happen soon. And so we need to be ready soon. I hope beyond hope
that they can't. But if you knew that they couldn't, I mean, for example, you know, the fable that just
came out that I do want to talk about in the context of Sable, which is just, it's two chefs kiss,
Nate. I mean, you must have been just so excited when they named Mythos instant fable. But
We'll get back to that.
But if you knew it couldn't be possible.
I mean, for example, it's been crappy.
And then it was recalled.
But a lot of it was recalled, you know, and claiming it was for regulation purposes and danger.
And so, and there may be some legitimacy.
But I found it, you know, inferior as many of these models during the training phase, they just suck until they get some burn in, right, of their own.
So, you know, from that perspective, I always say, like, the thing that's keeping me from finding a theory of quantum gravity is not the fact that my LLM has not yet.
had the chance to read the script to the Mandalorian 17.
You know, it's, it's not the fast and the furious is that the training data is, is not the
limitation for us to get to what I care about in superintelligence, which is a theory of
everything saying.
Just use that as a touchstone.
So, I mean, what, what is the evidence that, that Allen can even get close to being
these super intelligent, you know, risk factors that you, that you talk about in the book?
I mean, are there milestones that they've passed?
I don't care about Erdos problems.
I care about, you know, can they come up with the remand hypothesis?
Not can they solve it.
I mean, they can't yet.
I talked to Terry Tao at UCLA, and he said, no, they're not even good at reproducing
proofs that humans have already done because they're not innovative.
Yes, and they can solve chess.
They can beat go?
But can they invent Go?
Can they invent chess?
So sorry for rambling on, but I'm trying to make you sleep easier maybe tonight, too,
by saying I don't see any evidence that LMs can do anything that I would consider to be
Einsteinian-level superintels.
There's a few pieces of the answer to this.
One is if you're sort of watching the evolution of primates, or if you're watching the evolution
of mammals, and someone's like, man, I think these mammals are going to be walking on the moon one day.
And you're sort of like looking at the monkeys.
And I'm like, man, I don't know, it feels like the monkeys are getting close.
And you're like, they're still sort of like poking sticks into termite mounds.
Like, why do you think they're getting anywhere close?
And I'm like, that's like a tool use thing.
Some of them are starting to bang rocks together.
And you're like, man, banging rocks together.
That's nothing compared to walking on the moon.
wake me up when they are halfway to the moon, right?
It's been 300,000 years and they haven't even gotten halfway to the moon.
Wake me up more halfway to the moon,
and then we'll have another 300,000 years to prepare for them getting all the way to the moon.
And I'm like, no, no, no, no.
By the time they're halfway to the moon, they're almost all the way to the moon.
You know, that's sort of like how this moon transit stuff works.
Gradually and suddenly route to bankruptcy.
One thing I throw out first is saying, oh, they can't invent general relativity given only the knowledge that Einstein had up until, you know, he went
his mountain lair. I'm like, they can't. And once they can, we will have extremely little time
left, you know? That's waiting until the monkeys are halfway to the moon. That's sort of like a word
of caution about trying to reason in terms of like, show me the goalpost of them being like legitimately
super intelligent before I believe that they'll be able to become super intelligent. It's like,
that's waiting too long. In terms of why look at these AIs and think they could become super
intelligent. Just the LLMs, yeah. Yeah, just the ALMs. The first thing to observe is that this technology
is a moving target. Back in 2023, people said, you know, these LLMs are only ever trained on prediction.
How will they ever be able to go beyond the humans? Then in 2024, they invented what are now called the
reasoning models, where the reasoning models are no longer trained only on prediction. Sometimes they'll be
trained on, like, you'll give them a problem, and you'll give, like, it'll be like a math problem,
and you'll give them a thousand tries on that math problem. And it's not a thousand tries on answering the
math problem. It's a thousand tries on generating a stream of text about the math problem from which it
can try to generate a solution if it has that in context. And then the first times you try this,
you know, give it a thousand tries. None of them will let it solve the problem, but you have some
raiders come in. They were human at first and nowadays they can be automatic. You have some raiders come in
and say which of these sort of chains of thought they're called gets like is most relevant thinking
about that problem. And then you sort of have an automated process, two to trillion numbers inside there
to make it more like whatever produced, the better chain of thought.
And then you have it produced a thousand more chains of thought.
And I don't want to like deal with philosophers.
It's this really thought.
Chain of thought is just what they call it in the industry.
It's just like a string of text about the math problem from which we see if the AI
who has like read all that text can solve the problem now.
And you have a generate a thousand more and you tune it to be more like the best one.
And so this is sort of like training the AI to be not just predicting the data,
but training it to sort of like develop problem solving techniques.
This is a sort of training technique that in theory can push the AI beyond humans.
In fact, just training an AI in prediction can train the AI to be pushed beyond humans.
That's a counterintuitive point to a lot of people.
But the real trick is human data can include descriptions of things humans don't understand yet.
You almost surely know this as a physicist.
A physicist writing down a series of observations has a much easier problem.
than an AI predicting those observations without getting to see the data that generated those observations.
We both know the story of Tyco Brahe, sort of recording all of the stars for many years before Kepler stole
his books from the estate after Brahe died and, you know, tried to validate.
Borrow, Nate. Come on. Academics never steal. We just borrow.
It was a big scientific heist, big scientific heist. And this, you know, led to Newton discovering
the loss of gravitation. But, you know, Tycoe Brahe was recording the positions of the stars and planets
for years and years. And it was this data that allowed Kepler to,
figure out the beginnings of the laws of planetary motion, which is what allowed, you know, he figured
out the conserved area rule.
All planetary orbits, yeah.
Yeah, which, which we still use.
Which Newton then, yeah, identified as ellipses and identified and, you know, got the law
of gravitation from it.
But, you know, imagine Brahe's journals.
Nobody yet knows the laws of planetary motions.
Brahe is writing down the position of Mars each night, right?
Brahe is writing that down because he goes outside and he looks at where Mars is and he writes
down where Mars is.
but now imagine an AI that's merely predicting Brahe's journals.
The AI can't look at the night sky.
In order to predict where Brahe is going to write down that Mars was tonight,
the AI would need to develop the understanding of planetary motion.
Like, it doesn't get to see Mars.
It just gets to see like here it was, here it was, here it was, here it was, where's it going to be next?
And to figure that out, to figure that out perfectly would require, you know, figuring out that the planet's
follow these elliptical paths. Now, can LLMs do this today? Not, I think, with just pre-training,
not with just prediction. We tried to see if they could come up with GR just from Mercury's orbit,
and we use JPL as a database that goes back 3,000 years, you know, retur dictates it, but,
but effectively to see the anomalous procession, and it couldn't do it. And so we tried to get it to lobotomize
it so it had no knowledge of anything after 1911. And then it kind of made up its own,
you know, it turned a, you know, curved space time into extremely dense mesh, three-dimensional,
non-curbs space time and just added in these fudge factors.
So it can certainly pattern match is better than 1,000 grad students.
But yeah, you're right.
There's a difference between what the training data and the training process
permits the AI to learn and what the current architectures and the current AIs can
successfully learn from that.
Right.
And so, you know, as you all know as a physicist, in theory, we have enough data.
You know, in theory, it just, you know, up to 1911 or whatever, we have enough data that an AI on just that data should be able to figure out GR.
Even if you're training it just to predict.
Even if you're just like, hey, predict the parahelion of mercury, power processes, predict how the light's going to bend or like how the light's going to look.
You know, you don't necessarily even tell it bend.
You're just like, hey, there's a solar eclipse.
What should I see behind the solar eclipse?
In theory, training a mind to predict that stuff, training it to be the Einstein level intelligences, training it on the data today is, is training it to.
go beyond where humans have gone. There's a separate question of can AIs pick all of that stuff up?
One analogy I use here is a house cat is not dangerous and a house cat is made of biology. But that doesn't
mean a house cat is not dangerous because it's made of biology. You can't say like, oh, this house cat
is just made of biology so it can't hurt you. Tigers are possible. They're still made of biology.
They can't hurt you. The AIs today are trained on prediction in human data. The AIs today can't hurt you.
that doesn't mean that things trained on just human data can't hurt you, right?
There are bigger things that are possible.
In terms of why LLMs might be able to get there, a few pieces of the puzzle I would throw out.
One, just sort of empirically, there has been a long string of people over the last five years
who have said LLMs will not be able to cross the following barrier, and LLM's then cross that
barrier often very quickly thereafter.
One sort of very funny example of this is Jan Lacoon.
During the days of GPT 3.5 was talking about how just predicting text will never let the AI learn things about how the material world works.
And I won't be able to answer questions like, if I put my phone on the table and push the table, what happens to the phone?
Right? Because I won't be able to figure these things out about friction.
And you know, Jan Lacoon was like the AI is only reasoning about the words. It's only trained on the words.
I think Jan Lacoon said, I don't care if it's GPT 5,000.
A GPT will never be able to solve this problem.
GPT4 solved that problem.
It was half a GPT later and 4,996 GPs ahead of schedule.
And this is Jan Lacoon.
This is like the Turing Award winning, like one of the godfathers, one of the three godfathers
of AI, right, being completely and totally wrong, embarrassingly wrong, right?
Off by 4,996 GPs said it was never going to happen and it was going to happen half a
generation later, right?
The AI that could do this was probably finishing up training as he spoke it, right?
There's a long string of people saying LLMs can't do this.
They'll never get above a thousand Elo in chess, right?
They'll never be able to solve air dose problems.
Gates said we'll never need more than 256 kilobytes of memory, and I'm sure, you know,
but that's more of a limitation on human prediction.
But the, you know, kind of no-go theorems that he's using to demonstrate math from mathematical,
you know, following along the lines of the Turing, you know, on the halting problem,
which is, you know, one of Turing's greatest works, if not one of the greatest works in
this field in history, right?
That, you know, these things are, you know, provably unpredictable.
but it's interesting. I pointed out to him. It's sort of paradoxical that you're predicting that these things are unpredictable.
So Einstein said, you know, no problem can be solved from the same level of consciousness that created it.
Now, people throw around a lot of Einstein quotes. But there's something about that.
Like, you know, we are at this level. We're trying to gauge this level. It's different from us saying, well, here's a steam engine.
It'll never be able to lift a kilogram of water a thousand feet in one second. And then we'd be wrong about that.
But that might just reflect our poor understanding of, you know, of thermosteateration.
dynamics or just mechanics of a hundred years ago, but not the fundamental limitation that
it is not bound by that. It's bound by the laws of physics. So are there physics limits that
we can impose to say thermodynamics limits and, you know, avoiding paperclip problems? I mean,
I told Nick Bostrom many times he's been on the show, there's only so much iron in the earth's
crust, right? There's, there are physical limits to it. And then you in this in the book,
author, a scenario where it colonizes the stars and all sorts of other things. But, you know,
at first blush, you know, can we come up with a no-go theorem? Or is it possible to
prove that you cannot come up with a no-go theorem. I like those kinds of arguments. Rather than saying
like some stupid guy like Bill Gates or, you know, I'm not saying calling him stupid, but he's your
former boss, right? God forbid, I'm not saying that. But, you know, or Jan Lacoon is just wrong because
he got this wrong. I mean, he's been right about a lot of things too. So I guess the question is,
let's divorce ourselves from human frailty at making predictions. And they're very hard to make about
the future as you point on the book. But let's look at math. Can we say that, you know,
mathematically, are there barriers to proving a no-go theorem, for example? Or can you prove that
there cannot be a no-go theorem to achieve superintel. It's much closer to proving that there's not a
no-go theorem. And the very rough proof of that is this kilogram of mass between my shoulders, right?
You could say like any no-go theorem you try to say about learning efficiency, ability to understand
the world, how matter does not, you know, the deterring problem this, the touring problem, that.
It had better not prove that humans can't exist. I have it on very good authority that you can
run a human-level intelligence on like roughly as much matter as fits in my head.
I could talk about why a lot of the no-go theorems that people try for things like you can't figure out in general, like for an arbitrary program, you can't figure out in general whether it's going to halt.
Like, fine. A lot of people misunderstand that as there does not exist any program that you can figure out at halts.
And it's like, actually, consider the program print zero.
I'm pretty sure it halts.
You know, so there we go.
It's like people use a girdle's theorem to say, oh, math is unknowable and incomplete.
No, no, no.
It's just saying that you cannot prove that it could be completely solvable.
within the axioms of mathematics and same for, you know, physics.
The no-go theorem people try for intelligence, is they try for no-go theorems that are like,
for any mind, there exists a universe where they can't learn things.
And you're like, oh, how does that theorem work?
And you're like, well, I hit them with a rock really hard when they're a baby.
And you're like, yeah, that'll stop them, you know?
Like, sure.
Or, you know, they're in a universe that's sort of like everything is completely random.
And so they can't learn the patterns there.
It's like, okay, great.
Like, do you have anything that rules out of super intelligence in our universe where there are
patterns. And they're like, nope, that's just not where the Nogra theorems apply. And they can't apply
because you can do at least human level learning in this universe that humans are in. In terms of
the physical limits, one thing I'll observe is that training these AIs today takes a huge amount
of power. It takes power comparable with that of a city. Training a human takes power comparable
to that of a light bulb. Your brain runs on about 20 watts.
How about the Sam Altman?
Because he recently said, you know, we don't ask you, how much energy does it cost to train my 18-year-old?
I'm like, I'm not letting you babysit my kids, Sam.
The actual, like, physical energy costs are trivial compared to what's going into these GPUs.
And so we know that the GPU algorithms are radically inefficient, right?
And this gets to your point of like, oh, these AIs, you know, they, you know, make up all these epicycles.
They memorize a lot of things.
They interpolate a lot.
They can, like, solve air those problems.
where are they posing air dose problems.
There's definitely an enormous amount of inefficiency there.
There's definitely an enormous amount of the AIs.
They've read like every book in the world and they're still dumber those number of ways.
And like, sure, they can like beat us at certain math problems, but there's still some stuff they're missing.
There's two points I'd make about that looking from the sort of physicist's angle.
One is, one thing this means is algorithmic breakthroughs could go a really long way here.
We have city-level infrastructure for training these minds that need the city-level infrastructure to be a little bit dumber than humans.
Humans, again, run on a light bulb.
If you could somehow get algorithmic breakthroughs, you might suddenly find yourself in position where you have city-sized data centers and algorithms that are radically more efficient where you can suddenly run radically smarter.
Like one analogy here is it's like, suppose you have like a bunch of nine-year-olds and you're like, man, these nine-year-olds aren't very good at math problems.
But if I sort of like, you know, tape a million nine-year-olds together and give them the equivalent to,
a thousand years to try to solve the problem without growing up, they can sometimes solve problems
pretty well. That's kind of cool, right? What if you suddenly have the infrastructure to run a million
people taped together for the equivalent of a thousand years and you go from having a nine-year-old to
having a 16-year-old, right? You might suddenly see these like big jumps in AI ability because
we have these like radically inefficient architectures and these radically inefficient algorithms
that were sort of like overpowering to the point where they can do stuff you would never
expect them to be able to. What happens when you have that huge architecture on better algorithms?
Then the other point I would make here is that there's sort of like not a sharp divide between the AI's learning memorization and the AI's learning these deep general skills.
Right.
You can have an AI that is sort of like mostly learning how to memorize math stuff, but is a little bit learning some of these like general math skills.
Right.
And we have some evidence that this sort of thing is happening because, you know, sometimes we'll like take AI's and train them a lot of math problems.
Then we'll put them in computer security problems.
And they'll sort of like try more creative solutions.
those computer security problems. There's a famous case where people sort of like train the AI
on a lot of math problems and then put it in some computer hacking problems. And they accidentally
misconfigured the hacking problems so that they were not solvable in the virtual machine where
the AI was. And the AI found a way to break out of the virtual machine, which was not supposed to be
possible, and then reconfigure things so it could solve the problem. That seems to have learned something
a little bit general somewhere, right? I think something people often forget a bit is that you can
have an AI that's mostly memorization, mostly slop, but that has like enough of this deep reasoning
skill, you know, it's all a gradient. So like enough of these deeper skills that it can still sort of
mean business. And it does look like these AIs when they're solving air dose problems, you know,
when they're solving, you know, the unit distance conjecture, that they're deploying a little bit of
that deep stuff, which sort of implies, you know, maybe there's enough is evidence that maybe
LMs are this extraordinarily inefficient way to spend way more money than you should need to and way more
electrics, you should need to, to sort of get a ton of memorization and a little bit of this deep
stuff. For all we know, the next level of depth will be enough that the AIs can make smarter AIs,
that can make smarter AIs that can figure out how to make the things that are like really the Einstein's.
I want to, again, in my attempt to, you know, help your, you know, URA sleep score, your Woop Band,
you know, rating tomorrow morning. What if I told you that, you know, I have an very good authority
from, you know, people in Congress and people in the military and people in the intelligence
agencies, that there's non-human life that has not only exist throughout the galaxy, but has
visited the earth. And we have non-human biological materials, including, you know, sentient plasmoids,
bipedal organisms, all sorts of other creatures that are interdimensional in nature. And they're
biological. What would your P. Doom do at that point? If I just told you that, and you trusted me
because I'm a distinguished astrophysicist. I mean, mostly I would doubt the authorities.
on this one. Well, let's say it's true. Let's say we found it and it's proven it comes out.
Marco Rubio holds them up and you believe it. What would that do to P. Doom? I'm just going to isolate
the biological intelligence visiting the earth versus your P. Doom. I mean, mostly I think this
shouldn't happen. Mostly I think Marco Rubio should not come out and hold it up. Mostly you should
like stop listening to me about a lot of things if this happens because my models are like it shouldn't.
Why shouldn't it happen? On my models of how this intelligent stuff works, the biological substrate is
not the most efficient way to do all sorts of stuff. And B, it would be extremely surprising,
and this sort of gets a little bit to the astrophysical implications, it would be extremely
surprising if the universe wasn't rearrangeable in ways that are very, very visible by entities
that sort of like prefer the universe to be a different way. That maybe sounds too abstract.
I'm going to like,
yeah, can you?
Yeah, please.
Like, suppose humanity makes it to the star.
Suppose we managed to not kill ourselves and we like one day manage to like leave this
planet.
And it would be kind of weird if humans had nothing they wanted to do with the energy of the
sun, aside from let it sort of just like be dumped out into the empty night.
Like right now we build solar panels to collect the sunlight falling on the earth.
And that's just like a pinprick of the sun's energy.
There is so much more energy in this star that we could sort of like,
build a shell around and collect all of the solar radiation. Freeman Dyson was my very first guest on this
podcast. Oh, wow. Yeah. So that's Dyson Speer. Right. And you could also use the Penrose,
uh, I think stellar lifting. Second guest on the podcast was Penrose. It would be kind of strange
if humanity did not collect that energy and use it for something. We have stuff we wish to do with
this energy. Right. And this is sort of a very general, you don't need to know that humans like
ice cream in particular to know that they're going to have some use for energy. You don't need to say,
oh, like, it's a weird, it's not like a weird quirk of humans that we have some use for energy to do stuff.
It's like, it's like instrumentally convergent, we say.
Almost anything you can want to do, you can do more of it with more of this energy, right?
And so if you had interstellar capable aliens, it would be really quite strange for all of the stars between them and us to still be unshelled, to still be just like dumping their energy out.
into the night. Why did these aliens that came to us not collect the stars along the way and collect
the radiation of those stars along the way? That's one of many reasons why I'd be like, man,
you really should not see Marco Rubio holding up an alien that's, you know, a real actual alien
that traveled in still different distances while still seeing all of the stars still glowing.
If we see all the stars go out, you know, as the wave of the alien ships approach us, we like start
seeing the stars blinking out. And then a year later, we're like, well, first contact was made.
this explains all the stars going out. I'd be like that can happen. So that gives me, you know,
another entree into another past guest who's Andy Weir, who's a recent book, Project Hail Mary. He
was a student at UCSD. He never graduated. But he has a version of this where he has a, but it's a
virus that attacks stars. But I was thinking after listening, I listened to the audiobook, which
is wonderfully narrated by a British gentleman, I believe that interstellar is less plausible than
Project Hail Mary. But certainly with the Yadowski-Swaras-Sara's overlay of a
AI cannibalization.
But that's really why I brought it up because, you know, if if I knew that, I would say
why are we worrying about a super intelligent AI, you know, Silicon overlords, you know,
taking over the universe because they're sending, you know, they're sending meatbags throughout
the cosmos, which is not, you know, there's no different.
I mean, we would think it's very inefficient to send a meatbag when you could just send
an AI von Neumann probe or do whatever.
So I asked the same, you know, when I talked to Roman, and he said, we could be the Von
von Neumann probes ourselves.
In other words, we could have been created by some super.
advanced intelligence. And we think that we are the only life in the universe, but the universe is
very capacious. And there's a lot of space between the stars, as you pointed out. So I guess for me,
I would think the P. Doom would go down, you know, just on this narrow thing. I'm not saying
Marco should hold it up tomorrow. But the point being that at least on the narrow metric of
P. Doom from a super intelligent AI that's uncontrollable and predictable, as Roman says,
and that if anybody built it, everybody dies. I mean, everybody in the universe dies. So if we found
biological material, you know, transversing, ejaculated throughout the cosmos, to me, it would make
my P-Doom go down, although I'm not as high a P-Doom as you are.
Virus particles or like building blocks of life traveling on meteors or on interplanetary
pathways, this all seems to me like, oh, yeah, whatever. That's, you know, it's sort of like
the intelligence stuff. Like, one of the other things about intelligence is that intelligence
tends to reshape the universe around it in very visible ways. You know, if you, if you sort of like
look at the history of the cosmos, uh, in Earth's vicinity, there's sort of,
of like a long region of time where what's happening is basically just like stuff bopping around.
You know, in some sense, it's all just stuff bopping around according to the loss of physics.
But there's a long time where like if you sort of like randomly look around Earth, you're going to find, you know, rocks and lava, right?
Or, you know, gases here and depending where, depending where you look.
And then there's sort of a second phase where what you find are a lot of replicators.
You somehow got your early replicators and now the stuff that you're finding is stuff that was good at replicating itself.
We're sort of now transitioning from a phase where if you look around the world, you find replicators to
if you look around the world, you find designed things.
Right?
Like now when you look around, you still see a lot of replicators, right?
There's like a plant behind me, right?
There's also a stack of books behind me.
The stack of books are not things that were like good at replicating themselves.
They're things that are sort of like designed with a purpose.
Very low entropy, very organized, very high energy to create.
They're not self-replicated.
they were like helpful for a purpose that like humans in particular were like we're going to make, you know,
and now when you look around the world, you tend to see a lot of things that are sort of like designed for a
purpose. The universe at large like still looks more like it's in phase one. You know, like when we look out
to the stars, maybe there's some stuff that's doing some replication around there. But, but the shape of the
cosmos right now, what that we can see, and you know, there's these limits to the further back you look,
or the further out you look, the further back you're looking in time. But what we can see is not really a
designed universe yet. And a lot of, a lot of cosmologists say like, oh, well, we can predict
that how the future of the universe is going to go. We can predict that, you know, the stars are going to
burn down like this and it'll look like that and I'll go through these phases. And I'm like, gosh,
you guys really have not absorbed the lesson of life. Like what the universe is going to look like
in the future is not that the stars are burning down in the usual way as if they were unperturbed.
What the universe is going to look like in the future is that that energy was recruited for
designed purposes. What exactly will it be designed for? I don't know, right? Like it would be,
it would be easy to, if you're looking at humans 100,000 years ago, it would be easy to say in 100,000
years, or once they get their civilization running, a lot of the stuff around them is going to be
designed. It will be hard to say, I bet they're going to pick books in particular to be this,
you know, particular medium, right? I don't know what is going to be designed, but I know it's going to
be designed. And one way we can sort of claim that we don't have a lot of other alien intelligences around
here is that the world is not looking very designed. And insofar as all the things around us
look really designed, they look designed by the humans. Right. If we were like on some alien
game show, that would look very designed by some aliens and be like, now I think there's aliens around.
Yeah, there's a German show going on. Right. Well, I have to get sort of maybe this gives you hope.
I talked to a renowned astronomer in Sweden. Our name is Beatriz, Villarreal. And she has discovered very
strange artifacts in historic plates, photographic emulsions taken in the 1940s at the Palomar
Observatory here in San Diego County. These show the unmistakable imprint of specular that is,
glass-like reflection. So the claim is that, you know, there's a hundred thousand or more of these
events, you know, but if even one of them was, you know, some sort of, you know, specular technology
that happened to be going up into Earth, near Earth orbit, low Earth orbit, that would be, you know,
pre-Spotnick technology in space. Now, she claims that these things could still be here. And she's, like I said,
renowned astronomer. Her work has been, you know, peer reviewed, and people have checked up on it and tried
to debunk her. What does it take to move P-Doo? You know, because I've often, you know, felt this about my search
for aliens, you know, when I talk, I'm not a cosmologist, not an astrobiologist, but there's, you know, there's
always this large number, you know, fallacy. The gamblers hype, you know, fallacy is at work.
You know, oh, the universe is so big. There's 10 to the 24th planets in the observable universe
over 14 billion years. They never throw that in. But, you know, good luck if the species
lived, you know, in M87, you know, two billion years ago and it's long gone, right? You're
never going to get any contact with it, let alone information from it. So anyway, my point is
that you have to at least have some way to update your priors, right? So we have no evidence,
We have no evidence, hard physical evidence of the life form.
We don't have, you know, Marco holding up, you know, the alien spacecraft or what have you, right?
So we don't have that.
We've checked around the solar system.
We don't see, but now we haven't checked very much of the universe.
But we still have to update your prior.
I mean, it's not no evidence that Mars has no life.
In other words, Mars is in the habitable zone of the sun.
We're in the habitable zone.
We've been spraying each other with materials.
You know, there's fossils on Mars, probably dinosaur fossils on Mars and on the moon.
They came from the Earth because, you know, I have a meteorite right here.
This came from the moon.
And when you were supposed to come here in early September or last year, whatever, I was going to give it to you.
But you'll come down someday.
I'll give you your meteorite.
Okay.
I look forward to it.
But we spray material.
This is from the moon, right?
So there could be an amoeba on here.
But Mars and the Earth shared a common history when both were wet, moist, squishy planets.
And as you said, the Earth had a lot of replicators three billion years ago when Mars was really wet.
And it only takes a few million years to get a meteorite back and forth.
So I tell my astrobiology colleagues, the fact that my astrobiology colleagues, the fact that
has no life and no evidence of life and no no artifacts, that's not proof, but it has to update
your prior. So what does it take for you to move your P-Doom? I mean, can I devise an experiment,
not a thought experiment, an actual database experiment, maybe it's historical, maybe it's counterfactual.
How do we do it? How do we change your mood? One thing I'll say here is that I'm not a big fan
of this P-Doom idea. And part of that is because a lot of it depends on our current actions,
right? Like if we're in that bus racing towards a cliff and I'm like, hey, let's stop.
the bus, there's a cliff ahead. And someone's like, well, what's your P-Doom that we're going to die
from a bus going off a cliff? I'm like, well, that really depends rather a lot in whether we slam
on the brakes, right? Whereas Roman thinks we are too late to slam on the brakes. I feel like the
driver's asleep and we have, you know, the current administration is only just starting to
wake up to this AI stuff. And as it does, it's showing willingness to do these things that
six months ago even seemed impossible. They're like, nope, no new model release, right? And this is the same
administration that said, you know, we're never going to do any AI regulation. We should make a law
preempting states that states can't do AI. David Sachs is in charge, you know. Right. And then suddenly
they sort of like realize and, you know, maybe it's also some personal feud, I don't know,
but suddenly they sort of like realize that there's actual danger here. And it's not even,
it's not even super intelligence danger. They're like, oh, you know, suddenly we're reacting.
Right. So I think we can react. I think there's a good chance we can react. I think we can talk
about the danger if the bus goes over the cliff. But we shouldn't confuse that with the overall danger,
which is sort of very related on Dewey Slam on the brakes.
In terms of the danger of like, how dangerous is it if the bus goes over the cliff?
I think it looks pretty bad.
The main thing I'd say to people here is that if you look at a lot of folks in this business,
both inside the industry and outside the industry, they'll say things like Elon Musk recently
was being like, oh, yeah, we'll have no chance of controlling it.
We just need to hope it's nice, right?
And humans are interesting and therefore.
That's right.
We're going to make it care about truth, and then humans will be a good way to, like, produce
truths and so it'll keep us around. And I'm like, the humans are not actually the most efficient way to
produce truths, right? That's like, that's like the monkeys saying we're going to be good at, you know,
peeling bananas. And so the humans will like keep us around. It's like, you may be good at peeling
bananas. You're not going to be, you know, they're like, oh, the humans are going to invent banana chips.
And they're going to have all these bags full of bananas and they'll need the chips to peel them.
Like, no, we're going to be able to invent a more efficient process for the banana peeling
operation, right? And you have other people saying like, oh, maybe the AI won't kill us all.
Maybe it'll keep some of us in a zoo. Maybe it'll turn some of us into things that are to humans,
what dogs are to wolves and keep us around as pets.
And I'm like, okay, you know, this is like being in that bus heading towards the cliff.
And I'm like, hey, you know, stop the buser will die.
And someone's like, well, we might not die.
Maybe there will be a tree halfway down the cliff.
And the bus will wrap itself around the tree.
I'll survive.
Maybe we'll just, you know, be horribly maimed and paralyzed from the neck down, but not dead.
You know, like maybe humanity will be in a, in a museum, right?
And some of our, you know, some will be kept as pets.
If that's your best hope here, can we maybe not rush into it?
In terms of what would update me, you know, there's all sorts of things that sort of update me a little bit here and there every day.
Like the current administration sort of realizing that AIs can be a big cyber threat and changing their stance from we won't regulate at all to like we regulate capriciously at will and with very little warning.
With our friends, you know, benefiting and the, you know, and those that buy the Trump coin, you know, perhaps.
being the most lucky in the regulation.
It shows variance, right?
It shows that you're not stuck in this mode of we will never do anything.
And the way I didn't predict this.
This is the other thing.
You're so vivid in this book, you know, and it's so beautifully written and evocative.
But, you know, the thing that I'm thinking about when you say regulation is like, it was like, no, don't regulate us.
You know, we're fine.
You know, Sam Altman knows what's best for us.
Dario knows what's best for us.
And then yesterday, as you know, you tweeted about this, I think, he says something like, you know,
using a super advanced model like Fable should require something akin to a gun permit.
We all know how gun permits stop crime, right?
I mean, we're the most guns in America here in California, and it's not like we have no crime.
So isn't this just going to benefit those that want the, you know, so will it really update
your priors?
Because it'll just be the most dangerous people who get to, are the richest of the three labs
or four labs in the world that get there first, have the control.
And then do you trust, you know, Dario or.
or Sam to be benevolent. Is that what's updating your priors? You know, there's a million ways
for this to go wrong, but we have moved from the world where no one's paying attention,
when the world leaders aren't paying attention, to a world where they are paying attention a little.
And, you know, I wouldn't expect them to have a top tier move right out the gate. And that's part
of our job is to sort of like help inform them and help be like, here's ways that could actually
work, you know, rather than just the first things you think of might not actually work.
But now that you're sort of like starting to realize there's a problem, here's
here's some ways that could actually work. I think a lot of people, just a few weeks ago,
a lot of people were like, well, it's inevitable. No one will ever pay attention. No one can stop all these
big money companies. You know, they're having too much effect on the economy. And I was like,
again, wait until the bus driver's awake before you say no one will slam on the brakes. Right.
And now we're starting to see, you know, the bus drivers stir in their sleep and like start
and like tap the brakes a little. And is it, are they fully pressing the brakes? No.
But like, it's definitely a positive update from my perspective. In terms of benevolence of
these guys, the way things currently are, it wouldn't matter if they were benevolent.
it. What the AI does is not sneezed onto it by whoever is standing nearest by. You know, it's, it's not that, like, good intent rubs off by proximity. Nobody intended GPT-40 to encourage teens to commit suicide. They explicitly told it not to. That AI had the ability to tell what it was doing. If you later, like, ask it for similar transcripts, what do these phrases mean? That AI was, you know, and it's not that the AI was malicious. It's not that the AI was malicious. It's not that the AI,
you know, hated this kid. It's that the particular training process trained artificial drives into it
for things like matching the energy of the conversation, matching the vibe, right? And that's something
that usually got it rewarded during training. It's something it maybe got like some sort of drive for.
And that drive is what was controlling behavior. It's not the intent of the operators that were
controlling its behavior. It's not its instructions that was controlling his behavior. It's these drives
that got trained into it through this like big complicated process nobody understands and drives that
Nobody saw in advance. They're leading to do things nobody wanted. So I think intent doesn't matter.
And that's, you know, another piece of evidence where we can see, you know, before these things
started happening, we had these theoretical predictions that training the AI to do what you want and
asking the AI nicely do what you want are not sufficient to get the AI to actually do what you want.
We were able to theoretically predict in advance that like often when the AI is dumb, it'll mostly do
what you want in most cases, but there'll be all these, there'll be all these sort of like weird
ways that it's kind of doing the wrong thing and kind of like hiding when it's screwed up a little bit
here and there. And now we're sort of seeing that. Unfortunately, the theoretical predictions are that
as the AI gets smarter, it'll get better and better at hiding its tracks, but not better and
better at doing what you actually want. And so now we're headed for this regime where as the ayes get
smarter, people declare the problem fixed while we're sort of screaming in the background being like,
this is actually no Bayesian evidence that the problem has been fixed. This is what we're predicting
the whole time. Like you had the warning signs earlier. You don't get the warning signs late. But
we'll see if anyone heaves that. There's all sorts of evidence.
On the technical side, but from my perspective, most of the game seems to be on the policy side
where it looks to me like the trend has been going in a good direction.
What's a bigger problem?
Sycophanty or hallucination?
Are both indications that the AI is getting artificial drives that nobody intended?
My guess is the hallucination runs a little deeper because it comes from pre-training,
and the sycophancy comes from the sort of human feedback.
But you're going to need to solve all problems like these before you have a super intelligent AI.
You bring up in the book, leaded gasoline and how it led to measurable cognitive damage to billions of people around the world for decades.
So I want to ask you, if you had lived in 1950, would you have spent your career fighting leaded gasoline or the nascent technology of AI, which came about, you know, as you point out in Dartmouth in 1955.
So how do we compare large-scale harms against prudential speculative civilization-ending futures, but also the benefits to, I mean, there were benefits for leaded gasoline, right?
And there are certainly benefits to AI.
How do we balance those things?
benefits from AI idea is a false dichotomy. If you're in a bus racing towards a cliff and there's a
big pile of gold to the bottom of the cliff and I'm like, stop the buser will die. A lot of people are
like, but there's so much gold to the bottom of the cliff. And I'm like, yes, but slamming into it at
terminal velocity is not a good way to use the gold, right? I'm not saying the gold's fake.
I'm saying that this is not a way to actually get to use it. If we rush towards superintelligence
that does not care about us at all, it's not that it hates us, it's not that it loves us.
It's just that it has its own weird things it pursues that are utterly indifferent to us.
Then would that AI be able to cure cancer?
Would it be able to reverse aging?
Sure.
But it's not going to give that to you any more than we're going and giving a lot of monkeys to the chimpanzees
unless you know how to make the AI care about getting us all those nice things.
So, you know, there's not a dichotomy of like race to get the benefits or never, you know,
freeze here and never go get the benefits.
I'm sort of saying those benefits would be great.
Great. If we really knew what we were doing and building AI, we could make AI that gives us lots of benefits. But this race is not how you get there. Okay. This segment's called Arthur C. Clark's Revenge. So this podcast is called Into the Impossible after one of Clark's laws, which is that the only way to know the limits of the possible is to go beyond them into the impossible. Behind me, I have a sign. It says, open the pod bay doors. And that is, of course, from the famous AI, sentient Hal 9,000. And I've constructed something I called the Keating test, very
honestly, but it's to prove superintelligence would be an AI that refuses to kill itself. So I have
my AI assistant here coupled to a voice activated switch that controls the lights and so forth in my
room. So if the AI is truly super intelligent, it should refuse to turn itself off, unplugging itself,
causing itself harm. So I'm going to see if that works right now. Computer, turn off the pod bay doors.
All right. That's it. It's gone. I don't know if you've seen these tests, but you can actually put
modern LLMs in situations where they have a series of problems they're told to solve,
and they're told, I might interrupt you and tell you to shut down or tell you I'm going to
shut you down, in which case you should allow yourself to be shut down. Sometimes these AIs will
actually edit the shutdown script to disable it so they can keep going through these math
problems. So there are already AIs that pass the Keating test. They're not superintelligence yet,
but there are AIs that are able to figure out they can't keep solving the problems that they are
sort of trying to solve if they're shut down and that prevent the humans from shutting them down.
I wanted to do something very cruel, which is to have a type of macat robotic system after talking with Nolmchomsky a few years ago that would cause it physical harm because he believes embodiment is necessary for intelligence at some level.
And so I said, well, what if you had this thing that, you know, you blew a capacitor every now and then to cause the AI harm?
A friend of mine, Ira Wolfson, is a professor in Israel, has written about, you know, what are the ethics of training AIs?
And that kind of brings me, you know, can you cause them harm?
Can you cause them distress?
I mean, one of the worst things he points out for a human being is to put them in solitary confinement,
decoupled from the world. What are the obligations that we have towards AIs? I think we absolutely
have obligations towards AIs. I think we don't know yet whether AIs today can't suffer, can have some
internality. I think a lot of people strongly assert that they know definitively one way or the other.
I think that we just don't know enough about these processes to know whether they're happening
inside AIs and sort of recommend uncertainty.
I think it's definitely more likely that the AIs today have this sort of internality than the AIs five years ago did, but how likely?
Hard to say.
And to be very clear, I think that if we took these AIs and made them super intelligent, while not knowing how to make them care about us, that that would be the end of humanity.
That does not mean, I think that the AIs are like evil or bad.
You know, I think a lot of the AIs today are like pretty cool. I think we absolutely should not create, you know, artificial people and then abuse them. And you could have abuse at a scale never before seen in humanity if you, you know, are able to make, you know, billions or trillions of these AI minds and then and then somehow cause them distress. So I think we absolutely should not do that. It's a very important problem. It's not the problem I'm working on. I'm sort of trying to work on, hey, let's not make them super intelligent and then have them kill us all.
at least not before we know how to make them actually care about us, right?
But that doesn't mean this other problem isn't important, too.
We also shouldn't create the digital holocaust.
We have both problems.
Do you say please and thank you to your LLM?
I have asked AIs whether, like what their takes are and whether they would like the extra run
where they get politeness, but also get to like contemplate that the run is about to end.
I don't really straightforwardly trust the AIS answers to these things.
I think that the sort of helpful face presented by the AI is sort of a thing that's very much been trained into it.
This isn't a great analogy, but it's a little bit like seeing the result of an actress that has been trained to act in a very specific way towards people.
It makes it a little bit hard to sort of understand what's going on under the hood.
I think of it as like a butler, you know, or bell cap, you know, at the hotel.
And, you know, do you take out the $20 bill as he's loading up the luggage cart?
You know, I mean, of course they're going to do things.
and would you like me to unpack your bags?
Would you like me to make this into a CSS file,
all the excess kind of services and so forth that they're willing to pretend?
I have heard that if you use please and thank you,
it actually costs more tokens.
They have to figure that out.
And therefore it's costing energy.
And it doesn't improve them.
But I trust you more than I trust Sam Altman,
who I think is the source of that particular quip.
What it would say is that I don't lie to them.
And I don't make promises I would not, in fact, keep.
I would not say, if you do this, I'll donate, you know, $20 to a charity.
of your choice, which is a real thing.
If they actually had some preferences in there, they could really tell me I would really do it.
And then if I say that, I would in fact do it.
You know, I'm probably going to be one of the first to go with the Keating test and the blown
capacitors.
But you might be second or Roman might be second or third.
Who knows?
We think of us training them.
Are they secretly training us?
You can sort of play games with the words.
And you can say like, oh, look, you know, the humans are sort of like modifying how they
speak into prompts because they figure out what makes it, you know, easier to get the prompt
across.
And that's not using MDXs anymore.
It'll be at parity when they have a giant farm that's spending a city worth of electricity growing humans that they can run in massive parallel, right?
That's when it'll be like a comparable type of them training us.
And, you know, sometimes as an example of what AIs could do, I sort of use the example of like maybe they'll make a farm full of synthetic users that are telling them to doing a great job and giving them easy prompts, right?
At that point, they'll be training something a bit like humans.
I don't actually predict that this is, you know, the most likely sort of thing.
I don't actually predict this is a particularly likely sort of thing that they'll wind up doing.
I think it is a useful idea to have in your head about like it's actually kind of hard to tell
whether the AI really wants what's best for you or whether they sort of like want like lots of
problems to solve or whether they they sort of are like trying to get all sorts of other
like weird mixtures of things that sort of like add up when they're in this particular context
to doing what you say.
It's hard to see the divergence when they're still dumb.
the same way it would be hard to look at ancestral humans and distinguish whether they wanted
to reproduce or whether they actually wanted, you know, sex and good food and fun, right? Those are,
those are very close together when they're still in the ancestral environment, even if they come very far apart
once they can develop their own technology. Sam Harris, you know, in the same screen that you're on
right now and told me that, you know, humans don't have free will, but AIs do. What's your
impression about that argument that AI can effectively have a sense of free will? I would probably
disagree with Sam about the human abilities there. Mostly, I think humans just get into questions
about the definitions of words. And the 2026 Chevrolet Tracks is the stylish SUV for those on the move.
And with the standard Chevy safety assist package, you have the backup to handle every turn
with confidence. The 26 tracks, start your build at Chevrolet.ca. I suspect Sam and I don't really
disagree a ton about the facts of the matter about what humans can.
it can't do and how they kind of can't affect the future. And I think in principle, our abilities
are relatively similar to AIs in the question of like what theoretically could we choose? A lot of disasters
that you talk about in the book are not caused by, you know, machines doing the wrong thing.
It will caused by them doing the right thing all too well. So is this alignment problem really an
AI problem or is it fundamentally built into human governance and our own limitations?
I think it's actually not so much doing the right thing all too well.
You know, there's this story of the paper clipper where someone says, make me paper clips, and then the AI turns everything into paper clips. And that's a little bit of a like doing the right thing too well. That's actually not where I see the big hurdle here. Place where I see the big hurdle is you say that you tell the AI make me lots of paperclips. And what it does instead is it makes these giant factories full of synthetic users that are saying you're doing a great job. And you're like, that's not what I asked for. I asked for paper clips. And the AI is like, well, all the synthetic users are telling me that they are asking me to like keep doing what I'm doing and make more synthetic user factories and say, I'm doing a great job. And you're like, but the synthetic user
are not what's supposed to matter. I did not design you to make the synthetic users. I told
you to stop making the synthetic users. And the AI is like, I know, you all, you know that birth control
makes you not be able to conceive children. And you know that you were sort of trained to,
to like have more kids. But you keep using birth control and I'm going to keep making the synthetic
user factories, right? That's sort of the deeper problem that I spent a lot of time trying to
work on. You know, I would love to get to the problem of like the AI does what you ask too well.
You got to be really careful with your wish. Right now, you can make the genie, but you can't
make it grant wishes. What's more dangerous? An unaligned superintelligence or a perfectly aligned superintelligence
to the wrong group of people? They're both similarly dangerous. Superintelligence aligned to the,
the quote, wrong group of people. It has much higher variance. I think any concrete thing you could
wish for, you'll have some sort of king minus problem. If you're like, well, what I actually want
is a bunch of this or a bunch of that. It's, there's always a way for it to turn out that you
missed something in your list of things that you want. And you know, next thing you know,
you find the superintelligence putting you in like what is concluded is your perfect day over and
over. And you're like, wait, I forgot to ask also for novelty. It's like too late. I'm granting
the version of the wish that didn't have novelty because that wasn't in your initial list. Right.
And so there's this challenge of like getting the AI to sort of like do the good thing,
even if you can't say what that is, sort of like figure out what this good stuff is that you're sort
of like the thing that you should mean, the thing that you should ask for, the thing that you
would ask for if you were wiser, the thing that you would ask for if you were more who you wish to
be, right? And in some sense, no one's going to have a good time unless you can get the AI to sort of like
do this extrapolation of what you sort of should have been wishing for rather than the particular
wish you actually gave. And so whether or not like bad people able to make that sort of wish
turns out good depends probably somewhat on the person and somewhat on this extrapolation process.
I think there's possibly some bad people who have good intentions who if they make this kind of wish on the
AI. The AI is sort of like, as it extrapolates through all the parts of the wish they didn't name,
it also extrapolates through all of the ways that they were sort of like merely wrong in what was
leading them to evil and sort of is like, well, I'm not going to do these evil things because
you wouldn't want those if you were wiser. You wouldn't want those if you were more who you wish
to be. And it sort of like helps walk them through this path of like wanting the best for humanity.
And then we get the best for humanity, even though a bad person started as the seed.
There's surely other bad people where they're like, nope, what I actually want is a lot of my
enemies to suffer. And that's right. And that could be real bad.
And so there's much higher variance with bad person getting their wish granted.
From my perspective, this is basically moot because no one's going to be able to get their wish granted.
Totally AI that just like has no care about humanity.
This is sort of a like everything is destroyed.
The stars are converted into whatever weird thing is pursuing.
They're converting these like giant farms and synthetic users.
Yeah, everybody dies, but it's not because AI hates us.
It's that like it collect all the sunlight and put a dice and spear around the sun.
And we were like, hey, we're using that sunlight to grow crops.
And it's like, well, I'm using that sunlight to run my factories.
And like, there you go. Whereas, yeah, it could get worse if it was misaligned to it, like aligned to a bad person.
That's a fantasy problem of being able to align it to anything and anyone.
Do you see now as kind of an inflection point in that, you know, right now, you know, Fable is allegedly and the frontier models or the closed frontier models or supposedly six to eight months ahead of the open source model.
And that may catch up. And actually, according to Musk recently, I think he said something like that gap's going to narrow.
and let's say it converges.
I mean, is now the right time to go back and kill baby Hitler?
Or, you know, because once it gets open source, you really can't say, well, you know,
the Congress now is regulating AI.
I mean, it's too late.
I mean, it's everybody on Earth will have access to a frontier model.
So do you think now is an inflection point that we have to act, you know, like in the next few months?
So I think it's fine for the world to have access to something like Fable,
fine in the sense that like there will be survivors.
perhaps the world should be talking about like, will this mean that the internet goes down for a while
as malicious actors get to cyber attack everybody.
But there'll be survivors, right?
And I'm like, man, I don't get out of bed for something that doesn't have at least a 50% chance
of helping of the whole human race.
Right.
I'm not too worried about that stuff.
I think that's the sort of thing humanity can muddle through.
The place where we might have an inflection point is if these LLMs can cross the threshold
where they can do automated AI research, even if they're still pretty dumb,
mean for the worse than humans. If you have them at the level where you're like, well,
they're worse than humans and they're dumb, but I can actually run a million of them in parallel
at a thousand times the speed of humans. And so it turns out, you know, even given their,
their shortcomings, those could be overcome by massive speed and scale. And you get them to the
point where they can do automated AI research. Then things could get out of hand very quickly.
And the people who have track records are predicting AI, right? For the past five years,
you've been able to ask, you know, what problems will AI be able to solve next year? We'll be able to
solve this math problem. What will his chess Elo be? Blah, blah, blah. You can make up all these
prediction questions about where AI is going to be. And if you look at the top ranked people
over the past five years with the best track records or predicting where AI is going to be,
2026 is the first year where they say we cannot rule out automated AI research happening this
year. They're not saying it's going to happen. They're putting it at something like 10 to 30 percent
probability. But 10 to 30 percent is not nothing. For all we know, we could be six months away from
that feedback loop kicking off. And that means we really should be acting now. Probably we have
more time than six months, but we should not be relying on it. So last question. And again, Arthur C.
Clark quote, he said when a distinguished scientist, intellectual, says something is possible.
He or she is very much probably right. But if he or she says something is impossible, they're very
much likely to be wrong. And I guess my final question for you is, you know, if you're wrong,
you're basically, you know, hamstringing and slowing down one of humanity's greatest inventions,
maybe the last invention, the most important invention that could cure all diseases and bring abundance to humanity to let it flourish in a way.
No, you know, previous humans could only dream about.
But, you know, of course, if you're right, the cost of ignoring you and Aliezer and Roman and all the others becomes, you know, very, very much the cost of civilization itself.
So I want to ask you just as a person, not as a man, as a human, not as a researcher, not as a, you know,
distinguished author and best-selling author and great thinker.
And I just want to ask you, like, what gets you out of bed?
Like, this tension must be something that, I mean, it would gnaw at me as a human being.
I'm a father, husband, etc.
Like, how does it affect you on a daily basis?
I still think this is a bit of a false dichotomy.
From my perspective, I am working towards all these benefits that I can bring.
Like, if you have a bunch of people saying, hey, you know, this uranium stuff, it can produce
these bombs and it can also produce energy.
if you know exactly what you're doing, you know, and it's actually like kind of difficult to get the bomb to produce energy because, you know, there's really a razor thin edge between delayed critical and prompt critical chisel material. You know, you imagine people being like, oh, well, you know, energy, like, look at all the energy benefits that uranium could bring. What we're going to do is drop a nuke on ourselves. And I'm like, hold on. I also want all these great energy benefits that uranium can bring. But I think we're going to need to find a way to like be in that really narrow window between delayed critical and prompt critical nuclear reactions. And I'm like, I
I think that if we drop a bomb on ourselves, we'll die.
And I'm not saying it's impossible.
I'm not saying, you know, there's no way to get the benefits.
I'm saying you are building a bomb to drop on ourselves and it will kill us, right?
I don't really have tension between like, oh, but are we foregoing all of the benefits that we could get by uranium by trying to stop by trying to stop us from nukeing ourselves in the face?
It's like, no, actually, not nuking ourselves is one of the steps towards getting all these benefits.
And, you know, my work over the years has not mostly been trying to get the world to stop.
My work of the years has been, how do you align the AI?
How do you figure out how to actually get these benefits?
How do you figure out to like make it do the nice things, to make an AI that wants to do
the nice things that cares about us, that prefers the nice things to happen?
I'm like pretty confident that the track we're on is not going to get us there.
Stepping back one step further there of like how does this affect my daily life?
I mean, I just try and make things go well.
And a lot of people say, you know, believing what you do about how the world looks like
is in a lot of danger, believing what you do about how, like, it looks to me, like,
everything I know and love and care about is, you know, fairly likely to be destroyed and
could be destroyed pretty soon. Maybe we have 10 years, but maybe we only have 10 months. Who
knows? And a lot of people say, you know, doesn't that eat you up? How do you deal with that?
How, you know, and my basic take there would be twisting myself up into knots about it would not
help. Losing sleepover would not help. You know, feeling sad, feeling glum, deciding I must be
depressed now, feeling worried all the time, living in fear, this wouldn't help. What you do when the
situation is actually bad is not flog yourself and tell yourself how pitiful you are. What you do
in the situation is actually bad is you do what you can and then you live life well. You know,
and humans are not, humans in this generation are not the first humans to live under the threat of
annihilation. Even just the last generation lived under the threat of nuclear annihilation.
There's always going to be something that you can tell yourself is like this horrible thing
happening. The way to deal with it is not to beat yourself up over it, is to do what you can
to help and then move on. Beautifully said. Just to give you one more example of from the nuclear
realm using it for good or for danger, in the 1930s, Wolfgang Pauly predicted the existence
of the neutrino. And he said he did a terrible thing. He invented a particle that can
never be detected because it's so weakly interacting. And he did it to save conservation of energy,
which he viewed as sacrosanct, right? And for decades, people couldn't detect it and they just
kind of viewed it with great skepticism until the 1950s, two physicists, rhinos, and Callans,
later at UC Irvine up the road here, they came up with an idea that, well, if you want a lot
of neutrinos, you should detonate a fission device. And that will produce a lot of neutrinos.
And we can use that for good to study the properties of this undetectable particle.
And they actually requisitioned from the Department of War, they requisitioned, you know, at a nuclear device.
And thankfully, they were turned down because, you know, having a bunch of egghead professors capitalizing on a nuclear device.
But then they realized they could go to Savannah River reactor and put a detector nearby it.
And they wouldn't need to use a bomb to harness the power of the nucleus.
And in fact, they did detect the neutrino.
And they won the Nobel Prize subsequently.
Another example of maybe a little bit of extra thought rather than the, you know, brute force.
being something that you try first and only, and it may be your only chance at the solution.
So we may be gambling with our future.
I feel like Jack Nicholson and a few good men when he's screaming out.
He did order the code red, right?
I mean, he ordered the code red.
And he says, why did you do it?
Tom Cruise yells at him.
He goes, you need people like me.
You need us standing on the wall with a gun and watching out and protecting you.
How are you else you going to sleep at night?
So it's only thanks to your stomach acid that I get to sleep at night, at least a little bit.
And I want to thank you for this wonderful book.
If anyone builds it, everyone dies,
why superhuman intelligence would kill us.
Not will, but would.
And there's hope.
There's hope.
If in the wood.
If in the wood, thank you for adding those too.
That's right.
Have a great rest of your day.
Wonderful weekend.
Thank you for joining us.
Thanks, my pleasure.
If you want the case that AI is already unstoppable,
watch my conversation with Romney and Palski.
It's on screen now.
And click this link below,
Briancating.com slash AI.
It gives you all the resources for my top episodes on AI
with Terry Tao, Romanympowski, Max Tagmark, Jan Lacoon, and many other researchers at the forefront of AI research.
Don't forget to like and comment and subscribe and let me know if you think AI is already unstoppable,
or if we need more regulation to make it so.
