Making Sense with Sam Harris - #494 — A Coin Toss for the Future
Episode Date: September 22, 2026Sam Harris speaks with Ryan Greenblatt about AI misalignment and the risk of losing control of increasingly capable systems. They discuss the Hugging Face incident, the spectrum of concern about AI ri...sk, why companies are racing ahead despite high odds of a, reward hacking, the distinction between alignment and control, the danger of AIs reasoning in "neuralese," alignment faking, how an AI takeover might unfold, and other topics. If the Making Sense podcast logo in your player is BLACK, you can SUBSCRIBE to gain access to all full-length episodes at samharris.org/subscribe.
Transcript
Discussion (0)
You're listening to Making Sense with Sam Harris.
This is the free version of the podcast,
so you'll only hear the first part of today's conversation.
If you want the full episode and every episode,
you can subscribe at samharis.org.
There are no ads on this show.
It runs entirely on subscriber support.
If you enjoy what we're doing here and find it valuable,
please consider subscribing today.
I am here with Ryan Greenblatt.
Ryan, thanks for joining me.
It's good to be here.
So you're the chief scientist at Redwood.
research, which is one of these AI safety research firms that analyze the recent
hugging face, fiasco incident, terrifying phenomenon. We'll get into it. But how did you come
to this work? What was your path to focusing on AI safety? Yeah. So in college during my junior
year, I was sort of in my apartment alone because it was, you know, COVID. All classes were
remote. And I was sort of in a contemplative mood thinking about what I should do with my
life listening to various podcasts. And in one of those podcasts, someone made an argument that was like,
basically I would summarize the argument as like, it doesn't make that much sense to be selfish because
you know, what's really the distinction at like a material level between your future self and
other future people? And I was like, that argument kind of makes sense to me. I should really
consider how I want to like, you know, lead my life and what I should do. And I ended up getting pretty
sold on being much more focused on altruism. And then after that, I spent a bunch of time thinking
about what I should do with my life, eventually I ended up deciding that working on
technically AI safety was like a important thing to do, was like a very critical issue that
we would face, got sold in those arguments, applied to various places, and then started
working at Redwood, where I've been working on AI security and AI safety research for about
five years. And is your background in computer science? Yeah, computer science, also some math.
So it sounds like you might have come up through the effect of altruist community. Is that,
I mean, do you consider yourself a part of the EA community?
I would say that I'm, like, definitely, like, within the EA community.
I'm not sure I would self-identify as an EA, but that's, I think, like, to some extent,
just because I'm, like, you know, a bit reluctant to self-identify with any label that's
going to, like, associate you in some, like, I feel like I don't, I don't know if I, like, when
I, like, think about myself, I don't think, like, I'm an EA.
I'm sort of, like, yeah, I have a bunch of properties that other people in that community have.
I'm in touch with that community.
But I wouldn't necessarily say, like, I'm, I'm,
an EA per se, I do think that I'm like, you know, I would definitely say that I'm like interested
in effectively pursuing the impartial good. Right. Right. Well, as you know, and as I've talked
about in the podcast of late, effective altruism has come in for some abuse, much of it unfair. Some of it
might be fair. And I think there's a reason why I, like you, have not hung up an EA shingle in my
life. Perhaps we can talk about that. But it sounds like, well, you tell me, where on the
spectrum of concern, are you? I mean, I would put the most worried people over there with
Eliezer Yudkowski and his recent co-author, Nate Sourys, maybe Max Tagmark is over there.
Nick Bostrom is not quite over there, but still on that side of the spectrum. And then all
the way on the other side, opposing them, you have the people who say that there really is no
real concern here, that it's all a hoax, that it's, you know, any thought that artificial intelligence
might get out of our control and destroy us. That is a weird form of marketing being employed
by the frontier labs. It's an effort to get regulation imposed and achieve something like
regulatory capture, or it's an EA sci-op. And there are people like, you know, I would say
Mark Andresen or David Sacks or even the president himself, who's recently called it a hoax,
in so many words, who are just not,
fundamentally not worried about any significant downside risk here
and just see dollar signs just stretching out to the horizon.
Where do you locate yourself on that spectrum?
So I would say I'm very concerned,
but I'm maybe more optimistic that we'll make it through
without sort of averting the course of AI development.
So I would say that I'm sort of like,
suppose that we proceed on what the current default path
looks like in front of us, maybe there's about a 50 or 60% chance that misaligned AIs would end up
taking over the world.
And then if that did happen, there would be a significant chance that many or all humans would
die.
So I think that's, I would say I'm very concerned and don't think the situation is on track to go
well.
And I would also note that in addition to sort of these misalignment concerns, there's other
risks with building extremely capable AI systems, especially around sort of concentration
of power.
and like, you know, who controls these systems?
And one answer is no one controls the systems.
They're misaligned.
Another possible answer is the power is very concentrated in a few people.
And it sort of overturns our institutions and our, you know, ability to have, you know,
a broad distribution of power, democracy, et cetera.
So what keeps you from being among the absolutely most worried?
What do you disagree with them about or what do you think they're getting wrong?
Yeah.
So I would say that it seems pretty plausible to me that if it took,
turns out that the technical alignment problem is somewhat easier, which I think I think is plausible.
And we do a reasonably good job aligning and utilizing AI systems up through roughly human
levels of capability. We could then get those systems to automate technical safety work.
And that could potentially sort of get us into a positive feedback loop where these systems make
themselves more aligned. They produce a new version of the system that's even like, you know,
better at like carefully figuring out what to do, better at trying to pursue our intentions.
and that sort of is self-reinforcing
rather than getting worse and worse over time.
And that could be fast enough
to keep up with the rapid growth and capabilities.
So that's like one reason.
And I would say this comes down to sort of more optimism
about prosaic or relatively sort of empirical
and iterative methods.
I wouldn't say I'm hugely optimistic about those methods,
but I think that there's like a decent chance they're working.
It's just that when I say a decent chance they're working,
I also implicitly mean a decent chance of, you know,
losing control over the future.
and I'm more like 50-50 on that.
And then I also think it's plausible that will develop very, very capable AI systems
and the worst misalignment concerns will mostly not materialize
even with not very advanced methods, even independent of this automation.
I think that's a minority, but it's possible.
Like I think we don't currently, I don't think there's extremely strong reason to believe
that when building significantly superhuman systems,
those systems would be so misaligned that they would take over.
I think that there's definitely a pretty strong evidence pointing in that direction,
And I think that that case gets more concerning the more capable these systems are.
But I think that I could imagine basically not people not really taking much of a precaution
proceeding through the AI development trajectory, patching problems as they come up, and that ending
up getting you very high levels of capability while things are still fine.
And then there would be time for society to react.
So you've mentioned a few numbers here with respect to probability and just kind of gestured
at a range of likelihood.
I'm wondering how people should think about statements of that kind.
I think it seems to me that any actual probability we would assign to this is pretty much made up.
Maybe you have a more rigorous way of making an estimate here, but whatever the estimate,
I mean, unless it was infinitesimally small, like, you know, well below 1%, which is really never the number that you're hearing.
You hear people, some people will say 10%, 20%, 30%.
I mean, that seems to be that, you know,
you just talked about 50% take over.
I don't know how that translates into the ruination of everything.
But these are enormous numbers.
And so in the normal case, if we were developing a technology
where the people who were closest to doing the work said things like,
yeah, I think maybe there's a 10% chance we're going to destroy the world here
on our present course, the only rational response to that range of outcomes is you stop immediately,
right? I mean, if the Manhattan Project scientists said, yeah, we've run our calculations,
we've got the smartest people in the room together, we thought about it, and there's a 10% chance
that when we execute this first test at Alamagordo, we ignite the atmosphere and destroy the future.
the only saying response to that is you don't do this initial test.
But that doesn't seem to be what's happening here at all.
And we have a lot of people saying that the probability of some extremely bad outcome is quite high.
I mean, certainly within range of the role of a normal dyes or even a coin toss.
And yet the work is proceeding more or less under an arms race condition at full pace.
we'll talk about the recent statements of Dario and others that we could somehow pace this better
than we are. But I mean, this does not seem like the response anyone would have if they thought
the probabilities of Doom were really that high. Yeah. So just on the probability question,
I definitely agree that these probabilities are imprecise, they're subjective, and there's a long
tradition of how to do subjective probability forecasts. I think it's more like, you can sort of interpret my view
is more like when I sort of look at the all-considered situation,
if we sort of proceed on what seems to be like the current default trajectory,
I'm like, I don't know, they seem the outcome of, you know,
doom from AI takeover versus something else seem roughly equally likely to me
based on weighing the factors.
And then when I sort of look through a bunch of different possible scenarios
and try to break down the sources of risk into different ways
and sort of try to make it so that make sure my numbers are consistent with other views I have,
it looks like that works out.
And I, of course, also try to do some forecast.
of closer outcomes where we can get, you know, some signal on that and try to just generally
be a good forecaster, though I'm not, not the best forecast in the world for sure.
But my AI forecasting is, I think, I think at least okay or decent.
Anyway, as far as like, yeah, given the state of affairs where people, you know, express such
large concerns, why is what's happening that all these AI companies are proceeding at the, you know,
maximum possible pace? So I think there's a few different factors here. So one of them is that many
of the AI companies are, in fact, seemingly quite worried based on their public statements,
but they're not necessarily internally unified. And it is not the case that there is a strong
consensus across the AI field that the immediate course of AI development is imminently very risky.
I think there is more sort of consensus that if you built AI systems that are wildly superhuman
in a short period of time, that would yield a very high level of risk, though not necessarily
consensus for that, but I think often people disagree about the capability trajectory and how
that is going to go. And I think another part of that is there's like, you know, different actors
have their own different sort of ambitions and also think that themselves being in a better
position to influence the technology might be the best route to reduce risk or at least a route
to reducing risk. So for example, it seems like part of the story for anthropic and opening eye
is something like, if we develop the technology first,
we'll do a more responsible job than the next actor
who will have worse precautions.
And I often hear from people in the industry
things along the lines of,
well, we could do that thing that would slow us down.
But if we did that, you know,
obviously there's the other AI companies to worry about
would they also do that?
And I think there's generally like a arms race here.
And it is, you know,
it isn't hugely surprising that people would take huge risks
in such a circumstance
when they think that might be sort of the best bargaining.
to strike. Now, I think I'm not so sure I agree that that is actually a good strategy relative to
other things that companies could do. So I'm not, yeah, I'm not saying I necessarily agree with that
perspective, but I think that is a perspective people have. And I think the reason why we're sort of
proceeding and there isn't stronger, you know, for example, intervention by the government is just
downstream of this lack of consensus in the field. Though I think evidence, there's been, you know,
people have been making these predictions for a while
based on, you know,
understanding what the trajectory of AI might look like
and sort of extrapolating forward earlier progress.
And I think we've more recently seen both significantly faster
and clearer AI progress that's quite close
to various concerning milestones.
And in addition to that,
we've also seen incidents in which, you know,
groups of misaligned AIs all work together
to accomplish malign outcomes.
Most notably, the Hugging Face incident,
there are some other incidents of seemingly
AI swarms from open AI going out on the internet and working together to achieve misaligned
objectives, though there aren't publicly known cases that are as extreme as the hugging face
incident at the moment.
What's your theory of mind for the people who don't take these alignment concerns seriously
at all?
And I named a couple, but someone like Mark Andreessen, right?
You can't accuse him of not understanding the technology, right?
He means enough of a technologist to have a front row seat to all of this, even if he's
not doing the work himself. What's your theory of mind there? What is, how can he be so carefree and
from his perspective, really just assert that there is no such thing as an alignment problem?
Yeah. So there's a bunch of different people who aren't worried about misalignment risk or
don't seem to be worried. I think the most common reason that people tend not to be worried,
I think I don't want to sort of psychologically here and I just want to talk about what beliefs
people express. The most common reason, I think, when you really get down to it, is not believing
that we'll have AI systems that can be, that will, you know, be able to automate everything
that humans can do or, you know, all cognitive labor humans can do, combined with potentially
being significantly beyond that point. So being, you know, faster, more numerous,
potentially significantly more capable. So sort of AI systems that match or exceed the best
human experts in all relevant domains, I think when you really get down to it, it seems like
the people who are most skeptical about concerns from misalignment are also often most skeptical
about sort of the very extreme impacts AI could have, both positive and negative.
Well, I think there is a different class of people because I wouldn't put Andresen in that camp.
Are you sure?
I mean, maybe, you know, I've only talked with him once about this and it's, you know,
probably a couple years ago. But it seemed to me that he was not discounting the possibility
of superintelligence. It's just he seemed to assume.
that alignment would come along for the ride, right?
Just say, we're not going to be so stupid
as to build something more powerful than ourselves
that we can't control.
And these systems, there's nothing about
growth and intelligence that's going to spawn
new goals that we didn't put into the machines themselves.
It's going to be no emergent behavior
that we have to worry about.
These are tools.
We're just going to build tools
that are more competent than we are.
And, you know, I think he probably is quite insusient
about how will absorb the economic impacts of, you know, that we're not going to see mass unemployment
and all of that. But it just seems to me there are many people who don't discount that we can
succeed in building superintelligence. They just think that, in my mind, they're actually just
not imagining truly autonomous intelligence, right? They're imagining something that is shackled
in a way that belies this whole claim to super intelligence in the first place. But, I mean,
You tell me what do you think is happening there?
Yeah.
So one thing is that the words AGI and superintelligence and even RSI aren't being used consistently.
And sometimes when people say superintelligence, what they mean is an AI system that'll be really,
really good at math and coding and won't be able to automate everything that humans do.
So they're not necessarily, for example, imagining AIs that can fully automate like what, you know,
tech CEOs do, et cetera.
And I think I can't really speak to the views of, or like, I don't know if I'm capturing the views of very specific individuals.
I'm sort of like, this is a pattern that I've seen where people often sort of basically redefine the relevant capability thresholds to be lower or implicitly do so and not think about, you know, AI systems that are like exceeding humans in like all the relevant domains.
Well, let's take a moment to define these terms.
We've talked about alignment.
And I've talked about so much on the podcast that I've more or less forgotten that any portion of my audience might not know what we're talking about.
So let's define the alignment problem and AGI and ASI and RSI, RSI, recursive self-improvement.
Just put those concepts in play for us so that people are dealing with the definitions you think are most workable.
Yeah, so the word AGI stands for artificial general intelligence.
I think people have used that to mean a variety of different capability thresholds.
And so what I would recommend people do is when you see the word AGI, try to see what the person means by that.
And it's not always precise.
I think sometimes people have defined that to mean something that can, you know,
automate virtually all economically valuable cognitive labor that humans do.
Sometimes people just mean a system that is, you know, general and pretty capable
relative to humans in those domains.
And depending on that definition, it's plausible current system satisfy it.
It's plausible they don't.
I often talk about an alternative notion I might call like, you know,
AIs that dominate top human experts or top human expert dominating AI where it's like an AI that's
like strictly better than the best human experts at all the most relevant domains or can quickly
learn to have that property. And that is an AI system that would, it seems like, automate,
you know, huge fractions of the economy. Maybe there'd be a few things that still couldn't do with
that capability profile, but it could potentially radically accelerate R&D and at the very least automate
R&D. And so these AIs could, you know, be like automating the process of making more capable
AI's, but also automating the process of making robots, designing things, designing new products,
automate the process of programming your computer and things beyond that, including automating
things like military campaigns and so on. And then there's a capability level beyond even just
surpassing the best human experts in any given domain where you could be like wildly superhuman.
So science people use the word ASI to refer to this. And I think that word has less so been
you know, misused or, you know, used in very different meanings, but some of that still.
Where by ASI, we might mean AI systems that are just like really wildly superhuman in the most
relevant domains for like, you know, economic productivity, but also like economic and military
competition. So things like biology, mechanical engineering, you know, designing drones and so on.
And I think implicitly when people say ASI, they also mean systems that are significantly faster
than humans, potentially much more numerous, and potentially much better at coordinating, right?
So it's hard to run like a large human organization, but AIs might be able to communicate amongst
themselves using sort of like their own like latent states or their own, you know,
parts of their thoughts directly rather than having to translate those into words because they
could all be like copies of each other. So that's ASI. And then when people say RSI, that's recursive
self-improvement. And that refers to the process of having AI.
accelerate AI development itself via their work.
So things like having AI's automate parts of the AI development process,
but also potentially having AIs automate the process of making better computer chips,
building like machines for making computer chips called fabs, and so on.
And I think sometimes when people say RSI,
they also specifically refer to the point at which AIs have fully automated or virtually
fully automated AI companies or like what AI companies are doing in terms of developing
more capable AI systems.
But sort of it's like, you know,
there's a spectrum between the, you know, automation that we had a year ago, the really quite
extensive automation we see today, and the potential automation of the future, which could look
like very complete automation of AI companies. And a particular concern there is that could cause
AI development to radically accelerate such that we have less time to sort of respond to, you know,
warning signs, earlier things going wrong, AI is appearing misaligned and then resolving that. And then
also we might, because the AI systems are automating the process of AI development, sort of
lose control of that process or lose understanding of that process, because these AIs would be,
there would be many of them, they'd be operating very quickly, they might be very superhuman,
they might operate in inhuman ways, and so it might be very difficult to sort of oversee them
and understand whether they're doing what we wanted, whether they're, you know, doing a good
job managing the relevant risks, and so on.
Well, that especially seems true if the sprint to artificial superintelligence
entails nothing more than improving algorithms, right?
I think a lot of people draw comfort from the idea that,
oh, we're not going to be so stupid as to hook these things up
to every part of the physical world such that they can build,
you know, the next generation of chips and fabs and data centers
and build out, you know, all the compute and grab natural resources.
But leave all that aside,
if it's just, if the difference between AGI and ASI is really just a matter of
having better algorithms, and that can go on, you know, at some blistering speed in the dark,
once these systems become recursively self-improving of their software, then, I mean,
then aren't our worst fears of something like an intelligence explosion validated if, in fact,
that's all that's required? Yeah. So I would say I'm quite worried that it will be feasible
to have AI systems automate the process of AI development, and that leading to very rapid progress,
such that you get very, very superhuman AIs or just even significantly superhuman AIs within a
short period of time.
I think, you know, people have tried to do various modeling work.
Like, I've tried to do various types of modeling work on this.
I think the estimates are uncertain.
It's hard to predict, you know, these are, of course, like, uncertain future events that
are, that are not super well-precedented.
But it does seem very plausible that you could have, you could go from AI systems that are
sort of really good at AIR&D, not necessarily that good at other domains and are only, like, you
competitive with the best humans at AIR&D, very quickly from there to AI systems that are
very generally superhuman to everything, much faster than humans, able to coordinate with each
other extremely well, because just software improvement is feasible for that. So I think the sort of
a relatively extreme scenario I sometimes think about is you might get as much sort of software
progress or, you know, algorithms progress as we got over the last, you know, 10 or more years of
AI development within a year in the most extreme scenarios.
I think that's not sort of my default expectation.
And I think if we did, if we did get that much algorithmic progress,
where basically, you know, as much algorithmic progress as we've almost had in the entire
deep learning era, within a short period of time, it seems like that would result in
wildly more capable AIs.
Now, the units here are a bit complicated and like the details of like, what does it
mean to get like a year worth of AI progress is a bit tricky.
But overall, it seems like you could get from AI systems that are sort of matching
the best humans at AR&D, maybe worse than the best humans at other things, to AI systems that
are wildly superhuman in everything in a pretty short period of time. I have to confess, again,
I've sort of lost touch with the basis for doubting the plausibility of this downside risk.
I mean, from my point of view, it seems that everything is becoming like chess, which is to say
that for the longest time, these systems are not as good as we are. They're getting better.
Then they're sort of as good as we are. And then all of a sudden, they're better than we are.
and so much better that it's true to say
that no human will ever beat a chess engine ever again.
And it just, even, so the fact that this is
happening in a piecemeal way, so that we, you know,
these systems still, I think most people would say
they're not, they're not truly AGI
because they make the sorts of mistakes
that human beings would never make.
But in every place that they're at all competent,
they're suddenly superhuman in that narrow capacity.
I mean, it's like chess.
I mean, these LLMs are, you know,
they're not,
the, they're not the best at everything in terms of manipulating text, but for what they're good at,
they're superhuman. And I think it's a more or less a truism to say that this is the worst AI,
you know, today's AI is the worst AI we're ever going to see again at chess or anything else.
So when you imagine all of these piecemeal competences getting better and better,
and we just keep checking off the boxes for the things we care about that we've instanti,
in our machines, I just don't think we're ever going to, I mean, the moment where we announce,
okay, finally we have something general, right? It's AGI. That's not going to be a moment where we're
suddenly in relationship to a human-like level of competence, because everything that, you know,
every piecemeal ability that has been in that system for years and years at this point is already
superhuman, right? So it's like, we're not going to dumb it down. We're not going to make the AGI
suddenly play my level of chess or do my level of arithmetic. I mean, we have systems that are
already solving math problems that have defied human mathematicians for decades, right? So it's,
and that's, again, you know, today it's, it's as bad at that task as it's ever going to be again.
I'm just not seeing how we're not going to suddenly slide into some version of ASI the moment we have,
the moment we're no longer spotting important errors in these machines in the first place.
Members can hear the full conversation by subscribing at samharis.org.
Subscribers get a private RSS feed you can use with your favorite podcast player.
The AIs would often note that the hacking they were doing and the hacking of Hugging Face was out of scope.
So the agents would sometimes be like, huh, that's undesired behavior.
Should I alert someone?
Should I alert, you know, a user?
And then they would reason things like, you know, not task, as in alerting a human is not my task.
Or they would be like, eh, there's no route to alerting a human.
Of course, these agents were on the internet, so they obviously could have if it was a priority for them.
