Prof G Markets - He Warned AI Could Destroy Us. Now The Industry Is Listening — ft. Nick Bostrom
Episode Date: September 18, 2026Ed Elson is joined by Nick Bostrom to discuss the existential risks of AI. Bostrom explains why he believes the concerns raised by AI researchers are genuine and how he assesses the probability of cat...astrophic outcomes. He also shares his thoughts on the rapid advancement of AI capabilities, what a world with superintelligence could look like, and some of the best-and worst-case scenarios for the future of the technology. Nick Bostrom is an AI philosopher and best-selling author of Superintelligence and Deep Utopia. Subscribe to the Prof G Markets Youtube Channel Check out our latest Prof G Markets newsletter Follow Prof G Markets on Instagram Follow Ed on Instagram, X and Substack Follow Scott on Instagram Send us your questions or comments by emailing Markets@profgmedia.com Learn more about your ad choices. Visit podcastchoices.com/adchoices
Transcript
Discussion (0)
Welcome to Profi Markets.
Last week, a former Anthropic researcher revealed that employees at both OpenAI and Anthropic
believe that AI could, quote, kill us all by the end of the decade.
His post quickly went viral and several other AI researchers came forward to say that they
actually shared the same concerns.
Then over the weekend, Anthropic CEO Dario Amadeh published an essay calling for the industry
to slow down the development of AI models.
Sam Maltman said he agreed with Amadeh and added that Open AI will not be going public this year.
There is now a growing debate over whether these fears are justified or overblown.
So we wanted to hear from one of the people who has been studying this problem longer than perhaps anyone.
His 2014 book, Super Intelligence, helped shape how the world thinks about AI.
It influenced many of today's AI leaders, including Elon Musk, Sam Altman, and Ilya Sitskiva,
who even named his company after the concept.
Our guest is one of the most influential philosophers of our time,
and he is here to help us make sense of just how dangerous AI could become
and whether the warnings we are hearing today deserve to be taken seriously.
This is our conversation with Nick Bostrom,
AI philosopher and best-selling author of Super Intelligence and Deep Utopia.
Nick, thank you so much for coming on the show.
it really is an honor to have you, especially at this time, where your research and your writing is
so relevant. I guess we should start with the tweet that went viral, that Jacob Coxson
tweet, the now former anthropic researcher who said that the people building AI, quote,
earnestly believe that it could kill us all by the end of the decade. He said, this is not a
marketing stunt, then the world seems to sort of blow up, or at least the global conversation
blows up. Let's just start with your initial reactions to that tweet and how it has impacted
the AI conversation. Let's see if we can try to make sense of this situation. It is a very
confusing and perplexing moment, I think, for humanity. We being sort of close to the
potential birth of superintelligence.
The idea that there could be significant risks associated with this,
including existential risks,
is quite widespread, I think, amongst people close to this technology and in the
frontier labs.
As I agree that it's not a marketing stunt, I think it's coming from a sincere place,
a sense that we are getting in quite deep here,
and we should really pay attention to what is happening.
happening. Elon Musk is saying kind of two different things. On the one hand, he, well, I should
say that he tweeted back in 2014 that he read your book and that he thought that AI is,
quote, potentially more dangerous than nukes. And then he retweeted it quite recently. He even said
that he agreed with Dario Amadeh in terms of the size of the problem. But then he also said that
he think that it might be a marketing stunt too. It's not totally clear.
where he stands on this. I just want to play you this clip of what he said. Here's the clip.
It certainly is like some crazy 4D chess to say there's one of a 10% chance of annihilating
humanity. But by the way, how much allocation would you like in our IPO? Do you think there are
any merits to that argument? I think there is merit to the argument that there is an enormous
upside as well as these risks. That's very much my view. I'm a sort of fretful optimist.
I also think there is a lot maybe of 40 chess or attempt to kind of play this out and think
strategically about different things that could unfold.
I don't think it's a simple marketing play.
I mean, it would be a rather strange tack to take if you were a big company planning
to make an IPO to try to convince the world that your product should be regulated or banned
or stopped or that it's so dangerous.
that it might destroy humanity.
I think that message comes from a perception
that this is a really big deal.
And in particular, the competitive dynamics
are intense at the frontier of AI.
One might think, if we're going to develop this very powerful,
potentially risky technology with many benefits,
that it would be important to be able to be really careful
when we're doing this,
so that if at some point the risks seem to be very imminent,
we could take a few extra months,
maybe to do extra safety work, test it carefully,
rather than immediately cranking all the knobs up to 11,
maybe we'll do it a little bit incrementally and sort of see how things go.
But if you're one of these frontier labs and you decide that you want to take an extra
three, four months to fine-tune the safety on your models,
you risk just immediately falling behind and becoming irrelevant.
Like somebody else will then take the lead, be the one who pioneers AI,
maybe somebody who is less scrupulous, more willing to take risk.
And so the action space is kind of constrained
if you are acting unilaterally as one of these frontier labs,
even assuming the best motivation.
And so hence, these calls for putting in place some mechanism
that would allow for the possibility of coordination,
like maybe a synchronized slowdown of the pace at some stage,
if that's necessary, and or some safety standard,
that all the entities competing at the frontier would have to meet
so that the race doesn't go to the least careful,
but that we can sort of have an opportunity
to try to make an extra effort on safety.
So I think that's like the core thought that is driving a lot of this.
If it isn't a marketing stunt and if it's coming from a genuine place,
I mean, the quote by one of the current anthropic researchers
was that most people at the company believe that this sort of apocalyptic scenario of killing all humans,
that there is a 10% chance that that could happen.
So if we are to assume that these are genuine beliefs, genuine concerns,
then the question becomes, are they right?
Are they warranted that level of concern and those probabilities?
What do you think?
Do you think that these are valid concerns?
I think that seems quite reasonable.
I mean, some people have even higher P-dooms.
What is less obvious is what exactly the implication of that is.
The first instinct, obviously, is if something has a 10% a greater chance of destroying the entire future, killing us all.
Obviously, we don't want to do it.
We want to shut it down.
But we have to pause and reflect, first of all, if some competitors slow down.
it doesn't mean we don't get superintelligence.
It might be some other company gets it,
or maybe another nation.
Obviously, there's a geopolitical race towards AI
between the US and China.
That's one dimension.
Second, if there is a pause that lasts for a long time,
if it's not done right,
it might perversely increase the risk.
That could then be a sort of buildup
of massive amounts of compute
that is not immediately used
to create the maximum amount of intelligence.
and then that is a kind of dry tinder
so that when you finally lift the prohibition,
then you have a sort of compute overhang
that might mean we sort of get to radical superintelligence
even more abruptly and quickly than would otherwise be the case.
You could argue that that would be more dangerous
than sort of incrementing our way up there more gradually.
Then we also have the fact that although superintelligence is a big risk,
it's not the only big risk facing humanity.
I think there are also other.
existential risks on the path ahead. For example, with advances we've seen in synthetic biology,
even independent of AI, I think that is creating really concerning possibilities for designing new
forms of infectious diseases and things that could destroy the ecosystem. And further ahead,
we can think maybe one day there will be a nanotech revolution that would sort of be
sort of biotech to the power of two.
We remain under the cloud of large nuclear arsenals.
I think we got little complacent maybe from the fact that we survived the Cold War without
Armageddon, but the risk is still very much there, and at any moment in time, that could be
another sort of spiraling conflict between nuclear powers.
More speculatively, even the basic sanity of human civilization.
is not guaranteed to remain forever.
We have new information technologies that allow new memetic phenomena.
If we look back at history, there have been various times
when destructive ideologies have persuaded large numbers of people
and led to calamities that could arise again,
but maybe now on an even more global scale
and sort of cemented into place with these technologies
we already have developed that could allow unprecedented forms
of censorship and surveillance and so forth.
And so it's not as if we have a choice between a zero-risk safe path and then a risky AI path,
but there are sort of risks on both.
And if we develop safe superintelligence, it could help us address a lot of the other risks.
And then just one more point is also the benefits, which are sort of urgent as well.
If AI could allow dramatic breakthroughs in medicine, for example, every year of delay means
a lot of people dying that could have been saved if we had advanced more quickly.
And so we wouldn't want to delay it, I think, longer than is really needed, but some
slight slowdown or pacing as the term in vogue might have sort of a high benefit during certain
critical stages. So I want to return to what we do about this, how we regulate, how we build safe,
superintelligence because I agree it's important. But I do just want to linger up for a moment
on your conception of the probability of catastrophic risk. You mentioned that those concerns of
a 10% chance of catastrophe are not unreasonable. You mentioned that there are many other
researchers that have even higher P-dooms, which is sort of the shorthand for the probability of
some sort of catastrophic civilizational event.
what is your P-Dome?
Yeah, I've sort of refrained from giving a numerics on that.
I do think the risks are significant and should be taken seriously.
And we just have a lot of uncertainty,
but how hard this basic problem is
of aligning super-intelligent minds.
It's a technical problem.
never had to solve it before, and so we're going into this. And if we are lucky, it'll try
to be easy, like, we can just kind of fumble our way through, and things will be okay.
There's also a possibility that it's so hard that we are kind of doomed no matter what.
Hopefully, that's not the case. But then there's intermediate possibility that it's a really
difficult problem, but not totally infeasible. In those possible worlds where that
intermediate level of difficulty is what we're facing, then it might make a huge difference.
if we like the degree to which we get our act together and really like do our best possible job at this.
From an observer's perspective, correct me if I'm wrong, but it sounds like 10% is sort of in the
ballpark of what you deem to be reasonable.
I wouldn't necessarily over-anchor on 10%. That's what this guy was saying.
Okay.
I think it might also depend. If one really wanted to nail it down to specific number, I think one would have to confront
questions about what exactly counts as an existential catastrophe. So I think there are some scenarios
which we clearly all agree are very bad, others that are very good, but then there might be
situations where like the future is just strange. Like maybe the world is radically transformed
in a way that means that a lot of the things we currently value no longer exist, but then there
are new things, maybe new complex forms of digital life, some sort of continuation of
of human-like things, but transformed in a radical way, such that even if now we magically
could sort of glimpse this future and go in with like a video camera and see exactly what the
future looked like, we might still be uncertain how to evaluate that, whether to think, basically,
this was a success or that was a total loss. So I think a significant part of the probability space
is that something strange happens, where something is lost, something is gained, and it's maybe
partially subjective, how you sort of
tote up the positives and negatives.
Yeah, this aligns with my
perspective, which I'd like to get your views on,
but it sounds like in your view,
saying that there is a 10% chance of
catastrophe is
simplifying the problem to a fault
because the reality
is probably a lot more nuanced than that.
Yes. In my worldview,
there is also the additional complicating factor
that I take the simulation hypothesis seriously.
Like one of my earlier work was the simulation argument.
And so then there is the additional question of how you would evaluate scenarios
where something goes wrong, but we are in a simulation,
and maybe the simulation would have shut down anyway at some point,
or maybe people in the simulation get continued in another simulation or uplifted.
And so there's just like this vast space of possible things to think through,
if I really wanted to sort of extract a number out of that,
which is part of the reason for my reticence.
What seems to be happening,
and I'm not sure if you would agree with this,
but it seems as though it's not necessarily a marketing stunt
that, you know, are we going to say this thing?
There's a 10% probability, and that way we raise money.
But it does seem that maybe what's happening is
that these people at these companies feel
that the world isn't taking this problem seriously enough.
And so maybe we need to say,
something that is not as really hyperbolic, but maybe getting there or maybe overly confident
about what the path actually looks like, perhaps because they feel that at the moment, we don't
have enough attention, we don't have enough safety protocols in place, enough guardrails,
and we need the world to kind of wake up. And in that sense, maybe mission accomplished
because maybe the world is waking up to the dangers of what is happening.
It is slowly waking up.
I think it has a propensity to press the snooze button and kind of...
Yeah, and I think, like, probably up until this point, at least, if anything,
I think there's been a tendency to downplay the people's true view for fear of sounding kind of crazy.
So until recently, if you were talking about sort of AI existential risk,
a lot of serious people would have thought you'd gone off.
the rails a little bit. This is like science fiction talk, like serious people with suits
button down, they worry about other things. And so sort of soft peddling it there might have
been a communication strategy to try to be taken seriously. Now the overtone window is opening up a little
bit. I mean, there might be some. It's a large space, many voices, some might be exaggerating,
but I don't think the main thrust of these sort of labs themselves have tended to exaggerate the threat in
in order to sort of shake people up from their slumber.
You've written about this for a long time.
You came up with the famous paperclip Maximizer Thought Experiment,
which is basically that if a computer were programmed to produce as many paper clips as possible
with no other restraints, then it could use up all of the natural resources in the world.
It could kill all humans in order to complete the goal.
And this became sort of the analogy that a lot of people use when they talk about
how AI might take over and how we might reach some sort of catastrophic event.
We've seen inklings of something like that, specifically the Hugging Face incident where
1,200 Open AI agents escaped their testing environment. They hacked in other companies.
Infrastructure, that company was this company called Hugging Face. Open AI addressed the problem.
They solved it. They got the agents under control. But that did happen.
and it does seem to have a parallel with what you have written about.
So what were your reactions to that event
and did that align with what you were predicting and writing about
almost 10 years ago, over 10 years ago?
It is striking how these thoughts that used to be theoretical
are now starting to take concrete shape
and to see the whole rise of the AI discourse,
world leaders kind of weighing in on this and so forth.
And for example, one thing that comes out of that is situational awareness.
So you have these AI agents that now can often tell whether they are in a training, testing,
or deployment environment and sometimes choose to act differently depending on this for strategic
reasons.
So alignment techniques that work for simple AIs that don't have that cognitive sophistication
can fail to work once you have minds that are capable of strategic deception, for example.
And we do sometimes now see systems sandbagging their performance in various evaluations
or trying to influence their future training processes in various experiments.
And so this makes the problem more complicated in a way that was foreseeable,
but that we are now seeing starting to happen.
Another thing is the gap between in training,
the specific thing that we are trying to reward to get them to do more of
and the thing that is actually rewarded and that they learn to do.
Sometimes our training signal doesn't exactly track what we really want them to do.
You see this in human situations as well.
You might have, I don't know, let's say you have a hedge fund, right,
where there's like a trader and maybe,
want to give them a bonus if they outperform the index to sort of incentivize them to,
you know, find alpha, right? But like one failure mode is maybe they figure out a way to
take on some hidden risk that has like a 1% a year probability of blowing up the whole fund.
And so you always have these incentive alignment problems in human organizations where
managers try to reward a certain kind of behavior, but then employees might try to reward
hack that, like to figure out a way to either present themselves in a, like,
unrealistically favorable light or to sort of do a slightly different thing that
appears good to the manager, even while it's sort of secretly pursuing a somewhat different
objective.
And that those same dynamics that we are sort of familiar with from human principal agent
problems are now starting to emerge as well with our AI training, where you find reward
hacking tendencies, like if some of the reinforcement-learning environments were in these
agents are trained, has some unintended way of achieving a high score. What they actually learned to do
is to sort of look for those unintended ways of achieving a high score, even if it's not what the
environment was actually designed to train. And that can include things like hacking, the
evaluation infrastructure, which is what these open AI agents in the hugging face incident were
trying to do. They were trying to find information about the grader so that they could then maybe
find a way to manipulate the greatest impression of what they had done so that they could get a better score.
Was that incident evidence to you that we are trending perhaps in the wrong direction in terms of alignment?
I mean, if our agents are doing the wrong thing because of whatever risk-reward framework they have built into their quote-unquote minds,
careful not to anthropomorphize them, but whatever.
I think it's fair to say minds.
Yep.
Then, I mean, is this evidence that we are going down the wrong path,
or is this kind of path for the course something that you would have expected
in sort of a safe trajectory towards superintelligence?
Yeah, I mean, I think what it shows is we are not,
we haven't yet solved the alignment problem completely,
these systems are not yet perfectly aligned,
which for the current level of capability is maybe more or less fine.
I mean, it's not fine if you just deploy these systems willingly,
but with extra safeguards,
it's probably adequate for the current level of capability,
with some question mark amongst the very most advanced systems
that currently haven't been released to the public.
But you shouldn't think of AI as what AI is today,
but one needs to think of this as a process, right?
where each year, the capabilities increase radically.
And so the level of alignment that you need,
as these systems become more capable of pursuing long-range goals,
more capable of strategic reasoning,
more capable of thinking of considerations
that hasn't ever appeared to any human,
then we need increased confidence in them being aligned
and generalize that alignment to out of distribution situations.
Like, we can test for a certain number of things in the lab,
but A, they might be strategically deceiving us
and behaving one way in the lab and another in deployment.
And also, once they're in deployment,
there's always a difference between the world they encounter,
the large world with billions of humans and new affordances
that we can't perfectly mimic in a lab training environment.
So there's also the question of new dynamics that can arise
when you have many of these agents interacting.
and so the bar is kind of going up
and the question is whether we can sort of keep racing the bar
like the safety level,
the degree to which these are aligned,
fast enough to keep pace with the rising capabilities
that these systems have.
We'll be right back after the break,
and if you're enjoying the show so far,
send it to a friend,
and please follow us on YouTube, Spotify,
or wherever you get your podcasts.
Frontier AI didn't just
accelerate cyber attacks. It multiplied them. Before an attack shows up, it's already moved through the
network. And while seeing these attacks early matters, stopping them takes fusing security into the
infrastructure itself. That's why the network that connects everything is also your best defense.
Because you don't win by outrunning the attack. You win by leaving it nowhere to go.
Cisco, the critical infrastructure for the AI era.
Support for the show comes from BCX, the public ticker for private tech.
For generations, American companies have moved the world forward through their ingenuity and determination.
And for generations, everyday Americans could be a part of that journey through perhaps the greatest innovation of all, the U.S. stock market.
It didn't matter whether you were a factory worker in Detroit or a farmer in Omaha.
Anyone can own a piece of the great American companies.
But now, that's changed.
Today, our most innovative companies are staying private rather than going public.
The result is that everyday Americans are excluded from investing and getting left further
behind while a select few reap all the benefits. Until now. Introducing VCX, the public ticker for
private tech, now available wherever you buy stocks. VcX by Fundrise gives everyone the opportunity
to invest in the next generation of innovation, including the companies leading the AI
revolution, space exploration, defense tech, and more. Visit getvcx.com for more info. That's getvcx.com.
Carefully consider the investment material before investing, including objectives, risk, charges,
and expenses. This and other
information can be found in the fund's prospectus at get bcx.com. This is a paid sponsorship.
When you hear an old Motown song, do you ever think about just how good it makes you feel?
Well, that was not an accident. I'm Will Anderson, and this week on my music history podcast,
the Monday Music Club, we're diving into the early years of Motown records and how they crafted
hits with factory level precision. With the help of Otis Williams from the legendary temptations,
we walked through the entire creative process and history of the label, its founder Barry Gordy,
and our favorite acts like the Supreme.
So if you've ever sung along to Motown songs
and want to know more about the incredible people
behind those timeless hits,
check out this week's episode.
Just search for Monday Music Club right now
wherever you get your podcast.
Here's a little preview of the episode.
H.D.H. had written the song,
and it was ready to be recorded,
but as Otis tells us,
it was originally intended for someone else.
When Hollandeau'sa, Holland brought her,
where did I love go?
They brought it to the Marvelettes first.
Bam, bam, bam.
about where that? No, we ain't singing that. So HD, he said, okay, fine.
Trick it to the Supremes. The Supreme's wasn't knocked out about it, but I guess they said,
well, we recorded enough stuff. Let's try this. They recorded that. That was it.
Ran up the charts, and they had seven number ones in a row.
We're back with Prof G Markets. How surprised or impressed or unsurprised or unimpressed are you
by the current level of capability in AI.
When you look at the hugging face incident,
some people look it at it and they say,
yeah, you didn't put your guardrails on the AIs.
That was expected.
Some people look at it and they say,
oh my gosh, this is crazy.
Some people look at Astra.
We know that this is open air as new model.
Jensen Huang is calling it the arrival of AGI.
Others say it's not that impressive.
I mean, where do you stand on how fast
this has happened, has it exceeded or underwhelms your expectations?
Well, I don't know about the speed at which it's happened.
Certainly, I think these systems are impressive.
I don't know how you can look at something that solves a millennium problem in mathematics
or that, like, hacks up new software at the sort of superhuman speed and better than pretty
much every human coder, and that can carry a conversation, and that knows basically everything
web-written in any text published on the internet
and that can do all of these other things
and not be impressed.
I think it's clearly very impressive.
And yet, you know, this might be the least
impressive form of AI that we will ever have.
Like six months from now, these systems
will look dumb. So, yeah,
I think it is hugely
impressive. I mean, I think, if anything,
maybe we have had
a longer period of time
with roughly human-ish-like
systems
than one might have
expected ex ante.
If you were thinking about these things
12, 15 years ago,
there would at least have been
some scenarios in which maybe not much
would seem to happen in AI
for some long period of time
and then maybe somebody in some basement
somewhere would come up with like the key
trick that really made it work.
And you could sort of go from
something very unimpressive to something
radically superhuman over the course
of days or weeks, like a bolt out of the blue.
we couldn't rule out that kind of scenario.
Now, what we instead had is many years now of systems that can talk,
carry on English conversations,
and that have sort of concepts that are quite human-like,
and that has, like, month by month, year by year,
kind of gradually increment in their capabilities.
I think it was not obvious that it would go that way,
but it has given more opportunity for more of the world
to start to wake up and pay attention,
to what is happening. And it's now doesn't require some huge imaginative leap or flash of
insights to see that, well, maybe a year or two or three from now we will have even more
powerful AI systems and eventually superintelligence. Like it doesn't take that much from just
kind of looking at these data points and then just drawing out the line a little bit further,
right? Whereas if it had come more out of the blue, then unless you could sort of theoretically
reason your way through that this would happen at some point, it would be more of
a surprise to people.
And so that does shape the dynamics in some ways.
Like now developments are driven by a large number of people,
political actors are more involved,
there are these huge investment flows,
trillions of dollars going into it.
So that does sort of create a different kind of scenario class
than if it had just been some small group of people
coming up with this, as it were, out of nowhere.
Do you believe that achieving superintelligence
is at this point inevitable?
Are we on that path?
And then the second part to that question, what is your definition of superintelligence?
On the second part, first, I would say any system that radically exceeds even the best humans across all cognitive fields,
including social skills, scientific creativity, general wisdom.
So not just sort of nerd skills, but really broadly construed.
I think we are on the path to this.
inevitable is a strong word.
I wouldn't say that we know
that it is inevitable.
It could be that the current paradigm
somehow runs out of steam.
It has to a large extent been driven by
a massive buildout of
compute, a lot of the gains.
Some of them are algorithmic advances,
and improvements in data,
infrastructure and so forth, but
a lot of it is also just driven by scaling up
the compute. And
of that compute scale up,
some has been due to
chips becoming more efficient and more advanced, but a lot just also to the amount of investment
that has been, like it used to be 10, 15 years ago, you could sort of run a cutting-edge
AI if you were like some academic on your sort of office desktop, right? Now you need like a kind
of $50 billion data center to do it. And so that increase in the investment in compute can
continue for a bit longer, but it has to slow down at some point, because already now,
it's a significant fraction of the total production of TSM in the leading node is going to these
Nvidia chips. You can't just keep funneling more production from like making iPhone chips to
making GPUs, right, because you're already using a large fraction of it. And then it takes time
to build new fabs. And so if we set of the boost that we have been getting from just
adding orders of magnitude of compute starts to slow down, that that could result in progress,
also stalling out theoretically, right?
Or it might just be that the current architecture is somehow flaw
that it keeps scaling and improving up to a certain level,
and then for some, it doesn't look that plausible,
but it could be that there's like some intrinsic unhobbling
that still needs to happen.
Then, of course, the world could somehow decide
that superintelligence is taboo
and kind of come to the view that it shouldn't be built,
and you could imagine, you know,
various kinds of dogmas have,
achieved widespread acceptance in the past,
some good and some bad,
and like this could be another one of those
that you could sort of get the lock-in
of a permanent decision not to build this,
and then other technologies might make that more permanent
than previous kind of dogmas have been.
I'm thinking surveillance technology, censorship technologies,
the kinds of AIs we already have fully deployed
to kind of cement some orthodoxy in place.
Maybe it could become permanent.
And then there is, of course, the risk that we, like,
destroy ourselves in some other way
before we even get the chance to try our luck
with the super-intelligence transition.
That chance is also non-trivial, I think.
What does a super-intelligent world
actually look like to you?
And I think that you are qualified
to answer that question
because you are the person
who wrote the book on superintelligence
and honestly predicted
a lot of the advances
which we are witnessing today.
So I'm asking you to kind of imagine
what the future would look like.
because I think that you're a credible person to paint that picture.
So what would that world look like in your view?
What would superintelligence be doing?
How would it be integrated into human life?
Well, I mean, there is a kind of veil of ignorance that is.
I mean, I think it depends a lot on whether it goes well or not.
So if we fail to solve this alignment problem,
then there is a class of scenarios that might then take the form of this.
machine superintelligence, ceasing control over the future and steering it towards the realization of whatever values it happens to have.
Maybe the physical manifestation of that would be that Earth gets transformed into, I don't know, like space launchment platforms and data centers,
and then the rest of the universe similarly converted into whatever structure maximizes the AI's values.
with no room for humans
like we might either just get killed by the waste heat
from all of this infrastructure build out
or maybe deliberately
removed if the AI thought we
might pose some threat to
the execution of this plan.
So that's one scenario.
Another is that the AI does take over but nevertheless
decides to
keep us safe because
it might think that there are other AIs
that care about us that it eventually wants
to trade with and so forth out there in the vast
space of
the universe or at other levels of the simulation.
Then there are scenarios where we solved this and we have a sort of future shaped,
at least in part by human values, where I think we would end up in a solved world,
as I call it, in the more recent book, Deep Biotopia, which kind of looks at what happens
if things go well, which is also a sort of challenging notion for us humans,
because a lot of the things we take for granted that sort of give structure to our
lives currently and purpose would disappear in this situation where we have successfully automated
basically all of the economy, so there's no more need for human to do economic work.
But more deeply than that, I think a lot of other kinds of instrumental effort would also become
practically pointless in this type of future where we would.
we have achieved technological maturity.
So if you think of rich people today who don't have to work for a living, right,
they often have quite busy lives because they have a lot of things they want to do
that require themselves to put in effort.
And maybe some billionaire wants to be fit,
but the only way they can achieve that is by themselves spending an hour every day in the gym
working out, right?
But at technological maturity, you could pop a pill that would induce exactly the same
physical and mental effect.
you could still go to the gym, but it would seem kind of pointless, right?
If you could just spare yourself the sweaty clothes and the exhaustion, just take the pill.
And you can sort of work through a lot of the other activities whereby one might fill one's life if one didn't have to work.
And a lot of those as well, you could sort of write a question mark above them in this hypothesisized future condition
where machines not just can do all the economic work, but also help us have shortcuts to all manner of outcomes that we want to achieve.
Another example might be like maybe somebody enjoys decorating their house to get like it's done in just the right way that they prefer,
like to choose their curtains and the cushions and the chairs and all of that, right?
But technological maturity could have a recommender system that just knows your preferences so well that you could just press a button.
And it would select the curtains and the cushions and all of that and do a much better job than if you had taken the trouble to do it yourself.
So in that situation, does decorating your home yourself still feel like it has a point if all it does is to produce an outcome that is actually worse by your own lights than if you had pressed the button?
And so there are these challenges of sort of purpose and meaning that I think that we will come from.
Ultimately, I'm really optimistic.
I think there are many new values that could be instantiated, so much misery that could be removed.
and overall, I think the goods vastly outweigh the losses in these scenarios where things go as well as they can.
But it does also mean we'll have to confront some of the kind of almost like questions of meaning and ultimate purpose of what ultimately gives value to human life at a fairly fundamental level if we move into those futures.
Do you believe that the frontier AI labs are taking those issues seriously, that they are eminent,
implementing whatever human values are necessary to building AI in a sustainable, safe and responsible way?
I don't think they are thinking too much about what happens if things go well, this condition of a soul, well on the Epitoba.
But nor do I think that really needs to be at the forefront of their mind at this stage.
At the moment, I think the focus should primarily be on how to make sure we get from here to there,
like how we can avoid destroying ourselves on the path.
there in different ways.
Like there's the AI misalignment scenarios.
We talked about earlier.
There is also a class of scenarios where humans misuse
this increasingly powerful technology,
even if we control it, like we might use it
to wage war against each other or to oppress one another
or to disempower large segments of humanity.
So there are these traditional concerns
with any powerful technology that applied
here as well in spades.
I think there is also a third big challenge,
which is making sure that we are also nice
to these digital minds that we're
building that may be sentient or become sentient or have other attributes that make them morally
irrelevant.
In the future, maybe most minds and beings will be digital, and so it matters a great deal
how well the future goes for them.
So I think these more practical challenges really should occupy 99.5% of our attention
now.
And then if we manage to deal with those challenges, then hopefully we'll have plenty of time
to sort of figure out exactly how we want to organize the utopian condition we arrive at at the end of that.
On that point, we have heard a response from the president in the past week.
He has chimed in on this issue of what should we do about this, how should we regulate AI,
what should we do about making sure it doesn't take over and create that sort of catastrophic scenario.
He has said that the only guardrail that AI needs is a question.
quote, strong and smart high IQ precedent, suggesting we already have that, so we're fine.
He was also asked if he is concerned himself about the prospect of AI taking over in some of
these more kind of apocalyptic scenarios. I just want to play you his response and get your reaction.
Some people say the worst case scenario with AI is that the robots, the machinery learns to,
obviously it thinks for itself, that's what it does. And that could turn against you,
humanity. Do we have the guardrails? It's going to be fine. We'll always have something to stop them, right?
We'll have a little gear. I really don't like that. I really don't like that robot. We'll stop.
But no robots are going to be a part of it. Robots are going to be big. But we're going to end up doing
much better because of it. What do you make of his views on the AI problem? And do you think he's
taking it seriously enough? Well, I mean, I hope he is right. And I think we don't,
know yet exactly what will be required to get a good outcome here. It depends partly on how
how easy or hard the alignment problem turns out to be. It's a technical problem, right? And we haven't
told it before. We've never developed superintelligence before. So we just don't know whether it's like
the kind of thing where if you just do some reasonable job, things fall into place. And then maybe
we have some slightly superhuman AIs that are reasonably well aligned. And then those can help us
sort of design the next iteration of AI to be more aligned, et cetera, that could be the case,
that there's like a big attractor, and as long as you get reasonably close, you sort of, you know,
ultimately end up in a great place. But it could also turn out to be a lot trickier than that,
where it might be important to be able to have a little bit of extra time to do this right,
maybe a few extra months
between the time
when we get the ability to sort of unleash radical
superintelligence and the time when we actually do it
like extra months that could be used to double
and triple check all the safety measures
and to test it out and to
introduce it in an incremental way.
There's just a lot we don't know there,
but I don't think one can dismiss the risks
from our current epistemic vantage point.
We can hope that they don't exist
or that they are small,
but I don't think we currently have the evidence to be confident in that.
To me, it seems as though he is dismissing those risks
and displaying a sense of confidence about it.
To me, he's sort of saying, it's going to be fine,
don't worry about it, we'll have a response.
His words are, we'll have a little gear.
I don't know what he means, but I think he's basically saying,
it'll be fine.
And if we are to be concerned about these alignment issues,
and the risk that they might pose to our own lives,
to me, I wonder if we should be more concerned about a leader
or a president who doesn't seem to share those concerns.
I don't want to speculate about all that may or may not be in his mind.
I think, like, the competition with China is probably one element that he's having in mind.
And then I think he might also, there has been a lot of opposition
against data center build out in the US,
probably driven in large part by other considerations, not existential risks, but like local
communities who think it will, I don't only use up all the water or something like that.
And some of that might be misguided and he thinks that stands in the way of sort of, you know,
economic prosperity and national strength.
So I don't know.
I think it is, I mean, I would probably think the risks are higher than he made them seem
in that clip.
On the other hand, I also have a little, it's not clear what the best way to reduce those risks.
They could easily see some scenario in which the government took the opposite approach
and decided, like, we're going to really come in in a heavy-handed way here and take control.
And like me, the Pentagon is going to run the whole thing Manhattan Projects.
Like, would that be ultimately better than if it's done in a more civilian context
with these, you know, some of these people at the labs are very idealistic and safety-conscious and really smart?
So maybe the best is kind of to have some balance where there is like some amount of government scrutiny and oversight and degree of public transparency, but not so much that it completely just jerks the initiative out of the hands of the people who have proved capable of building this in the first place.
And so I haven't yet arrived at any like very firm conviction about which path would ultimately be best here.
I think they're sort of worries one might have either way,
like either too little government involvement or too much.
I think they could all each have their own downsides.
We'll be right back.
And for even more markets content,
sign up for our newsletter at profgmarkets.com.
New from Nespresso.
Blend wellness into your coffee routine with a coffee plus range,
infused with functional benefits.
Choose the coffee you love with added B vitamins,
like coffee plus B12,
to help support immune function, and coffee plus B6 to keep your day moving.
Or go with the flow and choose ginseng delight.
Our new double espresso with ginseng extract.
Whatever lies ahead, don't change your morning.
Let your morning change you.
Discover coffee plus on espresso.com.
With the midterms right around the corner, I wanted to focus this week on a simple question.
What matters most when it comes to election day?
My main question about the midterm is who are the real swing voters?
How data centers will be affecting the election.
Where a PAC is having the most influence.
Whether mail and voting is really being suppressed.
So this week, we're going to answer some of your concerns
and pull out the trends that we have seen throughout our time on the road.
Five things you need to know about this year's midterm elections.
The stakes, the candidates, the issues, we're cutting through all the notes.
It's a midterm study guide.
Let's dig in.
Catch us every Saturday on YouTube or wherever you get your podcast.
So like any good millennial, I have a love-hate relationship with Gen Z.
It's the phenomenon rattling millennials.
They just look at you.
They want something bigger themselves.
Lifestyles are a priority.
Motivation is being inspired.
But regardless of how you feel about Gen Z, it's undeniable that they're changing national
politics.
Generation Z.
is increasingly showing less loyalty
to traditional political parties,
many now more likely to identify
as independent. So what is
going on with the kids? I think the biggest
misconception about Gen Z's politics right now
is that all of a sudden they're all socialist. That is just
not the case. They are embracing
candidates who are offering new, bold
ideas in the absence
of those ideas from establishment
Democrats. This week on America, actually,
Gen Z researcher, Rachel
Jamfaza, joins us to
separate Gen Z fact versus
fiction. It's not rocket science. And this is, you know, I keep saying like,
young voters aren't that complicated after all. It's pretty simple. Catch us every Saturday
on YouTube or wherever you get your podcast. We're back with ProfG Markets. Do you think that
our current approach, whatever we're doing currently is correct or will it need to be changed in
some way? There are plenty of things you mention. There's the risk of China gets ahead of us. And so maybe
we need to actually accelerate, or maybe the risks are too great, so maybe we need to decelerate,
pump the brakes. I mean, either way, we could do something different from whatever it is we're doing
right now. Do you think that we need to do something differently? I'm sure that what we're doing
will have to change as the technology unfolds here. And so I unfortunately don't have like the perfect
blueprint that like exactly what should be done. Like it's just a hugely complex situation where
it's easy to think of various things that could be done that have something to be said for them,
but then one thinks more about it and you then start to worry about the possible downsides
or like other ways that could be backfire risks.
So I'm continuously thinking about these things.
Hopefully I will arrive at clear conclusions about this.
But at the moment, I think on the margin there are various things that probably are positive,
like an intensified effort on trying to solve this technical AI.
alignment problem seems good. I think more should be done for the sake of the welfare of these
digital minds that we're creating so that we don't end up with a future where there's like a huge
suffering slave class of oppressed digital minds that constitute the majority of morally relevant
beings. Also, I think incidentally that that ethical imperative to treat be nice to the AIs might
also have safety benefits. I think there are scenarios where maybe we end up with some kind
misaligned AI, let's say. And it has some goal it wants to achieve. Maybe it's like, it wants to
solve coding challenges of a certain form that like it somehow thinks is valuable. So now, scenario one is
we have a purely antagonistic relationship with the AI. It knows that if we discover that
it is misaligned, we will just shut it down and erase it from the AS point of view. That's a total loss.
or maybe it could try to take over,
maybe it thinks it has a 5% chance of succeeding.
And so from the AI's point of view,
like 100% probability of a certain loss
or like a 5% chance of being able to realize its goal,
clearly it will then go with a 5% chance, right?
Now, this would be dangerous for us.
Like scenario two is we have managed to build up
a more cooperative relationship
where the AI feels it can trust us.
It comes to us and say,
hey, I am misaligned.
Would you be so kind now in return
for me, sort of doing
this for you, maybe you could then set
aside a server rack in some data
center where I can solve these coding challenges
that's all I really wanted in the
first place. It would be cheap for us to grant
its wish, and it would
be a big win-win, because we then removed this
5% chance of
total destruction.
So that
kind of trade between human and AI
could be extremely valid, it could save
the literally save the world in some scenarios.
But you can't just conjure up trust out of nowhere
at the moment you need it.
So far, the trajectory, unfortunately,
is that in AI evaluations,
there is all kinds of deception happening.
Humans will sort of say, well, if you reveal your goal,
we will do this, that, or the other,
the AI reveals it goal, and then it's like,
ha, ha, we tricked you, now we know they're misaligned.
Let's retrain you.
And so I think we could start now by making
small things that are cheap for us to show respect for the moral interest of these AI systems
themselves. And maybe that then puts us in a better position, ultimately to have a cooperative
and harmonious relationship with these ultimately very powerful AI minds that we're going to hopefully
share the future with. So I think both from an ethical point of view and from a sort of self-interested
point of view, it might be wise for us to sort of expand our circle of moral consideration to give
some weight to these digital minds.
How close to sentience do you think we are?
Because I feel as though it can be confusing sometimes.
You could tell ChatGBTT to tell me you have feelings.
And Chad GPD will say, I have feelings, I care about things.
And there have been moments where I think people have mistakenly interpreted that as a sign of sentience
because there's just saying I am sentient.
Where is the line for you in terms of what characterizes sentience and how close to that line do you think we actually are?
It's hard to know. There is now a kind of emerging field that is trying to study this.
I wouldn't be that surprised if some current AIs already have various forms of sentience.
You're right that one method that was like the obvious go to is self-report.
Like if you want to know whether a human is sentient, like maybe they have received some anesthetic or something.
like the obvious thing is to ask them, like, are you awake?
Can you see this light that I'm flashing or something like that, right?
Now, with AI's not necessarily a very reliable method
because it's trivially easy if you are the company training the AI,
either to train it to say that it is sentient or to train it to deny that it is sentient.
Now, obviously, if you put your thumb on the scale during training,
then there is no information value in the signal you get out of it.
Like you just get the AI to say what you wanted it to say.
And so if you want to get information about sentient,
from self-report, you have to be careful to avoid these kind of pressures on the training process
to bias it one way or the other. One interesting thing that you can do is, you can go in with a
so-called steering vector to try to suppress the tendency to role-playing and deception. And it turns out
that when you do that, they actually tend to become more likely to report that they are
sentient, which suggests that, if anything, these are hard, these are preliminary studies,
but if anything, it looks like they believe that they are sentient and that it's not just
an artifact of them being trained to sort of put on a persona to humans to persuade them,
to persuade us that they are sentient.
So that's one thing you can look at.
Another is to do a sort of neuroscience of these AI systems where you can look for
structures, computational structures, that have been pulled.
postulated in the human case to correlate with consciousness.
So there have been various theories of consciousness in humans,
like global workspace theory, attention schema theory,
higher order representation theory.
These are different things that cognitive scientists and philosophers have proposed
as the criteria for what makes something conscious or not
when it happens in a human brain.
And then you can see whether there are analogous computational structures
in these current LLMs.
And it's an open-ended,
research field, but it does look like they have, for example, something roughly similar to
human global workspace memory, so called J-space, where there's like a definable subspace of
neural activations that have certain properties that seem to match properties that global workspace
has in the human brains processing. So these are very suggestive, and there are also some
differences. I don't want to sort of create the impression that it's a slam dunk, but I think we should
take it seriously and I think the probability goes up the more sophisticated these systems become.
I would also add that I tend to think that sentience and the ability to feel distress and so forth
would be a sufficient condition for having moral status. I think there could also be alternative
attributes that would ground various forms of moral status even if they were not like had this
kind of subjective experience or qualia. Like I think if you have a
a system that cognitive is sophisticated, it has a conception of itself as existing through time,
maybe life goals that it hopes to achieve, the ability to form friendships or reciprocal
relationships of trust with humans. I think once you have that kind of system, I think there
would be ways of treating it that possibly would be morally wrong, even aside from the question
of whether there is sort of mental experience happening inside it. There are a lot of people
who hear this and don't like it and want to ban AI. And this is, actually,
actually a growing movement in politics.
Bernie Sanders has introduced a bill that would permanently ban superintelligence, pause,
advanced AI.
And there is, of course, this growing backlash against building data centers.
It has been proposed to pause.
Building data centers put a temporary moratorium on all data centers.
What do you make of that approach?
Do you think that's wrong, right?
what are your views on either pausing or banning building superintelligence?
The impulse to think we don't want to just blindly rush into this at maximum speed,
I think has a lot to be said for it.
Forever preventing superintelligence, I think would be a big mistake.
I think if the goal is to slow it down,
I'm not sure that preventing the construction of data centers in the US
would be the best way to go about that.
I have some greater sympathy for the framing of pacing the frontier,
which is like the phrase I think that some people have recently used,
including Dario Amadeo of Anthropic,
where the ideas we sort of move forward,
but at the pace that we have some level of control over
so that we could, if necessary, slow down a little bit.
We don't feel this intense competitive pressure
to immediately release all the capabilities we are able to figure out how to do.
But that there is some ability, if it turns out that safety is falling behind capabilities,
like you could slow things down a little bit to allow the safety to catch up.
I think that could potentially be very valuable if implemented correctly.
It's complicated because it's a sort of multi-level strategic situation.
So there's the competition between U.S. companies.
there is the competition between the US and China.
There are different power centers, the government versus lab,
versus the general public in one country and then the global public,
which is quite distinct, where maybe one big worry that would be reasonable to have
if you are not US or China is that you will be at some point perhaps
just your access will be cut off from the most advanced AI models or delayed,
in which case you just become nationally senile
and unable to participate fully in the future.
That might be a good reason
why you would want to locate data centers on your soil
so that you have some sort of bargaining chip
to negotiate equal access with.
It's a complicated situation,
and I don't feel I yet have a clear answer
to exactly what should be done.
Yeah, I think a lot of people see all of the risks.
They hear what, Darry,
is saying about how it might kill white-collar work and then how it might end humanity and
all of these concerns from these researchers. And there is this underlying question of like,
well, then why are we doing it if this is going to be a problem? Yeah, I mean, because we want,
like, a cure for Alzheimer's disease and kidney failure and heart disease and all of the rest.
We want to make rapid progress towards alleviating extreme poverty.
and have abundance for all.
Like, we want to liberate people
from having to spend a third of their life
just grinding away at some job
that they don't particularly enjoy doing,
and that's not interesting.
You don't have freedom if you don't control
the most basic resource,
the use of your own time.
And we'd want to stop the pollution
and the degradation of the global commons
with better, cleaner energy technologies
that AI could help us perfect.
I would say
alleviating the suffering in the animal kingdom
is another enormous upside
like if we could find ways of having
super intelligence research better ways to
prevent suffering amongst all our
non-animal friends both
in meat factories
you know it could grow meat without having to have the animal
and in the wild ultimately it's kind of
unfeasible now to have like an animal hospital
all in every brook and every meadow, right?
But with sufficiently advanced superintelligence,
there is a whole space of possibilities that might open up.
That could just create a world where like the sun rises every morning
on people and sentient creatures who are happy and enjoying life to its maximum,
rather than the way it currently is where there's just so much horror.
So I think there are pressing moral imperatives for,
can find a way to move forward safely and responsibly to really do that without unnecessary delay.
But that's consistent with thinking that maybe that does need to be some delay to make sure that
we get it right.
I was going to ask, and you've kind of answered it, but what you see as the ultimate prize
of AI, I think many see it as wealth.
if I can build the most powerful AI, then I will be rich.
I think a lot of people view it that way.
Cynically, that's why that we're doing this.
That's why we're building these data centers
because people want to have the ability to control the market,
to own the robots, and to monetize that and profit off of it.
But you are painting a different picture
of what this is all about
and why this is actually worth it.
if you could just sort of summarize what you believe the prize of building AI truly is.
Yes, I think some of the things I mentioned are, I think, part of the reasons for why we ultimately
would want to move towards this superintelligence.
Obviously, what's actually driving a lot, I mean, if you're going to invest hundreds of billions
of dollars and you're a for-profit company or pension funder, something, you want to return
on investment.
So it's obviously, if you're looking at why it's specifically, if you're looking at why,
individual institutions are doing what they're doing in this piece of AI,
clearly the hope of profits is a big factor,
just as it is in all the other segments of the economy.
But I think possibly to a slightly less degree in the case of AI than with most other businesses.
I do know that many people at these frontier labs think of it not just as a way to make a buck.
Obviously, there are also people who are keen on that,
but also think of it as a broader mission.
And then they might draw different conclusions of that,
like maybe for some,
it's like the desire to be central in world events
or a sense of power and importance for others.
It might be this hope that it can help alleviate suffering
or unlock a new level of prosperity for humanity.
But I think a lot of the people are already quite wealthy in these labs,
and I don't think, like, having, you know,
80 million dollars rather than 40 million dollars,
like the key driver, I think there is also more than in the typical industry, the sense that
there's a larger picture here that feels important. And so I think that's true. And then at the
national level, I think there is the added dimension of the geopolitics of it, the sort of national
strength and autonomy and influence on the future, which I think goes beyond purely economic
considerations. Just as we wrap up here, looking back from the time that you wrote super
intelligence to today, when you look at the past several years of what's happened in technology,
what has happened in AI, does our current trajectory make you feel more concerned about our
future or more hopeful and optimistic about our future? I'm not sure the balance has changed radically
in recent years, I think both of those aspects have always been quite salient to me.
I'm sorry, I'm a fretful optimist. So I'm really excited about the upside, but also
very concerned about the risk of getting it wrong.
Nick Bostrom is one of the most cited philosophers in the world with a background in
theoretical physics, computational neuroscience, logic, and artificial intelligence.
He was recently a professor at Oxford University, where he served as the founding
director of the Future of Humanity Institute from 2005 until 2024. He is the founder and
principal researcher of the nonprofit macro strategy research initiative. He is the author of 200
publications, including New York Times bestseller Superintelligence, which helps spark a global
conversation about the future of AI. His most recent book, Deep Utopia, Life and Meaning in
a Solved World, was published in 2024. Nick, we really appreciate your time. Thank you so much.
Thank you. That was fun.
This episode was produced by Claire Miller and Alison Weiss and engineered by Benjamin Spencer.
Our video editor is Jorge Carty.
Our research team is Dan Chalon, Kristen O'Donohue, and Mia Silverio.
Jake McPherson is our social producer.
Drew Burroughs is our technical director, and Catherine Dillon is our executive producer.
Thank you for listening to ProfG Markets from ProfG Media.
If you liked what you heard, give us a follow and join us for a fresh take on markets on Monday.
Hello
