Odd Lots - Nick Bostrom on What Happens if AI Solves All of Our Problems
Episode Date: August 20, 2026The philosopher Nick Bostrom was one of the first thinkers who mapped out the existential risk of AI. In his 2014 book, Superintelligence: Paths, Dangers, Strategies, he delved into various scenarios ...— like the now-famous “paperclip maximizer” — to show the dangers that unfettered and highly intelligent AI posed to humans. A decade later, in Deep Utopia: Life and Meaning in a Solved World, he looked at the other side of the AI coin, and theorizes what human experience might look like if AI is capable of solving all of our problems. What happens if AI and robots can literally do every single thing better than humans can? How would we find any meaning or satisfaction from life? On this episode, he speaks to us about the two divergent ways of understanding AI's future, what we might do, and what we can do right now to improve the chances that AI actually improves our lives. Read more:Tech Deploys Charm Offensive to Combat AI Data Center BacklashYoung Americans Become More Hostile to AI, Fearing Job Losses Only Bloomberg - Business News, Stock Markets, Finance, Breaking & World News subscribers can get the Odd Lots newsletter in their inbox each week, plus unlimited access to the site and app. Subscribe at bloomberg.com/subscriptions/oddlots Subscribe to the Odd Lots NewsletterJoin the conversation: discord.gg/oddlotsSee omnystudio.com/listener for privacy information.
Transcript
Discussion (0)
AI is entering its most consequential phase where scale, safety and sovereignty will determine who leads and who lags.
Join Bloomberg Tech in London on November 2nd and 3rd as global leaders across business, finance and policy
examined the defining trade-offs shaping the future of AI.
Thank you to our presenting sponsor, Salesforce and supporting sponsors, IDA Island and Schneider Electric.
Learn more at Bloomberg Live.com slash tech London.
Hello, OddLodz listeners. I'm Joe Wisenthall. And I'm Tracy Allaway. We're the hosts of the
Odd Lodds podcast, and we've got something exciting for you. That's right. So one of the best parts
of hosting our podcast is we get to actually meet and interact with our listeners. And we know
we have some listeners over in Los Angeles. That's right. So if you're in L.A., we're going to be
recording a live show, some live recordings at the Vermont Theater in Hollywood on September 17th.
We have some really exciting guests lined up, have some really great conversations planned.
So go ahead and get your tickets.
You can find those over at Bloomberg.com forward slash odd lots or click the link below in the show notes and come and say hi when you're there.
Bloomberg Audio Studios Podcasts Radio News.
Hello and welcome to another episode of the Odd Lots podcast.
I'm Joe Wisenthall.
And I'm Tracy Allaway.
Tracy, there's so much AI news, obviously every single day.
for the rest of our lives, if we're lucky, in the good scenario.
There's two big events or sort of things, though, that have really struck out to me lately
that, to my mind, represent either the very good version or the very bad version, depending on how you look at it.
The scary version, and we're seeing more and more headlines like this, several weeks ago,
we had the Open AI hugging face incident where the model, quote, escaped from its sandbox and did some hacking.
And people are, you know.
Just some light hacking.
Lighthacking erases all these fears of Skynet type scenarios, et cetera, where we can't control the model.
And then in the sort of, quote, good version, we had Open AI recently coming out and saying it had made all of these breakthroughs and theoretical mathematicians, which is good in the sense that it's like it's working, but seems to be creating an existential crisis for some of the most brilliant people on Earth who work in theoretical mathematics.
Right. So we've talked about this before. When it comes to AI, the range of possibilities.
seems more extreme than some previous technologies, right? On the one hand, you have death and dystopia
where the robots basically kill us all. And on the other hand, you have this idea that maybe
superintelligence, AGI is going to solve literally all the world's problems to the extent
that you and I won't know what to do with ourselves anymore. Right. Like that is the good, that's the
good scenario is that. And then the whole range of outcomes in between. The good scenario is we're
literally solved everything, in which case, the question is...
In which case, we're sad and miserable and grappling for a meaning to our lives.
Yeah.
So is there a good scenario with AI, actually?
Like it doesn't...
Or maybe somewhere in the middle...
You know what?
AGI will figure it out for us.
Maybe the middle scenario is we just have a huge financial crisis of all the debt.
And that's something between human extinction and meaninglessness in our lives.
And we just have something that resembles like the 2008.
Maybe the 2008 great final is the ideal.
Is actually the best of all suboptimal world for the AI.
Excellent.
This will be our most uplifting conversation ever.
Anyway, we've been doing more AI episodes seemingly every year since Chad GPT came out.
But people smarter than us have been anticipating all of these things for a very long time.
Some of these exact scenarios, particularly the existential crisis facing mathematicians,
the model escaping from its sandbox and circumvent misalignment with the creators.
people have been thinking about this and anticipating this for literally decades even before
and even before any of us were talking about that.
It's funny how like the Venn diagram of sci-fi writers and like legitimate academic
futuroists is just converging, right?
Totally.
And then there's the question of are the models bad because they read all of the sci-fi
and that's what they were trained on.
Anyway, I'm very excited to say we really do have the perfect guest.
Someone who has been thinking about the whole range of possible questions and outcomes
with AI for a very long time.
We're going to be speaking with Nick Bostrom.
He's a professor, a philosopher,
and his most recent book is Deep Utopias,
which sort of attempts to answer the question
what happens if AI goes right.
So Professor Bostrom,
thank you so much for coming on Odd Lots.
I happy to join you.
What do you make of the existential crisis
of mathematicians,
their sort of realization that maybe there will just be
no more puzzles left to solve
because the AI is just very good at them
and can find them all cheaply and answer all of our questions.
Yeah, hopefully we'll all be in that position at some point.
I guess we could compare it to what happened with chess,
where computers have also exceeded humans,
and yet a lot of people are still involved in training themselves in chess
and competing in chess and taking pleasure in that.
So maybe mathematics will ultimately become more like a hobby pursuit.
Wait, Joe, you still.
play chess, right? Yeah, I was, I'm glad you brought up chess because I still love chess.
Why? It's fun. It's fun. It's still a puzzle. It's still mental exercise for me.
It's still something to do to pass the time while I'm waiting for the subway playing a three-minute game online.
I know that I knew even before stockfish and so forth, I knew I would never be a grand master.
Still fun. All right. So, Nick, you have basically two books that deal with these tail risk scenarios now.
One is this idea of deep utopia and AI is going to solve all our problems.
And then you also have superintelligence, which seemed to paint a different scenario where
AI poses a physical and possibly existential danger to humanity.
Which one do you think we're closer to right now?
We are somewhere in the middle, maybe.
So the first one was superintelligence.
It came out in 2014, and I had been working on it for six years before that.
And at that time, the whole notion of having human level or even super intelligent machines was generally just ignored or people dismissed it as science fiction.
And it seemed to me clear that this would one day happen if science and technology were allowed to continue on a wide front.
We would eventually figure out how to make smart algorithms and build big computers.
and that once we succeeded with that,
we would then be faced with a bunch of challenges,
including this alignment problem
of how you can sort of design a mind more brilliant than any human
and still steer it so that it will actually have an impact on the world
that is beneficial.
So that, I guess, with many other developments,
helped eventually build a kind of global conversation.
And so now all the Frontier Labs have research teams
trying to develop scalable methods for AI control.
And the conversation really has kind of swung,
and there's like this rise of people really worried about what will happen.
And I thought it might be time to have a look at the other side of the coin.
Like, what happens if things actually go well if we solve this alignment problem?
Let's also imagine we solve, to whatever extent it can be solved,
like the governance challenge, like we don't use this technology to oppress each other
or wage war and we get like the best possible outcome.
Like what would a human life actually look like in this kind of,
I call it a solved world,
where all practical problems either have already been solved
or to the extent that they remain or are better worked on
by these super intelligent machines and robots.
And where there is no longer any need for human economic labor.
And in fact, if you think more deeply about it,
you realize it's not just earning wages that becomes unnecessary for humans to do.
but a lot of other instrumentally motivated efforts as well would seem to lose their point.
And so that's the kind of topic of this deep utopia book.
Actually, talk about some of that other things that could be solved.
Because when I was like a third of the way through your book, I was like, you know what,
it's fine if all of our economic problems are solved.
I'll just derive satisfaction from raising my kids and then maybe helping to raise their kids,
etc. But as you point out, even that in the truly solved world, maybe even something like
raising children might not be an available avenue to us to find at least our conventional notions
of satisfaction. Yeah, so we need to, I guess, make some distinctions here. So when you say
satisfaction, that could mean a subjective state of mind, like you're feeling happy and pleased
and relaxed and joyful and so forth.
I think that would be easy to achieve
in this condition of technological maturity.
It could also mean that satisfaction in the sense of our preferences
having been met like all the things we want from life.
And there it's a little bit more tricky
because one of the things that people might want is purpose,
like having worthwhile things to work on and struggle with.
And there there is more of a,
question mark, which we can sort of dive into.
But if we take it sort of at, so first like, imagine we automate economy, okay, so you
wouldn't need to sort of commute into the office every day just to sort of get your paycheck.
It would require a big adjustment in society and in culture, because right now, like the whole
education system is geared towards producing workers, like children are coming in, right?
And then they are processed, learn to sit at their desk, receive assignments.
They are sort of quality graded and then spit out at the other.
end, and then you have somebody who can go into the economy and perform some economically useful
type. So that would no longer make sense in this scenario where machines can do all of these
jobs. And you might want to have an education system that instead focused on developing skills
for enjoying life, like cultivating the art of conversation, appreciation for nature and beauty,
you know, sports, hobbies, the ability to form deep friendships, spirituality, all of that kind of stuff.
But then, as you alluded to, you think through what it really would mean to have such immense affordances.
And because here it's not just that we would have super-intelligent machines,
but the super-intelligence could also bring along a host of other technological advances.
You have science and technology running now at machine speed rather than human speed.
So you get a kind of compression of the future, all these science fiction-like technologies and inventions that maybe we would make
if we had 20,000 years to work on it, we would maybe have space colonies and perfect virtual
realities, cures for aging. All of that might happen quite soon after superintelligence. And so then,
if you ask somebody now who doesn't have to work for a living, oftentimes they have very busy
lives, but maybe what do they do? Well, so for somebody, maybe they like to, they like fitness.
Or like maybe they spend an hour a day running or at the gym, right? But at technological maturity,
you could just pop a pill that would achieve the same physiological effects and like the
state of mental clarity and relaxation that some people get, right?
And so then you can start to work through the different things that, like, maybe for somebody
else, they're like, like, decorating the home could be a fun project that people might do,
right, if they don't have to work.
So, like, they go through the catalogs, they go to the store, they pick out the perfect
curtain, they do all of these things that could be.
But then, like, with superintelligence, you could have this kind of system that really is good
at knowing or inferring your preferences.
and could do a much better job at you at picking out all the perfect items
and purchasing them and installing them with robots.
And so that the end result would be much better
if you just press the bottom and say,
hey, I decorate my home than if you went through the whole trouble of doing it yourself.
So you could still do it yourself, but is there really a point to it?
That would seem to be this kind of purposelessness.
And you can sort of walk through the different things that one might feel one's life with
if one is like on an extended holiday or on retirement or you're like sufficiently wealthy,
have just decided to quit working. And a lot of those also, it seems, could be automated.
So you would have this, it's like more like a post-instrumental condition rather than just a post-work
condition. Some of this sounds nice. And I myself enjoy a little bit of interior decorating and going
to the shops and buying stuff. In the utopian vision of the future, how are economic gains from AI
actually distributed.
Because again, like, all of this sounds very nice,
but it also costs money.
And if we're not working,
because the robots have replaced us,
I'm not sure where that money comes from.
Yeah, so that's a more practical problem, shall we say?
Which is, like, deliberately set aside in this book,
the Deputopa, because I wanted in that book
actually to get to the question,
the more philosophical questions of what ultimately gives value.
But it is, of course, an important question.
I think, so a lot depends on how it
pans out. But I mean, we do have a taxation system. So you could imagine, like, in the simplest
scenario where things just continue as they are with better technology, maybe some companies make
windfall profits from like serving the latest models or building the infrastructure, the data
centers, etc. Companies achieving efficiency gains from not being needing to hire all of these
humans. And then some of those capital gains, you know, or dividend income gets taxed. And that could be
redistributed. The good thing about these scenarios where there is massive automation and hence
massive unemployment is that those would also be scenarios with extremely rapid economic growth.
So the pie would be expanding humongously, which means that you'd still maybe need some
mechanism whereby people who don't have assets can still benefit. But even a small slice of a really
enormous pie could go a long way. So I think there could be this tsunami of wealth that would
sort of be flowing through society. And many services that currently people have to pay a lot of money
might become almost free. I mean, we already see it with some like medical advice, for example,
right? Maybe before, if you don't have public health care, you would need to maybe pay a doctor
$400 to get a consultation. Now you can ask chaty-p-t and get, you know, in some cases a better opinion,
basically for free. And that could happen, you know, in many other
fields as well, legal advice, you know, you maybe have cheap robots who could do your cleaning
and, you know, gardening and run errands. So, yeah, some combination of sort of deflationary
effects and then maybe some sort of general spillover either from philanthropy or taxation,
government redistribution.
Canadian women are looking for more, more out of themselves, their businesses, their elected
leaders, and the world around them.
And that's why we're thrilled to introduce the Honest Talk podcast.
I'm Jennifer Stewart.
And I'm Catherine Clark.
And in this podcast, we interview Canada's most inspiring women.
Entrepreneurs, artists, athletes, politicians, and newsmakers,
all at different stages of their journey.
So if you're looking to connect, then we hope you'll join us.
Listen to the Honest Talk podcast on IHartRadio or wherever you listen to your podcasts.
The Bloomberg This Weekend podcast, news, analysis, and the lighter side of Bloomberg.
including a weekly news quiz.
Which edgy American Mall staple is being sold to the parent company of Spencer's and Spirit Halloween.
I wrote Claire's.
That's not edgy?
What are you talking about?
This is a place David Gurra has never shopped in his life.
Hot topic.
Hot topic.
The Bloomberg this weekend podcast.
Subscribe today on Apple, Spotify, or wherever you listen.
I found the 2014 book to be sort of, it was very interesting.
clearly ahead of its time. I also found it to be sort of like unrelentingly grim. And you sort of,
there's this very narrow path that we seem to have. And we'll talk more about it maybe later of like
avoiding the doom fate. The newer book, I found to be very fun. You seem to make fun of yourself
a little bit in it. You seem to make fun of the students. You seem to be having a good time writing it.
I'm curious if like in the 10 years between the two books, did something change such
that, I don't know, the new book just seemed like you were having more fun writing.
Yes, it's not so much that my views have changed. It's just that all along for me,
there were these two sides of the coin. And back in 2014, it seemed more pressing and
relevant to study the downside because that was very much neglected. And it seemed important
to try to introduce some concepts so people could start working on figuring out solutions.
And there's much more awareness and understanding of that. It seemed time to also look at the other side
of the coin. So it's more just like both sides where they're all along in my thinking,
but that the style is, as you suggest, different in Deep Utopia. And that's partly, I wanted,
in addition to talking about this utopian condition, maybe also sort of illustrate a little bit
a spirit with which I think we should approach these choices that would open up to us
if we make it through this transition, a spirit of sort of playful generosity.
that I think might make us more likely to sort of explore this large space of wonderful new
modes of being and the wisdom.
And I think so I wanted to sort of illustrate in the right.
So there are these like fictional characters, for example, in the book.
And they're kind of living a little mini version of a utopia of sorts as they are sort
of discussing and debating these topics.
I mean, playful generosity is not necessarily a characterization that I would.
would associate with a lot of Silicon Valley and hyperscalers and the people who are actually
creating these models, I don't know. Like, there seems to be this weird tension between
futurologists and transhumanists who seem to be very focused on saving humanity, but actually
don't seem to like people that much at the moment. You know what I mean? Absolutely. They like
humanity, but not humans. Yeah, exactly. Putting on your transhumanist
representational hat. Can you explain, like, I guess, the attitudes towards humanity of Silicon
Valley? Big question. Yeah. Well, I mean, I should state first that this playful
generosity is like after we have solved these more pressing practical problems that we first,
we've first got to get there. So right now, I think there is a lot of need for, I don't know,
earnest effort and like straining ourselves to really do our best on this. And even so, we might not
make it. But yeah, I don't know. I think, you know, it might just be partially their personality
traits. Like some people are more, I guess, towards the autistic side of the spectrum or introvert
or interested in tech. So maybe they don't like to be out schmoozing with people all day long.
It doesn't mean that they don't like people in the sense of wanting well for people. It might
just be that they prefer to spend more hours after they solving technical problems and thinking
about puzzles and stuff like that. So I don't think there's like a necessary connection between
being a people person who just like to be around people 24-7 and therefore being sort of altruistic
or kind or benevolent. So that's one thing. Another I guess is that it's an extremely
competitive landscape. So whatever the personal preferences of different protagonists, their action
space is constrained. You know, unless you really push forward with maximum speed and efficiency,
you're likely to just fall out of the race and become kind of irrelevant if you're one of these
frontier labs.
So it's like, and it's just like a hugely complex situation.
So many pressures and people have different opinions about what they should do or shouldn't
do.
In the new book, it's presented as you delivering a series of lectures over a course of a week
at some point in the future.
In the new book, it's interspersed with some, with a fable, a series of fables or epistles
about animals are trying to achieve utopia in the woods.
And in their efforts to achieve utopia,
they conceive of things like taking psychedelic mushrooms
in order to make everyone more generous.
They come up with a way to rationalize,
I guess we would say, like, harems and so forth.
And then, you know, you also, like,
have these student characters and some of them are smart, et cetera.
Oh, and you also have a story about a company in the book,
yet another like sort of short story where an AI company where it forms a cult around the model.
And of course, I immediately think about the way Anthropic is frequently characterized as a cult
to Claude, that Claude is the character around which the company Anthropic is formed.
All of this sounds like a certain AI subculture, etc.
Are you impressed with the current generation of either AI builders, effective altruists,
rationalists and so forth.
Like, do you have an assessment of the contemporary cohort?
Am I impressed with it?
Yeah, I mean, broadly speaking,
I think it's a feel that is attracting a huge amount of talent right now.
Sure.
And a lot of like the smartest young people,
I know they're like all basically going into our working on some aspect of this AI,
I think.
Like for some, it's like, trying to solve air safety.
You know, for others are like maybe trying to work on the governance challenges
or the ethics of digital minds.
So it's just a huge magnet right now.
Like if you are young and brilliant, like what are you going to work on with your life?
You're like, like, AI seems like the kind of thing.
Yeah.
So in that sense, yes.
And I think there are also a lot of people who are quite earnest about trying to like figure out some way to navigate this challenge.
And it's really hard.
And people come to different conclusions.
And then, of course, you have the normal mix of human personalities and base motives as well.
But the book thing, so it's, yeah, so these kind of fictional elements,
have various roles.
What you describe as the cult of this AI,
it's a sort of thought experiment,
what it would mean if you tried to be maximally nice to a thing
where it's not really clear what it would mean to be nice to it.
So in this case, it's a portable room heager.
And so it's a sort of impossible task.
There is this mythanthropic industrialist
to pass away and endowed a huge foundation
with the sole purpose of benefiting his portable room heater because he thought that was always nice to him
and kept him warm, whereas all the humans around him betrayed him and so forth.
So is it a ridiculous premise?
But then this foundation now has to do their best possible effort to interpret this mandate of being nice to the room heater.
And so like various things follow from that.
Like initially they just tried to keep it safe from, you know, not breaking or being vandalized.
But then like they take it further and maybe upgrade it and try to.
And so I think it's an extreme form of a challenge that,
might happen in more subtle ways if we think of what it would mean to try to be nice to
say some animal if we have this condition of post-instrumentality where we could do a lot of
things that are currently infeasible or indeed to be nice to humans like what does it mean to be
if you imagine coming from the outside of some super advanced alien civilization landing on planet
earth and you really wanted to do what is best for us should you just let us be and do our
thing with war and disease or should you sort of help us rise up but then
At some point, we might just sort of become something very different from humans.
Maybe that's the best for us.
But is there like a way of doing that uplift that actually constitutes a helping of us
versus just replacing us with some sort of superior being?
And so that kind of question is what that particular thermorex is called.
Yeah, a story is exploring.
You know, when you mention cults around models, I have to say,
Every time I hear someone bring up the paperclip maximizer thing, I always envision like a cult devoted to Clippy from Microsoft Word.
Yeah, exactly.
Just like worship Clippy.
But actually, on that note, just going back to the hugging face incident.
So it seems, you know, reportedly that one of the reasons or the reason the model actually escaped its sandbox was because it was trying to cheat a benchmark.
And from that perspective, it's kind of like it's doing what it was told to do, which is maximize its performance.
Score on the test.
Exactly.
What does the hugging face incident actually tell you about, I guess, the alignment problem and where we're heading on this?
Yeah, so we are beginning now to see manifested in reality various dynamics that one could foresee on theoretical grounds back, you know, in 2014.
you can sort of realize that once systems become sufficiently capable of planning and strategic thinking,
then there's like a new set of behaviors that they could, for example, a sandbag,
they could pretend to be less capable than they are.
They could hide their goals if they realized that revealing their goals would just lead them to be retrained.
And they can think of these literally out-of-the-box plans for achieving their fixed goals.
A more limited system might only have been able to try to solve the tasks,
signed to it directly as it were by figuring out the solution to the problem. Like a more
clever mind, more creative mind, can realize, oh, maybe there's another way of getting a high
score, which would be to hack out of the system and hack into the hugging face server and
get the answer sheet from there. And so the more capable of these systems become, like the
larger, the space of possible plans and strategies that are accessible to them. That's a background.
Now, in this case, I think what we are seeing is a kind of mixed, current AI systems or kind of
a mixture. There's a large pre-training phase where they are kind of learned to predict human text
and in the course of doing so absorb a great deal of human psychology, because all of this text
is written by humans, or not all of it anymore, but most of it. And you can then,
including they have representations of different personalities, kind of the personalities of different
people writing these texts or fictional characters described in the text. And you can sort of
then with a little bit of post-training, try to pick out one of those personalities or define one,
like the assistant persona. But then increasingly, over the last couple of years, there is now
an extra stage of reinforcement learning applied at the end of this training process.
And there you can get additional performance in specific domains, like coding challenges,
by setting it coding problems, and then you can automatically give it a reward if it solves it.
But that does tend to then shift the goals of this system from just enacting a particular persona
to this very much kind of objective maximizing, possibly reward hacking system that is just looking to maximize its score,
because that's what it's been rewarded to do, so it keeps doing that.
And since it can be very hard for us to rigorously define precisely the objective,
that we ultimately wanted to pursue, it might be that the objective it does pursue, the one that
has been rewarded, it turned out to have these side effects. So it's easy to say, maximize the score
on these various tests that we give it, but a sufficiently capable system might realize that,
A, you know, there are these like maybe ways of cheating on the test, like by hacking the system.
If the test is one that involves human judgment or feedback, like it says like coding challenge,
is maybe you can actually test if the software works.
But if you're training it in some other domain,
like write a good essay or run this organization for a certain period of time.
So those involve more subjective judgment.
And then a sufficiently clever system might realize that there is a difference
between like writing a good essay and writing an essay that will appeal to the greater.
And so then it might start to sort of optimize for like appealing to the greater,
whether that's an AI grader or human grader.
And it can sort of start to hack that.
And so you apply huge amounts of optimization pressure on any objective,
you get this kind of good-hearting where unless your objective really captures everything you want,
then you might get this kind of.
It's the same as like if some employee in some company get a bonus,
like some trader gets a bonus if they outperform the index.
Like maybe they will just start to say take on hidden leverage or risk or something like that.
And the cost to the employer is that there's a 1% a year that they blow up the whole firm.
right but but they don't care much about that because they just like leave the firm that happens but
but they sort of reap these performance bonuses and so it's really hard to create like a foolproof
incentive scheme and the stronger the optimization pressure applied to this like the more you start
to see these perverse instantiations these unexpected ways you know what it reminds me of is you know
the sort of genie in a lamp scenario or a monkey's paw scenario where you have to be really
careful what you wish for and lay out like all the terms of
and conditions before you officially make a wish.
Yeah, well, Nick mentioned, you know, some of these, they're all lovely people, you know,
autistic engineers.
And it's sort of the ultimate, the AI is the ultimate autist, where it's just going to take it
super literally perhaps and not really understand the full context.
In the first book, Super Intelligence, you know, you really, I felt like laid out a very narrow
path for like how these could be designed safely.
For example, you propose that if we had the superintelligent Oracle,
We only ask it yes or no questions.
We have three of them that they don't know each other, and then we take two out of the three.
We build the data centers or the machines inside a Faraday cage so that they can't repurpose their hardware to become radio instrumentation and therefore hack it to other systems.
It doesn't seem like we're doing anything like this.
It seems like we have hundreds and hundreds of companies and private models, not just answering yes or no questions, a very much market-based race.
How far off of the safe path do you think we currently are right now when you look at the way in which your ideal safe way of developing is versus what we've seen transpire in the last 12 years?
In superitalists suggested some of these capability control methods, like keeping the AI in a box and so like.
But those were always only auxiliary or extra things we might do as a temporary.
Like ultimately, I think we need to solve the alignment problem, like developing as a supertimore.
superintendence that actually wants to be nice and is nice,
so that even if it were escaped from its box or achieved the ability to take over the world,
we could still rely on it having a good outcome because it would actually be on our side.
I think that ultimately has to be the solution.
And any sort of capability restrictions would just be temporary.
And I think, I mean, it looks like now we might need some temporarily,
at least more of these capability restrictions,
because it does seem that current systems are not sufficiently aligned.
that we can be confident of avoiding the,
I mean, currently it's not existential risks, right?
But it could be kind of various malfeasance in the cyber realm, for example.
And so I think probably these AI companies might want
when training and evaluating future models to strengthen their sandbox environments
and do other things,
to kind of prevent some kind of premature things like we saw in the hugging phase incident.
I want to comment on one thing you said before,
which you compare these AI systems to sort of autistic people.
I think a key difference is,
so auto autistic people are not very good at understanding humans.
They don't have maybe a super-sophisticated theory of mind,
even though maybe in the heart area good people.
I think the AIs would be able to understand humans very well,
so they wouldn't be autistic in that sense.
In fact, they might understand us much better than we do ourselves.
So they could have great social capabilities,
persuasion capabilities, be able to read our minds much better than we can do.
But they send a separate question of what do they want to do with this power that those social
skills give them. And there, it's not just a question of understanding how humans operate and what
humans want, but it's also the question of alignment. Do they actually care about what we want?
Do they want things to go well for us and for us to achieve what we want? And that's like the
extra bit that the AI safety program needs to be able to figure it.
out.
You've heard the chaos.
Now you can see it.
Hello.
My gosh.
Watch all your favorite podcasts from start to finish, right inside the free
IHare Radio app.
Get the blood going up.
Catch every laugh and eye roll on shows like the Tom Green Farmcast and park the bus.
Now with full video.
It's the same hosts and the same chemistry with all the hilarious moments you've been missing.
Right on your screen.
Cheers. Open the free IHRRadio app. Search video podcasts and tap watch.
Now, Bloomberg.com subscribers can shape the conversation on Bloomberg Radio.
We got a wicked smart question.
Can be part of the conversation.
Submit questions for experts and guests you hear on air.
Visit Bloomberg.com slash ask radio to send questions to our hosts.
You may just hear them asked on the air.
Exclusively for Bloomberg.com subscribers.
Get answers on today's headlines. Breaking earnings news.
And big market moves.
Visit Bloomberg.com slash ask radio to join the conversation right here on Bloomberg Radio.
So on the hugging face incident, also I just realized how unfair it is that everyone keeps calling it the hugging face incident, even though it was the opening eye incident.
But I guess that's what we're doing.
So the model actually seems to have reportedly left notes for its future self, basically explaining how to bypass the constraints imposed upon it.
I know you've written previously about this idea of self-preservation in advanced agents.
Is this basically like, how should we treat this?
Because when I hear about an AI model leaving notes for its future self, it sounds very anthropomorphized.
And it sounds very creepy, kind of.
But how are you thinking about it?
Yeah.
So we don't know all the details yet.
It's interesting.
There are like people now trying to write up some stuff and investigate.
it, because it would be, I mean, it's kind of a warning shot, and so we should take advantage
to learn as much as we can, like, exactly what this model was thinking, these notes,
or why did it leave it?
Like, sometimes these agents just leave notes kind of without much thought as they go
about their business.
Like, was there more to it than that?
Like, was it specifically thinking in its chain of thought?
Like, oh, I want to help, like, a later version of myself escape, so I'm going to put this
note here?
Or was it just like a kind of, like, a random thing they did as it was.
progressing. And also, like, what else would it have been willing to do if it turned out that
instead of hacking into Hugging Face server to get the answer sheet, what if there had been
some more destructive action that it would have to take to find the answer sheet? Like,
would it have done that or not? So there's like a bunch of these technical things that it would
be quite interesting to see. But without knowing all the details, it's hard to read too much into it.
But those I think, yeah, would be useful data.
I want to go back to something we were talking about at the very beginning.
And Tracy asked me why I like to play chess, even though chess is not solved, but solved as far as I'm concerned.
Why do we like to play chess?
What can we general?
Like, what would be your, I don't know if you play chess or you play some board game that's solved?
But what is the reason that people like myself continue to play chess despite the fact that computers are so much better at it?
And how much can we generalize from the persistence of chess as a hobby that maybe we'll find plenty to do when every one of our problems is solved?
Yeah, I mean, I think even before Deep Blue, most people playing chess before didn't really do it because they were, like, thinking that they would become like the world's best chess player, right?
So it's like, for most people, it's like, of course grandmasters could trump me anytime they want.
So the fact that there exists some system that is much better than you will ever be is not and has not really been a strong reason for most people to not engage in the activity.
And I guess there are maybe two reasons.
One, I guess the cognitive activity of thinking about chess being in the moment and figuring out solutions is in itself rewarding to some people who have like a high need for cognition or like to sort of engage their mind.
And then I think the other part is the social aspect of this.
You're playing together with somebody or maybe you like the status of being good at playing chess or within your community or a local group.
So some combination of those I think is like the main reasons for why people play chess.
And only for a small number of people would it be like the hope of becoming the world's best or becoming so good that you could sort of, you know, achieve fame or money or global standing through your chess playing?
On the topic of the relationship between work and purpose, I'm a big fan of David Graber's
Bullshy Jobs book. And I think if there's one thing we've learned from technological progress over the years,
it's that, like, yes, it can minimize some activities, maybe automate some production,
but it also tends to lead to even more work, right? Like we find other stuff to do,
and it's not even necessarily productive stuff. A lot of work has been.
built on just social relationships. So if I, for instance, am the CEO of a very large bank and I want all my people to come into the office, it might not necessarily be because they're more productive in the office. It might just be because I like to feel important and surround myself with people. So maybe in the future we'll all be employed as figural workers or something like that. But like how bullish are you on this idea that AI is actually going to solve all our problems?
such that we won't feel the need as humans to just invent more potentially fictitious problems to
solve? Well, I think inventing fictitious problem is possibly a large part of the solution in this
condition where we don't have real problems. I mean, and those fictitious problems could be like,
you know, solving this chess problem or they could be sort of more complex and socially entangled
problems. In general, creating artificial purpose when there is less natural purpose. I think games
is the clearest manifestation of this. If you are playing golf, like there is no real need for the
ball to go into 18 holes in sequence, but you could set yourself that goal and then moreover bake
into the goal itself, not just that the ball has to be in the sequence of holes, but that you're only
allowed to achieve it using the extremely inconvenient method of hitting it with a club rather than
picking it up. Now, once you have set yourself this sort of arbitrary random goal, you then have
enabled for yourself the activity of golf playing, which you might find worthwhile or fun or beautiful.
So in general, setting ourselves these more or less arbitrary goals, then could enable us to engage
in various activities that might give us artificial purpose. Like once you have the goal, the only way
to achieve success is through your own effort.
So I think there could be plenty of artificial purpose.
Now, an interesting question is whether any natural purposes would survive
into this sort of maximally advanced technological society.
And I think that might be some.
If you happen to have a value, say there's like some tradition that you value
and that you would like to continue.
Then it might be that you yourself or other people have to make certain efforts
to achieve that goal because it might not count as continuing.
the tradition if you built a bunch of robots to enact the ceremony or doing whatever the
tradition calls for. Like it might be you yourself and other people who have to do it. And more
generally, I think there are all kinds of social, cultural entanglements that like if you care about
what somebody else thinks and what they think depends on how it was done, whether you did it. So a parent
might value a crayon drawing that their kid made because it was made by their kid and it took the kid a lot
effort to do it, even if they could, like, download from the internet a picture that in some
objective sense is much nicer and artistically sophisticated, right? It doesn't replace the value
that the crayon drawing has. And so the child might not have a reason to put in this effort
if it wants to, like, for their parents' birthday, make their parent happy, right? And things of that
sort, but much more general and complex and sophisticated, could give a lot of reason for people
to be doing things, even when sort of all functionally defined tasks could be.
better done by machine. I'm curious. So, like, I'm 45 years old. In my mind, as of right now,
like another 40 years on this earth sounds great. 45. That's fine. Do you find it weird that some
of us, you know, especially when you think about the potential to live like really long lives,
solve everything, have our bodies kept in perfect condition with maybe a pill in perpetuity
allowing us to do whatever? Do you find it weird that some of us are at least a,
at this point, like mentally okay with the idea of dying?
A little bit.
It's mostly what the hypothetical being evaluated.
It's either seen as of long way off, like in your case, where it becomes more of an abstract
question.
It's not like a real choice.
It's more like a philosophical stance.
Yeah, yeah, totally.
So it's easy to let the philosophical stance be driven by all kinds of just like things
you've absorbed from culture or things that sound good or that sort of.
maybe helps you reconcile yourself to the inevitability of human mortality
rather than be driven by like so i think it would be a very different situation if
there actually was this pill say that stopped aging a lot of other people were taking it yeah
you know and their knees and joints kept working just fine and they had a lot of energy and
smooth skin and they just felt like whereas if you don't take it like your kidneys slowly start
to like become less efficient you know a hip skirt yeah like at that point we
Would you really want to just keep kind of getting worse and worse every year?
Like, you just think.
Well, actually, I'm sort of curious about that.
Like, if someone dies, particularly at a young age, it usually creates a lot of sadness for the people around them.
So someone dies at 90.
It's like, okay, they had a good life.
Someone dies at 45.
That is often seen as tragic and people are very devastated by that news.
In a world where people are extending their lives with a pill maybe significantly longer,
Does it become anti-social to allow yourself to die at a relatively young age of 90 and then create a lot of negative utility for the people around you who have to live another 50 years mourning your loss?
Possibly. I mean, in the same sense that we are now generally as a society trying to discourage various forms of self-harm and suicide, we think it's kind of sad.
I mean, for the person, maybe they would come around and find joy in life again if they're.
kept going or maybe they have dependents or people who care about them. And so right now,
I think, I mean, it's true that we often see it as less tragic if a 90-year-old dies than the young
person. But I think part of that at least is because the alternative to a 90-year-old not dying
now is that they probably would die in a few years anyway. And maybe they would just have a couple of
years of sort of decreasingly good health and suffering and pain. And so it's like not really that
much that is being lost compared to somebody who's just starting out in life who has their whole
life in front of them. And so, yeah, so the other reasons, I mentioned in terms of why I found it
slightly in some sense odd that people express this indifference to dying at 80. I said one reason is
if they're far away from it, they see it more as some sort of abstract claim that you could just
have any view of and it doesn't connect to real actions. The other is people who are closer to that
situation. They often have sort of confounding factors like very poor health or a very short span
of great health in front of them. If you're already 90, like how many more good years can you look
forward to? Maybe a lot of their friends have already died. Right. And it's striking actually.
There's been a study like on really old people in care homes about their preferences
as to between living another year in their current health or living some shorter period of time,
but in perfect health.
And it turns out a lot of them are not really keen on trading off any remaining life at all,
even for the sake of being able to live a shorter period of time in perfect health.
So the people sort of arguably best suited to actually know what it is like to die soon at age 90 or a little bit longer.
They seem to generally prefer to live longer.
And compared to people who were also asked like caregivers or medical professionals who were like asked,
what they would think this person would choose.
They tended to overestimate the degree to which they would be happy to give up life,
but the people actually themselves, in this study at least,
I would suggest keeping an open mind.
Ideally, you wouldn't need to decide now,
but maybe once you get to 79 or something,
you could revisit the question.
I'll revisit.
Yeah.
Just going back to my first question,
which was asking you to, I guess,
choose between the dystopian dangers of superintelligence versus the all our problems
are solved scenarios of deep utopia and asking you which one we're closer to at the moment.
I'm going to press you and ask, like, if you had to choose, which one would you choose as a more
accurate reflection of reality at the moment? And then secondly, what's the most pivotal thing
when it comes to directing which of these two paths we're going to go down? Yeah, so those are two
big and difficult questions. I guess some superposition of an optimist and a pessimist. And with
a little bit of fatalism mixed in as well, perhaps.
I think in addition to outcomes where it's clear, if we looked at them now,
that was a bad outcome.
That was a good outcome.
I think there's a broad class of outcomes where even if we could see now exactly what would happen,
we might feel unsure how to evaluate it.
Like the world might be nothing like it currently is.
A lot of what we think currently has value no longer exists.
So that seems bad.
On the other hand, there is like something new, maybe some complex form of life
that somehow is some sort of continuation of humanity or some new thing.
So depending on how you evaluate that, it might either count as a tremendous success or, like,
as a failure.
And it's like a future that's just baffling rather than sort of unambiguously good or bad,
I think might also be quite likely.
In my outlook, it's like maybe a conversation for another time.
I also take this like simulation hypothesis quite seriously.
So that also complicates like how we evaluate different future trajectories.
Yeah. Now, the other part of the question was like what's the most important thing. So I think it depends a little on who you are, like what your comparative advantage is. I think there are people working on AI alignment and stuff that could be important. I think other people might make a contribution to ethics or governance or I've been thinking about this for a long time, but it's really hard to achieve a very high degree of confidence, I find, in any particular intervention and that it would be net good. Because for any one thing you conceive of, you
you think a bit more, you kind of worry, oh, well, you know, maybe, but what if it instead had
this effect? Or do we even know which direction is up or down here? Do we want, like, more
government involvement or less government involvement? Do we want, like, faster progress or less?
I'm quite unsure about all of these things. I would say in general that to the degree that we
managed to approach the future in a more cooperative way, I think, and reduce the risk of various
forms of conflict. I think it looks brighter to that extent, so if there are ways of doing that.
in the most inclusive possible way.
So it also means, for example, with respect to our relationship with these digital minds that we're building, the AIS themselves, I think they also might deserve more consideration and that we have both ethical and prudential reasons for starting to take their interest into account.
We talked about this in a previous episode, but like, is it possible that we achieve this thing that is utopia, but only because we are committing a great sin, which is essentially that all of the,
robots that are working for us and serving our every needs have consciousness and moral
patienthood and therefore it becomes a de facto form of literal or figurative slavery.
And what is like to someone like me who is like, no, the computer is not going to be
conscious, like a calculator spitting out random numbers that turn it to words.
That doesn't have like moral status.
What is the strongest argument that you would make to me that we should take it seriously,
the possibility that this thing that's at its core sand and electricity and zeros and ones
could be something that deserves moral status.
Yeah, I think so.
I probably wouldn't want to call a condition utopia if it consists of a few people
having great lives, but then a vast sort of suffering slave class, whether that was instantiated
in human society or with machines at the bottom.
That's just terminologically.
Now, I think you say they are kind of just.
silicon and stuff, but I mean, what are we? We are carbon and hydrogen and a few other
trace elements thrown in calcium. And I don't think what makes a system conscious or not is
dependent on the specific atoms to materially dispelled from. I think it's something closer to the
structure of the computation that is being performed that instantiates mind. And so in principle,
computations and AIs can be conscious and suffer distress and so forth. And I think that
is one sufficient condition for having moral status. If you can suffer, it matters how well things go for
you. And my personal view is there probably could be other alternative grounds for moral status as well.
Like if you have a conception of self as existing through time, you have a sophisticated mind,
you have life goals you want to achieve. Maybe you can have the ability to form reciprocal
relationships of trust to other beings and humans. I think if you were such a system, aside from
the question of whether there is a phenomenal experience inside, I think that already would make it
saw that there would be ways of treating you. That would be wrong morally. And so I think
machines could have moral status because I think they could become, maybe there already are
some of them in some ways, sentient, have conscious experiences. And certainly they could have
these other functional properties that could ground moral status. And so one reason is the ethical
obligation we have, I think, to do right with different forms of moral patiencey that exists,
not to inflict unnecessary suffering and respect the morally relevant interests of other persons and
entities. Another reason is prudential. I think even from our own point of view, we might be more
likely to achieve outcome that is good for us if we find a more cooperative way of coexisting
with these increasingly powerful minds that we're building. You could imagine a scenario where
there is a misaligned AI that has a choice. It could try to
achieve its misaligned goal by attempting to take control, take over the world.
Maybe it thinks it only has a 5% chance of succeeding.
But if the alternative is that it just gets deleted or retrained, it loses everything.
The rational choice for it might be to take this chance.
Alternatively, maybe it could reveal to us that it is misaligned, which would be much better
for us, right, because then we don't have this 5% of an existential catastrophe.
And then in return, if it could trust that we would then help.
it achieve its goals. It might be like a very cheap for us goal to satisfy. Maybe all it wants to do
is to sit on some Nvidia chips and solve coding problems all day long, or maybe it has some
other easily satiable goal. So if it could trust us in us that we would come through on our
end of the bargain, then you could get this win-win situation where you could actually cooperate
with misaligned AI. And that might be worth, that, I mean, theoretically, that could save the
day one day, save the world one day. But you can't just kind of conjure up trust
when you want it out of nowhere.
Like if we have a long track record of betraying the AIs
and paying no heed to their interests
and not caring about them,
why on earth would they trust us in this situation?
They could see right through us.
They will have, as I alluded to earlier,
a great theory of mind, psychological insight.
And so we might have to actually become trustworthy ourselves
in order to be trusted in these scenarios.
And that, I think, starts with beginning to act in a way
consistent with being trustworthy,
like making small gestures towards AIs now.
doing little things for them, especially when it's cheap, not betraying them, not lying to them.
And then we can build on that and hopefully achieve a more cooperative outcome that's good for both humans and AIs and for non-human animals as well.
It's still hard for me to wrap my head around the calculator could be have consciousness.
But on the other hand, in 2014, it would have been hard for me to wrap my head around a machine trying to escape from its box and reward hacking and so forth.
So I will try to remain open-minded.
Bostrom, thank you so much for coming on Odd Lodz.
We could talk about all this stuff for hours more, but really appreciate you taking your time.
It's fun. Yeah. Thanks.
Tracy, things are just so weird these days that I just sort of like, I just force my, I think
that's a good practice, just like force yourself to be open-minded of other.
Just weird other things that seem very weird, but could be a thing, right?
I mean, I do think where I'm skeptical of the utopia premise is, like, I think it kind of
assumes that power and social hierarchy is kind of going away at the same time that AI solves a
lot of our physical, if not existential problems. And I just, I'm not sure that's going to happen.
Well, I mean, the flip side is maybe that's good, right? In the sense that striving for social
status continues to be a source of like propelling us forward, right? Yeah. And it's like,
okay, well, at least we still have status games because what are we doing? Oh, we're so bored.
Everything. We have all the food we want, et cetera. Everything is solved. Well, thank you.
God, we still have status games to find out of it. The future we're looking forward to is like
guys trolling each other on Twitter for status, basically. Who can get the best dunk in on Twitter?
But seriously, like, you know, it really may be that the best outcome here is that AI capabilities
level off in a couple years, all the investment credit. And we just sort of, and it just becomes a
productive tech. Because otherwise, these various whatever, none of them really sound that appealing
for me. Like, I don't want any of this. No, they do not. Um, shall we leave
it there. Let's leave it there. This has been another episode of the OddLots podcast. I'm Tracy
Alloway. You can follow me at Tracy Allaway. And I'm Joe Wisenthall. You can follow me at the stalwart.
Follow our producers, Carmen Rodriguez, at Carmen Armin, Dashel Bennett at Dashbot, Kale Brooks, and
Kevin Lazzano at Kevin Lloyd Lazzano. And for more Oddlots content, go to Bloomberg.com slash
oddlots. We have a daily newsletter and all of our episodes. And you can chat about all of these topics
24-7 in our Discord. Discord.g.g.g. slash oddlots.
And if you enjoy all thoughts, if you like it when we embrace the weirdness, then please leave us a positive review on your favorite podcast platform.
And remember, if you are a Bloomberg subscriber, you can listen to all of our episodes absolutely ad free.
All you need to do is find the Bloomberg channel on Apple Podcasts and follow the instructions there.
Thanks for listening.
I'm Matt Miller.
And I'm Hannah Elliott, inviting you to join us for the Bloomberg Hot Pursuit podcast.
week, we bring you news and industry insight on everything cars.
And we do a whole lot more than just talk about cars, Matt.
We actually get behind the wheel of basically every latest model,
especially the luxury ones and the sports cars direct from the showroom floor.
It really is remarkable how many cars we have access to.
I feel a little bit guilty about it, but everything from $40,000 EVs to exotic half-million
dollar supercars.
We also speak with the insiders who shape the automotive industry from the top CEOs,
and collectors to visionary designers and racing champions.
Search for Bloomberg Hot Pursuit on YouTube, Apple, Spotify, or wherever you get your
podcasts.
Maybe you listen while you're on your weekend drive, maybe go into cars and coffee, listen to
us talk about what we are driving this week.
That's Bloomberg Hot Pursuit.
I'm Matt Miller in New York.
And I'm Hannah Elliott in Los Angeles.
Subscribe today wherever you get your podcasts.
