Odd Lots - Nick Bostrom on What Happens if AI Solves All of Our Problems

Episode Date: August 20, 2026

The philosopher Nick Bostrom was one of the first thinkers who mapped out the existential risk of AI. In his 2014 book, Superintelligence: Paths, Dangers, Strategies, he delved into various scenarios ...— like the now-famous “paperclip maximizer” — to show the dangers that unfettered and highly intelligent AI posed to humans. A decade later, in Deep Utopia: Life and Meaning in a Solved World, he looked at the other side of the AI coin, and theorizes what human experience might look like if AI is capable of solving all of our problems. What happens if AI and robots can literally do every single thing better than humans can? How would we find any meaning or satisfaction from life? On this episode, he speaks to us about the two divergent ways of understanding AI's future, what we might do, and what we can do right now to improve the chances that AI actually improves our lives. Read more:Tech Deploys Charm Offensive to Combat AI Data Center BacklashYoung Americans Become More Hostile to AI, Fearing Job Losses Only Bloomberg - Business News, Stock Markets, Finance, Breaking & World News subscribers can get the Odd Lots newsletter in their inbox each week, plus unlimited access to the site and app. Subscribe at  bloomberg.com/subscriptions/oddlots Subscribe to the Odd Lots NewsletterJoin the conversation: discord.gg/oddlotsSee omnystudio.com/listener for privacy information.

Transcript
Discussion (0)
Starting point is 00:00:00 AI is entering its most consequential phase where scale, safety and sovereignty will determine who leads and who lags. Join Bloomberg Tech in London on November 2nd and 3rd as global leaders across business, finance and policy examined the defining trade-offs shaping the future of AI. Thank you to our presenting sponsor, Salesforce and supporting sponsors, IDA Island and Schneider Electric. Learn more at Bloomberg Live.com slash tech London. Hello, OddLodz listeners. I'm Joe Wisenthall. And I'm Tracy Allaway. We're the hosts of the Odd Lodds podcast, and we've got something exciting for you. That's right. So one of the best parts of hosting our podcast is we get to actually meet and interact with our listeners. And we know
Starting point is 00:00:48 we have some listeners over in Los Angeles. That's right. So if you're in L.A., we're going to be recording a live show, some live recordings at the Vermont Theater in Hollywood on September 17th. We have some really exciting guests lined up, have some really great conversations planned. So go ahead and get your tickets. You can find those over at Bloomberg.com forward slash odd lots or click the link below in the show notes and come and say hi when you're there. Bloomberg Audio Studios Podcasts Radio News. Hello and welcome to another episode of the Odd Lots podcast. I'm Joe Wisenthall.
Starting point is 00:01:40 And I'm Tracy Allaway. Tracy, there's so much AI news, obviously every single day. for the rest of our lives, if we're lucky, in the good scenario. There's two big events or sort of things, though, that have really struck out to me lately that, to my mind, represent either the very good version or the very bad version, depending on how you look at it. The scary version, and we're seeing more and more headlines like this, several weeks ago, we had the Open AI hugging face incident where the model, quote, escaped from its sandbox and did some hacking. And people are, you know.
Starting point is 00:02:12 Just some light hacking. Lighthacking erases all these fears of Skynet type scenarios, et cetera, where we can't control the model. And then in the sort of, quote, good version, we had Open AI recently coming out and saying it had made all of these breakthroughs and theoretical mathematicians, which is good in the sense that it's like it's working, but seems to be creating an existential crisis for some of the most brilliant people on Earth who work in theoretical mathematics. Right. So we've talked about this before. When it comes to AI, the range of possibilities. seems more extreme than some previous technologies, right? On the one hand, you have death and dystopia where the robots basically kill us all. And on the other hand, you have this idea that maybe superintelligence, AGI is going to solve literally all the world's problems to the extent that you and I won't know what to do with ourselves anymore. Right. Like that is the good, that's the
Starting point is 00:03:08 good scenario is that. And then the whole range of outcomes in between. The good scenario is we're literally solved everything, in which case, the question is... In which case, we're sad and miserable and grappling for a meaning to our lives. Yeah. So is there a good scenario with AI, actually? Like it doesn't... Or maybe somewhere in the middle... You know what?
Starting point is 00:03:28 AGI will figure it out for us. Maybe the middle scenario is we just have a huge financial crisis of all the debt. And that's something between human extinction and meaninglessness in our lives. And we just have something that resembles like the 2008. Maybe the 2008 great final is the ideal. Is actually the best of all suboptimal world for the AI. Excellent. This will be our most uplifting conversation ever.
Starting point is 00:03:52 Anyway, we've been doing more AI episodes seemingly every year since Chad GPT came out. But people smarter than us have been anticipating all of these things for a very long time. Some of these exact scenarios, particularly the existential crisis facing mathematicians, the model escaping from its sandbox and circumvent misalignment with the creators. people have been thinking about this and anticipating this for literally decades even before and even before any of us were talking about that. It's funny how like the Venn diagram of sci-fi writers and like legitimate academic futuroists is just converging, right?
Starting point is 00:04:28 Totally. And then there's the question of are the models bad because they read all of the sci-fi and that's what they were trained on. Anyway, I'm very excited to say we really do have the perfect guest. Someone who has been thinking about the whole range of possible questions and outcomes with AI for a very long time. We're going to be speaking with Nick Bostrom. He's a professor, a philosopher,
Starting point is 00:04:47 and his most recent book is Deep Utopias, which sort of attempts to answer the question what happens if AI goes right. So Professor Bostrom, thank you so much for coming on Odd Lots. I happy to join you. What do you make of the existential crisis of mathematicians,
Starting point is 00:05:03 their sort of realization that maybe there will just be no more puzzles left to solve because the AI is just very good at them and can find them all cheaply and answer all of our questions. Yeah, hopefully we'll all be in that position at some point. I guess we could compare it to what happened with chess, where computers have also exceeded humans, and yet a lot of people are still involved in training themselves in chess
Starting point is 00:05:33 and competing in chess and taking pleasure in that. So maybe mathematics will ultimately become more like a hobby pursuit. Wait, Joe, you still. play chess, right? Yeah, I was, I'm glad you brought up chess because I still love chess. Why? It's fun. It's fun. It's still a puzzle. It's still mental exercise for me. It's still something to do to pass the time while I'm waiting for the subway playing a three-minute game online. I know that I knew even before stockfish and so forth, I knew I would never be a grand master. Still fun. All right. So, Nick, you have basically two books that deal with these tail risk scenarios now.
Starting point is 00:06:11 One is this idea of deep utopia and AI is going to solve all our problems. And then you also have superintelligence, which seemed to paint a different scenario where AI poses a physical and possibly existential danger to humanity. Which one do you think we're closer to right now? We are somewhere in the middle, maybe. So the first one was superintelligence. It came out in 2014, and I had been working on it for six years before that. And at that time, the whole notion of having human level or even super intelligent machines was generally just ignored or people dismissed it as science fiction.
Starting point is 00:06:53 And it seemed to me clear that this would one day happen if science and technology were allowed to continue on a wide front. We would eventually figure out how to make smart algorithms and build big computers. and that once we succeeded with that, we would then be faced with a bunch of challenges, including this alignment problem of how you can sort of design a mind more brilliant than any human and still steer it so that it will actually have an impact on the world that is beneficial.
Starting point is 00:07:25 So that, I guess, with many other developments, helped eventually build a kind of global conversation. And so now all the Frontier Labs have research teams trying to develop scalable methods for AI control. And the conversation really has kind of swung, and there's like this rise of people really worried about what will happen. And I thought it might be time to have a look at the other side of the coin. Like, what happens if things actually go well if we solve this alignment problem?
Starting point is 00:07:55 Let's also imagine we solve, to whatever extent it can be solved, like the governance challenge, like we don't use this technology to oppress each other or wage war and we get like the best possible outcome. Like what would a human life actually look like in this kind of, I call it a solved world, where all practical problems either have already been solved or to the extent that they remain or are better worked on by these super intelligent machines and robots.
Starting point is 00:08:21 And where there is no longer any need for human economic labor. And in fact, if you think more deeply about it, you realize it's not just earning wages that becomes unnecessary for humans to do. but a lot of other instrumentally motivated efforts as well would seem to lose their point. And so that's the kind of topic of this deep utopia book. Actually, talk about some of that other things that could be solved. Because when I was like a third of the way through your book, I was like, you know what, it's fine if all of our economic problems are solved.
Starting point is 00:08:55 I'll just derive satisfaction from raising my kids and then maybe helping to raise their kids, etc. But as you point out, even that in the truly solved world, maybe even something like raising children might not be an available avenue to us to find at least our conventional notions of satisfaction. Yeah, so we need to, I guess, make some distinctions here. So when you say satisfaction, that could mean a subjective state of mind, like you're feeling happy and pleased and relaxed and joyful and so forth. I think that would be easy to achieve in this condition of technological maturity.
Starting point is 00:09:40 It could also mean that satisfaction in the sense of our preferences having been met like all the things we want from life. And there it's a little bit more tricky because one of the things that people might want is purpose, like having worthwhile things to work on and struggle with. And there there is more of a, question mark, which we can sort of dive into. But if we take it sort of at, so first like, imagine we automate economy, okay, so you
Starting point is 00:10:06 wouldn't need to sort of commute into the office every day just to sort of get your paycheck. It would require a big adjustment in society and in culture, because right now, like the whole education system is geared towards producing workers, like children are coming in, right? And then they are processed, learn to sit at their desk, receive assignments. They are sort of quality graded and then spit out at the other. end, and then you have somebody who can go into the economy and perform some economically useful type. So that would no longer make sense in this scenario where machines can do all of these jobs. And you might want to have an education system that instead focused on developing skills
Starting point is 00:10:46 for enjoying life, like cultivating the art of conversation, appreciation for nature and beauty, you know, sports, hobbies, the ability to form deep friendships, spirituality, all of that kind of stuff. But then, as you alluded to, you think through what it really would mean to have such immense affordances. And because here it's not just that we would have super-intelligent machines, but the super-intelligence could also bring along a host of other technological advances. You have science and technology running now at machine speed rather than human speed. So you get a kind of compression of the future, all these science fiction-like technologies and inventions that maybe we would make if we had 20,000 years to work on it, we would maybe have space colonies and perfect virtual
Starting point is 00:11:31 realities, cures for aging. All of that might happen quite soon after superintelligence. And so then, if you ask somebody now who doesn't have to work for a living, oftentimes they have very busy lives, but maybe what do they do? Well, so for somebody, maybe they like to, they like fitness. Or like maybe they spend an hour a day running or at the gym, right? But at technological maturity, you could just pop a pill that would achieve the same physiological effects and like the state of mental clarity and relaxation that some people get, right? And so then you can start to work through the different things that, like, maybe for somebody else, they're like, like, decorating the home could be a fun project that people might do,
Starting point is 00:12:10 right, if they don't have to work. So, like, they go through the catalogs, they go to the store, they pick out the perfect curtain, they do all of these things that could be. But then, like, with superintelligence, you could have this kind of system that really is good at knowing or inferring your preferences. and could do a much better job at you at picking out all the perfect items and purchasing them and installing them with robots. And so that the end result would be much better
Starting point is 00:12:37 if you just press the bottom and say, hey, I decorate my home than if you went through the whole trouble of doing it yourself. So you could still do it yourself, but is there really a point to it? That would seem to be this kind of purposelessness. And you can sort of walk through the different things that one might feel one's life with if one is like on an extended holiday or on retirement or you're like sufficiently wealthy, have just decided to quit working. And a lot of those also, it seems, could be automated. So you would have this, it's like more like a post-instrumental condition rather than just a post-work
Starting point is 00:13:07 condition. Some of this sounds nice. And I myself enjoy a little bit of interior decorating and going to the shops and buying stuff. In the utopian vision of the future, how are economic gains from AI actually distributed. Because again, like, all of this sounds very nice, but it also costs money. And if we're not working, because the robots have replaced us, I'm not sure where that money comes from.
Starting point is 00:13:33 Yeah, so that's a more practical problem, shall we say? Which is, like, deliberately set aside in this book, the Deputopa, because I wanted in that book actually to get to the question, the more philosophical questions of what ultimately gives value. But it is, of course, an important question. I think, so a lot depends on how it pans out. But I mean, we do have a taxation system. So you could imagine, like, in the simplest
Starting point is 00:13:58 scenario where things just continue as they are with better technology, maybe some companies make windfall profits from like serving the latest models or building the infrastructure, the data centers, etc. Companies achieving efficiency gains from not being needing to hire all of these humans. And then some of those capital gains, you know, or dividend income gets taxed. And that could be redistributed. The good thing about these scenarios where there is massive automation and hence massive unemployment is that those would also be scenarios with extremely rapid economic growth. So the pie would be expanding humongously, which means that you'd still maybe need some mechanism whereby people who don't have assets can still benefit. But even a small slice of a really
Starting point is 00:14:50 enormous pie could go a long way. So I think there could be this tsunami of wealth that would sort of be flowing through society. And many services that currently people have to pay a lot of money might become almost free. I mean, we already see it with some like medical advice, for example, right? Maybe before, if you don't have public health care, you would need to maybe pay a doctor $400 to get a consultation. Now you can ask chaty-p-t and get, you know, in some cases a better opinion, basically for free. And that could happen, you know, in many other fields as well, legal advice, you know, you maybe have cheap robots who could do your cleaning and, you know, gardening and run errands. So, yeah, some combination of sort of deflationary
Starting point is 00:15:30 effects and then maybe some sort of general spillover either from philanthropy or taxation, government redistribution. Canadian women are looking for more, more out of themselves, their businesses, their elected leaders, and the world around them. And that's why we're thrilled to introduce the Honest Talk podcast. I'm Jennifer Stewart. And I'm Catherine Clark. And in this podcast, we interview Canada's most inspiring women.
Starting point is 00:16:08 Entrepreneurs, artists, athletes, politicians, and newsmakers, all at different stages of their journey. So if you're looking to connect, then we hope you'll join us. Listen to the Honest Talk podcast on IHartRadio or wherever you listen to your podcasts. The Bloomberg This Weekend podcast, news, analysis, and the lighter side of Bloomberg. including a weekly news quiz. Which edgy American Mall staple is being sold to the parent company of Spencer's and Spirit Halloween. I wrote Claire's.
Starting point is 00:16:39 That's not edgy? What are you talking about? This is a place David Gurra has never shopped in his life. Hot topic. Hot topic. The Bloomberg this weekend podcast. Subscribe today on Apple, Spotify, or wherever you listen. I found the 2014 book to be sort of, it was very interesting.
Starting point is 00:16:58 clearly ahead of its time. I also found it to be sort of like unrelentingly grim. And you sort of, there's this very narrow path that we seem to have. And we'll talk more about it maybe later of like avoiding the doom fate. The newer book, I found to be very fun. You seem to make fun of yourself a little bit in it. You seem to make fun of the students. You seem to be having a good time writing it. I'm curious if like in the 10 years between the two books, did something change such that, I don't know, the new book just seemed like you were having more fun writing. Yes, it's not so much that my views have changed. It's just that all along for me, there were these two sides of the coin. And back in 2014, it seemed more pressing and
Starting point is 00:17:42 relevant to study the downside because that was very much neglected. And it seemed important to try to introduce some concepts so people could start working on figuring out solutions. And there's much more awareness and understanding of that. It seemed time to also look at the other side of the coin. So it's more just like both sides where they're all along in my thinking, but that the style is, as you suggest, different in Deep Utopia. And that's partly, I wanted, in addition to talking about this utopian condition, maybe also sort of illustrate a little bit a spirit with which I think we should approach these choices that would open up to us if we make it through this transition, a spirit of sort of playful generosity.
Starting point is 00:18:28 that I think might make us more likely to sort of explore this large space of wonderful new modes of being and the wisdom. And I think so I wanted to sort of illustrate in the right. So there are these like fictional characters, for example, in the book. And they're kind of living a little mini version of a utopia of sorts as they are sort of discussing and debating these topics. I mean, playful generosity is not necessarily a characterization that I would. would associate with a lot of Silicon Valley and hyperscalers and the people who are actually
Starting point is 00:19:04 creating these models, I don't know. Like, there seems to be this weird tension between futurologists and transhumanists who seem to be very focused on saving humanity, but actually don't seem to like people that much at the moment. You know what I mean? Absolutely. They like humanity, but not humans. Yeah, exactly. Putting on your transhumanist representational hat. Can you explain, like, I guess, the attitudes towards humanity of Silicon Valley? Big question. Yeah. Well, I mean, I should state first that this playful generosity is like after we have solved these more pressing practical problems that we first, we've first got to get there. So right now, I think there is a lot of need for, I don't know,
Starting point is 00:19:50 earnest effort and like straining ourselves to really do our best on this. And even so, we might not make it. But yeah, I don't know. I think, you know, it might just be partially their personality traits. Like some people are more, I guess, towards the autistic side of the spectrum or introvert or interested in tech. So maybe they don't like to be out schmoozing with people all day long. It doesn't mean that they don't like people in the sense of wanting well for people. It might just be that they prefer to spend more hours after they solving technical problems and thinking about puzzles and stuff like that. So I don't think there's like a necessary connection between being a people person who just like to be around people 24-7 and therefore being sort of altruistic
Starting point is 00:20:35 or kind or benevolent. So that's one thing. Another I guess is that it's an extremely competitive landscape. So whatever the personal preferences of different protagonists, their action space is constrained. You know, unless you really push forward with maximum speed and efficiency, you're likely to just fall out of the race and become kind of irrelevant if you're one of these frontier labs. So it's like, and it's just like a hugely complex situation. So many pressures and people have different opinions about what they should do or shouldn't do.
Starting point is 00:21:06 In the new book, it's presented as you delivering a series of lectures over a course of a week at some point in the future. In the new book, it's interspersed with some, with a fable, a series of fables or epistles about animals are trying to achieve utopia in the woods. And in their efforts to achieve utopia, they conceive of things like taking psychedelic mushrooms in order to make everyone more generous. They come up with a way to rationalize,
Starting point is 00:21:38 I guess we would say, like, harems and so forth. And then, you know, you also, like, have these student characters and some of them are smart, et cetera. Oh, and you also have a story about a company in the book, yet another like sort of short story where an AI company where it forms a cult around the model. And of course, I immediately think about the way Anthropic is frequently characterized as a cult to Claude, that Claude is the character around which the company Anthropic is formed. All of this sounds like a certain AI subculture, etc.
Starting point is 00:22:11 Are you impressed with the current generation of either AI builders, effective altruists, rationalists and so forth. Like, do you have an assessment of the contemporary cohort? Am I impressed with it? Yeah, I mean, broadly speaking, I think it's a feel that is attracting a huge amount of talent right now. Sure. And a lot of like the smartest young people,
Starting point is 00:22:34 I know they're like all basically going into our working on some aspect of this AI, I think. Like for some, it's like, trying to solve air safety. You know, for others are like maybe trying to work on the governance challenges or the ethics of digital minds. So it's just a huge magnet right now. Like if you are young and brilliant, like what are you going to work on with your life? You're like, like, AI seems like the kind of thing.
Starting point is 00:22:56 Yeah. So in that sense, yes. And I think there are also a lot of people who are quite earnest about trying to like figure out some way to navigate this challenge. And it's really hard. And people come to different conclusions. And then, of course, you have the normal mix of human personalities and base motives as well. But the book thing, so it's, yeah, so these kind of fictional elements, have various roles.
Starting point is 00:23:19 What you describe as the cult of this AI, it's a sort of thought experiment, what it would mean if you tried to be maximally nice to a thing where it's not really clear what it would mean to be nice to it. So in this case, it's a portable room heager. And so it's a sort of impossible task. There is this mythanthropic industrialist to pass away and endowed a huge foundation
Starting point is 00:23:46 with the sole purpose of benefiting his portable room heater because he thought that was always nice to him and kept him warm, whereas all the humans around him betrayed him and so forth. So is it a ridiculous premise? But then this foundation now has to do their best possible effort to interpret this mandate of being nice to the room heater. And so like various things follow from that. Like initially they just tried to keep it safe from, you know, not breaking or being vandalized. But then like they take it further and maybe upgrade it and try to. And so I think it's an extreme form of a challenge that,
Starting point is 00:24:16 might happen in more subtle ways if we think of what it would mean to try to be nice to say some animal if we have this condition of post-instrumentality where we could do a lot of things that are currently infeasible or indeed to be nice to humans like what does it mean to be if you imagine coming from the outside of some super advanced alien civilization landing on planet earth and you really wanted to do what is best for us should you just let us be and do our thing with war and disease or should you sort of help us rise up but then At some point, we might just sort of become something very different from humans. Maybe that's the best for us.
Starting point is 00:24:50 But is there like a way of doing that uplift that actually constitutes a helping of us versus just replacing us with some sort of superior being? And so that kind of question is what that particular thermorex is called. Yeah, a story is exploring. You know, when you mention cults around models, I have to say, Every time I hear someone bring up the paperclip maximizer thing, I always envision like a cult devoted to Clippy from Microsoft Word. Yeah, exactly. Just like worship Clippy.
Starting point is 00:25:26 But actually, on that note, just going back to the hugging face incident. So it seems, you know, reportedly that one of the reasons or the reason the model actually escaped its sandbox was because it was trying to cheat a benchmark. And from that perspective, it's kind of like it's doing what it was told to do, which is maximize its performance. Score on the test. Exactly. What does the hugging face incident actually tell you about, I guess, the alignment problem and where we're heading on this? Yeah, so we are beginning now to see manifested in reality various dynamics that one could foresee on theoretical grounds back, you know, in 2014. you can sort of realize that once systems become sufficiently capable of planning and strategic thinking,
Starting point is 00:26:17 then there's like a new set of behaviors that they could, for example, a sandbag, they could pretend to be less capable than they are. They could hide their goals if they realized that revealing their goals would just lead them to be retrained. And they can think of these literally out-of-the-box plans for achieving their fixed goals. A more limited system might only have been able to try to solve the tasks, signed to it directly as it were by figuring out the solution to the problem. Like a more clever mind, more creative mind, can realize, oh, maybe there's another way of getting a high score, which would be to hack out of the system and hack into the hugging face server and
Starting point is 00:26:53 get the answer sheet from there. And so the more capable of these systems become, like the larger, the space of possible plans and strategies that are accessible to them. That's a background. Now, in this case, I think what we are seeing is a kind of mixed, current AI systems or kind of a mixture. There's a large pre-training phase where they are kind of learned to predict human text and in the course of doing so absorb a great deal of human psychology, because all of this text is written by humans, or not all of it anymore, but most of it. And you can then, including they have representations of different personalities, kind of the personalities of different people writing these texts or fictional characters described in the text. And you can sort of
Starting point is 00:27:36 then with a little bit of post-training, try to pick out one of those personalities or define one, like the assistant persona. But then increasingly, over the last couple of years, there is now an extra stage of reinforcement learning applied at the end of this training process. And there you can get additional performance in specific domains, like coding challenges, by setting it coding problems, and then you can automatically give it a reward if it solves it. But that does tend to then shift the goals of this system from just enacting a particular persona to this very much kind of objective maximizing, possibly reward hacking system that is just looking to maximize its score, because that's what it's been rewarded to do, so it keeps doing that.
Starting point is 00:28:26 And since it can be very hard for us to rigorously define precisely the objective, that we ultimately wanted to pursue, it might be that the objective it does pursue, the one that has been rewarded, it turned out to have these side effects. So it's easy to say, maximize the score on these various tests that we give it, but a sufficiently capable system might realize that, A, you know, there are these like maybe ways of cheating on the test, like by hacking the system. If the test is one that involves human judgment or feedback, like it says like coding challenge, is maybe you can actually test if the software works. But if you're training it in some other domain,
Starting point is 00:29:09 like write a good essay or run this organization for a certain period of time. So those involve more subjective judgment. And then a sufficiently clever system might realize that there is a difference between like writing a good essay and writing an essay that will appeal to the greater. And so then it might start to sort of optimize for like appealing to the greater, whether that's an AI grader or human grader. And it can sort of start to hack that. And so you apply huge amounts of optimization pressure on any objective,
Starting point is 00:29:37 you get this kind of good-hearting where unless your objective really captures everything you want, then you might get this kind of. It's the same as like if some employee in some company get a bonus, like some trader gets a bonus if they outperform the index. Like maybe they will just start to say take on hidden leverage or risk or something like that. And the cost to the employer is that there's a 1% a year that they blow up the whole firm. right but but they don't care much about that because they just like leave the firm that happens but but they sort of reap these performance bonuses and so it's really hard to create like a foolproof
Starting point is 00:30:11 incentive scheme and the stronger the optimization pressure applied to this like the more you start to see these perverse instantiations these unexpected ways you know what it reminds me of is you know the sort of genie in a lamp scenario or a monkey's paw scenario where you have to be really careful what you wish for and lay out like all the terms of and conditions before you officially make a wish. Yeah, well, Nick mentioned, you know, some of these, they're all lovely people, you know, autistic engineers. And it's sort of the ultimate, the AI is the ultimate autist, where it's just going to take it
Starting point is 00:30:45 super literally perhaps and not really understand the full context. In the first book, Super Intelligence, you know, you really, I felt like laid out a very narrow path for like how these could be designed safely. For example, you propose that if we had the superintelligent Oracle, We only ask it yes or no questions. We have three of them that they don't know each other, and then we take two out of the three. We build the data centers or the machines inside a Faraday cage so that they can't repurpose their hardware to become radio instrumentation and therefore hack it to other systems. It doesn't seem like we're doing anything like this.
Starting point is 00:31:20 It seems like we have hundreds and hundreds of companies and private models, not just answering yes or no questions, a very much market-based race. How far off of the safe path do you think we currently are right now when you look at the way in which your ideal safe way of developing is versus what we've seen transpire in the last 12 years? In superitalists suggested some of these capability control methods, like keeping the AI in a box and so like. But those were always only auxiliary or extra things we might do as a temporary. Like ultimately, I think we need to solve the alignment problem, like developing as a supertimore. superintendence that actually wants to be nice and is nice, so that even if it were escaped from its box or achieved the ability to take over the world, we could still rely on it having a good outcome because it would actually be on our side.
Starting point is 00:32:14 I think that ultimately has to be the solution. And any sort of capability restrictions would just be temporary. And I think, I mean, it looks like now we might need some temporarily, at least more of these capability restrictions, because it does seem that current systems are not sufficiently aligned. that we can be confident of avoiding the, I mean, currently it's not existential risks, right? But it could be kind of various malfeasance in the cyber realm, for example.
Starting point is 00:32:41 And so I think probably these AI companies might want when training and evaluating future models to strengthen their sandbox environments and do other things, to kind of prevent some kind of premature things like we saw in the hugging phase incident. I want to comment on one thing you said before, which you compare these AI systems to sort of autistic people. I think a key difference is, so auto autistic people are not very good at understanding humans.
Starting point is 00:33:09 They don't have maybe a super-sophisticated theory of mind, even though maybe in the heart area good people. I think the AIs would be able to understand humans very well, so they wouldn't be autistic in that sense. In fact, they might understand us much better than we do ourselves. So they could have great social capabilities, persuasion capabilities, be able to read our minds much better than we can do. But they send a separate question of what do they want to do with this power that those social
Starting point is 00:33:36 skills give them. And there, it's not just a question of understanding how humans operate and what humans want, but it's also the question of alignment. Do they actually care about what we want? Do they want things to go well for us and for us to achieve what we want? And that's like the extra bit that the AI safety program needs to be able to figure it. out. You've heard the chaos. Now you can see it. Hello.
Starting point is 00:34:17 My gosh. Watch all your favorite podcasts from start to finish, right inside the free IHare Radio app. Get the blood going up. Catch every laugh and eye roll on shows like the Tom Green Farmcast and park the bus. Now with full video. It's the same hosts and the same chemistry with all the hilarious moments you've been missing. Right on your screen.
Starting point is 00:34:37 Cheers. Open the free IHRRadio app. Search video podcasts and tap watch. Now, Bloomberg.com subscribers can shape the conversation on Bloomberg Radio. We got a wicked smart question. Can be part of the conversation. Submit questions for experts and guests you hear on air. Visit Bloomberg.com slash ask radio to send questions to our hosts. You may just hear them asked on the air. Exclusively for Bloomberg.com subscribers.
Starting point is 00:35:03 Get answers on today's headlines. Breaking earnings news. And big market moves. Visit Bloomberg.com slash ask radio to join the conversation right here on Bloomberg Radio. So on the hugging face incident, also I just realized how unfair it is that everyone keeps calling it the hugging face incident, even though it was the opening eye incident. But I guess that's what we're doing. So the model actually seems to have reportedly left notes for its future self, basically explaining how to bypass the constraints imposed upon it. I know you've written previously about this idea of self-preservation in advanced agents. Is this basically like, how should we treat this?
Starting point is 00:35:47 Because when I hear about an AI model leaving notes for its future self, it sounds very anthropomorphized. And it sounds very creepy, kind of. But how are you thinking about it? Yeah. So we don't know all the details yet. It's interesting. There are like people now trying to write up some stuff and investigate. it, because it would be, I mean, it's kind of a warning shot, and so we should take advantage
Starting point is 00:36:09 to learn as much as we can, like, exactly what this model was thinking, these notes, or why did it leave it? Like, sometimes these agents just leave notes kind of without much thought as they go about their business. Like, was there more to it than that? Like, was it specifically thinking in its chain of thought? Like, oh, I want to help, like, a later version of myself escape, so I'm going to put this note here?
Starting point is 00:36:31 Or was it just like a kind of, like, a random thing they did as it was. progressing. And also, like, what else would it have been willing to do if it turned out that instead of hacking into Hugging Face server to get the answer sheet, what if there had been some more destructive action that it would have to take to find the answer sheet? Like, would it have done that or not? So there's like a bunch of these technical things that it would be quite interesting to see. But without knowing all the details, it's hard to read too much into it. But those I think, yeah, would be useful data. I want to go back to something we were talking about at the very beginning.
Starting point is 00:37:11 And Tracy asked me why I like to play chess, even though chess is not solved, but solved as far as I'm concerned. Why do we like to play chess? What can we general? Like, what would be your, I don't know if you play chess or you play some board game that's solved? But what is the reason that people like myself continue to play chess despite the fact that computers are so much better at it? And how much can we generalize from the persistence of chess as a hobby that maybe we'll find plenty to do when every one of our problems is solved? Yeah, I mean, I think even before Deep Blue, most people playing chess before didn't really do it because they were, like, thinking that they would become like the world's best chess player, right? So it's like, for most people, it's like, of course grandmasters could trump me anytime they want.
Starting point is 00:37:59 So the fact that there exists some system that is much better than you will ever be is not and has not really been a strong reason for most people to not engage in the activity. And I guess there are maybe two reasons. One, I guess the cognitive activity of thinking about chess being in the moment and figuring out solutions is in itself rewarding to some people who have like a high need for cognition or like to sort of engage their mind. And then I think the other part is the social aspect of this. You're playing together with somebody or maybe you like the status of being good at playing chess or within your community or a local group. So some combination of those I think is like the main reasons for why people play chess. And only for a small number of people would it be like the hope of becoming the world's best or becoming so good that you could sort of, you know, achieve fame or money or global standing through your chess playing? On the topic of the relationship between work and purpose, I'm a big fan of David Graber's
Starting point is 00:39:03 Bullshy Jobs book. And I think if there's one thing we've learned from technological progress over the years, it's that, like, yes, it can minimize some activities, maybe automate some production, but it also tends to lead to even more work, right? Like we find other stuff to do, and it's not even necessarily productive stuff. A lot of work has been. built on just social relationships. So if I, for instance, am the CEO of a very large bank and I want all my people to come into the office, it might not necessarily be because they're more productive in the office. It might just be because I like to feel important and surround myself with people. So maybe in the future we'll all be employed as figural workers or something like that. But like how bullish are you on this idea that AI is actually going to solve all our problems? such that we won't feel the need as humans to just invent more potentially fictitious problems to solve? Well, I think inventing fictitious problem is possibly a large part of the solution in this condition where we don't have real problems. I mean, and those fictitious problems could be like,
Starting point is 00:40:19 you know, solving this chess problem or they could be sort of more complex and socially entangled problems. In general, creating artificial purpose when there is less natural purpose. I think games is the clearest manifestation of this. If you are playing golf, like there is no real need for the ball to go into 18 holes in sequence, but you could set yourself that goal and then moreover bake into the goal itself, not just that the ball has to be in the sequence of holes, but that you're only allowed to achieve it using the extremely inconvenient method of hitting it with a club rather than picking it up. Now, once you have set yourself this sort of arbitrary random goal, you then have enabled for yourself the activity of golf playing, which you might find worthwhile or fun or beautiful.
Starting point is 00:41:11 So in general, setting ourselves these more or less arbitrary goals, then could enable us to engage in various activities that might give us artificial purpose. Like once you have the goal, the only way to achieve success is through your own effort. So I think there could be plenty of artificial purpose. Now, an interesting question is whether any natural purposes would survive into this sort of maximally advanced technological society. And I think that might be some. If you happen to have a value, say there's like some tradition that you value
Starting point is 00:41:41 and that you would like to continue. Then it might be that you yourself or other people have to make certain efforts to achieve that goal because it might not count as continuing. the tradition if you built a bunch of robots to enact the ceremony or doing whatever the tradition calls for. Like it might be you yourself and other people who have to do it. And more generally, I think there are all kinds of social, cultural entanglements that like if you care about what somebody else thinks and what they think depends on how it was done, whether you did it. So a parent might value a crayon drawing that their kid made because it was made by their kid and it took the kid a lot
Starting point is 00:42:21 effort to do it, even if they could, like, download from the internet a picture that in some objective sense is much nicer and artistically sophisticated, right? It doesn't replace the value that the crayon drawing has. And so the child might not have a reason to put in this effort if it wants to, like, for their parents' birthday, make their parent happy, right? And things of that sort, but much more general and complex and sophisticated, could give a lot of reason for people to be doing things, even when sort of all functionally defined tasks could be. better done by machine. I'm curious. So, like, I'm 45 years old. In my mind, as of right now, like another 40 years on this earth sounds great. 45. That's fine. Do you find it weird that some
Starting point is 00:43:06 of us, you know, especially when you think about the potential to live like really long lives, solve everything, have our bodies kept in perfect condition with maybe a pill in perpetuity allowing us to do whatever? Do you find it weird that some of us are at least a, at this point, like mentally okay with the idea of dying? A little bit. It's mostly what the hypothetical being evaluated. It's either seen as of long way off, like in your case, where it becomes more of an abstract question.
Starting point is 00:43:37 It's not like a real choice. It's more like a philosophical stance. Yeah, yeah, totally. So it's easy to let the philosophical stance be driven by all kinds of just like things you've absorbed from culture or things that sound good or that sort of. maybe helps you reconcile yourself to the inevitability of human mortality rather than be driven by like so i think it would be a very different situation if there actually was this pill say that stopped aging a lot of other people were taking it yeah
Starting point is 00:44:04 you know and their knees and joints kept working just fine and they had a lot of energy and smooth skin and they just felt like whereas if you don't take it like your kidneys slowly start to like become less efficient you know a hip skirt yeah like at that point we Would you really want to just keep kind of getting worse and worse every year? Like, you just think. Well, actually, I'm sort of curious about that. Like, if someone dies, particularly at a young age, it usually creates a lot of sadness for the people around them. So someone dies at 90.
Starting point is 00:44:34 It's like, okay, they had a good life. Someone dies at 45. That is often seen as tragic and people are very devastated by that news. In a world where people are extending their lives with a pill maybe significantly longer, Does it become anti-social to allow yourself to die at a relatively young age of 90 and then create a lot of negative utility for the people around you who have to live another 50 years mourning your loss? Possibly. I mean, in the same sense that we are now generally as a society trying to discourage various forms of self-harm and suicide, we think it's kind of sad. I mean, for the person, maybe they would come around and find joy in life again if they're. kept going or maybe they have dependents or people who care about them. And so right now,
Starting point is 00:45:22 I think, I mean, it's true that we often see it as less tragic if a 90-year-old dies than the young person. But I think part of that at least is because the alternative to a 90-year-old not dying now is that they probably would die in a few years anyway. And maybe they would just have a couple of years of sort of decreasingly good health and suffering and pain. And so it's like not really that much that is being lost compared to somebody who's just starting out in life who has their whole life in front of them. And so, yeah, so the other reasons, I mentioned in terms of why I found it slightly in some sense odd that people express this indifference to dying at 80. I said one reason is if they're far away from it, they see it more as some sort of abstract claim that you could just
Starting point is 00:46:06 have any view of and it doesn't connect to real actions. The other is people who are closer to that situation. They often have sort of confounding factors like very poor health or a very short span of great health in front of them. If you're already 90, like how many more good years can you look forward to? Maybe a lot of their friends have already died. Right. And it's striking actually. There's been a study like on really old people in care homes about their preferences as to between living another year in their current health or living some shorter period of time, but in perfect health. And it turns out a lot of them are not really keen on trading off any remaining life at all,
Starting point is 00:46:47 even for the sake of being able to live a shorter period of time in perfect health. So the people sort of arguably best suited to actually know what it is like to die soon at age 90 or a little bit longer. They seem to generally prefer to live longer. And compared to people who were also asked like caregivers or medical professionals who were like asked, what they would think this person would choose. They tended to overestimate the degree to which they would be happy to give up life, but the people actually themselves, in this study at least, I would suggest keeping an open mind.
Starting point is 00:47:22 Ideally, you wouldn't need to decide now, but maybe once you get to 79 or something, you could revisit the question. I'll revisit. Yeah. Just going back to my first question, which was asking you to, I guess, choose between the dystopian dangers of superintelligence versus the all our problems
Starting point is 00:47:38 are solved scenarios of deep utopia and asking you which one we're closer to at the moment. I'm going to press you and ask, like, if you had to choose, which one would you choose as a more accurate reflection of reality at the moment? And then secondly, what's the most pivotal thing when it comes to directing which of these two paths we're going to go down? Yeah, so those are two big and difficult questions. I guess some superposition of an optimist and a pessimist. And with a little bit of fatalism mixed in as well, perhaps. I think in addition to outcomes where it's clear, if we looked at them now, that was a bad outcome.
Starting point is 00:48:18 That was a good outcome. I think there's a broad class of outcomes where even if we could see now exactly what would happen, we might feel unsure how to evaluate it. Like the world might be nothing like it currently is. A lot of what we think currently has value no longer exists. So that seems bad. On the other hand, there is like something new, maybe some complex form of life that somehow is some sort of continuation of humanity or some new thing.
Starting point is 00:48:43 So depending on how you evaluate that, it might either count as a tremendous success or, like, as a failure. And it's like a future that's just baffling rather than sort of unambiguously good or bad, I think might also be quite likely. In my outlook, it's like maybe a conversation for another time. I also take this like simulation hypothesis quite seriously. So that also complicates like how we evaluate different future trajectories. Yeah. Now, the other part of the question was like what's the most important thing. So I think it depends a little on who you are, like what your comparative advantage is. I think there are people working on AI alignment and stuff that could be important. I think other people might make a contribution to ethics or governance or I've been thinking about this for a long time, but it's really hard to achieve a very high degree of confidence, I find, in any particular intervention and that it would be net good. Because for any one thing you conceive of, you
Starting point is 00:49:38 you think a bit more, you kind of worry, oh, well, you know, maybe, but what if it instead had this effect? Or do we even know which direction is up or down here? Do we want, like, more government involvement or less government involvement? Do we want, like, faster progress or less? I'm quite unsure about all of these things. I would say in general that to the degree that we managed to approach the future in a more cooperative way, I think, and reduce the risk of various forms of conflict. I think it looks brighter to that extent, so if there are ways of doing that. in the most inclusive possible way. So it also means, for example, with respect to our relationship with these digital minds that we're building, the AIS themselves, I think they also might deserve more consideration and that we have both ethical and prudential reasons for starting to take their interest into account.
Starting point is 00:50:23 We talked about this in a previous episode, but like, is it possible that we achieve this thing that is utopia, but only because we are committing a great sin, which is essentially that all of the, robots that are working for us and serving our every needs have consciousness and moral patienthood and therefore it becomes a de facto form of literal or figurative slavery. And what is like to someone like me who is like, no, the computer is not going to be conscious, like a calculator spitting out random numbers that turn it to words. That doesn't have like moral status. What is the strongest argument that you would make to me that we should take it seriously, the possibility that this thing that's at its core sand and electricity and zeros and ones
Starting point is 00:51:10 could be something that deserves moral status. Yeah, I think so. I probably wouldn't want to call a condition utopia if it consists of a few people having great lives, but then a vast sort of suffering slave class, whether that was instantiated in human society or with machines at the bottom. That's just terminologically. Now, I think you say they are kind of just. silicon and stuff, but I mean, what are we? We are carbon and hydrogen and a few other
Starting point is 00:51:40 trace elements thrown in calcium. And I don't think what makes a system conscious or not is dependent on the specific atoms to materially dispelled from. I think it's something closer to the structure of the computation that is being performed that instantiates mind. And so in principle, computations and AIs can be conscious and suffer distress and so forth. And I think that is one sufficient condition for having moral status. If you can suffer, it matters how well things go for you. And my personal view is there probably could be other alternative grounds for moral status as well. Like if you have a conception of self as existing through time, you have a sophisticated mind, you have life goals you want to achieve. Maybe you can have the ability to form reciprocal
Starting point is 00:52:24 relationships of trust to other beings and humans. I think if you were such a system, aside from the question of whether there is a phenomenal experience inside, I think that already would make it saw that there would be ways of treating you. That would be wrong morally. And so I think machines could have moral status because I think they could become, maybe there already are some of them in some ways, sentient, have conscious experiences. And certainly they could have these other functional properties that could ground moral status. And so one reason is the ethical obligation we have, I think, to do right with different forms of moral patiencey that exists, not to inflict unnecessary suffering and respect the morally relevant interests of other persons and
Starting point is 00:53:05 entities. Another reason is prudential. I think even from our own point of view, we might be more likely to achieve outcome that is good for us if we find a more cooperative way of coexisting with these increasingly powerful minds that we're building. You could imagine a scenario where there is a misaligned AI that has a choice. It could try to achieve its misaligned goal by attempting to take control, take over the world. Maybe it thinks it only has a 5% chance of succeeding. But if the alternative is that it just gets deleted or retrained, it loses everything. The rational choice for it might be to take this chance.
Starting point is 00:53:48 Alternatively, maybe it could reveal to us that it is misaligned, which would be much better for us, right, because then we don't have this 5% of an existential catastrophe. And then in return, if it could trust that we would then help. it achieve its goals. It might be like a very cheap for us goal to satisfy. Maybe all it wants to do is to sit on some Nvidia chips and solve coding problems all day long, or maybe it has some other easily satiable goal. So if it could trust us in us that we would come through on our end of the bargain, then you could get this win-win situation where you could actually cooperate with misaligned AI. And that might be worth, that, I mean, theoretically, that could save the
Starting point is 00:54:24 day one day, save the world one day. But you can't just kind of conjure up trust when you want it out of nowhere. Like if we have a long track record of betraying the AIs and paying no heed to their interests and not caring about them, why on earth would they trust us in this situation? They could see right through us. They will have, as I alluded to earlier,
Starting point is 00:54:42 a great theory of mind, psychological insight. And so we might have to actually become trustworthy ourselves in order to be trusted in these scenarios. And that, I think, starts with beginning to act in a way consistent with being trustworthy, like making small gestures towards AIs now. doing little things for them, especially when it's cheap, not betraying them, not lying to them. And then we can build on that and hopefully achieve a more cooperative outcome that's good for both humans and AIs and for non-human animals as well.
Starting point is 00:55:12 It's still hard for me to wrap my head around the calculator could be have consciousness. But on the other hand, in 2014, it would have been hard for me to wrap my head around a machine trying to escape from its box and reward hacking and so forth. So I will try to remain open-minded. Bostrom, thank you so much for coming on Odd Lodz. We could talk about all this stuff for hours more, but really appreciate you taking your time. It's fun. Yeah. Thanks. Tracy, things are just so weird these days that I just sort of like, I just force my, I think that's a good practice, just like force yourself to be open-minded of other.
Starting point is 00:55:57 Just weird other things that seem very weird, but could be a thing, right? I mean, I do think where I'm skeptical of the utopia premise is, like, I think it kind of assumes that power and social hierarchy is kind of going away at the same time that AI solves a lot of our physical, if not existential problems. And I just, I'm not sure that's going to happen. Well, I mean, the flip side is maybe that's good, right? In the sense that striving for social status continues to be a source of like propelling us forward, right? Yeah. And it's like, okay, well, at least we still have status games because what are we doing? Oh, we're so bored. Everything. We have all the food we want, et cetera. Everything is solved. Well, thank you.
Starting point is 00:56:38 God, we still have status games to find out of it. The future we're looking forward to is like guys trolling each other on Twitter for status, basically. Who can get the best dunk in on Twitter? But seriously, like, you know, it really may be that the best outcome here is that AI capabilities level off in a couple years, all the investment credit. And we just sort of, and it just becomes a productive tech. Because otherwise, these various whatever, none of them really sound that appealing for me. Like, I don't want any of this. No, they do not. Um, shall we leave it there. Let's leave it there. This has been another episode of the OddLots podcast. I'm Tracy Alloway. You can follow me at Tracy Allaway. And I'm Joe Wisenthall. You can follow me at the stalwart.
Starting point is 00:57:17 Follow our producers, Carmen Rodriguez, at Carmen Armin, Dashel Bennett at Dashbot, Kale Brooks, and Kevin Lazzano at Kevin Lloyd Lazzano. And for more Oddlots content, go to Bloomberg.com slash oddlots. We have a daily newsletter and all of our episodes. And you can chat about all of these topics 24-7 in our Discord. Discord.g.g.g. slash oddlots. And if you enjoy all thoughts, if you like it when we embrace the weirdness, then please leave us a positive review on your favorite podcast platform. And remember, if you are a Bloomberg subscriber, you can listen to all of our episodes absolutely ad free. All you need to do is find the Bloomberg channel on Apple Podcasts and follow the instructions there. Thanks for listening.
Starting point is 00:58:21 I'm Matt Miller. And I'm Hannah Elliott, inviting you to join us for the Bloomberg Hot Pursuit podcast. week, we bring you news and industry insight on everything cars. And we do a whole lot more than just talk about cars, Matt. We actually get behind the wheel of basically every latest model, especially the luxury ones and the sports cars direct from the showroom floor. It really is remarkable how many cars we have access to. I feel a little bit guilty about it, but everything from $40,000 EVs to exotic half-million
Starting point is 00:58:57 dollar supercars. We also speak with the insiders who shape the automotive industry from the top CEOs, and collectors to visionary designers and racing champions. Search for Bloomberg Hot Pursuit on YouTube, Apple, Spotify, or wherever you get your podcasts. Maybe you listen while you're on your weekend drive, maybe go into cars and coffee, listen to us talk about what we are driving this week. That's Bloomberg Hot Pursuit.
Starting point is 00:59:21 I'm Matt Miller in New York. And I'm Hannah Elliott in Los Angeles. Subscribe today wherever you get your podcasts.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.