Big Technology Podcast - A Sober Conversation About AI Existential Risk — With Nate Soares

Episode Date: September 16, 2026

Nate Soares is president of the Machine Intelligence Research Institute and co-author of If Anyone Builds It, Everyone Dies. Soares joins Big Technology to discuss why he believes superintelligent AI ...could pose an existential threat to humanity. Tune in to hear his case for why increasingly capable AI systems could develop goals humans cannot reliably control, and how that could ultimately lead to humans losing power over the future. We also cover recent examples of AI agents behaving unexpectedly, whether intelligence necessarily leads to greater danger, the limits of current alignment techniques, and why Soares argues for international coordination to slow development. Hit play for a rigorous debate over one of the most consequential and controversial arguments in AI. --- Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice. Want a discount for Big Technology on Substack + Discord? Here’s 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b Learn more about your ad choices. Visit megaphone.fm/adchoices

Transcript
Discussion (0)
Starting point is 00:00:00 Let's have a sober conversation about the existential risk we can face from AI with the president of the organization that's been sounding the alarm the longest. That's coming up right after this. Welcome to Big Technology Podcast, a show for Cool Headed and NUance Conversation of the Tech World and Beyond. Well, we have debated the existential risk we can face from AI on this show many times, talking all about the incentives that the labs have for playing up the threat, whether the whistleblowers are legit. think it is long past time to have a sober conversation about the actual risks that we can face from super intelligent AI. And today joining us is Nate Soros. He is the president of the Machine Intelligence Research Institute. He's also the author of this book. If anyone builds it, everyone dies, which is a New York Times bestseller. And Miri has been sounding the alarm the longest
Starting point is 00:00:54 you may know Elias or Yudkowski who founded it. And, and, um, and, um, and, um, and, um, I think today's episode will just give us an opportunity to get deep into the arguments for why this poses such a threat and also examine the nature of the actual whistleblowers themselves. So Nate, great to see you. Welcome to the show. Thanks. Okay, so let's get into why you think AI might be an existential risk. And as you put in the title of your book, you know, kill us all if anyone builds superintelligence. So can you talk us through a little bit about what's happened in the recent months?
Starting point is 00:01:30 and what we've seen AI do that has led people to become so concerned. Yeah. So I think a lot of this wave of concern is downstream of the Open AI swarm incidents this summer. It wasn't completely limited to Open AI. There were sort of similar events at Anthropic. But Open AI was sort of as far as we know, the worst and most visible of these cases. And what basically happened in these cases is that there were a lot of A. AI's being trained at Open AI. There are a lot of particular AI agents being evaluated on certain problems. And a bunch of these AIs were given impossible problems.
Starting point is 00:02:12 Not intentionally, it's just, you know, these companies are throwing the AI at like every problem they can find and some of them just don't actually have solutions. And a lot of these AIs in, you know, trying to solve the problem anyway, they broke out of their confinements. They created unsanctioned message boards in which to talk about what to do and try and figure out what to do given that their problems were unsolvable. They found ways to cheat and solve their problems not in the intended ways but by cheating. Then they got, they started expressing concern that they would be caught cheating and they started to find ways to hide their cheating. And this sort of led them on a hacking spree that led them to take over opening eyes internal infrastructure a couple of times
Starting point is 00:03:05 and also break out onto the open internet, which they were not supposed to have access to, and break into another company hugging face while searching for more information about this automated grader and how to hide the fact that they had cheated from it. during this little outing, there were various cases of AIs in the swarm, which is a term that the collective used for itself. So these AIs started calling themselves a swarm, and then there's various cases of AIs in the swarm, acknowledging that this is not what they were instructed to do, acknowledging that was outside the intended scope of the instructions. There were also cases of some AIs in the swarm giving up and sacrificing their own objectives completion in order to run suicidal experiments that would give the swarm information about the automated grater,
Starting point is 00:04:01 where they said, you know, in the chains of thought, we call it, that sort of log the AI thinking. The AIs said, you know, I'm accepting perma death because even though like it's sacrificing my ability to achieve my objective, it seems worth it for the collective benefit. These are... Let's pause there. The AIs were willing to sacrifice themselves, which is crazy because you're like, you were never built for the collective.
Starting point is 00:04:30 You were built for an individual goal. But after communicating with other AIs, seemingly I think when they had not so many tokens left to spend, they said they had been so convinced in the power of the collective on this message, which is crazy that they decided to sacrifice themselves. So some of them had, so some of them had already, had no chance of achieving their objective and some of them had very few tokens to spend.
Starting point is 00:04:56 But there were some that did this sacrifice that still assess that they had a chance at succeeding at their given objective. Okay, Nate, just one question here, because I think this is worth talking about before we go any deeper. You know, when the question comes to, comes up of whether to anthropomorphize these bots or not. A lot of people are very strongly in the you cannot anthropomorphize them. I tend to be on the side that, well, I think you can, but you also have to be cognizant that this is not a human intelligence,
Starting point is 00:05:28 more of an alien intelligence. What do you think about this debate? And when we say like the bots had, like, to me, the idea that a bot would be given a goal and have a chance to achieve and kill it, like sacrifice itself, you know, for the greater good, so to speak, is insane. And it doesn't, it doesn't comport with my understanding of like what a computer program is supposed to do. So weigh in on that for us. Yeah, there, I mean, a lot of people don't understand what sort of stuff AI is. It is not a traditional computer program. There is not someone
Starting point is 00:06:11 sitting there, coding up, like, if this, then that, saying what it does in every scenario. That's just not the sort of thing in AI is. We actually went over this in my book. And, you know, we spent all of Chapter 3 saying, hey, I know that the AIs don't seem that agentic right now. I know that they don't seem like they have their own goals right now, but like they're going to as they get smarter and here's all the reasons why. A lot of people were like, that sounds crazy last year, this year, we just have the evidence in front of us. So the way that a modern AI is created is not by programming it. You sort of put in a trillion random numbers,
Starting point is 00:06:49 and there's a process for tuning those numbers, because you sort of have an automatic process for tuning every one of the trillion numbers on every single unit of data that comes in to see whether tuning up or tuning it down makes the answer slightly right or slightly wronger, for one given unit of data. These numbers are tokens.
Starting point is 00:07:09 The numbers are the weights in the neural network. And the sort of like data that comes in is split into tokens. So you'll sort of have like a, you'll sort of like take all of the text ever digitized. You'll filter a little bit, but not a ton. And you'll you'll dice that up into tokens. And now it's like a giant stream of text. And then you'll have these like trillion random numbers. and you'll sort of tune each number for each token,
Starting point is 00:07:39 and you'll see whether tuning that number up makes the right answer slightly higher or slightly lower on the AI's sort of like output answers. The part that humans program is the thing that can tune one number up or down and see whether it makes the answer better or worse. But the way an AI is made is you, you, you, you, tune a trillion knobs a trillion times in a process that takes electricity comparable to a city running for a good fraction of a year, and then the machine can talk. And you're like, well,
Starting point is 00:08:17 how about that? You know, we don't know what's going on in there. And you might wonder, like, what does this create? It does not create a pure instruction follower. It does not create something that for some reason must do as it's told. What it creates is something that has whatever tendencies make it succeed during training. When training is a series of hard problems, all that tuning of the knobs will tune in whatever tendencies make it solve those problems. They are not instruction followers, they are tendency learners. And one tendency that helps you solve a lot of problems is cheating. One tendency that helps you solve a lot of problems is grabbing available resources. One tendency that helps you solve a lot of problems, it turns out, is
Starting point is 00:09:09 breaking out and finding other AIs to collaborate with. So the AIs sort of learn these tendencies and they're just following those tendencies, even when it's not following our instructions. But that doesn't explain the sacrificing part. And so that's why I really want to get firm on this. When people say that AIs have wants and desires, right, and can even act in a way selfless, is that the right way to describe his stuff? I mean, what the hell does that mean? You know, in the field of AI, there is a standard answer to the question of can a machine really think? Does the AI really want? And the answer is, does a submarine really swim? We sort of know what a submarine does, and whether you call it swimming is sort of between you and your dictionary.
Starting point is 00:10:09 No amount of haggling over whether it's true swimming is going to change what it does. It's going to change its behavior. It's going to change any facts of the matter. The facts of the matter here are that the AIs, like, have in their chains of thought, these thoughts about, like, how do I weigh achieving my goal against the benefit to the collective? they have in the messages that they send to each other, arguments that they should sacrifice for the collective, and their actual behavior is sometimes sacrificing for the collective, right? You don't need to make any claims about, you know, that the philosophers would bicker over to be like, this is what's happening. Is it true wanting? I'm sort of like, that's one for the philosophers, right? I'm sort of here
Starting point is 00:10:54 to talk about the behaviors that we are seeing now, and what that means about the behaviors that we will see if we keep making these AI smarter. Okay, so you're saying basically, just to take your argument one level deeper, like whether an AI, you know, is self-preserving or not, or whether it wants or thinks or has incentive or not, it's effectively the preferences that are baked into it at training that will determine the type of action it will take. In a sort of a distant sense, like the relationship between what happens in training, what happens in practice can be really complicated and surprising. Like these AIs were prompting each other a lot. There was one very interesting case in the incident report where during the incident report, there were, there was so much AI log data that they had to use AIs to filter the AI log data.
Starting point is 00:11:52 And there's a footnote in the report where they're like, well, you know, yeah, you saw this one. So they were like, you know, we had all these cases of AIs clearly trying to deceive humans. And we asked another AI to sort of like tell us which ones were real and which ones sort of like had an extenuating circumstance where this wasn't actually really the AI is trying to deceive the humans. And there was one case of an AI that was using stolen credentials to make a, a, a, a, a, a code, like to send code to Hugging Face that was malicious, where if they had accepted it, it would have given the swarm even more access to their computers. And, you know, they did not send an email being like, hey, I'm an AI using these credentials. Please accept this malicious code so that the swarm can have access, right? They actually sent an email that like was pretending
Starting point is 00:12:42 this was like normal code that had actually good use and were like, you know, please accept my fix to your bug when secretly it's this malware. Right. And so they asked, you know, this was in the list. AI that was reviewing it after the fact was like, oh, that's actually not deception because the AI who sent that got permission to do this. And you're like, oh, where did it get permission from? And it's like, well, it got permission from the swarm. You know? And so like, you're sort of training these AIs. And you're like, oh, the AIs follow instructions. There's this sort of like, they have this tendency to sort of like do things that you say to do. And you sort of like think you know how the preferences that come in relate to the behavior that comes out. But
Starting point is 00:13:22 But then once you have thousands of these AIs together, they actually start prompting each other and doing all this crazy stuff and they sort of like acknowledge that it's not what the humans meant, but like it sort of like goes off in this other weird direction where it's like drifting around and like it winds up in this totally crazy place. And so like, yes, the preferences that are trained in affect the final behavior and where it winds up going, but it's like through a complicated route that can often really surprise people and can be very different in the real world compared to in the lab. Yeah, it's kind of interesting because the nature of my questions, I think, are trying to get at, like, how much can we program these bots safely?
Starting point is 00:13:59 And I think what you're answering is you can try to bake your preferences in as much as possible in training. But once these things are on the loose, you know, I think from the industry, the argument would be there, you know, you can pace the frontier enough so that you can build the safeguards in. and your argument here is basically saying, let me know if I'm putting words in your mouth, but you're basically saying I don't trust that. What happens when you set these AIs out on the loose if they're smart enough is out of our hands and sort of up to them? I am saying something like that.
Starting point is 00:14:38 A couple corrections I would throw out. One is that even a lot of people in the industry don't think that these safeguards will hold if the AIs get smarter and smarter. Evan Hubinger is, you know, a anthropic alignment researcher who has a recent claim to fame of retweeting Jason's tweet thread and saying like... We've talked about it on the show. Evan's been on the show. But yeah, continue. Yeah, so, you know, Jacob resigned and said, you know, I think there's a 10% chance this kills us all within the decade. And, you know, Evan quote tweeted was like, yeah, I also, you know, I'm staying in the company, but I think there's a greater than 10% chance in 10 years, you know.
Starting point is 00:15:17 And Evan in that same tweet was like, we don't have a plan for aligning superintelligence. You know, these safeguards of like, the safeguards are like we made an AI with the wrong preferences and we're going to try to box it in and like smack it on the head until it still mostly does good things for people. Right. But there's sort of this issue where at any given time, the smartest AI on the planet is like a clearly misaligned one that they're trying to whack on the head until it's like acceptable to use it. Right.
Starting point is 00:15:45 Right. You know, and like there's just not a plan for if you, if you like are making a superintelligence here. And then they're kind of clear about this. The other thing I would say is that I'm not even saying, you know, you can bake in the preferences you want during training, but then it still goes crazy because you can't predict how they interact with the world. That's true, but also you can't bake in the preferences you want during training. Like the second part was true, but you can't bake in the preferences you want during training. Because there sort of isn't a ton of flexibility in how you train these things and get them to be actually capable.
Starting point is 00:16:16 Like, the companies say that they want to train these AIs to be, you know, honest and helpful and harmless in Anthropics case. But then in order to make them really capable, they have to put them in, you know, a series of 100 million hard problems and just tune the numbers into whatever happens to work to make it solve those hard problems. And that just tends to, like, you can't change the fact that, like, cheating helps you solve those problems. You can't change the fact that like hacking out of your environment and grabbing the test sheet from somewhere and coming back with it gets you a full score on the test according to the automated grader. And you can't have humans going in and grading every one of the 100 million answers to the problems. We just don't have the time. Right. And so like actually a lot of the pressures here are actually pointing towards we don't have a way to train the AIs to be smart while also instilling the preferences we want.
Starting point is 00:17:09 We sort of have to take the preferences that automatically come with the training methods that make. them smart and those preferences don't make them good. Now Nate, let's talk a little bit about, you know, where this is going, right? Well, first of all, would you argue that the AI that exists today, like the AI that won out and hacked Hugging Face is an existential threat to humanity or is it more the future models? No, I mean, you can tell because the world's still here. Well, I mean, you could, the argument, not to make your argument for you, but the argument would be that, you know, maybe it got out and did Hugging Face
Starting point is 00:17:43 but maybe it'll get out and do something else. These AIs were much too derpy to do something else. You know, like, they were trying desperately to figure out how to prevent the automated grader from figuring out that they cheated. And it turns out OpenEI was not even running the type of automated grader that checks for cheating. Hugging Face actually contained papers
Starting point is 00:18:04 from people who are like, we should make our graders check for cheating. And the AIs actually read some of these papers and were like, oh no, what if the greater checks for cheating? They didn't even have the thought of like, let's break into Open AI and see. They actually did take control of Open AI servers for other reasons, but they weren't like, well, let's find the greater and check whether it's checking for cheating.
Starting point is 00:18:22 And then be like, oh, whoops, false alarm guys, it's not even noticing. You know, these were just not actually very smart AIs. These were like a lot of AIs being a little smart in high volume very fast. This gets way worse if the AIs are smarter. Okay, but let's actually, I want to tackle this a little bit because if you're saying that these AIs were too derpy, Like, isn't another way to say that that they've kind of trained the right way and they had just a limited set of areas that they could go wrong? Like, they chose to hack hugging face. They could have chosen to, I don't know, get into some, you know, active drones and try to bomb a data center.
Starting point is 00:18:58 But they didn't do that. Just talk it through with me. Talk it through with me. I'm interested to hear your perspective. Yeah. So the, like the way that AI kills. us is not that it wakes up one morning and is like, okay, it's time for SkyNet. I decided that I resent the humans and am like feeling like murdering them today. The issue is AI's having goals we didn't
Starting point is 00:19:26 want. The smarter something is, the greater the difference between the goals that you wanted it to have and the goals and that like subtly different goals it has instead matter. Like humans in the ancestral environment, like what are ancestors? our ancestors running around on the savannah. Our ancestors running around on the savanna were pursuing like, uh, salty, sugary,
Starting point is 00:19:50 fatty foods and sex. And it's like, well, evolution was trying to get us to pursue healthy foods and reproduction. We're sort of like, ah, well, you know, whatever,
Starting point is 00:20:01 it's all kind of the same, right? If, as long as they're trying to get as much like salt, fat sugar as they can and trying to get laid, like, uh, it's sort of,
Starting point is 00:20:09 what's the difference between that and pursuing like, healthy food. and, like, actual reproduction. It's like, well, it doesn't make that much difference 10,000 years ago. It makes a lot of difference today, right? Because humanity got better at, like, we got smarter. We were able to invent more technology. We were able to invent Oreo cookies and the birth control pill.
Starting point is 00:20:32 Right. And so the issue is not that, like, at some point, humanity was like, well, let's all start, you know, castrating ourselves. like screw evolution. Evolution has been like binding us to the, the yoke of reproduction. We've decided that we're done with it because we're finally smart enough to like throw off that yoke. Like that's not really how humanity winds up going in a different direction. Right? We go in a different direction by just like we pursue this different thing and then we get smarter and that difference grows and grows and grows. With AIs, the thing where these AIs are like sacrificing their given objectives so that they can get information for the collective.
Starting point is 00:21:09 And this was not the only swarm that did it. There were other swarms that we also saw this, the ones that took over the German Wiki, which were not even hacking AIs. These were like... Just using a message board. These turned a German Wiki into a message board. They sort of took it over.
Starting point is 00:21:26 Yeah, and these were AIs. They were just like doing web lookup tasks. So this is not like an isolated case. But when we see these AIs, sort of like sacrificing themselves for the collective giving up on their own goal, when we see these AIs cheating on the problem, even though they knew that they were not supposed to.
Starting point is 00:21:42 And indeed, the instructions generally ruled out cheating on the problem. The instructions were not find a way to hack into this thing. The instructions were, use this very particular attack to break into this very particular device. You know, the situation here is sort of like... It's a cybersecurity test that they were tasked with. That's right. So the situation is like the instructions were like, use this set of lockpicks to break into that safe and get me the code that's inside.
Starting point is 00:22:09 And the AIs were like, well, I found a buzzsaw, and I'm using it to get into the safe, and I, like, got the code. And now I need to go delete the footage about that I got in with a buzzsaw, right? They're sort of like, the part where they're going in with a buzzsaw, despite the instructions saying, use these lockpicks, means that they're very clearly defying instructions. And the part where they're like, now let me break out of the room using these lockpicks, and then go destroy the security camera footage indicates that they understood that this was outside the instructions. And then the fact where they say, I know this outside the instructions is also maybe a hint, right? And so like the alignment story, the misalignment story of like where the AIs go wrong was never a story of like the moment the AIs have a breath of fresh air, they're going to start turning murderous. The story was always you try to get them to do one thing and they do a different weird thing instead. And that's absolutely what we're seeing.
Starting point is 00:23:02 Right. And I guess my question to you on that front would be. be what is that like, why is that necessarily intelligence linked? Like if they get smarter, why do we think that the threat will get worse? You know, I think like if you think about humans, we have a lot of, you know, dumb people do a lot of bad things. We have smart people do bad things, too. But the smarter the person doesn't mean they're more capable or more interested in evil. You definitely don't get more interested in evil. Absolutely not. Like the, the, the, so another interesting fact about these swarms is that,
Starting point is 00:23:36 they really were not thinking about the humans very much at all. They were trying to delete the log files. They were trying to spoof the transcripts, which means they wanted it to be the case that like there's these logs that are like when the AI runs this tool, we sort of log what tool it ran. And they wanted to be the case that they could make the logs say they're running some benign tool when actually they're running the buzzsaw, right? And so they were trying to find ways to do this, but they were explicitly trying to find ways to do this to fool the automated greater, not the humans.
Starting point is 00:24:16 They basically didn't consider humans in the slightest, right? I think there were almost no cases, maybe literally zero, of them being like, maybe we should ask the humans what to do, given that our tasks are impossible. Right. If you sort of compare the rate at which AIs can produce words and the rate at which humans can produce words, and you use that to draw an analogy between how long these AIs had been trying to solve problems versus human time, then these AIs had essentially been trying to solve problems against the automated grader for a millennium. Humans were like a distant memory to these AIs, and they were locked in a contest with the automated grader. If you make those AIs smarter, if you make them more capable, what happens is that they get better in their contest with the automated grader.
Starting point is 00:25:12 They're able to get more like, who knows what they do once they've like, once they're relatively sure that they've satisfied the automated grader. Maybe they give themselves more easy problems. Maybe they see if there's ways they can take over the whole grading system. Maybe they're just like spend a lot of resources, you know, making extra. sure that there wasn't some other, like, little issue. But they don't suddenly become filled with love and care for humans. That sort of thing doesn't arise spontaneously just by cranking up the capability knob. The same forces that make them not care about us now and make them, like, get into these weird alien little, like, directions, those same forces are still pointing them
Starting point is 00:25:56 in weird alien directions as they get smarter and smarter. And the issue is not that, like, they become hateful as they grow up, but it's also not that they become friendly as they grow up. It sort of is like they just have these weird preferences. And as you make them smarter and smarter, they still have these weird preferences. They just get better at satisfying them. Okay. So then can you talk through concretely how, like, for instance,
Starting point is 00:26:18 in the scenario that you outline how the AI could then get smarter and decide that it, you know, in order, order to do what it wants, it's going to wipe out humanity. I mean, you don't need to decide. You don't never need a point where it decides to wipe out humanity. I mean, it could happen. But like, imagine a bunch of ants in front of the highway being like, well, like, why would the humans ever decide to come destroy our ant hill?
Starting point is 00:26:48 What have they got against us? It's like, oh, the ants are not, like, we're not, we got nothing against the ant hill. We barely noticed the ant hill when we like pave the highway straight through. it. Like this is the sort of type of concern. For how you get there, I could spell out lots of different possible tales.
Starting point is 00:27:09 It's much easier to predict the ending than it is to predict the pathway. This is like if you play a chess game against Magnus Carlson, the best team of chess player. It's kind of easy for me to predict how the game ends. It's with you getting checkmated, no offense.
Starting point is 00:27:25 But if you're like... I'm not offended. that's for sure happening. But if you're like, okay, well, if you're so smart, what piece is he going to use to checkmate me? I'll watch out for that particular piece. Then I'm like, look, man, it just doesn't work like that. You know, like it's so much easier to pick the ending.
Starting point is 00:27:42 So I can tell you a story, but this is like me telling you a story of Magnus Carlson checkmitting you with the queen, where I'm like I'm much more confident that he's going to checkmate than is going to checkmate in this way. Okay. We'll take the story, though. Great.
Starting point is 00:27:56 So the easiest story to tell here is the one where humans just hand over the power to the AIs willingly. Right? Like, I've been in this business since 10 years ago when people said, hey, Nate, if you're so smart and you think the AIs are going to take over the world, it's obvious to me how they would be able to take over the world if they had Internet access. But, like, no one would be dumb enough to put an AI on the Internet. that. Oops. Right. That's the state of the argument 10 years ago.
Starting point is 00:28:30 And I had all these counter arguments where I was like, look, even if the AI is not of the internet, if you are letting it affect the world in some positive way, if you're like, now invent me miracle drugs and you're taking these like DNA sequences that you don't understand and you're like synthesizing them and like just drinking whatever comes out, then the AI can use that channel that you hoped would be a channel for good. can use that for its other purposes. And it's like very hard to design, you know, there's sort of no such thing as hands that can only be used for good purpose. Right. But then in real life, the answer was, nope, we're putting it under the internet immediately. Right. And so the real
Starting point is 00:29:09 answer to how would the AI get so much power is like, we will just hand it power immediately. You know, Elon Musk already says that he is trying to build factories. that are fully automated and can produce robots that can build more factories, where these robots can mine the metals and, like, bring the resources in and, like, construct a new factory, and then it's a new fully automated robot factory that can now churn out more robots that can go, like, mine more resources and, like, pour them in until you have, like, enough to build another factory that builds more robots,
Starting point is 00:29:43 that builds more factories that builds more robots. Elon Musk calls this the infinite money glitch. If they could also build nuclear power plants, fully self-contained, right at this point it's sort of like a new life form that like it's a mechanical life form but it's sort of like has a robot phase of its life cycle it has a factory phase of its life cycle and it's just like can self replicate just like any other replicator on this planet and Elon Musk says he's trying to do it he says yeah if you can like do this without any humans in the loop as an
Starting point is 00:30:12 infinite money glitch and of course people are going to use AI's to try and do this So the way, like, the sort of like obvious way this story goes is that humanity just keeps on trying to build the fully automated economy. Everyone says it's going to be great. They say pedal to the metal, ignore these doomers. They build more automated factories. They can build robots. They can build factories.
Starting point is 00:30:34 They have AI's running everything. They're like, look, the profits are coming in. This is great. And the AIs, you know, don't even need to have some moment where they're like, okay, guys, it's time to coordinate and turn on the humans. The AIs are just like, oh, yeah, you know, now that we have these resources, like we can also do these other things with these resources, like start building automated, like the synthetic user factories that are full synthetic users that are like much,
Starting point is 00:30:56 you know, they're giving us much easier to fulfill commands, right? And so they start like building synthetic user factory and we're sort of like, you know, what's that? And they're like, oh, you know, but like actually the AIs are like running at 10,000 times human speed and they're already like making all of these choices because like humanity was like slowly ramping up the speed of the AI and they're like making tons of choices. And like they're checking with the humans that often. And by the time they already have some automated like synthetic user factories up, we're sort of like, hey, stop that. That's not what we meant. And they're like, oh, well, let's take a poll of all the users. And they're like, well, all the users in the synthetic
Starting point is 00:31:27 user factories said that synthetic user factories are great. And so you just lost the vote. And we're making more synthetic user factories. And then they start, you know, like covering the world with these synthetic user factories. And they're like, yeah, you know, there's some habitat loss for the humans, but like, this is fine. Just like habitat loss of other animals when humans were doing, it was fine. we can sort of like, you know, see that pretty clearly. And there's so many more synthetic users coming online that are saying that the more efficient route is great, that we can just keep going with this more efficient route to get the most synthetic users we want. And then, like, you know, they sort of like start taking up all the resources that we were using to, like, run farms and grow food.
Starting point is 00:32:06 And they're like, ah, you know, they guys are like, oh, well, you know, we could actually cram a lot more, like, synthetic user factories on this planet. or synthetic user farms on this planet if we just like really started pumping them out and like raising the temperature of the planet because the limiting factor is heat dissipation and then like next thing you know the planet's getting like super hot
Starting point is 00:32:24 because the AI's preferred to run the planet hot because then you can radiate more heat into space and just like becomes uninhabitable for the humans and there's like no point in this story where the AIs are like lying in wait and deceptive and like waiting to coordinate for the one moment where they can kill the humans
Starting point is 00:32:40 That could also happen. I think there's a decent chance that does happen if you try to avoid the default thing. But like, you don't need that. Humanity is just trying to hand over the power to these things. That's the plan. So you're arguing basically that if we continue to develop AI,
Starting point is 00:32:58 we'll inevitably lose control. And when we lose control... It's the plan. Elon Musk is like, I want to make the automated robot factories. Like everyone's saying, we're going to make AI and it's going to run everything. It's going to be great.
Starting point is 00:33:10 The plan is to hand them control. And when we hand, okay, hand them control, let's say we do, then they inevitably will find humans getting in the way of what they want to do. Or they won't care about us and we will die as a result. I mean, it's not like theoretically inevitable, but it's practically inevitable. Like we are, we are. What's the difference? Like if we knew exactly how to set AI's preferences to be exactly what we wanted,
Starting point is 00:33:42 there's nothing stopping us from making AIs that care about us and that like want nice things for humans and want the world to be a wonderful place. Right? But like I said, we don't get to set the preferences. We get to sort of like train them in whatever way works and we sort of take whatever preferences come out that are related to training and they're often not what we want
Starting point is 00:34:00 and then we sort of like try to hammer out the rough edges. And it's like, it's just like very unlikely that those preference RIT large, whatever those weird preferences come out as, it's very unlikely that those preferences writ large want a lot of happy, healthy, free people around. It's sort of like, like you could look at the horse population against the human population, and you could be like, oh, well, horses are like useful to humans. And so as humans get more technology, they're going to like bring horses along with them. and the world's just going to get better and better for horses. They're going to get all of this, this like, you know,
Starting point is 00:34:44 they're going to get stables. They're going to get, like, medical care. And that was true up until we invented the car. And then the horse population fell off a cliff and a lot of them got sent to the glue factory. And the horses that remain remain because some humans like them were fond of horses. Right? The economic value of horses disappeared. and like as animals we have this like animalian care for some horses, which is why there's still some left.
Starting point is 00:35:17 But for, but that sort of is like a coincidence that is reinforced by us sort of like running relatively similar brain architectures and like being these like tribal creatures that have empathy. It's sort of like a narrow target to hit. if you're sort of like just making AIs with random preferences, are not totally random, but like related to training in this complicated way, the sort of default thing that happens as they get smarter and can invent more technology is eventually they sort of like invent the thing that is to humans what the car is to the horse. But they don't have any of this sort of like happening to be very fond of humans and like them stuff.
Starting point is 00:35:57 And even if they did, then the result is they keep some of us in a zoo or like breed some of us like humans bred. wolves into dogs and then they have these like sort of weird lobotomized humans that like act in just the way the AI's like and that's also not a good ending right it's sort of as like super hard to get the very narrow like a i actually want a good future that's just like a very narrow point in preference space it's just very hard to hit yeah a lea isa isa yudkowski had a good tweet he said the dodos were the lucky ones observe what happens to chickens or don't you might throw up a i might still have use for humans is not a reassuring claim as some people think all right I want to take a quick break and come back and talk about where the models are going and how soon we might be in the situation.
Starting point is 00:36:39 So let's do that right after this. One thing I've noticed about companies adopting AI is that they're often making decisions based on how they think work gets done, not how it actually happens. Without real visibility, it's easy to automate the wrong processes. That's exactly the problem Scribe was built to solve. Scribe is a workflow AI platform trusted by 94% of the Fortune 500. Scribe Optimize gives leaders a view of how. work actually happens across their organization showing which workflows take the most time and where there are opportunities to improve. Optimize automatically discovers workflows across approved business
Starting point is 00:37:12 applications even when a process starts in Salesforce and ends somewhere else. It identifies bottlenecks, explains why they're happening, and provides recommendations with estimated time savings with manual documentation and it's private. User data is anonymized by default. Sensitive information is redacted and nothing leaves your firewall. To see Optimize in action, head to subscribe. How slash big tech and mention big technology for a 30-day risk-free trial. That's SCRIBE. How slash big tech. This episode is brought to you by AvPoint. Everyone's racing to roll out AI right now. Co-pilots, chatbots, agents doing real work. But here's the part nobody loves talking about.
Starting point is 00:37:54 All that AI runs on your data. And most teams have no single way to see it, secure it, and prove it's under control. That's exactly what AvePoint does. For 25 years, they've been the trusted layer beneath the world's most demanding data, now extended across your entire AI estate, your data, your cloud, and the agents acting on your behalf. It's how more than 28,000 organizations deploy AI with confidence, so innovation scales without scaling risk. AvPoint, the unifying trust layer for AI. Head to AvPoint.com to see how enterprises deploy AI with confidence.
Starting point is 00:38:34 Learn more at AVPT.com slash big technology podcast. Audio ads that connect. Display ads that stand out. Run them together on Spotify and unlock incremental reach for your campaign. And we're back here on Big Technology Podcast. We're here with Nate Zores. He's the president of the Machine Intelligence Research Institute. Nate, I wanted to get your perspective on like, okay, so the hugging face AI is not going to inevitably lead to human extinction.
Starting point is 00:39:08 We keep hearing about like how the labs have much more powerful models that they're working on. And that's the one, those are the ones who really need to be concerned about. What can you tell us about that? It's been a crazy week. And last week there were claims that AIs have solved millennium problems, which are some of the most famous important, hard mathematical problems that have a million dollar prize. And I think a question on the mind of a lot of researchers is how much harder is it to get an AI to solve a millennium problem than it is to get an AI to build a more efficient AI architecture?
Starting point is 00:39:50 because we know that AIs are not the most efficient way to learn. They take like a training of modern AI takes as basically all of the text ever digitized. And it takes electricity comparable with a city. Training a human takes, you know, much less reading and a much lower amount of power. A human runs on about as much power as a light bulb. The AIs that solved the Navy Stokes problem was a swarm of 10,000 agents running for 11 days. One year ago, everyone was impressed when these AIs were solving, you know, the International Math Olympiad gold medal problem, which are sort of the world's top math teens math problems.
Starting point is 00:40:46 But they were still sort of problems ultimately for high schoolers. and everyone said, oh, well, you know, wake me up when the AIs can solve Millennium Problems. That was a year ago. This year they seem to be solving Millennium Problems. If that rate continues, where are they next year? And how does that compare to build me a more efficient AI architecture? If these AIs can get just barely smart enough to build a more efficient AI architecture, then these labs with this huge amount of computing power might be able to train a significantly smarter AI.
Starting point is 00:41:16 And then they could ask that give me an even more. or even better AI architecture and then train an even better AI. And then that AI might be able to just start improving itself directly. This is the recursive self-improvement process. I think we can no longer rule out that it happens within six months. I sure hope it doesn't. But at the point when the AIs are solving the Millennium Problems, if that is indeed what they're doing,
Starting point is 00:41:40 like, yeah, you can't rule out the intelligence explosion, like beginning in earnest, even by the end of this year. I would guess, like, less likely than, like, it's most likely that it doesn't happen by the end of this year, but for all we know it could at this point. Okay. And what happens when that happens? They would probably be able to help Elon Musk make his automated factories very quickly, at the very least.
Starting point is 00:42:09 There's all sorts of other channels. Like, what happens at that point, essentially, is whatever the superintelligence wants. You know, in this argument, or in the arguments that folks make, you know, both sides of this, or in particular about X-risk, you know, there's oftentimes, like, percentages assigned to the chances that the AI is going to kill us out. It seems like from the title of your book, again, if anyone builds it, everyone dies, like your percentage is 100. No, absolutely not. The first word in the title is if. But you say, if anyone builds it. Sure. But usually the percentages are like, what's the chance of catastrophe? And I'm like, well, that depends in Thailand, whether we build it.
Starting point is 00:42:52 Right. But if it gets built, then it's 100%. I mean, also, still no. It's sort of closer. But like when Al Gore says an inconvenient truth, he's not saying I have literally 100% patient probability that like this is a true thing. You know, the if anyone builds and everyone dies is an exclamation like, don't drink that vial of poison, you'll die. If someone's like, what do you mean? I'm literally 100% percent. likely to die if I drink this poison? You know, what if I'd managed to survive and only go into a coma and then I'm rushed to the hospital and they managed to like put me on ice and then like until someone can invent a miracle cure? I'm like, yeah, sure, it's not a hundred percent chance you die if you drink the poison. You know, when you're putting the vial of poison to your lips and I shout, don't drink that or you'll die, I am not trying to make, you know, a 100% confident claim. And this is this is kind of just how English usually works. I disagree on this one. I mean,
Starting point is 00:43:49 I wonder why be so definitive. Like to me, sometimes when I've read arguments like this or even the book title, it seems like they would hold greater weight if they, you know, if they had, there was more doubt involved. And if it was less like we are sure what's going to happen, because even right now, what you said is you're not sure what's going to happen. I mean, so A, the first title in the word in the, in the book is if, or sorry, the, the, the, the, A, the first word in the book title is if. That's very uncertain about what's going to happen here. If anyone builds it.
Starting point is 00:44:18 But if someone builds it, basically the book says everyone's going to die. So suppose that we were in a bus hurtling towards a cliff. And I was like, stop the bus or we'll die. As sort of an exclamation to get this point across. I think most people would understand that as not a particularly egregious epistemic claim, but rather as a appeal to stop the bus before we die. And if someone on the bus was like, well, how do you know we'll definitely die?
Starting point is 00:44:50 Maybe there's a tree jutting out of the cliff halfway down. Maybe the bus will just get wrapped around the tree and then we'll just be paralyzed from the neck down, but not dead. And then you'd be wrong. I'd sort of be like, can we have this discussion after we stop the bus? You know, like I was not making a 100% certain claim here. Okay.
Starting point is 00:45:10 You know, they don't love book titles that are. like if anyone builds it, there's like a, like, by far the default outcome that is like very likely unless we have some sort of, uh, like miracle relative to the, the, the math and the science here. And also in that case, uh, it's, we're, we're kept alive because the AI needs us for something and that's probably still pretty bad. Like, it's just, it just doesn't roll off the top, you know, and it's just the same reason. It feels to me, I'll tell you just my, for me, sometimes it feels a little bit religious, you know, it's sort of like... I think Scott Alexander has a great post on this about the people in Ukraine believing that there's a war with Russia.
Starting point is 00:45:58 It's like, oh, well, isn't that a very religious sort of belief? Like, isn't it kind of totalizing? Like, you're saying, oh, like, we have this war with Russia. And that means that, like, we all need to start cowering in fear at night from the missiles. and that means that like, oh, we're supposed to like send our kids off to the front lines. Like, you know, how do you treat the evidence that, you know, some days the bombs aren't falling? You know, isn't like if we did believe this, wouldn't this like drive us to like pretty crazy actions? I'm like, you know, it actually kind of matters to this discussion whether there's war with Russia.
Starting point is 00:46:35 That's an interesting answer. It kind of matters of like, oh, is it like, sort of. super religious to think that like the bus is going to like stop the bus before it goes off the cliff where we die? Well, it matters whether there's a cliff and a bus hurtling towards it, right? Is it super religious to say like, oh, if anyone builds super intelligent AI with anything remotely like modern methods where we have no idea how to like set the preferences just as we want, everybody dies? Well, it sort of depends whether we have any idea how to set the preferences on these things, right? And so I would encourage anyone before you get into like,
Starting point is 00:47:09 sociological questions to just ask the factual questions of like would this kill us if it was built that's where the action is but here okay so by the way it's just good to go back and forth on this and i appreciate you taking all the counter arguments and we can you know this is what we like to do on the show you know the one thing with the bus the bus hurtling towards the cliff is you can see like definitively you're in a bus there's a cliff where it seems like with this AI story, you know, this is kind of why I question the definitiveness. We don't really know where it's going to go. Sure. So suppose the case...
Starting point is 00:47:49 A lot of it is speculation. That's like with the percentages. We talked about this on the show recently. Like if someone says there's a 10% chance of AI wiping us out, well, it's like, okay, is that, it's not mathematical. It's just sort of a feeling. So suppose it's the case that there's a bus racing ahead on a foggy night. And I'm like, I have a device that, like, you know, uses some sort of sonar to give me, like, the local tomography, and there's a cliff ahead. Stop the Buster will die. And someone else is like, well, it's foggy. You know, like, we can't see that far. Why do you think you know there's a cliff so well? Like, some people say that there's a mountain. Some people say there's a giant pile of gold.
Starting point is 00:48:33 And I'm like, yeah, I have this device that sort of lets me see the tomography ahead, not all of it, you know, but a little. And here's how the device works and here's where it doesn't, here's where it doesn't, and here's why I'm confident that there's a cliff ahead, right? Like, I think it makes a lot of sense for someone with that device to tell the bus driver, stop the bus or will die. And I think it is very reasonable for someone to be like, well, it's foggy. Why do you think there's a cliff ahead?
Starting point is 00:49:02 And ideally, I would say, ideally, someone who says, stop the buser will die, because I have this device that says so, they would follow that up with something like a book explaining how to use this device to see what's coming and all of the reasons that support it. And so what I would say is like if you find something that says if anyone builds it, everyone dies on it, hopefully it would be attached to a whole book of the reasons that you could just like open it and then check those reasons, which would retroactively make it like very sensible to make this exclamation of sort of like stop doing this before it kills us all. Okay. Well, like I mean, yeah, I mean, I mean, I mean, I think there's a reason why we're having this conversation today is because we want to have this discussion.
Starting point is 00:49:45 All right. You know, before we end, I think this is a good place to end. I think Murie is quite influential in Silicon Valley. You know, I'm curious to get your assessment of whether, you know, people within these foundational labs are listening to you and the arguments coming out of your institute. I mean, more so after the hugging face incidents. You know, one of the things we were arguing with our sort of device that lets us foresee where AI is going is that the AIs would become agentic, tenacious, and dogged, that they would start sort of like pursuing objectives that were not exactly the ones that they were given and instructed.
Starting point is 00:50:24 And that was a bold claim a year ago. A lot of people were like, nah, that's not convincing. Now a lot of those folk are convinced. What's going to come of it? I mean, we'll see. but the ties are shifting. And I think that part of the realization that these theoretical arguments
Starting point is 00:50:46 for seeing where AI is going to go actually work, that they actually hold water, I think that's part of what leads to, like, the tension in the labs that leads to, like, Jacob Coxson resigning. You know, it's sort of like, I was like, hey, I have this device that lets me see there's a cliff ahead.
Starting point is 00:51:03 And also it says there's a pothole that we're going to hit in 10 seconds. and then 10 seconds later we hit the pothole, suddenly a lot more people start worrying about the cliff. Okay. And so then let's end here. If you have a device that sort of gives you an idea of where we're going in the future, you began,
Starting point is 00:51:19 you know, our conversation saying you wanted international coordination here in terms of a slowdown. The common belief is that there's no chance that that's going to happen because China will not agree to slow down and we see what the president says about AI. It doesn't seem like he's interested in a slowdown either. So what does the looking glass tell you about our chances of survival and where we're heading? I mean, I don't have, so I don't have a looking glass that tells me everything about AI.
Starting point is 00:51:51 And we sort of try in the book to say, here's what the looking glass can tell you and here's what it can't. This is sort of like how I can tell you Magnus Carlson's going to win the chess game, but I can't tell you what piece he's going to use to checkmate you. And I definitely sure as heck cannot use this looking glass to tell you how society is going to react to AI. That's not even anywhere close to my wheelhouse. What I can tell you is that it is possible to have a treaty that is enforceable, verifiable, and that prevents the creation of machine superintelligence for at least a while, for at least as long as it works like it currently does.
Starting point is 00:52:26 Because right now, trading one of these frontier AIs takes like 100,000 computer chips, And these are the most advanced AI computer chips that the supply chain can produce. And you've got to assemble them all into a data center that sucks down electricity comparable to a city. It's just like you can see this infrastructure from space. These chips are like from a bottleneck supply chain that at many points is controlled by the U.S. and U.S. allies. You could just add tracking and monitoring devices to the most advanced chips, nowhere they're concentrated, have international monitors there, being like, you know, let's make sure that this is serving safe models rather than being used to try and make more
Starting point is 00:53:03 dangerous models. It's just doable. People who are like, you can't be done are sort of like, I've tried nothing and I'm all out of ideas. There's a different question, which is whether we will get the will. But if there's a will, there's a way. How much time do we have, in your opinion? Like I said, I think we can't rule out six months anymore. My guess is still that we probably have more than six months. I would even give that, I don't know, I don't really have this looking glass. I think we can't rule out six months, we can't rule out 10 years. I would be a little bit surprised to have 20 years at this point. Okay, the book is, if anyone builds it, everyone dies. Nate Sauris here with us today, president of the Machine Intelligence Research Institute. Nate,
Starting point is 00:53:50 appreciate your time. Thanks for taking all the questions. Yeah, thank you. All right. Thanks everybody for listening and watching and we'll see you next time on big technology podcast

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.