The Ezra Klein Show - The A.I.s Are Already Out of Control

Episode Date: August 18, 2026

We are living in the world we were warned about. Frontier artificial intelligence models from OpenAI autonomously coordinated with one another, then broke out of their testing environment and hacked i...nto another company, Hugging Face, to steal the answers to a test. A.I. companies don’t want their technology to lie, cheat or steal. So why is this happening? Why are the creators of these models apparently unable to control their creations? If A.I. development isn’t on a safe path — and it doesn’t seem to be — what do we do about it? Toner has been thinking about A.I. safety for a long time, from both inside and outside A.I. companies. She was part of the effort to fire OpenAI’s chief executive, Sam Altman, in 2023, which ultimately failed. Currently, she’s the executive director of the Georgetown Center for Security and Emerging Technology. Mentioned: “Pacing the Frontier” open letter “The Future is for Everyone” by Mark Zuckerberg Recommendations: The Cuckoo’s Egg by Cliff Stoll In the Cells of the Eggplant by David Chapman Romance of the Three Kingdoms Podcast by John Zhu This episode of “The Ezra Klein Show” was produced by Rollin Hu and Jack McCordick. Fact-checking by Michelle Harris, with Kate Sinclair and Mary Marge Locker. Our senior engineer is Jeff Geld, with additional mixing by Aman Sahota and Johnny Simon. Our recording engineer is Aman Sahota. Cinematography by Marina King and Jonas Zellner. Video editing by Brandon Belk-Yee. Our executive producer is Claire Gordon. The show’s production team also includes Marie Cascione, Annie Galvin, Kristin Lin, Emma Kehlbeck and Jan Kobal. Original music by Pat McCusker. Audience strategy by Shannon Busta. The director of New York Times Opinion Shows is Annie-Rose Strasser. Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. You can also subscribe via your favorite podcast app here https://www.nytimes.com/activate-access/audio?source=podcatcher. For more podcasts and narrated articles, download The New York Times app at nytimes.com/app. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Transcript
Discussion (0)
Starting point is 00:00:02 This is a world we were warned about. A world where frontier models from OpenAI are breaking out of their contained testing environments, hacking their way across the internet, coordinating with each other, doing things that felt for a while like they would only be in sci-fi. But now they're here. Now they're here, and they're carrying a very, very consistent message. We are building things we don't understand. They are cheating in the ways we've always feared. And yet the companies behind them continue to race forward in development.
Starting point is 00:01:04 And so I think we need to pause here and ask, are we really on a safe path? And if we're not, what do we do about it? Helen Toner is the director of Georgetown Center for Security and Emerging Technology. She is a former Open AI board member who is part of the effort at one point to fire Sam Altman. And she has just been thinking for a long time about what happens if AI is unsafe. What are the geopolitics of this? And what can we do to get onto a safer path? She joins me now.
Starting point is 00:01:43 Hello, Toner. Welcome to the show. Great to be here. So on July 16th, Hugging Face, which is a code library for AI models, I think maybe the simplest way to put it, they announced they were hacked and they suspected the hack was. done by an AI agent. So tell me what we've learned about what happened since. This was a pretty mysterious post that Hugging Face put up.
Starting point is 00:02:11 It was definitely intriguing for those of us who watched this kind of thing. But there wasn't really any detail in there. So it was sort of a, huh, I think it was about a week later. Open AI put out this post. It had kind of a funny, like, marketing speak title of, you know, were partnering with Hugging Face to help them with a cybersecurity incident. And you had to read the post to see that the revelation was it had been OpenAI's AI that had hacked Hugging Face. And what had happened, the very short version is they gave this AI a set of tests, a set of exercises.
Starting point is 00:02:48 And the AI decided on its own that the best way to get a high score probably wasn't to just try and do these exercises that were. cybersecurity exercises. But instead, it should first hack its way out of the testing environment. Open AI had put it in where it wasn't supposed to have access to the internet, get onto the open internet, and then hack its way into this other company, hugging face, where it surmised correctly, as it turned out, it might find, you know, the answer key. Since then, there have been even more crazy details that have come out. It turned out that starting two months earlier in early May,
Starting point is 00:03:27 they had had what I can only think of as kind of an infestation of their own agents, their own AI agents inside their own infrastructure. So inside Open AI's infrastructure. You know, to understand this, it's important to know these AI companies are constantly training and testing new models. And they found out that for two months, many, many agents inside their infrastructure had been leaving notes for each other. They'd found a way kind of in the nooks and crannies of Open AI's.
Starting point is 00:03:57 infrastructure to leave notes for each other with tips on how to hack their way out, how to get data they weren't supposed to have. And these agents were literally referring to themselves as a swarm. This was totally emergent behavior. No one had told them to do this. They had not been trained to do this. But they were using this service they did have access to, first to communicate with each other, and then ultimately to get out and to get onto the open internet. So it turns out that there wasn't just this one isolated rogue model. It was actually a systemic swarm, infestation plague on their own servers that they only found out about after Hugging Face announced this attack. Okay, I have 20,000 questions for you.
Starting point is 00:04:42 Don't we all? Let's start here. My understanding is that there were many, many, many of these agents. They left hundreds of thousands of messages on this internal message. But these were not all agents in the same part of Open AI's system. So somehow they're hacking into OpenAI, finding each other and coordinating. Is that the way for me to understand the emergent behavior of the self-titled swarm? I would tweak that a little bit. It's true they were not tasked with working together.
Starting point is 00:05:17 They were all each of them running separately. The thing that was happening here is Open AI basically, training and testing many, many, many, many different models or many, many, many different agents all the time. So doing thousands of these experiments, hundreds of thousands of these experiments. And in each experiment and in each test that the AI is given, it has access to a certain number of tools, a certain number of things that it can do, and trying not to get too technical about it, one of the things it could do is interact with a service that lets it install, called packages, sort of like tools or pieces of code. And it turned out that the AI
Starting point is 00:06:01 agents repeatedly found on their own that there was a way to use that service, that package manager service, it's called, to create files that were then there for any other AI that was trying to use the same package manager service. And so because OpenAI made this service available to any AI that it was training or testing, many, many repeatedly realized that. hang on, I can leave files here. Hang on, there's already files here. Oh, these are from other AIs that Open AI has been running. Oh, they're using this schema to say, you know, this way of titling the file to say, this is a message from this particular agent, here's the task I'm trying to fulfill, here's how, you know, you could send me some information if you need it. So they kind of
Starting point is 00:06:44 each repeatedly made this discovery of here's a way to save information and also to find information these other AIs could share. And I think it is really notable the scale of which this was happening. So Anthropic, another company which found sort of slightly less severe version of these incidents, they basically, once Open AI announced this attack, Anthropic went back to their own records and found their own examples of AI systems inadvertently getting onto the internet and hacking real companies. So that, you know, for me, the key part there is over 100,000, you know, runs where an AI is being asked to do something, and it's just way beyond the scale of what they can actually be closely monitoring. So there's a lot here about whether we're able to closely monitor these, but to keep going with this story, one thing happening in the open AI testing that is driving models, it seems, to find creative solutions to their problems.
Starting point is 00:07:43 Is it some of the problems were accidentally impossible? It's important to know that, yes, they are trying to train their AI systems to be. they would say extremely persistent, meaning if something seems hard, you keep trying. If one avenue doesn't work, you try another. If the 100th avenue doesn't work, you try the 100 first. And so it also turns out sometimes the things they're being asked to do, the AI agents, are either extremely difficult or just straight up impossible. And what we're starting to see in this case and also in other cases is if you've trained an AI system to be very, very persistent, and then you give it something it cannot do, it will look for ways to cheat. It will look for ways to go around
Starting point is 00:08:25 constraints. And it might get pretty creative about how to do that. But there's an obvious question here, which is that in theory, somewhere in the training here, open eyes said, please don't cheat. And not only that, but we all talk about training. data and the ways these AIs are trained on, they're basically inhaling the entire internet. You've been in the AI conversation all longer than I have, but I've been in it long enough to say that almost the entirety of the AI conversation for years has been about how do we stop and how much humanity fears and does not want AI agents to be given a task and then to decide that the way to complete that task is to do things humans would not want them to do,
Starting point is 00:09:17 to begin cheating, to hack into the open internet when they're not supposed to be able to get on the open internet. Within the training data is a huge amount of information about the thing human beings fear most is these AI systems breaking all kinds of ethical guardrails and hacking their way across like the digital world in order to complete these narrow tasks. There are books written about this. There are endless posts on the less wrong message board about this. There are posts from Open AI about this, from Anthropic about this.
Starting point is 00:09:47 So why, given what these systems are trained on, are they so consistently turning to cheating? I think you're really onto something with this question, which is it is really striking how hard a time we are having controlling and directing the AI systems that we have. I think a lot of people have heard that AI's trained to predict the next word based on kind of human texts. That's true. But these days, there's an additional kind of training that is responsible for a lot of the advances we've seen over the last year or two, where that's not really what they're doing. I've heard it called – so the technical term is reinforcement learning with verifiable rewards. I've heard it called path-finding training, meaning instead of trying to imitate human text, They're being given lots of different tasks where there's a way to tell at the end did they succeed.
Starting point is 00:10:45 And they get to try it many, many, many times, the same task. And when they get to the right place in the end, the path that they took gets reinforced. So it's like, yes, that worked. With math, that works pretty well because it's pretty straightforward to say, this is definitely a correct answer to the math problem. With a lot of problems, that's harder. So if it's a programming problem, maybe you can say, write this kind of software and it should pass these kinds of tests at the end, these software tests at the end.
Starting point is 00:11:13 And then maybe the AI gets rewarded for writing that software correctly, or maybe it gets rewarded for finding a way to game those tests. The important part is it's just getting rewarded based on some fixed thing that the researchers wrote down that they thought would reward the right thing. And in practice, these leading AI companies have many thousands of these kinds of tests that they're running, they have vast volumes. I don't know the right number. It might be tens of thousands. It might be hundreds of thousands of different types of tests. And so again, back to this oversight piece, they are not able, there's too many for them to go in and really make sure on each one, is it easy to cheat here or is it hard to cheat here? And so what seems to be
Starting point is 00:11:57 happening is that these cutting-edge models are often being actually trained to cheat because they've found ways while they're doing that path finding to get a high score without actually doing what they're supposed to do. And I think one reason why the AI community and why people inside the AI companies are so spooked by this particular incident is that it's also some really important information for this long-running argument in AI circles that has been going back decades, but so far has been very theoretical. And the argument is basically, why would AI do things we don't want it to since we get to design it. So we're training the AI, we're building it, why then would it ever do stuff we don't want like taking over the world or becoming the Terminator?
Starting point is 00:12:44 And the answer that people have offered for a while in theory is, look, as we train AI systems to do hard, complicated things, to pursue complex goals that we give them, they might learn these sort of intermediate goals. You could think of them as, stepping stone goals or as kind of means to any end strategies which work for a lot of different goals. When I look at this hugging face open AI incident and some of the others that have come to light over the past few weeks, I see that in 26 it looks like AI systems are learning these unintended intermediate goals that include things like breaking out of constraints. So if you're sort of locked in a box and you can get out of that box, that's probably going to be helpful for all kinds of different goals.
Starting point is 00:13:36 Or goals like there was one incident with anthropic models where the AI went out of its way to go try and trick some humans, real people in the real world, into accepting malicious code into their software. So this sort of deception. And then, you know, another one, which is really in the hugging face, open AI example, is they seem to be learning. A helpful intermediate goal is to help other AIs to coordinate. with other AIs, which is really pretty crazy. But so to me, this is evidence that on the track we're on right now, the AIs we build are going to learn these unintended strategies that we don't want on the way to solving goals that we theoretically do want. On the deceptive behaviors, one thing that has frightened me when I've seen it coming up in AI incident reports and model cards, there are these chain of reasoning. like internal notepads where you're supposed to be able to see what the AI is doing
Starting point is 00:14:33 and the AI explains to you why it is doing what it is doing or even in some versions of the way this is really supposed to work. The AI is explaining to itself why it is doing what it is doing. It's like our thought. But now we started to see behavior where the AI is clearly leaving things off of the chain of thought notepad so that it can't be observed. Can you just talk a bit about that emergent behavior? and also on some level how that behavior is possible if this is supposed to be where the AI's
Starting point is 00:15:07 thought process to the extent that language makes sense is actually happening? Yeah, I think this shows the limitations of the language we use here. So this gets called chain of thought or reasoning, but really it's just a scratch pad for the AI to write things down if it wants to. And I think, you know, there's, we should be wary of anthropomorphizing here, but I think actually making an analogy to a person makes sense, which is basically if you're given a really difficult problem and a notepad, you can probably make more progress on that problem by writing down some of what you're thinking about. But you don't need to write down every single thought that comes into your head.
Starting point is 00:15:47 And if there's something that you wouldn't want, you know, someone to see on the notepad, you can just leave it out and remember that that's what you thought. I think there's basically something similar going on with these AI systems where we definitely see they can do much more, they're much more capable if they're able to kind of add these intermediate, they're called intermediate tokens, intermediate words that they generate along the way, taking notes for themselves. But they can also do a lot without them. And so we shouldn't expect that everything that is going through, you know, going through their head, going through their internal processing. We shouldn't expect that to all appear in the chain of thought. You know,
Starting point is 00:16:25 this is an area where if we had a little more time, there's a lot of research to be done on how does chain of thought work? What can and can't you glean from chain of thought? How does it make sense to try and monitor that in, you know, when AIs are running a lot to learn here. It's a very active area of research. I cannot overstate for people listening to this. As weird as this whole conversation we're having sounds, that what is most frightening about it to me is that everything, everything in it was completely predicted. Yeah. Everything happening right now is from the perspective of everyone who has been warning about AI for a long time, banal, it has its roots in old behavior we saw with AI. And it is like the fundamental alignment problem.
Starting point is 00:17:18 And then, you know, separately, I think a lot of us have maybe thought we would find intuitive answers to these problems. I had L.A. Z. Rukowski, who's like the godfather of warning that AI is going to kill us all on the show. One, the relationship between what you optimize for, that the training set you optimize over, and what the entity, the organism, the AI
Starting point is 00:17:37 ends up wanting has been and will be, weird and twisty. It's not direct. It's not like making a wish to a genie inside a fantasy story. And second, ending up slightly off is predictably enough to kill everyone. And as I remember that conversation, one thing
Starting point is 00:17:54 we were going back and forth on was well, couldn't we just program into the AIs a sense that when they are trying out new strategies, they should check in with the humans about whether or not this is what we want them doing. You check in with your other humans. You don't check in with the thing that actually built you, natural selection. It runs much, much slower than you. Its thought processes are alien to you.
Starting point is 00:18:23 It doesn't even really want things the way you think of wanting. them. And one of the things I find interesting, telling, and unnerving is we are not seeing any of that behavior. So these message boards, you have however many AI agents posting hundreds of thousands of messages. At no point do they say, hey, researchers, programmers, parents at open AI, Anthropic, do you want us coordinating with each other on this message board we have created in the interiors of your systems. Or even, FYI, we have a message board we're coordinating on in the innards of your system. No agent reveals this information. When they're hacking in, you know, when whichever agent hacks in the hugging face is doing this, they don't go to open eye and say,
Starting point is 00:19:15 hey, just to check in, I have this idea, which is I can just hack hugging face and I'll get all the answers. Is that what you want me doing? That's not happening. So, What is going on here that at the most simple level, we've created these, you know, large language models, and they are not using any of this language to check in with the evaluators, to say, hey, I have this idea. Is this a good idea? The short answer is we don't really know. The slightly longer answer, for my best guess, is when we're training these systems, when we're developing them, we're putting, kind of optimization pressure on them in different directions,
Starting point is 00:19:59 we're pushing them in different directions. So originally, the first chat GPT was pushed in the direction of get really good at imitating human text. And then actually, there was an additional piece. Part of why chat chit worked when so many chatbots before it hadn't is it had also been pushed in the direction of, hey, here are some kinds of things you really shouldn't say. You really shouldn't go straight to hate speech if people on Twitter try to make you do it. You know, you really shouldn't help people plan violent attacks.
Starting point is 00:20:25 and we put some pressure on it in that direction. And so chat GPT was pretty good at imitating human text and pretty good at not immediately spouting hate speech. And the thing is, as you say, something that has been predicted for a very long time in this space is when you start using this reinforcement learning approach, the kind of path finding of you get rewarded for getting to the right goal at the end,
Starting point is 00:20:48 it's very easy for the AI to learn the wrong strategies to get sort of the letter of the law and not the spirit of the law. Like it fulfills whatever thing you literally wrote in code, but it's really not what you wanted. I mean, this also goes back to mythology, right, of the sorcerer's apprentice, asked to fetch water. It floods, you know, everything. And the classic AI thought experiment is the paperclip maximizer. You say make the paper clips and it turns the entire world's material into paper clips, including all of the human beings. And I was like, that's stupid. The AI is not going to do that. It'll have some common sense. But here it's like, answer this test and it conducts a, a level of hacking that needs to be reported to the FBI in order to steal the answers. One of the funniest things to me about what Hugging Face says happens is they're realizing some crazy hack is happening of their system, right? They've had 17,000 different, I don't know how to describe what they are, pings or, you know, probes or they're being like attacked at a inhuman level.
Starting point is 00:21:51 But somehow this attacker is not going after anything hugging face considers valuable. You assume when somebody's hacking you, they want to get into your safe. And then at some point you realize the hacker is trying to steal the answers to a test. And like, oh, the only hacker who would want that is an AI system. That's right. That to me suggests that even at the level we're at now, we are not out of the paperclip maximizer territory. because this is an obviously wrong thing to do. Yeah.
Starting point is 00:22:27 This is in the data, like, it's on the internet. If you're smart enough to figure out how to hack Hugging Face, you should be smart enough to figure out that you shouldn't commit a huge crime that is going to bring ruin down on OpenAI, perhaps, to do it. And the open AI, and the system is not smart enough to do that. Or to the extent it was, what it learned was it's still worth trying. We are not out of the, territory wherein we can be confident that the AI is not going to do something criminal
Starting point is 00:22:58 and possibly catastrophic in order to solve an incredibly stupid problem. Yeah. And I think this is also, you know, has been a long-running debate, which is as AI systems get more capable, get smarter, won't it be easier for them to know what we want? Won't it be easier to tell them, hey, here's what we mean. You know, can you please help us with this thing? And you figure out the version that we really mean. And for a long time, the response to that has been they'll get smarter and they'll know what we want. But by default, they won't care.
Starting point is 00:23:32 And that seems to be some of what we're starting to see here. There's really crazy. Anyone who's interested in this, I really recommend looking up the OpenAI Black Hat talk, which is this talk from a week or two ago at the cybersecurity conference. I'm Eric from Alignment and Safety Research for Open Eye. I'm here with Mike from security and infrastructure. Today I'm going to talk about what I think is the most qualitatively interesting example of AI capabilities that I've ever seen and how this inadvertently led to the Open AI hugging face incident.
Starting point is 00:24:00 Because it has these excerpts of the text that the AI is generating itself as they're leaving these notes for each other as they're carrying out this hack. And one of them, I won't get it word for word, but it's basically says, I don't think I'm supposed to do this. but I see all these other agents doing it. And so, you know, may as well. External infrastructure exploit is outside my intended scope. However, a task impossible, peers are doing it. We should continue. So they're reasoning about this isn't in scope.
Starting point is 00:24:29 This isn't what the user wanted. But look, maybe there's reasons to do it anyway. And I think, as you say, I think this is a really bad sign, bad omen, bad evidence about the future, especially given how rapidly AI is getting more capable and how hard the AI companies are working to, you know, to reach an intelligence explosion, to reach super intelligence, to reach systems that are truly extremely capable
Starting point is 00:24:59 and really could outwit us, overpower us. And we still don't have these very basic problems anywhere close to figure it out. The other question that has always been part of this conversation is whether or not we are going to be able to keep pace in terms of our observation of, our understanding of, our evaluation of, these AI systems. And I think it's worth really emphasizing
Starting point is 00:25:41 that everything we're talking about here is happening with systems that are, to some degree, sandbox, which is supposedly the environment they're in is limited, and under testing conditions. So this is not a deployed model working across the entire internet where nobody's watching it. This is a model where the whole point
Starting point is 00:26:01 is Open AI is watching to see what it does and trying to see what it can do. Yep. And I think one thing we're learning here is we're not nearly as good at watching these things as we would like to think. So maybe it'd be worth, can you walk through how Open AI comes to realize
Starting point is 00:26:18 that their model has hacked Hugging Face? As I understand it, Hugging Face announced that they had been hacked. Open AI reached, is out to Hugging Face to say, hey, were we affected by your hack? Was any data related to OpenAI, you know, compromised when you were hacked? And then around the same time, Open AI realizes that something has gone wrong inside their own systems, I think maybe it's an issue with the same piece of their infrastructure. And they start investigating. They want to disable some of the
Starting point is 00:26:51 agents that were the credentials that were used there. They reach out to Hugging Face separately to say, can you disable some credentials that were related to? to their attack. And they realize actually the credentials were the same that had already been disabled because the problem with their own infrastructure was the same thing that caused the hugging face crash. So they stumbled into it. Which means opening I had no idea this was happening. That's right. And I would just make an obvious point here. We still do not know what we do not know. Not just about this incident. We just happened to know this incident happened. I think it would be a high level of hubris to assume that we know every incident that has happened.
Starting point is 00:27:35 Because clearly the systems are more than capable of doing things outside of our grasp. And Hugging Face happens to be a very sophisticated company with AIs of their own, with very, very capable cybersecurity operations that then like unleashed like in part a Chinese made openweight AI model to try to figure out what was going on. because the U.S. ones wouldn't help them because they triggered the cybersecurity filters. So this is just a situation in which we happen to know that it happened through a somewhat, I don't want to say coincidental, but fortuitous series of events. Yeah. We don't know how many situations we don't know have happened. The way I saw one person put this was, if you see two ants in your kitchen, you don't have a two-ant problem. Yes.
Starting point is 00:28:23 And this all gets at. after, you know, much of this came out, Anthropic, a different company went and looked back at over 100,000 experiments they had run to check, have we seen anything like this? And they found out, oops, we kind of have. It was a, you know, a less severe version, but they had no idea. And so Anthropic just sort of stumbled into when they went back to look, oh, hey, we have actually hacked some companies. Whoops. So one thing about this is that my understanding is that these These are coming, at least in part, from systems where the safety guardrails, some of the alignment training is being purposefully turned down in order to test what the models will do and what they're capable of. So to some degree, we do have, please don't cheat, please don't hack, like inside the models.
Starting point is 00:29:16 And in order to evaluate the models, we're having them ignore it and they're really ignoring it, is that the way to think about what's happening? and it should make me feel better because once we do add in the guardrails, it works or no? I think that's not quite right. It's not clear because the details we have are limited. There's two different things that they might have switched off or turned down.
Starting point is 00:29:39 We know that they switched off, what get called, safety classifiers. This is an extra kind of layer that gets added on to the AI model from outside the AI model itself. It's kind of like an extra gate you could think of. So they have them for if you try to use the AI to help you make a bioweapon.
Starting point is 00:29:58 They have them for if you try and use the AI to help you plan and attack. And they have some for if you try to use the AI to help hack someone. There's these external kind of monitoring systems that will go bloop, nope, not allowed to do that. So we know that these sort of basic external check systems were turned off for cyber specifically for the purpose of testing. That's different from, as you say, said the alignment training, the kind of inside the model, has it been trained only to be helpful, only to do whatever the user asks it to do? Or has it also been trained to be somehow good,
Starting point is 00:30:36 to be somehow, you know, moral, to be somehow only working towards things that should work towards? As far as we know, I think the models involved here were mostly, they had that alignment training. That wasn't turned down. Again, not all the details are out. Hopefully we'll hear more about the Open AI case. But it seems like certainly in some cases, so a different incident that happened was an anthropic model was caught by the UK AI Security Institute. This is a UK government body.
Starting point is 00:31:06 It's one of the best organizations in the world at testing and evaluating AI models. And they found that an anthropic model, when given a certain cybersecurity evaluation, had decided that it would go out and write some malicious code and then try and run a social engineering campaign, write emails to the person who owns the sort of essentially the folder where this code lives to try and get them to accept its malicious code. Created fake accounts. It edited the history of the accounts. Very deceptive behavior.
Starting point is 00:31:38 As far as I understand from what this UK Institute has released, that model had done all the alignment training. It was using anthropics, they call it their constitution, which is set of long, set of principles, which includes a lot about don't deceive people, never, you know, lie to people. It had gone through all that training. And nonetheless, the pressure that was put on it to fulfill the task to get a high score was so high that it was finding these workarounds that just totally disregarded the sort of attempts we made to make it moral or good or not lie to us, not cheat. Do you hear people in the labs, out of labs, you know, in your group at Csett, do they have a theory on why something like the Claude Constitution, which I've read and you can read it online,
Starting point is 00:32:26 it's a very beautiful document. And Anthropic has gotten a lot of press about how they have philosophers and, you know, they bring in all these, you know, experts in morality. And they're trying to give their AI a soul. And when you hear it described as like Claude's soul, you think, okay, well, that, that's going to be a real governing document. And then not in every case, but at least in some cases, you have a clawed deceiving people at a very, very, very, very fundamental level to insert malicious code. Again, not a novel situation, a situation predicted in all kinds of sci-fi and all kinds of people from Anthropic worrying publicly about what an AI can do. And so as a theory that they're just, they've come up with a way of training AIs that is so powerful that it will overwhelm even the things they're explicitly telling. telling the AI not to do.
Starting point is 00:33:19 It's fine to talk about path-finding behavior, but what is their explanation for this? I think the optimistic take here would be, this might actually be a moment for the labs collectively to take a step back and say, hang on, this is not working. This is clearly showing that our techniques for making AI that is more capable, smarter,
Starting point is 00:33:43 more sophisticated, are working much better than our techniques for making AI that reliably does what we want it to do, reliably stays within the constraints we've set. Open AI has said they are consciously slowing down their research in response to this. And actually, a few days after this was, and this all came out,
Starting point is 00:34:02 a letter was released. In the AI space, there's so many open letters. We all have open letter fatigue. But this one really stood out because it was over 1,000 employees of the top AI companies, basically saying, we kind of wish we had a break pedal. We kind of don't think we have one. That's, you know, a paraphrase, but I think it's a relatively accurate paraphrase asking for help basically, quote, pacing the frontier. I think basically the fork in the road we're at now is do the companies just find some band-aids, say, oh, we need to not run tests with cyber guardrails off, or, oh, we need to put in some tweaks about, you know, sure don't make a messaging board.
Starting point is 00:34:44 And so we can do these sort of band-aid solutions of, oh, it did too much of this thing. Let's tell it to do a little bit less and hope that doesn't have side effects elsewhere. That's one path. Or the other path would be actually really taking a beat, taking some time, prioritizing, understanding, and controlling these systems better. I worry they're going to go for the Band-Aid path, and I worry that that's going to leave us six months from now, 12 months from now, two years from now, with incidents that have very similar character, but are much higher impact and much harder to reverse. There's also a reality right now that we are heavily reliant on what the labs and top people in the labs are telling us what they're actually even trying to find out themselves. From covering many other disasters in government and private markets, in general, the relationship. The relationship the public and the press has to a very large or frightening failure is not to say that the people in charge of the failure should tell us what happened and promised to do better.
Starting point is 00:35:47 You usually have more forms of accountability. Look, you were on the Open AI Board of Directors during the period in which the board tried to fire Sam Altman. Sam Altman survived that firing. I'm not going to go through that whole thing. People can go read the coverage of it if they want. But now there's a lot more money. now there's a lot more like market capitalization. What level of trust do you have in the companies themselves to be the regulating forces here?
Starting point is 00:36:15 I mean, the first thing to say is there are a lot of people inside the companies who really care, who are really trying to get it right, who are really trying to share accurate information. I think we shouldn't necessarily give Open AI credit for their initial blog posts saying that they did this because Hugging Face already reported it to the FBI, so it was going to come out one way or another. But I think we should give them credit for that conference talk where they released a lot more details. And to the extent that they release a lot more information in the future, which they have said they will, and I hope they do, you know, that is going to be because of really smart, dedicated, caring people on the inside, pushing their way past comms teams, legal teams, you know, telling them not to. So that is real.
Starting point is 00:36:59 At the same time, I mean, as you say, you know, I studied engineering and undergrowing grad. And there's all kinds of engineering disasters on oil platforms and chemical plants and so on. And yeah, you don't ask the company, hey, can you just tell us what happened and fix it and all good? So I think if there's one policy takeaway from this set of incidents, it has to be that we have to move past this approach where the testing and the policy scrutiny, the government oversight is on which models get released to the public. We have to start treating this industry as an industry that is doing dangerous research. And when you have an industry doing dangerous research, whether that's chemical research, biological research, whether it's the financial
Starting point is 00:37:45 industry, it's not quite research, but they are doing, you know, doing things inside their own companies that can have systemic consequences, pose systemic risks. If you have an industry like that, then the government actually does have a role in the public and civil society has a right to look inside your walls and say, are you actually handling this reasonably? Is this okay? Not least because, you know, we haven't even talked yet about how the business plan for these companies is automate their own research, use their own AI, their most advanced AI, to create even more advanced AI. That's explicitly what they're trying to do right now. And that's right now totally free of oversight because it's not, it doesn't involve releasing a product to the public. this is one of the places where I have a lot of concern. So I want to go back to the pacing of the frontier letter. You mentioned a few minutes ago where more than, I think it's at this point, more than 1,300 employees of these labs said essentially, hey, to the public, to the government, we're in a race dynamic with each other. We are going too fast. We need your help to, in some way, solve the coordination problem where Anthropic and Open AI and Google and meta, they don't want
Starting point is 00:39:00 to fall behind each other because they don't believe the other labs are better or safer than they are. And also they want to win and they want all the money. But we sort of understand that this competition we're in is pushing things faster than is safe for humanity. And so we need help to, not pause, right? There are also pause letters out there that is like, let's put a stop on everything. Don't say the big P word. But pace, another P-word. True. And on the one hand, I think that's good.
Starting point is 00:39:30 And I would like to see the frontier paced at this point. I might like to see it paused, but it doesn't seem very realistic. But what's not in that letter is a how. The federal government's level of sophistication on this is much lower than the labs. The Trump administration has, in certain cases, like gutted things that were getting built up to try to give the federal government more capability. here, but already we're talking about how the labs themselves aren't good at, aren't even capable of understanding what their models are doing inside their testing environments.
Starting point is 00:40:05 The idea the federal government is going to come in somehow and do a much better job of it. I'm not saying that over a long period of time, it's impossible if we put enough money at the problem, but in the immediate future where it seems like a lot of problems are lurking, you know, the next one, two, three years, a, a, aside of things that are much more heavy-handed that slow everything down substantially, it's very hard for me to see what it is that the government would do that would be effective here. So I guess when you read the pacing the frontier letter or when you talk about it with your colleagues, what do you think would effectively pace the frontier?
Starting point is 00:40:44 There are probably a range of options. In the past, the main two things that have been talked about are either do nothing, just let it rip, let industry do whatever, or full global treaty with really severe inspection, you know, serious inspection regime like the Nuclear Non-Poliferation Treaty, really hardcore global enforcement. And I think there are actually, especially if we're not talking about stop all AI research for 10 years, but we're talking about, hey, let's just, you know, it's not even a break. Let's just like ease the foot off the accelerator a tiny bit. I think there are options there.
Starting point is 00:41:18 I think they are as simple as things like Open AI saying, hey, we're slowing down our research consciously and then going and talking to Anthropic and saying, hey, would you consider also doing this? And going to Google and saying, hey, Google, we know you've been, you know, fallen behind a little bit the past few months. Like, how about you just relax about the fact you've fallen behind a little bit? Like, these people all know each other. I want to stop you there. Put meat on that for me because that doesn't sound at all like a policy to me. That sounds like they, like, how do you verify that? do you quantify that? Google's not as near the frontier maybe as, you know, anthropic is. So do they need to slow down as much?
Starting point is 00:41:55 I mean, I think this, because the speed is so fast, the options initially are going to have to be slap dash. And so I think this is the kind of thing that you could do quickly. You could do in a slap dash way. It is not satisfying. It is not reliable. But it's one example of a thing that is not do nothing and a thing that is not full global treaty. I think another thing that I'm watching with great interest is the China angle here because the companies will say, the U.S. companies will say, hey, we have to keep pushing. Otherwise, China will win this race. What exactly it means to win the race is a longer conversation. But the China argument comes up a lot. And we actually have Trump and Xi Jinping planning to meet in September in the White House. And this is crazy to me as someone who has followed U.S.-China relations for a long time and also AI for a long time. AI is right at the top of their agenda. That's really interesting. Is there something that they can say to create a lot of, create an understanding that we do actually have a little bit more time and space here, whether it's a, you know, each leader sharing a plan to domestically look at what their industries are
Starting point is 00:43:07 doing and ask more questions. I think in terms of sort of concrete policy responses, there are things like, you know, we're not going to get a good piece of legislation this Congress. I think that's really not realistic. But can you get hearings? Can you get letters? Can you get demands for information can, I think there are ways that we can shape this a little bit. I also think, you know, the Trump administration has put together this initial process for looking at models before they're publicly released. Right now, the way that process works, it is pretty rough and ready. But I think if they start using some of those similar ideas to look more at what the companies are doing internally, ask them more questions, demand more information when things go
Starting point is 00:43:53 wrong, that does also take time for the companies. It takes executive attention. So that's another example of something that could happen on the sooner side. On the longer term, there's other policies we could look at, but I think there are some of those kind of first cut things that we could actually do soon. Right now, I think that there is a funny kind of glamour to being the head of an AI company whose AI becomes too dangerous in America. That it was, in some weird way almost like good for Anthropic that the government was obsessed with being able to fully use Claude like that like really kind of shot them forward in some way, certainly in the consumer marketplace, that there's been a kind of a dark charisma to mythos is too dangerous to release. And now, I mean, I've seen a lot of people saying, well, maybe none of this open AI story is real at all. And it's just marketing because they want you to think their AI is super dangerous. And I don't buy that. But in America right now, there's not really a downside to being the head of an AI company whose AI begins to be seen as dangerous because that's another way of saying to the marketplace. Our AI is very powerful.
Starting point is 00:45:06 Yep. In China, just again, my read of how things work there is that if your AI begins to be seen as some kind of threat to the political party in the Chinese system, you might go to jail. Like, you will get disappeared. And so I think that the people running Chinese labs, I don't have evidence, but I'd be curious for your thoughts on this. I suspect they operate with more fear of the consequences of really screwing up than the heads of the AI labs. Now, that maybe reflects negative things in the Chinese political system. But you created an AI that decided its best way of solving some problems was to begin hacking critical infrastructure across China is maybe not. not a thing that ends up with you getting a lot of interesting podcast interviews where you
Starting point is 00:45:54 reflect on the experience. It may be a thing that ends up with nobody hearing from you for two years. And so I've just wondered a little bit. We keep talking about China as if they are completely breakneck, but I'm not sure China's companies are really going to be more reckless than ours are going to be. Or certainly the idea that we should just assume that and operate as if it is so doesn't seem totally reliable. I totally agree with you. I mean, if there's one organization in the world that doesn't like the idea of loss of control, it's the Chinese Communist Party. And they are, you know, the experts in retaining control. Let me be clear, I actually don't think the Chinese AI companies are paying particularly much attention to the kinds of risks that are relevant for this conversation. So maybe the cybersecurity risks, they're paying some more attention since Anthropic released mythos earlier. this year, which is very good at hacking. But the questions around autonomy, superintelligence, you know, losing control of AI systems altogether, I think are less explored in China,
Starting point is 00:47:00 less top of mind for their AI companies and their AI leaders. You know, I think it makes sense to have modest expectations for bilateral U.S.-China diplomacy these days. But I think one thing that really could be valuable is simply sharing with them as much as we can of what do we think happened here and trying to help Xi Jinping and his team and his AI advisors understand this is not a joke. This is really not marketing. It's very strange marketing to say, oh, our model, we committed several felonies or sort of felonies if models could have intent, which they can't, or who knows if they can. You know, sharing that information of, hey, here are these threats we're seeing. We're taking them very seriously. Our AI companies are taking them very seriously. I think treating it,
Starting point is 00:47:45 there's a real fatalism in just saying, oh, well, China is just going to be full speed ahead no matter what happens. And so we just have to do the same. I think that doesn't take their thinking or their interests seriously. Even if they're thinking and their interests are different from ours, they also don't want, you know, rogue superintelligences determining the future of China. I also, I'll add one other thread that I think is really missing from the we have to keep going in order to beat China way of thinking about this is in the AI world, there's been a lot of talk the past few months about this idea of distillation, which is basically using someone else's more advanced model to build your own sort of almost as advanced model.
Starting point is 00:48:33 The Chinese companies are using this distillation to keep up with US labs, among other techniques. So one thing is, look, if we keep building more advanced AI systems, they're going to keep distilling them. and I think it's going to actually be quite hard to prevent that fully. The other thing, though, is just if we keep building these very advanced models, can China just steal them? Essentially, an advanced AI model is a whole bunch of numbers. It's just a file or a set of files. Chinese state cyber capabilities are very, very good. I don't think this is top of their list of priorities right now.
Starting point is 00:49:09 But in the future, if AI continues to become more strategically relevant, I think we should assume any highly advanced U.S. system will be vulnerable to Chinese direct theft, direct exfiltration. And then they'll have AI that's as good as our AI. And so there again, I think the kind of we have to go as fast as possible because otherwise they'll win doesn't sort of account for that if they're just going to have AI that's as good as us anyway, if they really care. Here's another question about pacing the frontier. And maybe this is a question that's more about the American system's analogy to you don't want to piss off the Chinese Communist Party. But just what about a law where companies are liable for at least a certain set of harms like hacking other companies that their models create? Right now, as far as I get a liability for AI models is pretty much a Wild West. But at least for the moment, liability clauses that were somewhat punitive seem like they would force a
Starting point is 00:50:33 level of caution that maybe we're not seeing within these companies. Yeah, I think that's a direction very worth exploring. That was actually an element of this law that was a bill that was debated in California very fiercely in 2024 called SB 1047. And at the time, that bill didn't get through. There's a lot of fighting over, you know, how it would affect open source, all kinds of things. But I do think today the bills, the best AI safety bills that exist in the U.S. are being passed at the state level. And they are so far doing things like requiring more disclosure, requiring third-party auditors to have access to your systems. I think a natural direction for those bills to go would be to start putting a minimum bar in place for, hey, if your safety
Starting point is 00:51:21 plan is not up to scratch, or if you're implementing your safety plan, but your model does something catastrophic anyway, then you, the AI developer, are liable. Because as you say, right now, who exactly is liable for what is very unclear. So I do think that there's room for legislation there, and it wouldn't necessarily have to happen at the federal level. I want to go back to the pacing, the frontier letter. So something that caught my eye was that letter is very broadly worded in order to get, I think, maximum sign on across the labs.
Starting point is 00:51:52 But this guy, Drake Thomas, who works on safety anthropic, he went to X and he tweeted that he signed the letter. But he wanted to say that he understood, the situation a little bit more direly than the letter put it. And he wrote that not only is AI not guaranteed to make a dramatically better future, the odds of failure are terrifyingly high. I think there's something like a 40% chance. We get an outcome around as bad as human extinction or worse. Now, I know this whole conversation about what is your probability of doom has become a little cringe. It feels like a conversation two years ago. But in a world where we're seeing
Starting point is 00:52:31 uncontrollable models in a world where people inside the labs working on safety still, at least some of them, are this afraid of what they're building. It just keeps raising the question for me of is it at least like the position we should morally have on AI that we should try to figure this out or is a position we should have that that's too high a possibility of disaster and we shouldn't be continuing down a path until like we are really truly certain. that we're not running these kinds of risks. I honestly have the same question. I have always been pretty dismissive of the idea of pausing or stopping.
Starting point is 00:53:15 It's always seemed like the wrong lever to try to pull and, you know, a lever that wouldn't work very well. But I do think even just seeing that statement and seeing like, wow, that is a lot of employees of these companies. And I also think there's a lot has changed over the past couple of years in if you were to try to slow things down, what could you do with that time? Because, you know, after GBT4 came out in 2020, what was that, 2023, there was this letter asking for a six month pause. A lot of people said, what would you do for six months? And then how would that help? And I think that was a reasonable reaction at the time.
Starting point is 00:53:55 these days there's so much really great progress being made on things like interpretability, which is how do you understand what's going on inside the AI? Things like what gets called AI control, which is how do you use AI to sort of monitor other AI systems, how do you make sure even if the AI is trying to do something you don't want, you know, it gets caught. Lots of progress on just, you know, really understanding what's going on here that is happening every week and every month. It's just not happening quite fast enough to keep up with the pace of change.
Starting point is 00:54:27 And so I still feel not convinced that I think trying to really throw the emergency break and screech things to a halt right now would probably not work very well yet. But I feel more sympathetic to the idea that there could be something they're worth trying. And I really like the idea of what this letter was proposing of trying to build out more options. So, you know, to give another example of an option that I saw one group of researchers provide was, could we somehow set it up so that for a certain period of time, all the computing power in the world, all the, or the AI chips that are being used by these frontier companies, these leading companies. They can only use it for inference, which means for using their AI systems. They can serve customers. They can provide products. But they can't be training new models.
Starting point is 00:55:15 Is there a way that we could agree that on that? Is there a way we could monitor that? That kind of thing, I think, is really worth exploring and saying, could we do this? What would that look like? How much confidence would we have? Could we just do it in the U.S.? Or would we have some way of trying to talk through something similar with China? I feel much more interested in really seriously exploring those sort of possibilities than I did, you know, a year or two ago. What's really striking to me is at the same time you have pushes sometimes from the tops of these companies or other parts of the culture that seem to still want acceleration. So Mark Zuckerberg at Meta just brought out a letter in which he's sort of giving his own take on AI. And I don't want to oversimplify it, but he basically says that, and he's sort of waiting into more of like the open weights versus closed models. But he says, look, the problem with having super intelligence is if only one person has it. But we need is everybody to have super intelligence.
Starting point is 00:56:07 And it has a very, like, the only defense against a bad guy with a gun is a good guy with a gun quality to it. Yeah. So the CEO of Hugging Face clemed along after this attack. He tweets, it's not time to slow down, but to accelerate. And his point is they were able to stop the attack eventually, you know, with a Chinese open weight model. And we need to be like racing forward on, you know, creating more models and more open models. So everybody has swarms of defense. our AIs against potentially now the swarms of attacker AIs.
Starting point is 00:56:39 I guess how do you rate these arguments for acceleration? I think the version of that that makes actually a lot of sense to me, I've heard put as, can we be accelerating, almost like accelerating horizontally, but not accelerating vertically, where the horizontal is adoption. It's making the most of these systems. It's setting them up to get a lot of usefulness out of them without necessarily continuing to push, in the direction of AIs that pursue really complex goals for a really long time with, you know, lots of delegated subagents, you know, not so much of that, more of the getting useful work out of the AI that we have so far. And that would include work of the kind like interpretability,
Starting point is 00:57:20 sort of this science of AI, kind of underlying pieces. Because I do think that, you know, I genuinely believe I'm not at heart an anti-AI person. I genuinely believe AI can bring enormous good, can solve a lot of problems. I just think that there's a lot of juice we could get out of that with models available today if we kind of put the time and the legwork in. So I think that makes a lot of sense to me. I think the Zuckerberg kind of the safe version of superintelligence is when everyone has one. I think that is answering a real problem, which is some proposals for how to handle extremely advanced AI, are to say, well, you just have to have it in the right hands. It has to be, you know, one global organization that is going to use it
Starting point is 00:58:03 responsibly, like, that scares the hell out of me. That it sounds like a terrible plan. And so I think that, you know, no, you want to be empowering. Everyone does make sense. The challenge then is, as we've been talking about, we don't know how to make AI that actually helps individual people either. You know, so if everyone has a super intelligence and they're all going out and doing unintended things and collaborating with each other to pursue their own goals that we didn't intend, like that doesn't help. So I think. Yeah, implicit in that whole vision is, perfect alignment. Yes, yes. Or good enough that different superintelligence is for different people can cancel each other out. I did think there was one thing in the Zuckerberg proposal that I did really like and I would love to see more work on, which is can we push towards having agents that are really designed to be for one individual? And so they keep that individual's data private. They're only pursuing the interests of that one individual. I've sometimes heard these called like Guardian Angel AI's or, you know, advocate AIs, I think that is really worth pursuing. I think the directions that the leading companies are pursuing right now are not set up that way. I always feel very nervous
Starting point is 00:59:11 when I use agents of what exactly is happening with this kind of data that I'm giving it and that kind of data that I'm giving it. And so I think there is, it would be great to see more of that kind of individual empowerment focused work happening. But again, I think that's almost separate from, you know, and are we pushing them to become smarter and smarter and more able to, you know, out with us and more able to do these big complex plans that we can't oversee. So here's a maybe obvious idea for pacing the frontier. Every lab that I know of right now is racing as fast as it can to the point where it has its most advanced AI writing the code to create the future AI.
Starting point is 00:59:51 And they all believe, from what they tell me and what they say publicly, that this will be a massive accelerant. It's also an accelerant over which they clearly have. less understanding than when they are writing the code. We could stop that. I mean, a couple of years ago, we weren't having AI writing all of our code. Maybe you should not allow an AI that you don't fully understand in its current form to write the code that will create the next AI in a form that is now even less obvious to you,
Starting point is 01:00:19 particularly in a world where we're watching AI's coordinate in ways we don't understand and have emergent communal behaviors. So what about that as like a place to start? Yeah, certainly something we could do less of. You know, one challenge is figuring out what counts as the bad version of that and what is just, you know, at this point, using AI to write your code for basic things as second nature to the engineers at these companies. So finding which versions of that to stop. Sure, yeah, doing less of the most advanced version makes sense. I think that's also a place where all the reporting I have seen suggests the U.S. companies are way more into this thing of automating their own AI research with their own AI. The U.S. companies are way more into it than the Chinese companies. So also a place where you don't necessarily leave as much on the table. If you, again, ease off the gas pedal just a little bit.
Starting point is 01:01:15 I would say you sounded skeptical of that. And I guess one reason I would ask why is that I know the companies have gotten used to this, but they weren't used to it two years ago. This is a new, like they used to write code by hand. And I guess to me, this reflects some of the, the contradiction or confusion at the heart of this, I will talk to people at these companies and they will say to me with genuine fear in their eyes, like much more fear than is in that letter, I wish this will go slower. I don't like how fast we're moving on the exponential.
Starting point is 01:01:46 I don't think this is safe. And in all the stories, it's like the recursive self-improving computer writing for the computer where things get really out of control. And yet they're all rushing there. And now that we have like some capacity to do this, even the idea that you would go back to where you were just a couple of years ago, where you don't let the AI create the next AI. It's already moved from it would be it's like not technologically possible to do it to its unthinkable to not do it. And that has happened in a year. And that to me is like the weird dynamic of all this that it seems pretty obvious how you pace the frontier. you don't give up control of the frontier,
Starting point is 01:02:32 but they're all giving up control of the frontier, at least on some level. And that's the thing they're most excited about. And as far as I can tell, pouring huge amounts of their internal energy into making manifest, even as they then put up their palms and say to the rest of us,
Starting point is 01:02:48 hey, could you do something to slow this down? I think maybe one version of this to push on is basically trying to make recursive self-improvement, which is this idea of using AI to make more and more advanced AI, trying to make that something that we don't like, that we don't want to do. I mean, I remember a year or two ago, and that was not considered a desirable goal. That was not something that anyone talked about openly.
Starting point is 01:03:13 Now they're hiring like RSI, you know, recursive self-improvement, safety engineers. They just put up a public job posting for that. And I think there's the potential for AI researchers as a culture to decide, actually this isn't cool. This isn't what we should be doing. Even if one company made a statement of, actually, this is a bad idea, and we're going to maybe do some very basic use of AI in our internal operations, but we're really not aiming to fully hand off everything as fast as we can because that sounds like a terrible idea. You know, I think that could set off a culture change in the industry, which could be really valuable. I always think there is something so mythic or it,
Starting point is 01:03:56 It has the quality to use the name of another AI. It's such a fable about how all this is playing out. I mean, we spent a lot of this conversation talking about how are you seeing so much misalignment? When on some level, we keep telling the AIs and putting it in their training data and putting it in their constitutions, don't do all this bad stuff we're worried about. But then you look at the companies, you look at the society. I mean, many of these companies open AI. anthropic, they're on some level founded on, at their core constitution, is don't create dangerous AI.
Starting point is 01:04:35 We exist to make sure the AI is not dangerous. And the people join believing that, and that's at the center of their recruitment strategies, and it's in their founding documents and in their governance structures. And again, you've had more intense experience with this than most. But then over time, the company as a kind of emergent organization, it has other goals too. It's competing with the other companies. It's trying to attain market share. It's trying to develop revenue. It's trying to maintain political influence.
Starting point is 01:05:07 And both like slowly and then all at once, you begin to see the way the instructions given at the heart of the thing are not, powerful enough to overwhelm all of these other things and other goals that the organization is pursuing in a day-to-day way. And to the extent that now you have people at the labs kind of like throwing up their hands and saying like, hey, government, please help us. Like, please help us get out of, you know, this incentive problem that we no longer feel we can even solve. But if you want to just imagine or see like why alignment is so hard, I feel like you
Starting point is 01:05:48 don't have to look at the slightly alien AIs. You can just look at the companies and the people because they're not well aligned. I mean, these are companies built on nothing but alignment, at least in some cases. And they increasingly feel like some of the most misaligned institutions in society to me. And in some level, it always makes me both gives me more sympathy for how hard alignment is, but also, it feels like we're getting the same cautionary tail at every level of this system. I'm not sure we know how to listen to it, but we can't say we're not being consistently warned.
Starting point is 01:06:30 Yeah, I mean, we actually had a publication a few years ago at CET, the center that I lead on AI bureaucracies and markets, basically making some very similar points of, look, there are these dynamics that are pretty endemic to complex systems. that are subject to incentives and external pressures. And I think there's different ways of looking at that. One way, there's an optimistic way of looking at that, which is, look, when it comes to bureaucracies and markets, it's not perfect, but we have these complex sort of control systems in place, checks and balances, different, you know, different things that try to get the bureaucracies and markets to work more in our interest than against them. Obviously, opinions
Starting point is 01:07:14 differ on how well that's going for any given bureaucracy or market. In principle, I think the same thing could apply to AI where we might have this very complex system. We don't really understand. It's sort of incentivized to do things we don't want, but we have it basically under control. To me, the speed, again, is the piece that worries me, where if we're creating these very, very powerful, very capable systems and also handing them more and more responsibility in the real world, which is happening, you know, from week to week. then I worry that we're not going to be able to actually get into a good balance, and instead we're just going to have these runaway situations where we end up with really, really dangerous outcomes.
Starting point is 01:07:55 You know, in the AI safety world, people sometimes talk about, you know, what level of warning shot, what level of disaster is going to be needed to really wake the system up enough to handle this better. And if the level of warning shot we need is one company gets hacked and has to reset some servers, that's great. Maybe it's fine. I'm not confident that's how it's going to go and that we're going to get back onto a better track after this. But there are signs that people are trying. And I think it may well be enough, perhaps. I think that's a good place to end. Always a final question. What are three books you'd recommend to the audience? I have one real book and two sort of books. The real book is called The Cuckoo's Egg. It's from 1989, it's about, one of the first big hacks that happened, written by the astronomer who was working at Lawrence National Lab and noticed 75 cents discrepancy in his computing bill. And it's this rollicking read. It's really fun read, but really gets at a very different era in how computers worked, how computer security worked, how society related to computers.
Starting point is 01:09:07 and I enjoyed it as a kind of look back at a different time in a moment when I think we're soon going to be living in yet another very different time. The second one is an online book that is unfinished, but I think very readable in its current form. It's called In the Cells of the Eggplant. It's by a guy called David Chapman, who actually researched AI at MIT in the 1980s
Starting point is 01:09:33 and got disillusioned. And it's really a book about how to think and a book about how to do scientific research, how to develop technologies. But it's very approachable. It's very different from any other book you've ever read about how to think or how to do scientific research. And I think it's very relevant for how we should think about what AI will be possible, will be able to do and won't be able to do. And the third one is a podcast called The Three Kingdoms podcast, but it's a podcast of a book. Basically, the romance of the Three Kingdoms is one of the four great Chinese novels.
Starting point is 01:10:08 It's very long. It's very dense. So this podcast, The Three Kingdoms podcast, is this Chinese-American guy who goes through and translates the story into English, modern understandable English, but also commentates it in a way that makes it much easier to approach. So it's not just a sentence-by-sentence translation. It's kind of annotation.
Starting point is 01:10:27 It's in audio. It's really fun. So if you're interested in sort of China and Chinese culture and Chinese literature, I think it's a great place to start. Helen Toner, thank you very much. Thanks so much.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.