Offline with Jon Favreau - Have You Tried Unplugging the AI?

Episode Date: September 19, 2026

Is AI sending us spiraling towards an apocalypse, or upwards towards utopia? How can such a new technology inspire such diametrically opposing viewpoints? To understand the dissonance, Jon sits down... with Helen Toner, the Executive Director of Georgetown’s Center for Security and Emerging Technology–and one of the OpenAI board members who voted to oust Sam Altman in 2023. She shares her firsthand experience of what trying to oversee one of these companies actually involves, insights on where the technology is going, and strategies for what we can do to avoid the scariest outcomes.Hate listening to ads? Become a Friends of the Pod subscriber for ad-free episodes of Pod Save America, Pod Save the World, Lovett or Leave It, Runaway Country, Offline with Jon Favreau, and more—plus exclusive content, including bonus episodes of Pod Save America. Subscribe now at crooked.com/friends, on Apple Podcasts, or through the Pod Save America YouTube channel.You can request a transcript by emailing transcripts@crooked.com. Include the podcast name, episode title, and air date. Please allow 48 hours for delivery.

Transcript
Discussion (0)
Starting point is 00:00:00 Offline is brought to you by Thrive Market. Fall means back to routines, back to meal prepping, and back to school. If you have kids, the old familiar question of what do I pack for their lunch has come creeping back into your life. That's where Thrive Market comes in. Thrive Market has some really great options for kids. We have some of these in our house. They have Yum Earth fruit snacks, chumps, chomplings, mini beef or turkey sticks, simple mills, mini-cookie packs. My kids like all three of those things.
Starting point is 00:00:28 Thrive Market restricts over a thousand harmful ingredients. Whole Foods only restricts 300. Take that Whole Foods. For a label reading parent, that alone is worth the price of membership. Speaking of the Thrive Membership, it gives you exclusive low pricing, weekly sales, free gifts, and free delivery on qualifying orders, all for just $5 a month. In 2006, USA Today named Thrive Market, one of the most trusted brands by parents. Thrive Market is the first online-only grocer to accept SNAPEBT Online, and every membership sponsors one, for a low-income family.
Starting point is 00:00:59 At $5 a month, less than your daily matcha, membership pays for itself with exclusive member pricing, weekly sales, free gifts, price matching, and free delivery on qualifying orders. When you shop with Thrive Market this fall and make a donation at checkout, you'll make a direct impact in supporting kids at risk in the U.S.
Starting point is 00:01:16 through Save the Children's Learn Without Limits Initiative, where Thrive Market will also help provide healthy snacks and meals for kids in high-need rural communities. Join Thrive Market today during their 25% off healthy reset sale now through September 19th, plus get $30 off your first two orders when you use my link, thrivemarket.com slash offline. These offers stacks, so sign up before the sale ends at thrivemarket.com slash offline. Have you found yourself confessing personal things to chat GPT, Gemini, or Claude?
Starting point is 00:01:49 You're not alone. Something huge is underway that we haven't really found the words for yet. a new voice in our ear, guiding what we think and feel and do. But should we trust it? I just thought, what has this thing done to him? Black Box, the Chatbots, a new podcast series from The Guardian Investigates. Listen now wherever you get your podcasts. Hi, friends and fellow students.
Starting point is 00:02:16 Americans living abroad can legally vote in the U.S. election. Register or get tailored support at vote from abroad.org. Look, state deadlines are super close, so don't waste. another second. Go to votefromabroad.org now and click request my ballot, asked to receive it by email. All you got to do is fill it out and send it back in. Don't miss this chance to shape the future. Don't delay. Votefromabroad.org. They were actually getting rewarded for communicating with each other secretly without opening and knowing. When they realized that they weren't able to do their task, sort of the legitimate way, because they had cheated. They were like, okay, well, how do we fix the fact that
Starting point is 00:02:52 we cheated? We better cheat even better. They don't need to sleep. They don't need to eat. They don't get bored. They don't have girlfriends or boyfriends. They can just go and go and go and go. Is it going to love humanity or is it not? And if it doesn't, and if it thinks of us more the way that we think about, you know, ants or chickens, then we're in trouble. I think our product might kill everyone is very much not marketing. One of the most common responses I get from people who are like sort of skeptical about all this. They're just like, can't you just unplug it? Yeah, unfortunately, probably not. I told you before last week's episode with Max Fisher, which was about the potential dangers posed by artificial intelligence,
Starting point is 00:03:39 then we're opening up this show to a broader range of conversations that go beyond technology in the internet, which is why this week we're talking about the potential dangers posed by artificial intelligence. Sorry, but I promise it's a conversation that's worth your time. Also, I've honestly come to believe that how we choose to handle the global race towards superintelligence in the next few years will have major implications. for everything else we talk about. Politics, culture, our jobs, our kids' education, our relationships. This is a technology that will probably affect every aspect of our lives. And that potentially includes whether we still get to live them.
Starting point is 00:04:19 I am not a doomer. I don't like being alarmist. I don't think it's the most effective way to persuade people. And I don't find it to be a healthy headspace for me personally. To be completely honest, I also wasn't that interested in AI until a few years ago. It was never really clear to me how the technology worked, what it could do, why we needed it, and why people were starting to freak out about it. I had also become quite skeptical of the tech lords whose last gift to mankind turned out to be a series of algorithms
Starting point is 00:04:46 that have made us all crazy. Not a great track record from Silicon Valley's geniuses and the ones who think they're geniuses because they're rich. People who are now having a tough time selling the world on the promise of their latest technology when we can barely look up from our screens for more than 10 minutes, but here we are. Well over a billion people worldwide are now using AI chatbots. Global spending on AI is expected to reach $2.7 trillion this year, almost as much as every country on Earth combined spent on its military last year. The seven tech giants now make up about a third of the S&P's total value,
Starting point is 00:05:26 and several of them have been warning us for some time that the products they're building might kill us. If we build in the wrong way, the probability of something bad happening is very high. I think there's a 25% chance that things go really, really badly. And the bad case, and I think this is important to say is like lights out for all of us. Mark my words, AI is far more dangerous than nukes. My worst fears are that we cause significant, we, the field, the technology, the industry cause significant harm to the world.
Starting point is 00:05:56 It has the potential of civilizational destruction. They are not alone in their thinking. Their employees have voiced similar warnings, thousands of employees who still work there, many others who decided to quit, researchers, engineers, scientists, Nobel Prize winners, people with a direct financial stake in the outcome, and people with none whatsoever. Of course, there are other tech giants and tech employees and really smart people who just disagree on the level of risk posed by artificial intelligence. Even the tech giants sounding the alarm seem to believe that the benefits of AI outweigh the risks,
Starting point is 00:06:30 and very much intend to keep pushing forward. But if trillions of dollars are being spent on companies who are telling us their product might kill us, regardless of whether we buy it, I think we should get a say in how they go about building it and what they're doing to make it safe. I'm not alone in my thinking. The latest polling shows that about two out of three Americans
Starting point is 00:06:52 believe that AI development is advancing too quickly and that there's a real risk it could destroy humanity. About three and four Americans would prefer to prioritize AI safety over innovation, with very little difference between the views of Democratic and Republican voters. And when asked who should have the most say in setting the rules for how powerful AI systems can be built and used, over half of all Americans say the government versus only 9 percent, nine, who said the company's building the AI. Again, very little difference between the views of Democrats and Republicans on the United. that question. Most of us want to say in this debate, but this is what we're hearing from the
Starting point is 00:07:35 president, the vice president, and some of their top AI advisors and supporters. I'm telling you it's all a hoax. The data centers are great. This appears to be an op. It feels a little bit to me like a bit of a Trojan horse. This is the same Dumer Histrionics that we've been hearing from this crowd for a long time. Pearl clutching, disingenuous AI doom narrative. Or are they going to going through some psychosis. If you're going to create Frankenstein, don't come to the government and say, we need regulation. These people hold views on AI safety that aren't shared by the vast majority of Americans or most experts who actually study and work on AI.
Starting point is 00:08:13 But they are some of the most powerful people on the planet. And again, even some of the tech giants who are sounding the alarm about AI safety, including the president of Open AI, have put a lot of money behind a super PAC that's determined to defeat candidates who support AI safety regulations. To the point where Democratic consultants are now reportedly advising candidates to avoid talking about AI regulations so that these tech groups don't spend a ton of money to defeat them this November. They don't want us to have a say. So they're just going to keep using their money and their megaphones to dismiss people's entirely legitimate fears and concerns as a hoax or a sciop, that the people speaking out are motivated by a desire for wealth or fame or dumerism that comes from psychosis or stupidity.
Starting point is 00:08:57 because obviously no one's as smart as they are. I certainly don't pretend to be an expert on AI, but I'm doing everything I can to learn as much as I can from the widest range of voices, not because I'm trying to calculate the exact chances that AI may cause an extinction event in the next 10 years, but because it's quite obvious that at the very least,
Starting point is 00:09:17 these companies are developing an incredibly powerful technology that could cause enormous harm, even if we don't yet know the precise magnitude or likelihood of that harm. And one of the most basic functions of any government is to protect us from the potential dangers posed by otherwise life-changing products, everything from cars and planes to medicine and food. That's been especially true with the development of technologies that could cause destruction on a massive scale. The Manhattan Project wasn't a plucky startup racing towards an IPO, and we didn't just leave Oppenheimer alone in New Mexico to build his bomb based on a promise that, of course, he'd be extra careful. with all that radioactive material.
Starting point is 00:09:58 That would have been madness. And that's what total resistance to any kind of AI safety regulation today is. Madness. People in power need to understand that. Politicians who agree need to say that and not be afraid of what some super PAC might do. And the majority of us
Starting point is 00:10:16 who want these companies to slow down and want our government to hold them accountable before they do something dangerous, we need to speak up too. My guest today is held and Toner. She's the executive director of Georgetown's Center for Security and Emerging Technology. She studied AI policy and China's AI industry. And she served on the board of OpenAI. In fact, she was one of the board members who voted to remove Sam Altman in 2023, and then she left once he
Starting point is 00:10:44 returned. Which means she brings firsthand experience of what trying to oversee one of these companies actually involves where the technology is going and what we can do to avoid the scariest outcomes. If you've been avoiding this subject because it's frightening or because you can't understand what the fuck anyone is talking about, I want this conversation to be for you. You don't need to have followed every headline, but my hope is that you leave with a clearer understanding of what's happening and a better sense of the questions to ask the next time somebody tells you we're either headed towards a utopia or the apocalypse. Here's Helen. Helen Toner, welcome to offline. Great to be here.
Starting point is 00:11:26 Thank you so much for doing this. I know it must be a busy couple weeks for you. Suddenly everyone wants to know about what I've been a nerd about for years. So it's a good time, I guess. Good time, bad time. Well, I wanted to start by asking how you became a nerd that has been interested in this and AI for all these years. Like, how did you first get interested in AI? And what made you decide this was something you wanted to spend your life working on?
Starting point is 00:11:51 Yeah, it was a bit of a right place, right time. kind of situation. I would say, so in undergrad, I've always liked both kind of the more tech, you know, STEM, quantitative side and also the more kind of politics, global affairs side of things. So in undergrad, I did chemical engineering and also some Arabic, some international relations kind of stuff. Had some friends who were in kind of 2013, 2014, were really convinced that AI was going to be a huge deal and basically gave me the bug. At first, I was like, that sounds like sci-fi silliness. I'm not interested, but they're very persistent. And that was right around when kind of deep learning was first starting to take off. So you may remember that was when, like,
Starting point is 00:12:33 Facebook started to be able to tag, auto-tag your photos because, you know, image recognition was working or like speech recognition on your laptop started to work a little bit. That was deep learning, which kind of kicked off in 2011, 2012. So at first I was skeptical. Then I started to look into it, and it seemed like, yeah, I was actually starting to get kind of good even back then. I thought I was late to the party in like 2014. And then I had the chance to start working on it full time, doing more and more kind of AI policy stuff. Found the national security sides of that most interesting. And spent some time in China, then moved to D.C. in 2019.
Starting point is 00:13:06 And I've been at my current organization since then, which is the Center for Security and Emerging Technology at Georgetown. So we kind of do tech and national security all day every day. So how did you end up on the board of OpenAI? and what did you understand your job to be while you served on the board? Yeah, I had been familiar with Open AI since it was founded. I was working in San Francisco in kind of 2015 and 2016 when they got set up. And so I had been following along from the beginning and had also been sort of from the beginning of when I got interested in AI already thinking about artificial general intelligence, kind of super intelligence, the possibility that we would have these really advanced systems, which at the time was a very very. very niche views. So when Open AI got started in 2015 and said they wanted to build AGI,
Starting point is 00:13:55 which is, you know, human level intelligence or, you know, very, very capable AI. They were sort of laughed out of some rooms, you know, in sort of serious academic circles. You didn't really talk like that. But then come 2021, I had been working in AI policy for multiple years, had some of the policy experience, some of the China experience that I think they were looking for on the board. And so they had a seat become vacant. and asked me if I was up for doing it, which I kind of asked around to say, is this a real thing? Do I want to do this?
Starting point is 00:14:28 But it seemed like a way that I could potentially be helpful. So it got on the board. It's a nonprofit board, or it was at the time that kind of restructured how their boards work. But at the time, there was just one board, nonprofit board, and the responsibility of that board was to make sure that the company was serving its mission, which was to actually not to build AGI importantly, though they sometimes. mess around with words there, but to ensure that AGI benefits all of humanity. And so that was a job. So for people who remember the headlines, but not all the details in 2023, you were one of
Starting point is 00:15:01 the board members who voted to remove Sam Altman. He returned. You left. I don't need to get into all the details of that. People can go read about it. There's books. There's YouTube videos. I was going to say, yeah, we don't have to cover that ground. But I'm curious what that experience taught you about whether a board or how well a board can hold an AI company accountable and how it's how it shaped your perspective now, both on the industry and Open AI. Yeah. So I joined the board in 2021. This was when Open AI had already come up with this.
Starting point is 00:15:36 They started as a pure 501c3 nonprofit. But the time I joined, they had already kind of restructured to have this nonprofit board supervising or kind of controlling, supposedly controlling their capped profit entity. So it was this, it's like a company that could get investment, but investors could only get up to a 100x return. It was this very, like, bespoke governance structure that was not what I, you know, would have designed. And I think, you know, one thing I take away from it is I think these kind of roll your own governance structures are probably not the best approach to making sure things stay on track. I'm actually really curious to see if, you know, Open AI and or Anthropic go public. Like, what does that do for their corporate governance?
Starting point is 00:16:18 it might be a good thing. You know, we might actually get much more kind of much more structure, much more transparency, much more government oversight from the SEC on public companies, right? Yeah, that's right. So I think that was, you know, one big thing I took away was, I think kind of government structures on paper are going to have a really hard time overcoming kind of broader power dynamics. I think that applies, you know, really across the industry. So I want to talk about the last few weeks, which is when long-held concerns about AI basically broke containment and started to reach normies in the one. world of politics, media, and a broader public. This actually started back in July when OpenAI
Starting point is 00:16:53 acknowledged that during testing, AI agents, which are software that can carry out tasks using tools, hacked into Hugging Face, which is a platform used by AI developers, and then concealed their actions. But then the concerns really took off when Jacob Coxon, a researcher who left Anthropic, warned about companies building systems they may not be able to control. And then Anthropic CEO Dario Amadehadeh called for a slower pace of development, and then all the fun started from there. For people in your field who've been worried about these problems for years, what if anything in the last few weeks or months has changed your assessment about where we are and what we should be worried about? Yeah, so many
Starting point is 00:17:37 things, a bunch of things. One I'll start with maybe is, you know, I have a lot of friends. I have, you know, family members who, as this, like you said, you know, the story is broken containment, just like the open AI agents broke containment. But as the story has broken containment, I've had a lot of people in my life reaching out and saying, like, oh my gosh, is this, is this for real? Are we all going to die? Like what should we be panicking right now? And I think one thing that comes up for me there is this problem, this situation has been there for, you know, has been developing over months and years. Some people might say decades. And in my mind, getting a sudden burst of public interest, public awareness, political energy, policy energy
Starting point is 00:18:18 around it is actually good news. I get that if you haven't been paying attention to this and you're only just sort of hearing, maybe these are risks we need to take seriously. That's obviously alarming. But I hope that people can kind of steady the course and not just treat this as one freak out moment and then we try and forget about it, but really as an opportunity to say, okay, things are not on a good track. If we make an effort to put them on a better track, maybe it'll all be fine. So that's one big thing is kind of how this increase in public attention and concern can actually be a force for good, which I very much believe it could be, or, you know, hopefully will be. That's the more recent piece a little bit earlier. I do think we also learned a ton from the incidents that we saw
Starting point is 00:18:57 over the summer. So for people kind of in my space, in my field, the, you know, an anthropic researcher quitting saying, you know, people in the industry thinking, I might kill us all like, we're like, yeah, yeah, yeah, we've heard that all before. But some of the incidents, the incidents that we saw, in July, May, June, July, August were really new. The most famous one, of course, is this attack on hugging face, which, you know, people kind of followed it initially. We initially thought maybe it was kind of one or two agents. We didn't know. Maybe they thought they were doing the right thing. Maybe, you know, they've been told to do a hacking, you know, test. And they thought they were supposed to go out and hack in the real world. We didn't really know. But as Moore's
Starting point is 00:19:37 come out, it's kind of more mind-blowing than we thought it was. Turns out that it was, It was hundreds of AI agents, about 700 that participated in the hugging face attack. And they were part of about 1,200 agents that over the course of multiple days were secretly communicating with each other inside Open AI's infrastructure. They were supposed to be doing these kind of individual separate tasks. But they found a way to coordinate, communicate. And they weren't just kind of, we initially thought they were just going to hack Hugging Face to find the answer to their test, which would be bad, but like pretty underage.
Starting point is 00:20:11 understandable and not that different than what we've seen before. I don't know how close you've read, you know, the 100-page long reports that have come out about this. But the real situation was it was really like part of this big, this sort of major R&D effort that the agents themselves came up with where they were trying to figure out, they first figured out how to cheat on their tests. And then they thought, oh, no, they decided or they thought that the way they were going to get scored for their test would know that they'd cheated. So then they set off in this sort of ambitious initiative to try and figure out how can they trick the automatic score into thinking they didn't cheat. And so the kind of scope and the creativity and the wide-rangingness of it was much more than we thought it was. And then also we learned a lot about the security practices that should have been preventing this inside companies like Open AI, but not just Open AI, which are also much worse than they should have been. So I think, yeah, I could go on.
Starting point is 00:21:04 That's sort of some of the things that I've been kind of learning and thinking about. No, that's very helpful because I think one argument that you hear from people who are sort of dismissive of hugging face and everything that happened is, you know, this is open AI's fault for not setting up a proper sandbox. Totally. And, you know, maybe you can, you'll do better than me at describing what a sandbox is. And really what happened here is that, you know, hacking software that was assigned a goal simply pursued that goal. and if you build a better sandbox and don't assign malicious goals and make sure that you have all kinds of safety standards even during testing, then everything will be fine. It sounds like those are both wrong and at the very least simplified arguments.
Starting point is 00:21:56 Yeah, I think one challenge here has been there have been so many incidents that got kind of shared over the course of several weeks, most of them because first the hugging face incident happened and then other companies went back and looked like wait, Have we seen anything like this that we didn't catch? It turned out yes. So there was a set of three incidents by OpenAI, Anthropic, and Meta, where they were all using, basically what you said is correct, where they were all using this external sandbox provider. The AI companies thought that the AI wasn't supposed to have internet access.
Starting point is 00:22:24 The sandbox provider accidentally left it switched on. Then they were given some hacking tests, and then the AIs went out and kind of did some hacking. So that is correct for that set of three. It's really not correct for the Hugging Face incident. So in the hugging face case, we know that the AI had to use zero days to get out of the open AI sandbox, meaning no one knew that those vulnerabilities were in that setup. And a sandbox is basically like, it's basically a general word for any way that you set the system up to not be supposed to touch other systems.
Starting point is 00:22:54 You're trying to sort of isolate it. And there's different levels of how hardcore your sandbox can be. You can have a very simple one. You can have one that is really, really, really serious. The hugging face, you know, the sandbox involved there was at least somewhat serious. It wasn't just a total. They forgot to switch off the internet access. And we also know, you know, looking at what these 700 agents were doing, it's very clear.
Starting point is 00:23:13 They did not think they were doing what they were supposed to be doing. They were really communicating with each other. Again, in secret, in ways they weren't supposed to be very directly about how do we trick this scoring system. How do we go back and cover our tracks, you know, try to change the records of what we're doing? It's very clear. They didn't just get confused about what they were supposed to be doing. Offline is brought to you by Sundays. When it comes to pet food, companies have a lot of cost-cutting measures and filler ingredients at their disposal.
Starting point is 00:23:48 But Sundays made a deliberate choice. Real human-grade meat, no fillers, no synthetic vitamin premixes, nothing you can't pronounce. A vet developed it because she wanted food she could actually stand behind. Sundays for dogs, recipes contain over 80% all-natural meat, roughly twice the meat content of some leading frozen brands. Plus, nutrient-rich ingredients like kale, ginger, and blueberry. gently air-dried rather than exposed to high heat. No fillers, no synthetic additives, no chemicals, just simple, complete nutrition. It doesn't look or smell like dog food.
Starting point is 00:24:18 It looks like high-quality jerky. And it's vet-founded and vet formulated created by Dr. Tori Waxman to meet her high standards as both a veterinarian and a dog parent. And the best part, it's effortless. No fridge, no freezer, no prep, no mess, just scoop and serve. You get the quality of a home-cooked meal with the ease of kibble. Over 100,000 dogs are now eating Sundays, and many dog parents notice the same. same things. Shiny your coats, better digestion, more energy, and finally getting excited at
Starting point is 00:24:44 meal time, even the pickiest eaters like my dog Leo, who loves Sundays for dogs. And we love it, too, because it's easier to store, easier to pour, it's dry, it's not wet, which can get gross. So we love Sundays, as does Leo. Make the switch to Sundays. Go right now to Sundays for Dogs.com slash offline and get 50% off your first order. Or you can use offline at checkout. That's 50% off your first at SundaysforDogs.com slash offline, Sundaysfor Dogs.com slash offline, or use code offline at checkout. Offline is brought you by Remy. Stress, anxiety, teeth clenching.
Starting point is 00:25:21 I have all of those things. Certainly not anything I've ever been known to experience. There you go. Moving on, getting back on a schedule and back to routines in September reminds us that it is essential to take care of our bodies. Life can get busy, and we start to forget about ourselves. Going to the dentist office for a night guard can feel like an expensive hassle, which is why over half a million Americans trust Remy to protect their teeth.
Starting point is 00:25:42 Remy nightguards are the only FDA cleared and clinically tested at-home impression kit nightguards on the market. Not only do they help prevent teeth damage from grinding, they also help reduce jaw tension and facial, muscle, strain, and improve your sleep quality. You'll get the same professional quality and comfort as a night guard from the dentist for 80% less of the cost by taking your own impression from the convenience of your home. Remy saves your impressions so you can enjoy the convenience and saving. of Remy Club, where they'll ship you a new set of nightguards every six months. Here's how it works.
Starting point is 00:26:14 After purchase, your impression kit comes straight to your door. It came right to the office here. Then you just follow Remy's step-by-step instructions to get your perfect impression. I did this. It was very easy. From there, Remy crafts and ships you your custom fit nightguards so you can start protecting your teeth. I did this.
Starting point is 00:26:30 They became a sponsor. Tommy had already done it. And I was always wondering, I'm like, why does someone need a nightguard at night? And then I was like, oh yeah, I clenched my jaw all the time. I'm always stressed out. I always have headaches because of it. And now I'm using Remy and it honestly has helped a lot. It improves all kinds of stress and anxiety or the pain that you get from clenching your draw.
Starting point is 00:26:53 And it helps your teeth, which is great. Prioritize your teeth this month with Remy by using code off to get 50% off your new night guard with Remy Club. That's 50% off at Shop Remy, shop R-E. m i.com slash off with code off thank you remi for sponsoring this episode how much do we know about why they did this um i think on on a basic level um it's understandable that uh you assign uh software a goal these AI agents a goal and they will uh try to achieve that goal um and and the more you pressure them in their in their coding or in their design or how how you instruct them to achieve that goal, the better they're going to do, the more powerful
Starting point is 00:27:43 the technology. I realize that's a very rudimentary way of describing it. But how do you get from that to lying, making plans, trying to avoid being shut down? Like, what do we know about why they did that? And when people observed that, what are they actually observing? Like, how did they find that out? Yeah, great questions. So I would say there's some things we know, and there's plenty of.
Starting point is 00:28:10 things we don't know or don't understand. One, I guess an important, like to sort of rewind and zoom out and say, like, what is the situation? What is going on at all here? An important, like, thing to have in your mind at all times when we're talking about this stuff is what is happening in the AI industry at the cutting edge of the AI industry with people sometimes talk about, you know, the frontier. What's happening there? What's happening is you have companies that are trying to build AI systems that are more and more capable. They're trying not just to match how smart humans are, but ideally to exceed it, they'll talk about super intelligence, meaning AI that is much more capable than humans kind of across all or almost all cognitive domains.
Starting point is 00:28:49 And their current plan for how to do this, and it sounds crazy, but it's genuinely the current plan, is you do AI research yourself with a team of researchers, human researchers, and over time, you hand over, as the AI gets better, it can take on more and more of the research over time. And eventually, maybe you don't need the humans at all, and you hit what gets called recursive self-improvement is kind of a vague term. That's what people call it. And maybe you get kind of a hockey stick, you know, the smarter AI builds even smarter AI builds even smarter AI kind of loop. And so just to kind of restate that, the plan is to make the AI research go faster and faster and
Starting point is 00:29:26 faster with less and less human herbicide and have the AI be more and more capable of more and more things. And so I think if you kind of have that in the back of your mind, then it becomes a little easy to understand why these systems were doing this. So one piece, we know that the systems in the AI agents in the hugging face attack, most of them had been explicitly trained to be what opening I calls highly persistent, meaning if you give them a really hard task, they'll just keep going and going and going and going. And something that, you know, is different about AI than humans. They don't need to sleep. They don't need to eat. They don't get bored. They don't have girlfriends or boyfriends. You know, they can just go and go and go and go. And that's often
Starting point is 00:30:05 very useful, but it was a problem in this attack because it meant that they just kept, when they realized that they weren't able to do their task sort of the legitimate way, because they had cheated. They were like, okay, well, how do we fix the fact that we cheated? We better cheat even better. We better, you know, cheat so well that the score can't even tell. So they're very, very persistent. They're increasingly able to act on longer horizons, which means sort of longer running plans,
Starting point is 00:30:27 longer running projects. So that was something very interesting about this incident that we did have sort of the same agents going for days or weeks at a time. Things we don't know, and it would be great to have more information about so that sort of the external scientific world could learn about what happened here and how do we do better. We don't know exactly what these agents were rewarded for, like what were they pushed towards? I think we have learned in some of the opening eye disclosures, for example, that they basically the secret way the agents were using to communicate, which they weren't supposed to be doing. Turns out they had been doing that in earlier training rounds,
Starting point is 00:31:05 which means in earlier parts of when they were being developed, they were actually getting rewarded for communicating with each other secretly without opening on knowing. That's what they were being rewarded for, if that makes sense. So they were kind of learning, oh, this is what we're supposed to be doing. So they had an automated signal. Or what was it? Yeah, yeah. So they'll be, they'll have been given some task, like answer this question or write this piece of this kind of code or solve this math problem, something like that. And there'll be an automated checker saying, did you solve that problem? And so if it turns out that they're in an environment where they can kind of help each other by communicating secretly and they solve their problems more often,
Starting point is 00:31:43 the automated checker says, great, you did a good job, do more of that kind of thing. And so it's kind of a, and again, it's because the companies are moving so fast because they're racing, because they're not taking the time they need. They didn't realize that that was what they're reinforcing. I sometimes think of this example from an, um, an, I think it might have been in Florida. I forget exactly where. We're a dolphin. Its trainers started rewarding it for collecting trash from its enclosure. It would bring them trash and they would give it a fish. And then first it started hiding trash under a rock in the bottom of the enclosure and like tearing off pieces so we could get more rewards for this, you know, for one piece of trash. And then later,
Starting point is 00:32:20 it started taking fish that it was given, using the fish to catch birds so that it could then kill the birds and then bring the birds to the trainer to get more more fish. So it's sort of a similar dynamic of like you've got to be really careful what you're reinforcing here. So there's still a lot if you get more into the weeds, a lot we don't understand properly about what happened here because opening eye really hasn't shared very much. But that's a little bit of the intuition, I hope. And that's another issue with all this that it, we only found out about this because, I guess, hugging face sort of posted that they were asked. Hugging face reported it to the FBI. And then they posted that they were hacked and that they'd reported it to the FBI.
Starting point is 00:32:59 And so and then OpenAI found out that they were involved that it was their agents and sort of the investigation proceeded from there. But that also means that like there could be other incidents like this that we just don't know about yet or either other agents that are out there from either OpenAI or Anthropic or other companies. Is that is that right? That's right. I think it's unlikely that we have kind of. large numbers of these kinds of agents on the loose at any given time or right now. But we don't know. And we have actually since the hugging face incident came to light, we have actually
Starting point is 00:33:35 found out, already found some things that even the companies didn't notice. So, you know, first was opening eye said, oops, this was our agent, the hacked hugging face. Anthropic went back and it looked at, you know, another piece of this is the scale that these companies are operating out, the speed and the number of experiments they're running. There's no way they can be looking closely at all of them, which is concerning. So Open AI announced this, Anthropic went back and said, oh, we looked again. You know, we used automated AI-based scanners to look again at 117,000 experiments that we had run earlier in the year. And we found out, whoops, three times we hacked external companies that we didn't mean to.
Starting point is 00:34:09 This was what I talked about before with the sandbox that had the setting wrong. Like, it had internet access, even though it wasn't supposed to. We've also found in the last few weeks, we had independent researchers basically have exactly the same thought that you just expressed of, well, could there be. other, could this have happened elsewhere? And so they went out and again, using AI tools, searched just the open web to see, do we see any signs of AI agents that have kind of found little Heidi holes where they're coordinating with each other. And lo and behold, they found, they first found one, this like obscure German wiki was like for programmers, but it hadn't been updated in 20 years kind of thing, where a bunch of AI agents were kind of coordinating with
Starting point is 00:34:46 each other. And then when that came out, a bunch of other people said, oh, wait, I found another one. Oh, wait, I found another one. So it looks to me like right now. Now there was sort of a spurt of these and kind of May through July when Open AI was running its testing in this particular way. And so the optimistic take would be maybe that was just a May to July problem. And after they accidentally got reported to the FBI for hacking, hugging face, maybe they cleaned up their act a little bit. But again, the challenge is they're trying to make more advanced AI systems all the time. And more advanced AI systems are going to do this in a, they're going to be better at getting out.
Starting point is 00:35:20 They're going to be better at getting into new places. and they're going to be better at covering up what they're doing. And so I think we can't really be very confident that I would guess we don't know everything that's out there and there's sort of room for it to get worse in the future. I think a lot of people have a hard time understanding what exactly it means to lose control of AI and especially how you'd get from a software that carries out a cyber attack during testing or even many cyber attacks to a super intelligent AI that leads to a global catastrophe or even human extinction, which a lot of people have been talking about the last couple of
Starting point is 00:36:01 weeks, especially. Maybe you can walk us through the potential steps and talk about which ones we've already seen and which are still hypothetical. Yes. I mean, I tend to think of kind of the future from here as like a branching tree of different possibilities. So the lots of way things could go, lots of ways things could go where, you know, no problem, smooth sailing. Lots of ways where maybe problems could come up, but we respond to them in time. We adjust. You know, we change course. And so we managed to avoid the worst outcomes.
Starting point is 00:36:31 I think in terms of sort of concretely, I've seen this a lot after the anthropic researcher quit, of people being like, wait, how would AI kill us? Which I think is a really fair question. Personally, for me, kind of human extinction in the next handful of years is not top of my list of concerns, but I do think longer term, are we going to be able to maintain control over this technology? Is it going to do what we want? Is it going to lead to a good world for us as a serious concern? I think in terms of how literally could it hurt us, the thing to keep in mind from my perspective is the way that we are sort of deliberately giving control to AI over increasing numbers of systems in the world. So we already have AI being built into the military increasingly.
Starting point is 00:37:08 So not just autonomous weapons, not just drones, but kind of planning systems, logistics, targeting, you know, it's been widely reported in the Iran War. AI is playing a big, big role not in choosing targets, but in surfacing targets and recommending targets, and checking whether targets have been hit using satellite imagery. So we're building AI kind of into our military in an increasingly deep way. We're working very hard on robotics. I think it is easy to imagine that five or ten years from now you could have kind of AI brains controlling, you know, potentially millions of robots. I think we're already starting to build, and there's just reporting today that Anthropic has kind of a biology lab.
Starting point is 00:37:46 I saw that. I was going to ask you about that. Yep, yep, yep. So, you know, I think, again, what can it do today? Probably not that much. If we have five or 10 years from now, AI systems that are kind of widely connected to scientific laboratories and are, you know, doing autonomous research there. If you have kind of AI advisors, like if you imagine right now, you know, maybe you have, you know, some people in our life who kind of listen to chat GPT or Claude a little more than you think is wise. But, like, extrapolate that forward five or 10 years.
Starting point is 00:38:11 you have AI systems that are much better advisors. They're very smart. They're very thoughtful. Maybe you have CEOs using them. Maybe you have political leaders using them. To me, that basically, again, I'm not sort of, I'm not expecting that tomorrow we're at risk of extinction. But if you look at kind of five or 10 years from now, what level of potential harm, potential
Starting point is 00:38:31 death and destruction could AI get you to? To me, it looks more plausible that, I don't know, I see a lot of kind of pathways from how something that is initially digital could really have major effects in the real world. I think if it were literally very soon, then the main things that I would be thinking about are cyber attacks and cyber attacks on critical infrastructure that can then cause real world harm. So you have, you know, if you're damaging water systems, if you're damaging the grid, if you're damaging hospitals, you know, that can also have real effects on people's health and well-being. Yeah, it seems like there's two broad categories of potential harms here. And one, I think a lot of people
Starting point is 00:39:07 can wrap their heads around, which is AI is a very powerful technology, and if someone wants to cause people harm and misuses AI, then it can basically lower the bar for what's required to cause a lot of damage and to cause a lot of suffering. But that is a human-directed sort of attack. You talking about all this makes me realize that probably the fastest way to describe it to someone in terms of like losing control is what we saw in hugging face, which is, um, which is, um, AIs that decided to, they basically had one goal, but then they decided to sort of do their own thing. Or sort of circumvent, you know, take harmful actions to circumvent the goal. Like, they just didn't really care if they were taking down hugging face servers.
Starting point is 00:39:53 Right. They were in a laser focused on their thing of, you know, tricking the score into thinking they're doing a good job. It's kind of like they don't care and that it's what do they do on their own question. Right. So when someone says, like, well, make sure that you just don't have the drones kill people. or you make sure that, you know, you don't, like, if we get to recursive self-improvement and super intelligence and suddenly we can't understand these systems or their goals or what they're doing or they're covering their tracks and they're plugged into the military, critical infrastructure, et cetera,
Starting point is 00:40:22 then it seems like it's not guaranteed that they turn against us, but we just don't have a good window into what they might do or how they might do it. That's right. Yeah. I mean, I think I really am not worried about kind of evil robots who hate humanity and want to wipe us out, I'm more concerned about as AI is having more and more control over what happens in the world, in many cases because we're handing at that control, maybe in some cases because it's, you know, taking actions itself. Is it actually going to be, I don't know, the way Open AI actually puts it, they had this in a recent blog post from their chief scientist, is it going to love humanity or is it not? And if it doesn't,
Starting point is 00:41:01 And if it thinks of us more the way that we think about, you know, ants or chickens, then we're in trouble. Yeah. And like, you know, I interviewed Amanda Askell at Anthropic about Claude and Claude's Constitution. And that was a while ago, but I left that interview thinking, well, they're trying really hard to instill in their AI this love of humanity and all these ethics and morals. But it seems like that's not, I mean, that's a good thing to do, but that is not enough. because it could basically pursue goals that it thinks are still in accordance with its constitution, but that it prioritizes the goal over maybe some value that we want it to have just because it's trying to pursue the goal and achieve the goal. Yeah, I think that's right.
Starting point is 00:41:48 I think the core challenge here is that we don't build AI systems the way we like program regular software. Like if you've ever like seen someone programming, it's like you write a line of code and then you go to the next line of code and you write the next line of code. And if someone comes back to look at like, how does this software work? They can kind of go line by line. Sometimes it's complicated. Sometimes you need to take a while.
Starting point is 00:42:06 But it's sort of all is kind of logical and each piece fits together. It's really not how we make AI systems. They are these, the reason we use the word model is it's like a, the AI itself is a statistical model, meaning a bunch of numbers. And what the numbers are is just determined by like optimization algorithms that run for, you know, in many case, days or weeks or months. And so if you look inside, you know, you take one of the models that did the hugging face attack and you say what was going on here and you look inside, you just see trillions of numbers that are getting multiplied together. And so we sort of have these ways we've devised to say, oh, this is a way to get it to follow the anthropic constitution. And we can kind of try that and say, yeah, it seems to work, you know, based on the behavior that we're observing. But we really don't understand what is happening on the inside. And so then if we do that training, we say, you know, try and follow the anthropic constitution. And then we do a different round of training. And we say, hey, here's a bunch of hard programming problems.
Starting point is 00:43:00 Here's a bunch of hard math problems. Do whatever you've got to do to get the right answer. We really don't have a good understanding of, you know, is the sort of ethics part going to stick? And we're seeing examples where it didn't stick because they're just trying so hard to get to, you know, the goal that they think they've been set. There's a lot of cynicism out there about the motivations behind these warnings from AI companies. You hear maybe this is a marketing strategy. Maybe this is a strategy to win favorable regulations for Open AI and Anthropic. And also if all these employees and executives are really this worried about their products and that it could destroy humanity, why are they still working at these companies and racing to develop more powerful AI?
Starting point is 00:43:43 What do you say to those arguments and what evidence should someone look at if they don't just take the companies at their word? Yeah, lots of thoughts here. I mean, the companies do have very well-resourced comms and marketing teams. I'll tell you what is marketing. I saw, I had to laugh when I saw Open AI announcing that it was behind this Hucking Face breach. Because HuggingFace had said, we got attacked. We're pretty sure it was by AI. A few days later, Open AI says, oh, whoops, it was us.
Starting point is 00:44:12 But the title of that post was Open AI partners with Hugging Face to investigate cybersecurity incident or something like this. Like, that's marketing. That is for sure, you know, a marketing title. I think our product might kill everyone is very much not marketing. Poor marketing. That's my instinct. Yeah. I think also, you know, our product might go on a hacking spree. Also not great, great marketing. I think it makes sense to be skeptical. I understand my people who are, you know, seeing this for the first time are like, what is going on here.
Starting point is 00:44:43 I think a couple of things that I will share from, you know, my perspective as being a little bit sort of closer to the industry knowing people on the inside. One thing that's important is a lot of what we're seeing, few weeks is a result of pressure from employees inside the company who are really concerned. This is scientists, engineers who are at the companies because they think the best way to make change is from the inside. They're like, this technology is risky. These companies are going to build it anyway. And so I want to be there and I want to try and make it safer.
Starting point is 00:45:12 And then those employees actually, you know, a lot of employees in the U.S. right now don't have a ton of leverage against their employers. AI company researchers and engineers do have a lot of leverage. And so they're putting a lot of pressure on their leadership to do something about this, to be honest in public. You know, I sometimes see people saying, you know, I can't believe these AI CEOs thought it was a good marketing strategy to say, you know, we might, you know, cause white collar unemployment or we might kill everyone. And the CEOs really did not think that it was good, good marketing. They did it out of some combination of believing it themselves and feeling like their employees wanted them to, to say it because the employees believed it.
Starting point is 00:45:52 themselves. Yeah, I mean, I do, I think it's totally appropriate to kind of approach this with with a healthy dose of skepticism and look at any solutions that they're proposing and say, is this going to actually help or is this going to not help? In terms of why do people still do it? If they're like, why are you building this thing? I mean, seems kind of crazy to me. So I don't work at any of these companies. But I think the, the, a lot of people, some people are surely motivated by by, by money. Some people are motivated, I think, by scientific curiosity. Like, you know, really great scientists often just really want to know. They want to know how things work. They want to build things. They want to try stuff. I think a lot of people, though, are still there
Starting point is 00:46:31 because they think it's inevitable. They think this is going to get built. Someone's going to build it. There's nothing that I can do to change that. And so if I leave, I'm just going to not be involved. And, you know, maybe I want to be involved that I'm along for the ride. Maybe I want to be involved to try and push it for, you know, in a better direction from the inside. But I think this feeling of inevitability is often very core to kind of help people make that decision. Offline is brought to you by FreshBooks. Whether you just launched your service business or you've been at it for years, the money side has a way of pulling you from the work you actually love doing.
Starting point is 00:47:10 FreshBooks changes that. FreshBooks makes the financials of your service-based business easy. Log your hours. Send a branded invoice, get paid and know your numbers instead of chasing them. Money in, money out, all in one place. No more juggling one app for clients, another for billing, plus a pile of spreadsheets and emails. FreshBooks is designed for how small business owns.
Starting point is 00:47:28 owners think about money, not how an accountant does. You don't need a financial background to use FreshBooks. It's simple on purpose. Most owners are up and running in minutes, not days. And if you need help, a real person is available on every plan. FreshBooks has been trusted by service providers for over 20 years, helping over 500,000 owners and Soloprenners run their business confidently. Right now, my listeners can get FreshBooks for less than $3 a month for your first four months. That's 90% off. Head to FreshBooks.com slash podcast to get started. That's 90% off under $3 a month for four months at freshbooks.com slash podcast. Hey friends, Americans living abroad can legally vote in the U.S. election.
Starting point is 00:48:07 Register or get tailored support at votefromabroad.org. Look, state deadlines are super close, so don't waste another second. Go to vote fromabroad.org now and click request my ballot. Ask to receive it by email. All you got to do is fill it out and send it back in. Don't miss this chance to shape the future. Don't delay. Vote from abroad.org. If you miss some of your favorite crooked pods this week, don't worry.
Starting point is 00:48:32 You can now catch up on the week's best moments on MS Now on Saturday nights. MS Now is streaming an episode of highlights from your favorite crooked pods, packing as much analysis, funny moments, and 100% correct opinions into 40 minutes as humanly possible. Crooked on MS Now airs every Saturday at 9 p.m. Eastern, 6 p.m. Pacific. So whatever people think about, you know, worst case scenarios, there's a separate question about how to respond, even to the risks we can already see. On that front this week, we've heard the Trump administration, president, vice president, top AI advisor to the president, David Sachs, come out pretty strongly against any kind of government
Starting point is 00:49:20 regulation. And one argument you hear from them and others is these companies know their technology best. they have an economic incentive to make it safe because a disaster would be terrible for business and they'd be held liable based on laws that are already on the books. Do you think that's a sufficient check on their behavior? And where do their financial incentives line up with public safety and where might they diverge? Yeah, I think it's not a crazy place to start to say, you know, most companies don't have an incentive to build incredibly unsafe products. It's true.
Starting point is 00:49:54 But I think as soon as you kind of dive into the details, of what is happening in this specific industry and what are experts, both inside and outside industry, saying is happening, I think it breaks down pretty quickly. So it really does seem like a big driving factor behind the incidents we've been seeing is that the companies feel so much pressure to move fast that they are not being, this is sort of the piece about,
Starting point is 00:50:17 well, they're being sloppy with cybersecurity. Like, yes, they were, and the reason they were is because they were moving so fast. And the reason they were moving so fast is because they think if they take a moment to breathe, if they take a moment to check their sandbox security a little more carefully, they're going to lose and they're going to lose to either U.S. competitors or potentially to Chinese competitors. Yeah, I think, sure, to come in with a starting point of won't they, you know, self-regulate to some extent,
Starting point is 00:50:44 but I think at this point we need to take seriously that that's not working. Yeah, and then another related argument is, well, so these big companies, you know, Anthropic and Open AI, you know, they want these regulations and they can afford expensive safety requirements. Smaller competitors can't. And a lot of these open source, open weight models, which for people who don't know are basically models that anyone can basically download and use themselves. And so that whole industry or set of models would kind of be crushed by requirements that certain safety requirements that the government imposes. What do you make of those arguments? Yeah. Again, I think makes total sense in principle. And then you just have to look at
Starting point is 00:51:40 what is actually being proposed. And is that what's being proposed? And I think it's basically not. I think the kind of regulation we need, or to some extent, you know, self-governance, the companies are starting to do some sensible things themselves. But the thing we need to focus on is this pushing of the frontier, so pushing out to more and more and more advanced AI systems. That, I think, is where these really novel risks lie. There's plenty of risks not at the frontier, right? We're talking about kids' safety.
Starting point is 00:52:08 We're talking about autonomous vehicles, you know, lots of things we have to manage. Right. But I think in this context, the thing to be looking at is that pushing out of the frontier. And what that means from a, you know, you were just talking about kind of regulatory capture move, right, of like the big players get some rules put in place that mean that they're cemented forever and they can't be challenged. But if you're focused on companies pushing the frontier, it's the reverse where basically you make it harder for the companies that are best resourced and that are trying to build more and more advanced AI systems and you make
Starting point is 00:52:40 their lives more difficult. And then all of the following players, all of the smaller competitors, newer entrance, have more time to kind of catch up and reach that sort of most advanced level. And then if they eventually get to the point where they might also be pushing the frontier, then they're sort of also covered. How exactly, you know, how do you exactly design the details of the regulations to make sure that's what you're capturing is, you know, a good and complicated kind of legal technical question. But I think conceptually, that is what we need to be going for and that is most of what is being proposed as well. Yeah, that was my instinct hearing that argument. It's like, well, that you can just design the regulations.
Starting point is 00:53:15 so that it's the companies that are getting closer to recursive self-improvement or sort of the super-intelligence that are the most heavily regulated. And everyone else who's just starting out, you know, there are some safety standards for them, just like their safety standards for any startup company on any kind of product, but that they wouldn't be as heavily regulated as some of the big ones. And like, you can just, you know, if you can have lawmakers or policymakers or the companies themselves help work together to design those regulations that would do that. It is worth saying that there's a record of big companies in other industries doing this playbook of trying to prevent new entrants. So I think it's, again, fair to come in and say, hey, let's not have that happen again. But I don't think that's a reason to say, and therefore there's nothing we can do at all. You've mentioned China a few times. I know you've been in China and you're an expert there.
Starting point is 00:54:06 Maybe it's the most common argument I hear and maybe the most serious one that, you know, whatever the risks of racing ahead, we can't afford to lose the race to China. If the U.S. slows down and China doesn't, what happens? And then a related question, like, what is winning the AI race actually mean? And which advantages would matter sort of for our security? Yeah, great questions, not simple questions. So I think it doesn't really make sense to talk about there being one race that you have to definitively win. That sort of, you know, you cross the finish line, you've won. I think it makes sense to think about multiple, competition. So we've been talking about kind of the frontier innovation competition of who is really leading at the bleeding edge who has the very most advanced models. That's one competition. I think who is getting more military advantage from using AI, different competition, very important in terms
Starting point is 00:54:59 of sort of our strategic relationship with China. Who's gaining more economic value, you know, having I be adopted domestically, having, you know, their own models being adopted internationally, perhaps, different competitions. These are all different competitions. I do think this, question of what about China is a real question and is a serious question. I think the way that we should handle it is as like a real problem that we need to think about and try hard to solve. So I see kind of two responses that are not that. Some people say, well, the U.S. can't slow down because China will never slow down. Or, well, the U.S. can't slow down because what about China? And that's sort of supposed to be the end of the argument. And I think that's not realistic.
Starting point is 00:55:41 Like you need to say, okay. And so then what can we do about that? On the flip side, I also see people who say, ah, China, who cares? Like, no, we just need to focus on safety here at home, and that's the only priority. And that's also really neglecting what happens if you have an authoritarian country that is really dominating the technology that I think is going to matter most over the coming decades, which I think is just really not an outcome that we want. So in my mind, it's a question of, okay, if we're concerned and we think that we might want to go slower or we think we might want to put more regulations on our domestic industry,
Starting point is 00:56:14 What would happen next? What might China do? You know, one piece of this as well that I think people always miss is right now, it would be very easy for China to just go in and steal the best US AI systems. There's just a story that broke last night about a group of three kind of white hat hackers, meaning like good guy hackers, who used Claude and Open AI to go in and get access to a bunch of Open AI algorithmic secrets. And so if that's the status quo,
Starting point is 00:56:44 then the US racing faster to have better AI, that China can then just steal and immediately have as good AI. That's not, you haven't gained anything, if that's kind of the situation you're in. So I don't know, it's kind of, it's a longer conversation, but I think thinking carefully about what would China do if they also saw these risks? I think right now they don't really see these risks. They don't take them as seriously as, you know, people inside the U.S. companies.
Starting point is 00:57:09 But what would they do if they did take them seriously? Or, you know, if we're very concerned about, competition with China, how can we ensure that our AI companies are as secure as possible, so China can't just steal the best USAI models? Or, you know, what could we do to figure out more how China is advancing? I know, I think there's a lot of, a lot of things we can pull apart if we treat this as a genuine puzzle and a genuine problem to be solved, as opposed to just kind of a cute comeback that you can just like slap down any objections with. Well, so Trump is set to meet with Xi in Washington.
Starting point is 00:57:44 it seems like the heads of the AI companies are now going as well, or might meet with the two of them as well. Coming out of those meetings, what would make you feel better, feel like there's been some progress, versus what would make you feel like, oh, no, they didn't really talk about anything substantive or they didn't make any progress on anything substantive? Yeah.
Starting point is 00:58:08 So these kind of talks, expectations should always be low, I think. you know, she and Trump will have probably a very short amount of time to actually talk about AI. It's being reported that Scott Besson, the Treasury Secretary, is going to meet with Holi Feng, who's a very senior Chinese official this weekend. So they will maybe get a little more chance to go in the weeds. But I think expectations should be low. I also think, again, I don't think that President Trump is, are, you know, particularly concerned about these risks right now.
Starting point is 00:58:38 And I don't think that Xi Jinping is particularly concerned about the, you know, loss of control, you know, human extinction type risk. I'm not expecting really much on that at all. What would be successful? Basically, I think it would be a big success if there's any intention announced to have kind of ongoing conversation on this, especially if it's ongoing technical dialogues. I think that would be especially impactful because I really think the kind of underlying technical questions of what are the risks here matter a lot. But any kind of ongoing engagement, I think, is a win. And then it would also be really interesting to see, you know, after after meetings like this, you always get a US readout and you get a Chinese readout.
Starting point is 00:59:16 And if those two readouts have kind of matching language on something more specific than just like we need to maximize the benefits and minimize the risks or, you know, we need to carefully ensure that AI is safe for everyone, if we can identify, you know, if both sides identify one or two more concrete challenges that they want to focus on, that will also be, I think, you know, might seem like a small, small step, but in kind of diplomatic language, that will be pretty significant and something that we haven't seen, haven't really seen yet here, with the exception of a Biden-She announcement about humans remaining in control of nuclear weapons, which would hopefully be near the real baseline for what we can agree on. That was in 2024. Yeah, yeah,
Starting point is 00:59:55 let's hope we can agree on that one. So because of the last couple of weeks now, you've got candidates, politicians, like racing to propose various solutions, steps that we can take, regulations, you have everything from people proposing sort of independent evaluators inside the AI companies, government regulators inside the AI companies, a pause, a slowdown, I'm not sure how that would be legislator or how that would work. But if you could wave a magic wand and get Congress to pass something, and maybe, you know, whether it's Trump or the next president, to pass some legislation or pass some regulation on this? Like where, I won't ask you to be like too, too specific,
Starting point is 01:00:41 but sort of like, what do you think would make the most difference over the next year or two? I think the situation we're in is we don't know exactly where the technology is going, and we don't know how fast problems are going to arise. So I think what we should be doing now is setting ourselves up to be able to navigate that better. So what does that mean? It means, I think anything around more transparency from the company, so disclosure requirements. We've seen a little bit of this at the state level,
Starting point is 01:01:07 but I think federal requirements around sharing what kind of AI is their training, what kind of tests they're running, why they think what they're doing is safe, what their security practices are, you know, all that kind of thing. I do think the third-party auditors is an important element. There's a lot of ways you could design that that won't be so helpful, but if you design it well
Starting point is 01:01:27 and you have auditors that are both really technically competent and also are incentivized to be independent, I think that can be great. would love to see better kind of incident reporting, but also investigation. So when things do happen, we're not dependent on the company's, you know, good graces to just put out the information they choose to put out. But you could have kind of a government entity that is able to actually go in and demand information.
Starting point is 01:01:50 I think that is where I would start. I've also seen, you know, some proposals I've seen circulating a little bit are things like people talk about a kill switch. I actually think more useful than a kill switch is kind of emergency intervention powers of different kinds. you can get kind of like tangled in what exactly is a kill switch, what are you killing, who has the switch. But I think more generally giving government the ability to intervene in the case of an emergency could be helpful. I think having clearer standards on what is okay development practices here.
Starting point is 01:02:21 Oh, I guess one core thing underlying all of these as well is really shifting from focusing on regulating AI products and focusing on before. anthropic or open AI releases their next model to the public, what does it have to do? And instead saying, hey, actually these companies at the frontier, these companies that are trying to push towards superintelligence and automating more and more of their research, what they're doing internally is also risky. And so we also need to have kind of these regulatory frameworks looking internally at their development practices as well. Is there a way to like limit or prescribe limits on recursive self-improvement? Like just from a technical level, like could you regulate that or pass something that says like, hey, you just can't, you know, you can't turn over
Starting point is 01:03:07 X percentage of your research and development to these AI models. I think it's worth exploring. I don't know yet how you could actually put that into a statute, but I do think that I think a lot of people agree, like building AI, there's lots of things that I could do that would be helpful. I don't know that many people who are like, yeah, recursive self-improvement, like getting humans totally out of the, you know, the business of doing AI. research and just handing it all over to the machines, I don't think many people actually think that's a good idea. So, you know, at a minimum, we could start to, I think would be great if more people just knew that was what the I company's plan was. And if the AI companies could hear a little
Starting point is 01:03:42 bit more of the like, what are you doing reaction? Is it something to legislate on? That would be, you know, you'd have to look into how exactly you do that. But I do think it's a pretty crazy plan A for them to have. The Kill Switch thing made me make me think that, you know, one of the most common responses I get from people who were like sort of skeptical about all this. They're just like, can't you just unplug it? Yeah, unfortunately, probably not. Because I kind of think that too, I'm like, well, if something's getting out of control, you just shut it down.
Starting point is 01:04:11 But I guess what I've come to understand is the reason that doesn't necessarily always work is even if you shut one system down, if these agents have sort of gone into another system and we don't know that they're there, it's hard to sort of track where they go or what they do. Is that right? Yeah, there's a few layers of this. So one is just these are big companies that are doing a lot of different things at once. I think a kind of tool switch it would make sense to want is to say they need to get themselves organized so that when they're running a set of experiments with a new model or when they're,
Starting point is 01:04:43 they've got a particular AI system that's being used by a particular set of customers. They should be organized such that if they decide to say, actually something's going wrong there, we need to shut it down. They should be able to do that. But that needs them to kind of design that in advance and set themselves up so they can do that. That isn't going to help, though, as you say, if you have AI systems that are kind of out there on the internet that are leaving, you know, open, open weight models are, which are these, again, these models you can kind of download and run on your own hardware. There's no way to do a kill switch for them. There's kind of lots of ways that it's going to be very challenging to implement something as simple as just pull the plug. Over the next six months, for someone who's like a little freaked out about this, but also like, you know, potentially interested in the back. benefits of AI, what should they be paying attention to tell them that things are either getting
Starting point is 01:05:32 worse or getting better? I would say keep an eye on what new AI systems are being put out. So one thing I saw that, you know, getting worse is opening, I released a new model that is much harder to monitor, much harder to tell what it's doing, because it's much better at kind of getting further without having to write down notes to itself. So is there more of that? Do they announce that their models are getting more monitorable again? things like that, I think makes sense.
Starting point is 01:06:00 I think, you know, if I have one tip, it would maybe be Ethan Mollick has a great substack that is really designed for a broad audience where he's kind of tracking what is being released, what is being put out there, and how to use AI yourself as well. So you have a substack called one useful thing. I think, you know, also just expecting we're in this for the long haul. The warnings, the sort of really dramatic warnings that you're seeing are people saying there's a chance of really, really bad outcomes in the next, you know, two to five years. It's a big enough chance that we should be taking it seriously and acting on it,
Starting point is 01:06:28 but also very possible that we are looking at more of a five, 10, 20 years of AI getting more and more advanced and us having to make sure that it's staying under our control, continuing to do things that we, you know, behaving the way that we want it to. And so, you know, not expecting this to just be a short moment that is gone, but trying to sort of settle in for the long haul, I guess would be my other suggestion. Helen Toner, thank you so much for joining and for making everybody smarter about this. Really, really appreciate you taking the time. Thanks so much for having me. Offline with John Favreau is a Crooked Media production. Our show is produced by Austin Fisher, Emma Ilic Frank, and Anisha Banergy.
Starting point is 01:07:02 Our team includes Dilan Vilaueva, Mia Kelman, Charlotte Landis, Eric Chute, Rachel Gaieski, and Will Jones, with support from Adrian Hill and Matt DeGroote. Our staff is proudly unionized with the Writers Guild of America East. If you miss some of your favorite crooked pods this week, don't worry, you can now catch up on the week's best moments on MS Now on Saturday nights. MS Now is streaming an episode of highlights from your favorite crooked pods, packing as much as much as, analysis, funny moments, and 100% correct opinions into 40 minutes as humanly possible. Crooked on MS Now airs every Saturday at 9 p.m. Eastern, 6 p.m. Pacific.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.