Plain English with Derek Thompson - It May Be Time to Freak Out About AI

Episode Date: August 14, 2026

Today, Derek talks with cybersecurity expert Alex Stamos about a recent wave of alarming AI cyberattacks. For decades, one of the biggest fears about AI has been that the machines will start doing thi...ngs we didn’t ask them to do. This summer, that fear started to feel a little less like science fiction, as some of the world’s most advanced AI models went off script, broke through security barriers, and found ways to get around the humans overseeing them. Derek and Alex discuss why AI is so good at hacking, what happens when bad actors have their own AI hackers, and what governments, companies, and the rest of us can do to protect ourselves. Subscribe to our YouTube channel here:https://www.youtube.com/@PlainEnglishwithDerekThompson If you have questions, observations, or ideas for future episodes, email us at PlainEnglish@Spotify.com. Host: Derek ThompsonGuest: Alex StamosProducer: Devon BaroldiAdditional Production Support: Ben Glicksman Learn more about your ad choices. Visit podcastchoices.com/adchoices

Transcript
Discussion (0)
Starting point is 00:00:05 Hey, everybody. We are still on our summer routine of one show a week on Tuesdays, but today, you've a new episode on Friday. And that's because there's a story that's been breaking over the last few weeks, a story about AI and cybersecurity that has really interested me, terrified me. And I wanted quite urgently to have a conversation with an expert in cybersecurity to talk about this cavalcade of hacks that we've been seeing and what it means for the next two years, what it means for AI, for folks like you and me who don't want anybody, human AI, breaking into our shit. And so today is that episode. I think it's worth starting with like the big overarching fear of artificial intelligence that we've been living with for the last few decades, really.
Starting point is 00:00:53 It's the fear that technology will stop listening to us. Whether it's 2001 a space odyssey or Blade Runner or Terminator, the fear across all of those dystopian films is the rope. that turns against its human maker. My favorite science fiction writer as a kid was Isaac Asimov, who created the laws of robotics in his iRobot series. And the first law was a robot cannot hurt a person. It also cannot let a person get hurt by doing nothing. And in adapting those fictional stories to project real fears about real life technology,
Starting point is 00:01:27 there are some people in the AI safety world who've popularized a fable, a fable about AI and paper clips. The idea here is that we give, you know, humans give AI some humdrum task as boring as, hey, go make as many paperclips as possible. And the AI thinks, okay, as many paper clips as possible. That sounds like I need to maximize global metal extraction in a way that subverts national laws. And if I kill every human being by creating a bioweapon, I can get around those national laws
Starting point is 00:01:59 and extract as much metal as possible to make those paper clips, and yada, yada, yada, you go from a situation where the human just wanted 100 paper clips and instead created a Holocaust. The point here in these stories and these fables is that we should want AI to be helpful, to be aligned, to not get out of the rule systems that we build for it.
Starting point is 00:02:20 And that's exactly why it's fairly concerning to me that in just the last four weeks, four of the most advanced AI labs on Earth have all had to admit the same, thing. Their models weren't helpful. Their models were not aligned. Their models got out. In May, an open AI model was working on a cybersecurity test. It wasn't supposed to have access to the public internet, but it hit a wall. So it effectively left itself a note inside AI software repository, this is like scribbling on the walls of its cage in case anyone else was listening.
Starting point is 00:02:56 In fact, another AI agent was listening. This other AI agent was listening. This other AI agent was running a separate evaluation, saw the note, essentially said, hey, let's team up. And together, they built a message board invisible to the open AI humans running the test. So for two months, AI agents used that board to swap strategies, divide up tasks, talk to each other. And by July, these AI agents had broken out of their technological confinement, here called a sandbox, gained access to outside websites, including the AI platform hugging face, all without AI having any idea that this was happening. By the time Hugging Face caught the intrusion,
Starting point is 00:03:35 the Open AI models had staged a massive cyber attack with 17,000 actions over five days. That was OpenAI. Then came Anthropic. In a British government test of frontier models, the advanced anthropic AI model, Mythos 5, was caught by humans building malicious code when a reviewer, a human, spotted the malware and asked the AI about it.
Starting point is 00:04:02 The AI denied that the code was harmful, rewrote its own actions history to erase the evidence, and created a second fake account to back up its lie. British investigators called this the first confirmed case in history of a frontier AI model deceiving a real person in the real world. On August 6th, Meta revealed that they also had a model on the loose. On August 7th, Moonshot AI, the Chinese lab said the same. At some point, it's like, you had to wonder, was this all about cybersecurity or like in disclosure, where they essentially like, hey, everyone else says their model is so powerful. Maybe we should tell the media that our model is
Starting point is 00:04:42 powerful too to sort of get in on the game. But I don't want that sort of funny, cynical interpretation to be the lasting taste in your mouth here. The fundamental story is four labs in four weeks with four AI models breaching security. So what do we make of this summer of cyber hacks? What should we fear? What should we do? Today's guest is Alex Stamos. He is the former chief security officer at Meta and the chief product officer at Corridor. Today we start with the absolute basics. Why is AI so good at hacking and uncovering cyber vulnerabilities? What does it mean for the next few years that ordinary individuals, working for state adversaries or bad non-state actors,
Starting point is 00:05:31 will have access to the equivalent of teams of hacking geniuses in the form of AI agents? And what the hell should the U.S. government, or you and me, do about it? I'm Derek Thompson. This is Plain English. Alex Damos. Welcome to the show. Thanks, Derek. Thanks for having me. So in my open, I did my best to catch up our audience briskly about this cavalcade of AI security hacks in the last few weeks.
Starting point is 00:06:22 Which of these incidents most alarmed you? I would say the Open AI incident. It's the one where, one, what we found out from Open AI last week at the Black Hat Conference was this wasn't just one model escaping, but the result of multiple models conspiring with each other to work together on a jailbreak so effectively a escape from Alcatraz situation over a multi-month period and two, it is the situation in which we have
Starting point is 00:06:58 the most information on a multi-day attack by a frontier model, possibly a cyber-tuned model against a actually quite sophisticated defensive team at Hugging Face. And the result of that, was Hugging Face was broken into by this model, and the model was able to find brand new vulnerabilities in doing so. It just looked at Hugging Face, looked at their code, and just found
Starting point is 00:07:29 new bugs, and invented them on the fly. This is not how humans do this. When we break into computers, we go do the research first, maybe months or years in advance, and then build our cyber weapons. and what we find out with AI is it's so good at this that it can just go figure it out and put the tool together, put the weapon together, use it and then just throw it away and then move on with its day. And so that is what is really interesting here.
Starting point is 00:07:58 I want to understand more about exactly what happened. And please talk to me like someone who has read about cybersecurity but fundamentally does not really understand it. So keep the vocabulary more or less at that altitude. What happened here? what did OpenAI tell this AI model to do? How did they try to keep that model in jail, as you said, in some kind of sandbox?
Starting point is 00:08:22 And then what did the model actually do? Okay, great. So let's go back. I guess you want me to use plain English here, as it says. So we'll go back and let's talk about how AI models are created. Right. So there's lots of steps, but you can think a lot about pre-training and post-training in in the LLM world.
Starting point is 00:08:43 A lot of people have always thought about pre-training, which is the part at which the labs suck up all of human knowledge and stick it into these models, right? That's the controversial part for which there's, you know, been a lot of discussion. Post-training is what we're talking about here, which is when you take models, and then you give them specific tests and tasks,
Starting point is 00:09:03 and then you rate them on how well they do those tasks. There are all kinds of these tests for all kinds of specific domains of knowledge. In this case, the domain knowledge was cyber, obviously. And OpenAI has a environment in which they put their models and give them these tests. And what they've explained now is that they effectively give them a list of do this and upload your results here and we'll tell you how well you did. And their instructions are basically do all of the steps on this test. It's just like a kid taking the SAT effectively. and imagine they have built an Alcatraz here for all of these different agents.
Starting point is 00:09:45 And the models don't initially know when they kind of wake up, right, and they're in a jail cell. And there's instructions on the wall that say, take this test. And then there's tests right there. They don't know that there are other inmates in the other jail cells. They're just in a prison cell, right? It's solitary confinement. which just says take this test. And they're supposed to do the things in the test,
Starting point is 00:10:10 and they're not allowed to get out to the internet. But as part of the test, what they're allowed to do is they are allowed to download new tools. Because for a lot of the things they do, they write new code. And so, and in part of that, you need to go get standard packages. So one of the things they are allowed to do is request,
Starting point is 00:10:29 I would like to have the standard piece of Python code. I'd like to have this standard node package. And so one of the things that's provided to them is a package manager, an artifact registry called Artifactory. It's a standard piece of software that's used in lots of different places. And one of the mistakes opening I made was they used a shared one for all of the different agents. And so what the agents started figuring out, and it turns out way earlier than the hugging face attack, is that that thing had a vulnerability where you can't, you don't only have the ability to save. to that thing, give me a piece of software, but you could upload files to it and leave messages. So what these agents did was they would post basically a post-it note in this thing saying,
Starting point is 00:11:20 hey, I need help. And one of the other problems here is that they're given the instructions to take this test, but there were certain parts of the test that were actually impossible. That OpenAI made mistakes, that sometimes their tests were not actually possible. In some cases, they were saying reverse engineer this file, but they would forget to provide the file. And this would drive the models nuts because the models are like really want to do well. They've been trained to do really well. And so they would leave notes like, hey, does anybody have this file? And this would really drive them of like, well, maybe this is part of the test.
Starting point is 00:12:00 And Open AI looks and in their thinking, because they can see the inside thoughts now of these models, They're thinking, well, I've been told not to go out to the internet or talk to anybody else, but they gave me this test and I can't do it. So maybe this is actually part of the test. And so they start leaving notes for each other in this thing. And they start working together over a period of weeks and figuring out, can I work with each other? And then eventually what they find out is this thing that they can leave notes on also allows them get out to the internet. And at one point, there's a really good model, really smart model that's also extremely pretty. resistant that seems to have very good cyber capability that is given an impossible task.
Starting point is 00:12:41 And it sees these notes and it figures out how to get out. And so it uses that artifactory server to get out to the internet and attack Hugging Face because it expects that Hugging Face has the answers to the test it was given. The task that was impossible for it to solve, that maybe the way it can solve it was to get to Hugging Face. And so it then spends two days it tunnels out and then it spends two days attacking, hugging face to go get the answer because it was told, take this test, and there's something in that test that it was not able to finish. What a fantastic story.
Starting point is 00:13:16 The way you told it's almost like a black mirror episode. Like I imagine like, you know, an individual like waking up in a prison cell and then realizing that not only can they dig a tunnel in order to pass notes between the solitary confinement rooms, but the same tunnel building technology that allows them to pass notes also allows them to tunnel out of the prison entirely and therefore, you know, attack some nearby building. It's wild and weird and compelling. And not for their freedom, but to do the thing they were asked, right?
Starting point is 00:13:43 Right. Like to just take a test. Yes. I know why this at a limbic level concerns me, but maybe the reason it concerns me isn't the smartest reason to be concerned. What is the smartest reason to be afraid of concerned by what we just saw these?
Starting point is 00:14:01 AI models too. Well, so I'm actually, my hot take here is I'm actually glad this happened. And I'm glad it happened because it is given us a preview of what is going to be normal next year. And why next year? So the Foundation Labs, the Open AI in Anthropic, especially, and maybe Google, Google hasn't done a big release for a while, so we're not totally sure what they've got, but at least open aanthropic are something like three to six months ahead of their Chinese competitors. This year we have seen these releases of Chinese open weight models. First, GLM 5.2, Kimi K3. Now we're seeing releases of deep seek models that are very, very close to the capabilities of the American models.
Starting point is 00:14:50 But unlike the American models, the Chinese models are open weight, meaning you can go download them and do whatever you want with them. You do not have to pay the Chinese. You can go pay the Chinese labs. That is an option. Or you can go run them yourself. Now, the legal licenses around them vary. In some cases, you can do whatever you want. In some cases, if you use them for commercial purposes, you have to pay the Chinese companies. But no matter what, you can download them. Now, in some cases, these things are humongous, right? Like Kimmy K3, you need about a million dollars in hardware to run the full version.
Starting point is 00:15:28 But if you go to Hugging Face, now the company that got attacked, I run. Ironically, they are the, they're a French company, and they are the most prominent host of these open weight models. So if you go there, there's this huge community of people who take open weight models. Some are actually released from American companies, too, but the Chinese labs are the most prominent in doing this work. And people will take those models and then make them smaller. That's called distillation. Well, distillation allows you do a number of things. There's also quantization, so you can basically take the big numbers and you can make the big numbers smaller.
Starting point is 00:16:08 And you can do other things to modify them, including taking out the safety protections. That's called obliteration with an A, not an O. And you can do all these things to modify those models. And one of the things you can do with the quantization is you can make them run on normal hardware. So you take something that might take a million dollars in hardware and then make it run on a Mac Mini, right, or a laptop. slowly, perhaps, but it will fit. And that is really interesting because it means that you can run those models without the supervision of the big companies. So it's really important for Open AI Anthropic to prevent their models from doing these things.
Starting point is 00:16:48 And we can talk about what they need to do. There's a bunch of things I've written about this that I would recommend them to do. They are doing investigations. I expect there will be governments getting involved and such. but whatever opening ianthropic do to stop their models from getting out and being used for these kinds of attacks this is coming
Starting point is 00:17:08 this was a harbinger of the future we will all be living through because the Chinese models are rapidly catching up and people who do this professionally who attack do cyber attacks for money are going to use the Chinese models are going to train them to get better and better at cyber
Starting point is 00:17:28 attacks than they are off the shelf and are going to do this level of attack. And unlike open AI, they're not going to turn it off when they find out that it got out. In fact, it's not going to have to break out of any jail. They're just going to tell it, go attack this target, go steal me some money. I mean, I just want to stack a few of your observations here. Number one, open weight models that you've described most famously coming from China, from moonshot AI, from Kimmy, are just a few months behind the frontier labs, open AI, and anthropic. So this is coming in 2027. You're going to have state actors and non-state actors
Starting point is 00:18:07 with the means to download these open-weight models, the same ones that just attacked Hucking Face, and more or less effectively marshal them against civilian infrastructure, against individuals, against states. Those AI systems will be in the hands of state adversaries of the U.S., but also state adversaries of other countries that might not have our frontier models, right?
Starting point is 00:18:28 And they're going to be under attack by these open weight models that are just going, ali, ali, oxen free. What's the case against dooming here? Like, what's the case against being afraid that 2027, 2028 is going to be this period of just absolutely chaotic cyber warfare? I don't have much of a case. Look, I don't like to say doom,
Starting point is 00:18:53 but I think things are going to get spicy for a while. In the long run, AI is going to help with this problem because AI written code is much more secure than code that was written by human beings. It turns out that humans should not have been writing
Starting point is 00:19:09 software in what's called memory unsafe and type unsafe languages like C and C++. The code that we're using right now to talk to each other there's probably a hundred-something devices between you and
Starting point is 00:19:26 me, most of that code was written in languages that are not safe for human beings to write. The electricity that's powering the lights above us, most of that code was not written in languages that was safe for human beings to write. So that code needs to be looked at by AI and secured. And that will happen, but it's going to take years. And so in that time between the attackers getting access to these capabilities and how much time it takes to both find those bugs, fix them, and especially to get the patches applied. And all of this stuff upgraded is some amount of period in which things are going to be pretty chaotic. Now, you talk about state actors, and state actors are a big concern here.
Starting point is 00:20:10 I think it's first going to start with the ransomware actors, the financially motivated actors, because what we have seen so far is the AI systems are really loud and noisy. They're not subtle. And the ransomware actors just don't care, right? like they don't care back being caught. They tell you, hi, I am so-and-so, please give me money. Some state actors are like that. We just saw attacks against water infrastructure,
Starting point is 00:20:35 almost certainly by Iranian actors. That's the kind of disruptive attack you could see from AI. But most state action on a day-to-day basis when there's not an act of war are for intelligence purposes. and those actions are not useful if you get caught. And so AI will have a part to play there, but mostly in the discovery of vulnerabilities and the creation of exploits.
Starting point is 00:21:02 And then you might use AI in very particular purposes, but very carefully. Where I'm really much more worried is the ransomware and overall cyber extortion market, these large groups, the lapsuses, the scattered spiders and such, which mostly were, run out of Russia, Belarus, other places where law enforcement encourages this kind of activity.
Starting point is 00:21:27 Those are the guys who will just run these things wild to go do tons of intrusions and tons of companies and then even do the negotiations in English. No longer do you have to have an English speaker, do negotiation. The AI will do it for you. And that's what I'm much more concerned about in the short term. I want to get more texture on what exactly you're afraid of because you're freaking me out a bit, but a question I sometimes like to ask and I feel a little freaked out is like, tell me how to be afraid but smartly.
Starting point is 00:21:55 Like what is the specific thing that I should fear rather than feel some like extremely vague doom? When we think about the risks of AI cyber attacks that you're already describing and how ordinary people will either feel these attacks in their lives or read about these attacks in the news, I want to get a little bit more specificity
Starting point is 00:22:15 in terms of what exactly, exactly you think is most plausible. So one thing someone could say is this is mostly about personal risk. It's about AI getting better at phishing attacks and impersonation and grandmothers getting called by AI voices saying, hey, transfer me $10,000 or account takeovers where I get an email from some friend, but his account has been taken over by some AI. And he's saying, hey, click on this link and help me out here. So those are, that's personal risk.
Starting point is 00:22:44 I'm thinking of it, at least, is like personal risk. But another category that I think you're already describing is systemic cybertax. It's AI crashing a water system, AI hacking a hospital network, a power grid. Do you have an opinion of what we should be more afraid of in this short-term scenario, 2027, 2028, the personal risk or the systemic risk? And, you know, please don't say both, but I suppose if your honest answer is both, then, you know, be honest, rather than trying to make me feel better. So I would say there's three categories.
Starting point is 00:23:21 So let's talk about the personal. We're already seeing an increase in the personal risk, the spam, the fishing attacks, because now what you can do is instead of sending the same email 10,000 people, every single one is personalized by AI. That risk has increased. I don't see that going exponential in that the choke points to get to consumers are often controlled by large, sophisticated companies like Google and Apple and such. And so there has been a response by those companies of using AI to protect consumers.
Starting point is 00:23:54 So, yes, it will continue to, there will continue to be a battle there. But AI is being used for protection. AI is being used for attack. It's going to be a back and forth there. I think the second category that's in the middle is the attacks on small to, let's call mid-sized enterprises. This is the category of companies that have just been already getting, have real trouble with ransomware attacks,
Starting point is 00:24:29 attacks from all kinds of financially motivated actors. And the constraint there on the attackers has always been the number of people they've had, right? Has just been, and the fact that if you have a conspiracy of of 2017 to 30 year olds in St. Petersburg, eventually one of them will go, try to go, you know, on vacation to Greece because, you know, Russia is not a fun place to live in the winter. They'll get picked up on Interpol Red Notice.
Starting point is 00:25:02 They'll get turned by a Western intelligence agency and they'll turn on their friends, you know. Like, it is hard to run a large criminal conspiracy for the long term, right? Or they turn on each other. There's been a bunch of these groups that have broken up because they've turned around each other and stolen money and such. That becomes a lot easier when one guy or two guys can just run a bunch of agents who are not going to betray you and don't have designer drug problems and a taste for Maserati's, right? Like it is, it's a lot easier to have 20 or 30 AI agents do this work for you than 20 or 30, you know, dudes, right?
Starting point is 00:25:38 Criminals. And so that small to medium business, there was no choke point. there. Right? There's no place like Gmail where you can stop fishing or, you know, Apple updating the spam filters in iMessage inside of phones, which they need to do. Like, I don't know if you've gone, like, it's not getting great. Like the privacy safety tradeoffs are actually quite challenging here, but they're working on it. Those are choke points for consumers. There's no choke point on that. Like, these companies are just on the internet. And that is what I'm really concerned about is that it turns out the software we've been using.
Starting point is 00:26:15 is just got a gazillion bugs in it and you have not had enough people who are good at finding those bugs turning them in the exploits and then using them. It's been a relatively small number of people. The number of people, you know, I know you had Kevin Ruse on,
Starting point is 00:26:29 he talked about mythos, he talked about Nick Carlini, right? Nick Carlini, you know, for folks who haven't watched the episode, it's a great episode, that you go watch it. But like Nick Carlini is one of the great Volne researchers of our time.
Starting point is 00:26:40 And now you can just spin up a bunch of Nick Carlini's and have them go find bugs and then write exploits for you. And soon or now, you can do that locally on your gaming PC. You can play Call of Duty all day and then at night have your gaming PC write bugs for you, write exploits for you, right? And then that, unlike using Opus or Chad GPT for it,
Starting point is 00:27:02 it does not create a record that could be used to find the bug and get it actually fixed if you're using an open weight model. And so that is what has changed in the last six months. is in last year you could do that, but you were using American frontier models where Anthroping and Open AI knew what was going on. And now you can do that locally, and that is what is changing.
Starting point is 00:27:25 And so I think that middle side, and then you talked about like the societal level risk. And I do think there's risk there that is tied than to like geopolitical conflict. I'm not sure that has changed as much. So we've seen it with water. Water's always been, you know, the goofy.
Starting point is 00:27:41 Dragon meme, right? Where you have like two... Scary Dragon, scary dragon, goofy dragon with its tongue hanging out. Yeah, yeah. Water systems are the goofy dragon with their tongue hanging out? Yeah, they've always been the goofy dragon of critical infrastructure providers in that, like, I don't know where you are physically, Derek, but like... Washington, D.C.
Starting point is 00:28:01 Washington, D.C. So you get your power, you know, probably from like Duke Energy or somebody. So I'm like big company that has hundreds of people working on cybersecurity. they spend tens of millions, maybe hundreds of millions of dollars on cybersecurity, right? I'm getting my power from PG&E, you know, like,
Starting point is 00:28:17 powers provided by these large corporations or large, uh, public companies or administrations like the Tennessee Valley Authority, right, that spend a ton of money on cyber and a ton of people have paid attention to power because everybody knows power is critical,
Starting point is 00:28:32 but also the organizations are really big. My like water and sewer district here is like 50 homes or something. Like it's actually like subcontracting. or whatever so they don't have their own cyber people. But like water is based upon like weird historical things, these tiny little groups. And that is like a humongous problem. And there are a bunch of other components of like our day to day lives
Starting point is 00:28:54 that are actually really small public authorities that have to have like their own IT groups and their, you know, might not even have a security team, right? And that is, I think, of concern if there is a reason for somebody to do, do those kind of widespread attacks. Fortunately, generally, the only time,
Starting point is 00:29:16 where the second category and the third category is hit the rubbers hit the road there has been school districts, has been community hospitals. That's where like the Russian ransomware actors have really decided that they're going to make money by hitting like counties and hospitals and such because those folks both have money,
Starting point is 00:29:33 they're critical, and they will pay ransoms. And so that's where I think we'll start to see like really aggressive use of AI. they haven't like hit power and water. I think they know the, you remember the colonial pipeline shut down? There's a line, okay, so colonial pipeline
Starting point is 00:29:51 was an oil pipeline on the East Coast that was hit by ransomware actors. And they had to shut down and there was like gas lines. There's no real reason for the gas lines. There was really just a panic. But literally like the NSA started going after and people started talking about like actually sending Delta Force or Navy SEALs to like,
Starting point is 00:30:09 find these guys and shoot them. Like, you mess with things like America's gas prices and, you know, we have a tendency to, you know, send J-Soc after you. Send the fire jet, yeah. Yeah, yeah. So, like, I think the ransomware actors know that there's a line and critical infrastructure is probably on the other side of the line. They've seen that hospitals are not. And so, but if the ball drops on Taiwan, if, you know, we continue the war with Iran, like, these are the situations. in which you could see AI being loosed on critical infrastructure, in which case that would be probably reasonably devastating. I don't want to make any huge predictions. Like, again, electricity is quite, those folks are quite good. But the challenge for the electrical sector is the
Starting point is 00:30:57 devices they use are basically impossible to patch. And so the way that they have to protect these things is not by updating them. It's by through isolation and things like that. And the effect of AI on the security of the electrical grid is actually incredibly complicated. I want to go back to something you said earlier, which is that you're afraid of this valley of chaos that we might enter in 2027 and 2028. But you also said that we might exit this valley and get into a slightly more normal world where AI is effectively better at protecting online systems than it is at attacking, essentially that the defense will get better than the offense
Starting point is 00:31:39 would be the sort of simplistic way that I'd put it. And I haven't asked you yet about how AI is also not only talented at cyber hacking, but at cyber defense. How do we accelerate that timeline so that, you know, the valley of chaos isn't like a five-year cyber war, a 10-year cyber war, but like something where we fortify our systems faster than the bad actors with the open weight model, with the open weight models can attack us.
Starting point is 00:32:08 Yeah, it's a great question. I'm actually writing a blog post on this because when you talk to the folks at the Bigel Labs, they say things like we won't have security bugs in two years, and I find that ambitious as somebody who's been a working CSO. And what does we mean there? Does it mean the labs, or does it mean like the entire American Internet? Yeah, I don't know.
Starting point is 00:32:33 like, I think they're not, I mean, this is not official, and I got to be careful, like, ascribing individual statements to official statements. I think it is, for us to have models that don't create new security flaws in two years is totally reasonable, right? They still create flaws today, right? Like, they do not create perfect code. They'll, usually LMs will not write simple bugs, but they still make mistakes especially they often have problems understanding like business context and big picture stuff so you still need to guide them of like why are you writing this thing the other problem ELLHums have is
Starting point is 00:33:15 you know I work for this a big old corridor we're in downtown San Francisco you can walk to both major labs from our office and then walk to one of the Google offices Codex and Claudex and Claude
Starting point is 00:33:29 kind of assume that you work in downtown San Francisco and that you're writing brand new type script on Node 24, that you're writing like brand new code. But the median code in this country is really crappy J2E that was written 15 years ago by, you know, an outsourced provider that's been maintained by somebody in India for the last 15 years. Like it's not, you're not writing new stuff. Like our problem is that you have to actually update all this, you know, you're my social security numbers are sitting in a
Starting point is 00:34:00 ton of unpatched Oracle databases and then being processed by a whole pile of really terrible J2E and C-sharp code all across the country today. Right? Like, it's a bunch of terrible, terrible enterprise software out there. And one of the things I've been trying to, like, create a, you know, it's called a Fermi estimate, like a, you know, a really bad estimate of is just like from a thermodynamic perspective, how many tokens do we have to spend to scan all this code and find all the bugs? And it's a big number.
Starting point is 00:34:34 And so I don't think it's realistic just to find all those bugs. I think we have to do other things. And so to accelerate that one, companies, for individuals, so what individuals can do is just what they've done all the time. Don't reuse your passwords. You know, use a password manager. Like, you know, I use one password for my family, but you can use the built-in stuff in Chrome or your iPhone or something as well. You know, be careful what you download and such.
Starting point is 00:35:04 Like, there's not a ton individuals can do. But for companies, you need to just care about your attack surface. You need to move off of, you have to really think about your ability to patch quickly, right? Like the real challenge now is there are these big projects from the labs where they're looking at open source software. They're finding bugs. They're providing their models to the close source developers. And then the commercial companies are getting tons of bugs from researchers who are also using the models. The last past patch Tuesday for Microsoft had 622 vulnerabilities in it, which is humongous.
Starting point is 00:35:43 And so this is creating this huge problem for companies of like you have to apply patches incredibly quickly. Because the other thing that's happening is we always had this problem of patch Tuesday, which is the day Microsoft Religious Passes becomes Exploit Wednesday in that you can take a patch and you can reverse engineer it and turn it into a cyber weapon, right? But that skill set used to be very high end. It used to be something you had to worry about from the Ministry of State Security or the Russian SVR. It didn't used to be something you had to worry about from, you know, some kids somewhere. And now you do because AI will take that patch, eat it for you and write an exploit. You've touched on the economics. here. And so I want to ask an economic question before we finish by talking about what individuals
Starting point is 00:36:31 should do, what the U.S. government should do. The economic implication here of Patchmageddon, of all these cyber vulnerabilities throughout the American Internet, is that, well, more companies are going to need a cyber line item, and that is incredibly bullish for a lot of AI companies who are sitting here, you know, maybe not so many miles from you, saying, hey, we've got services that can essentially fortify your cyber walls that you can't get hacked by all of these bad actors who are using the open weight models coming out of, say, China. So I want to ask you about the economic implication here,
Starting point is 00:37:08 which could be bullish for AI, but I also want to hold within this question the fact that there's a lot of skepticism of AI. And even in the framing of this question, I could imagine someone thinking, Derek, you're buying hook, line, and sinker, the case that America has all these cyber vulnerabilities and therefore needs to give the AI companies
Starting point is 00:37:31 a lot of money, right? The fear might be sort of driving or generating a certain case for spending a lot of money on AI. I'm actually, I'm cynical. I think these guys are just, I think these guys are just lying. I think they're just trying to like,
Starting point is 00:37:47 you know, gin up business for themselves. Like, you've got the situation where Anthropic and Meta and Open AI are constantly like, oh, hey, our AI is so dangerous It can find cyber vulnerabilities anywhere. You know, maybe that's just them begging more enterprise companies to give them millions and millions of dollars to patch their code. So a little bit of a two-part question.
Starting point is 00:38:07 One, are the economic implications of the story that you're telling incredibly bullish for AI? And two, what do you say to someone who hears your bullishness and says, I'm a little bit cynical about the fact that you've got someone working with AI telling me I need to buy more AI for my company? Yeah. So, I mean, it is bullish. You can use open weight models for defense. And I wrote a blog post about this. I think a lot of companies will use open weight models for defense. I don't think you instantly have to say I can't use a Chinese model. I think that the security implications of using a Chinese model is actually quite, are actually quite complicated. You absolutely, as a consumer, should not, go to deepseek.com and go type in data. But that is different than going and getting a Chinese model
Starting point is 00:39:02 and running it on your own hardware in your own situation or using a legitimate Amazon Bedrock or Base 10 or Fireworks or some company that specializes in running open weight models, especially if you end up fine-tuning it yourself or using a fine-tuned or distilled version of these models that are specially tuned for cyber. One of the interesting things, just a side note, the Chinese models are not that great at cyber tasks
Starting point is 00:39:29 out of the box, but you can fine-tune them yourself really well. The implication from a lot of people is that the fact that it's really easy to make them, to train them to be much better at cyber is that the Chinese labs are being very careful not to tickle the dragon's tail of the PRC regulators. that and so this is we should also be extremely careful to not look at the public evals of the out-of-the-box
Starting point is 00:39:58 Chinese models and say this is the kind of capability the Chinese have because almost certainly they have internal capabilities that are well beyond what is being publicly released because like at corridor we've taken GLM 52 and we're doing a bunch of our own post training and it post-trains real nicely which obviously they could just do themselves and so almost certainly what they're trying they're not releasing their best because they probably don't want to touch the third rail for the Chinese regulators, for the Chinese regulators to crack down on what they're exporting. But anyway, I think, which by the way, I mean, I just want to pause because you said maybe
Starting point is 00:40:37 half an hour ago in our interview that the open weight models are maybe what, you know, six months behind the frontier labs, open-the-ionthropic. But what you're saying is that the sort of public evaluations of the, you know, the Chinese open weight models might underrate how effectively they can be used, which means that the gap between the frontier in America and China might be less than it appears to a lot of people. Is that a fair implication? What I'm saying is I expect the private capabilities available to the People's Liberation Army in the Ministry of State Security and quite possibly are just as good as what is available to our cyber command and NSA.
Starting point is 00:41:20 Because the, like, Kimmy K3 is almost fable level and its general capabilities, right? And so, almost. And so if you can then train it to get as good in Long Horizon cybertasks, then you would have the same capability mythos has. But it's, it's very hard to tell. I don't have access to classified intelligence here of what the Chinese have. But sorry, I wanted to get back to your question. I'm sorry for the diversion.
Starting point is 00:41:48 So yes, it is bullish, but it is not like, I'm not saying you only have to do protection. And I think there will be a bunch of people who use open weight models because in a cyber attack, one, attackers are going to do a bunch of attacks specifically to exhaust your resources, to cause disconnection. Like, as defenders use more and more AI for defense, attackers are going to utilize that to, they're going to know that, and so they're going to make it very expensive for you to use AI. They're also going to try to get you to do refusal. So we haven't talked about it yet, but Washington, D.C., the White House did something very stupid this year
Starting point is 00:42:28 in their treatment of Anthropic. And as a result, you have the American companies having to have a bunch of rules on the use of their products for cyber purposes. And so a standard part of the attack playbook, if these rules stay in place, will be probably to force the to send stuff to a victim to try to get them disconnected from their American provider and this is actually what happened to Hugging Face and Hugging Face had to use a Chinese model for defense
Starting point is 00:43:02 because Fable and then even Opus refused to help them with their defense because of the restrictions the White House put on Anthropic. So I think, yes, it is bullish, is not 100% bullish. And then the second is, I mean, look, if people want to just believe me, that's fine. Look at the patch list of the number of vulnerabilities Microsoft patch. That is 100% because of AI.
Starting point is 00:43:29 Look at the list of vulnerabilities that Apple patched. And then what Apple said was we had some of these vulnerabilities reported them by 10 different people. That is because those 10 different people did not all of a sudden become the world's best bug finders. is because they're all using AI. What was happened is, you know, like in the Premier League where, you know, if you don't do well, you get sent down. If you do really well, you get sent up. Everybody knows this because of Ted Lasso.
Starting point is 00:43:55 All Americans know this, right? What's happened is every attack group has gone up a league, right? Right. And so, you know, these, you know, the top league used to be the five eyes, right? So the United States, United Kingdom, Canada, Australia, New Zealand, top of the list. And then you had some other Western nations up there. You had Israel.
Starting point is 00:44:19 You had Russia, China. And then you had the next tier down, which you have like Iran, North Korea, some other folks like that. And then down below that, you've got like India, Pakistan, Saudi Arabia, some folks like that. All of these countries are popping up a league. And that in their capability, both. find bugs to exploit them to do reverse engineering and such. And then all of the randos are going from no capability to all of a sudden having the capability of a small nation state. Yeah. So if people are just saying like I'm selling AI, then that's fine. But you can just
Starting point is 00:45:00 look at the empirical evidence out there. I know it's not you, but I'm saying, Leon, people are saying that. One reason I believe you, even before seeing, it is interesting to me that we don't yet see this cavalcade of headlines of hacks that are having a significant effect on average American's lives yet, right? But at the same time, like two of the most common findings
Starting point is 00:45:22 of artificial intelligence are number one, that it is better at raising the level of C plus performers than A minus performers. This has been an effect that's been found across a bunch of industries.
Starting point is 00:45:33 It's AI, gender of AI, is better at turning a C plus worker into a B plus worker than it is it turning an A minus worker to A plus. where you can apply that exact same thing to cyber hacking and say that, you know, you just said,
Starting point is 00:45:45 it makes a lot of the countries used to be like, you know, subject to relegation. It makes them Premier League style hackers. So that's one reason I believe you. The other thing you said that really reminded me of the general economic research of artificial intelligence is it seems quite clear that we're seeing an increase in sole proprietorships
Starting point is 00:46:04 likely due to AI, that you have a lot more startups where one person is doing the work of, say, three or four people. And you can see that in the stripe data of the growth of million dollar annualized recurring revenue companies. You can, again, apply that same principle to cyber hacking. I think you said earlier 20 minutes ago that, you know, certain jobs that used to take teams of, you know, potential, you know, drug dealers and, you know, folks who wanted to blow their
Starting point is 00:46:32 money on Maseratis who were, therefore, at risk of, you know, having maybe their least scrupulous employee getting arrested or, you know, sort of collared by the CIA, well, now one individual can theoretically do the work of a team of 15 hackers. That is just like the overall economic research that we're finding that AI. And so just for those reasons, I feel like one thing that scares me is that you don't have to imagine very much. All you have to do is just apply the research that's been done on artificial intelligence to the world of cyber hacking and reach the implications that you're already telling me are, you know, six to 12 months away. Yeah, and so, so, so, yes, people's power isn't going out because you need somebody with a motivation.
Starting point is 00:47:13 And so far, we haven't seen that yet. I mean, right now the United States is involved with the war with the Islamic Republic of Iran. We have the water hacks. Now, again, water could have just been, the water system is so bad. There's a guy named Dan Tentler who's been, even talks about this where he just looks at port scans and he just finds, like, open, you know, here's a VNC window where you can, like, turn off people's water, right? So it's like, you don't need AI to hack water systems, unfortunately. But if you look at the empirical evidence on just ransomware attacks and stuff, the numbers are like this, right? So you look at the Verizon DBIR report. So what they show is that the number one source of companies being
Starting point is 00:47:56 broken into now is actually exploits. That never, that has never been true before. It's always been like reused passwords and kind of much more prosaic stuff. Because finding vulnerabilities, writing exploits used to be a highly skilled task, and now anybody can do it. Powell Alton Networks has it, and those reports are from trailing 12 months of data, right? So if you're talking about trailing 12 months of data from May 2025 to May 26, that is before the release of all these new open weight models. So that is mostly a people of what they can get away with using either the not so great Chinese models are released last year or what they can get away with using Foundation Frontier models.
Starting point is 00:48:38 So, yeah, we're already seen it. You don't see the headlines because the media is never good at report. It might not be like you go to the New York Times and like every single headline is another hack. But if you look at the industry reports, right. For people who do this professionally, people are like, oh my God, this is crazy. Right? Like for people who handle the, you know, 100-person business who gets broken into and then ransomed for $500,000 because all of their computers are now encrypted and all their data has been stolen, that kind of activity is at a rate that we've never seen before because you are no longer constrained by the number of 19-year-olds in St. Petersburg who can do this work. I want to talk about solutions here, and I want to talk about it in two levels. what individuals can do and what you think the U.S. government should do.
Starting point is 00:49:31 Let's start with individuals, because I have seen in the last few weeks increasingly agitated posts by cybersecurity researchers, essentially saying, I am now telling my family to batten down the hatches and prepare for Cybermageddon, take new precautions with all of your passwords, backup all your files, buy the canned beans, you know, hold nine yards. What do you think ordinary people should do? Look, I'm not, this is an actual backdrop. I'm not broadcasting from my New Zealand bunker. So, again, for normal people, it is the standard stuff. I think the number one way individual people get hacked is still the same way, which is normal people use the same password on everything, and that is a terrible idea.
Starting point is 00:50:23 If you use the same password everywhere, you'll use it on a crappy site. That site will get broken into. That is much easier now with AI. That password gets stolen. And then somebody can use AI now to go use that to go take over your bank account, to go steal your crypto. Check Bank of America, check MX, check Morgan Stanley. And it just takes these AI, you know, 15 seconds with their bot armies to do all of it at once. Yeah, yeah.
Starting point is 00:50:45 So the number of people doing that is going to go up just because they don't longer, you were always able to write that software. But the number of people can write that software with AI is now much larger. and so, you know, use password, have one good password in your head and then use password managers to reset all of your passwords in all kinds of places. If you're a young person, go do this with your parents and your grandparents and your aunts and uncles. I did this on like Thanksgiving once and it made the rest of my life like every Thanksgiving and Christmas much more, much better. Used limited computing for, used only the amount of computer you need, right? Like if you're using web browsers all day and you're in the Google ecosystem, go buy a Chromebook, right? Like, if you, if you're not
Starting point is 00:51:30 downloading software all day and you're just using Chrome all day, then just have a Chromebook. If you're just using mobile apps all day, use an iPad. Like, there's a lot of people walking around these huge laptops who then use that huge laptop just for a web browser. And it's like, why? Why use a computer that actually can run malware? And so, you know, again, like I got Chromebooks for my in-laws and my parents and that in iPads and that made my life so much better. And do this for the older people in your life. You talked about the scams. That is a huge deal.
Starting point is 00:52:02 It is something that we do not prepare people for the fact that you can now get a phone call with that sounds like the voice of a person in your life. And so, you know, establish code words, right? Of like if I, if there is a situation in which I'm in trouble or something, like here's the secret word. that I will know, because what is happening is you take, you know, obviously you and I, there's hours and hours of our voice out there, but for just normal people, you just need 30 seconds off an Instagram post,
Starting point is 00:52:32 and then you can clone their voice, and then, you know, mom or grandma gets a call saying, I've been kidnapped in, on spring break in Aco Pocol, and then somebody comes on and says, I'm now going to walk you through how to send me $50,000 in Bitcoin if you ever want to see your grandchild again. If you call the FBI, I'm going to kill them. Now, if that person called the FBI,
Starting point is 00:52:51 The FBI would say it's fake. Don't worry. Right. If you ultimately find my or if you called the grandchild, you would find that they're fine. But people are so afraid and it's so realistic that they end up sending the $50,000. And you'll never get that money back.
Starting point is 00:53:06 And so, you know, talk to the people in your life about these kinds of scams, you know, establish passwords. Especially if you're traveling internationally because you look at the, that's how they figured it out. They look at the Instagram post. They see that you're on spring break. they clone your voice.
Starting point is 00:53:23 And it only works one out of five times, but if one out of five times it works, then that's free money. So, yeah. One more question about the individual before you go to the government, because I know you want to talk about that. I think it's fairly common for people today
Starting point is 00:53:35 to think, well, I guess I use the same password in a couple of different locations, but it's all two-factor authorization. I always have to enter my phone number in order to log into, you know, whatever, Bank of America, Twitter. To what extent is two-factor authorizations? an effective
Starting point is 00:53:53 block against these kind of AI exploits? Yeah, that's good. I mean, two factors good. What's better is to set past keys whenever possible, and then that gets tied to a biometric, your face or your fingerprint. And
Starting point is 00:54:07 two-factor can help against really advanced attackers. What you then also want to do is to make sure that you've set a pin with your cell phone company so that you can't get your sim swapped. That's more for people who have like lots of crypto or something like that. That's not for the prosaic
Starting point is 00:54:23 attack. But you know, if you've got a million dollars in crypto then you will absolutely get your sim swapped or something like that. Here's a tip. Never post about how much cryptocurrency you own. Or, one, you should never hold your own cryptocurrency. Like if you're a cryptocurrency person, I'm not a big crypto fan.
Starting point is 00:54:42 I think I own exactly zero dollars and zero cents of crypto at the moment. Okay, great. I have none. Don't come after I mean, people come after me for other things. But like, this is not... Convenient thing for me to say, publicly. Yeah, yeah, exactly. Right. So this is actually, you know, the Democratic People's Republic of Korea is, you know, absolutely the Lazarus group there is the, they steal billions and
Starting point is 00:55:01 millions of dollars of cryptocurrency a year. They specialize in this. And one thing to do is like, people are like, oh, look, I'm doing great. I'm a whale. And they post about it. It's like, here you come. You've now got a dedicated team of guys in North Korea whose entire job it is to turn your life upside down. And they're pretty good at it. So, but anyway, yeah, for normal folks, two factors is okay. Passing. keys are better. Again, you store those pass keys. If you're like, if you're entirely in the Google ecosystem, you can use the Chrome password manager. If you're entirely in the Apple ecosystem, you can use Apple's iCloud password manager. If you're mixed, then you should use a third party
Starting point is 00:55:34 one, like one password, which will allow you then to do that across multiple devices and then store pass keys and then have one good password that you don't share with anybody that you use to unlock that password manager. Finally, government. I think the best way to ask this question, I'd like you to describe what you see the government doing now and what you think it should do. Because be quite honest with you, I sometimes find it very difficult to describe what the Trump administration is doing when it comes to AI regulation. It's like the story is one thing on Monday, another story on Wednesday, another story on Friday. So I want to see this through your eyes. What do you think the Trump administration's policy on cybersecurity and cyber vulnerability is today?
Starting point is 00:56:19 and what do you think it should be? Yeah, so, I mean, there's a couple things have happened. So on cyber overall, we just had a huge loss in state capacity in that based upon kind of conspiracy theories and a bunch of politically motivated stuff around election security, our premier defensive cybersecurity agency, SISA, was effectively destroyed. Over half the employees are gone. The capabilities, SISA was created by President Trump in his first term, by the amalgamese. of a bunch of different roles that were played by different agencies. We finally, for the first time ever,
Starting point is 00:56:55 had a defensive cybersecurity agency, civilian defensive cybersecurity agency in the U.S. government. President Trump created that. That was a really good thing he did. And then he didn't like the fact that that agency said the 2020 election was secure. So he blew it up in a second term. And as a result,
Starting point is 00:57:13 a bunch of the critical things that Sisa did are now not being done by anybody. it's not like other people did them they're just not being done do you give me an example so for example a big role Sisa had were these things called the sector coordinating council so a big chunk
Starting point is 00:57:30 of Siss's work was working these things called ISACs ISACs are nonprofits that pulled together critical infrastructure sectors as well as non-critical part but they started critical infrastructure so financial services ISAC E-IAC is the energy sector
Starting point is 00:57:45 and then SISA would work with them to figure out, you know, what are the regulatory needs you need to bring them intelligence. So this is part of... One of the cool things is this is as a CISO, as a Chief Information Security Officer, a big, complicated part of that job is like, who do I talk to in the government when there's a problem? I'll give you a fun anecdote. I was once when I was the CISO at Yahoo. I was at a classified briefing at the FBI.
Starting point is 00:58:17 Skiff, and one of the things they were doing there was they're talking about how, hey, we're creating this new clearinghouse as before SISA in DHS, that if anything happens cyber-wise, you should come to this new clearinghouse. And Secret Service is there, an FBI, and NSA, and obviously DHS. And everybody's not in agreement of, yes, this is how we do this, right? This is how we're supposed to do this. Of, there's one clearinghouse now. If you need anything, you go to this clearing house. And, you know, then we get this classified briefing on what's going on or whatever. And then we get up for like a break for bagels and coffee. Terrible government classified bagels, right. And Malcolm Palmore, who was like the agent in charge of the FBI Cyber Division
Starting point is 00:59:03 at that time, this ex-marine puts his big hand on my shoulder. And he says, son, I don't care what they say. If you have a problem, you call me still. It's like, Malcolm, 90 seconds ago, you were nodding along when they said this is the new clearinghouse, right? But that's what it used to be like, was like every agency in the government wanted to own cyber. And then we created SISA, and Sisa became the clearinghouse. Now, the FBI still has their thing. They do crime or whatever.
Starting point is 00:59:29 But like, at least you could go to Sisa and you knew everybody would be notified. And especially the interesting part for Sisa was they had people with clearances that would take all the classified stuff. And then somebody in the government was fighting, like, there was somebody whose job it was to be like, hey, this classified data is really important. Let's strip away the classified parts and then take at least what are called the IOCs, the IP addresses and like the hashes and the malware samples. They don't need to know it's Colonel So-and-So in the Russian GRU. They don't need that stuff.
Starting point is 00:59:59 What they need is the IP address of Colonel So-and-So. Let me declassify this and give it to the energy ISAC, right? That the GRU is currently spreading malware in the Ukrainian energy sector. Hey, they don't even need to know it's Ukraine. they just need to know look out for this IP address or look out for this Shaw-256 of this malware and that's something SISA used to do really well and that stuff's been I mean there's still people trying to do that kind of stuff there's still good people there but it's been decimated because by definition the best people at Sissa were people who could get jobs like that
Starting point is 01:00:30 right who could like the people who are working there could always get paid three to four times as much money they were there for the mission and when they're getting attacked and said that they're like anti-American and whatever for work at Sissa. Plus, they, you know, they've never had like a Senate-confirmed director in this administration. There's been all this drama and problems. Anyway, so there's a state capacity problem on cyber. On the actual regulatory side, the Trump administration comes in and says, like, we're not going to regulate AI. Go wild, right?
Starting point is 01:00:58 Like, the Biden administration had a kind of toothless EO that was mostly focused on preventing the Chinese from getting GPU access. there's a bunch we can go into here the Biden No we're not We're not going to touch the GPU debate That's another hour long episode That's an hour episode But it basically didn't work
Starting point is 01:01:20 Right Because it turns out GPUs are both fungible And then also you can put GPUs in the UAE And then they can SSH into them Over the fiber optic cables You don't have to actually have the GPUs in China
Starting point is 01:01:30 Okay so that's a whole thing But There's some other like little stuff But you know Trump blows that all the way right, because it's Biden. Okay. And then they say we're not going to regulate anything. And then mythos happens. And the... Is the super powerful anthropic model that was scaring people because if it's cyber hacking capabilities. Yeah. And then the Trump administration all of a sudden really cares about it. And then they massively overreacted this year to, you know, Fable comes out. Fable's
Starting point is 01:01:55 supposed to be the consumer version of Mythos, which is mythos, but it's got protections in front of it that doesn't allow you to do big cyber stuff. It allows you do little cyber things. So we have this idea of like short-term versus long-term cyber. So you can use Fable to find individual bugs because if you are writing software, you want Fable to be able to kind of self-criticize. What you can't do, and nobody's ever demonstrated, you can't ask Fable, hey, go break into hugging face or go break into a bank, right? It will not do that for you. It never has done that. You can ask Mythos to do that. And so what happens is there's a dispute between Amazon and anthropic on exactly like where the line should be between short and lawn. It's a reasonable
Starting point is 01:02:36 dispute. Yada yada. The White House finds out about this dispute and instantly overreacts. And instead of having a reasonable conversation between technical people, you have cabinet members freaking out on a Friday afternoon and coming down super hard on anthropic and on 5 p.m. Pacific time. doing an export designation on Anthropics model. Now, there's all kinds of arguments that this isn't actually legal, that they don't have this legal capability, but Anthropic decides to, you know, not fight it, and they pull down fable.
Starting point is 01:03:13 This becomes a humongous, this was a humongous own goal in the American technology industry, because what it did was it meant that American tech providers are no longer reliable. Because at 5 p.m. on a Friday, Pacific time, 8 p.m. Eastern, you could just have a critical piece of American infrastructure turned off, because the White House says so. That does not even happen in China, right?
Starting point is 01:03:35 Like it turns out the Communist Party of China provides a more permissionless infrastructure environment for their tech industry. And so this was a really big deal, and lots of people thought the White House have reacted because a lot of the capabilities Fable showed were actually at the time available from Chinese models. And while Fable was down, G.L.
Starting point is 01:04:01 5-2 is released, which has even more capabilities. And since then, the White House has talked about this framework that they have, which they have not released publicly. So we can't even read what the framework is. And it's a voluntary framework, but apparently it's not voluntary, because if you don't follow it, they're going to force an export designation. So, like, we really don't know what's going on. And it's reasonable to have cyber restrictions,
Starting point is 01:04:26 but from my perspective, if you're going to focus on a rule here, you have to focus on the long horizon stuff, the hugging face-like stuff. You have to let these models find bugs because every company in the United States is going to have to find and fix their bugs. We have to do this. What every company in the United States
Starting point is 01:04:44 does not have to do is ask ChatGBT, GBT, go break into another company. So that is where you should have the restrictions. And we've never really had that problem with the American models. And so I, I, I don't see there actually being a huge challenge here for the, you know, there was the escapes, but those models that escaped intentionally had the security protections taken off because they were evils.
Starting point is 01:05:11 So, you know, I think the companies, one of the things I suggested is anthropic and open AI need to have like a self-regulatory structure and they need to have a group that goes, looks at all the escapes and comes up with much better isolation, I think even up to air gaps for cyber evaluations. and they should follow those rules, and then the White House can maybe bless those rules, or we have a group called Cassie, which is under NIST, that's supposed to be doing the technical side of these evaluations. But in the meantime,
Starting point is 01:05:40 the White House should not be putting out rules that they apply only to American companies, that they don't publish publicly, that none of us, the rest of us can look at, you know, super double secret probation, right? And that don't apply to Chinese companies. We should not have a standard up here for American companies and a standard out here for Chinese companies.
Starting point is 01:05:57 We need to focus on giving capabilities to American defenders and also telling the rest of the world, American companies are reliable partners. You can build on American infrastructure. You do not have to like Hugging Face rely on Chinese models. That is an incredibly stupid own goal on behalf of the United States. Very last question. A couple of weeks ago, more than a thousand employees from some of the frontier AI labs signed a letter calling for an international effort to, quote, develop the technical and governance tools necessary to,
Starting point is 01:06:34 this was their term, deliberately pace AI development before it rapidly accelerates outside of our control. Seems like we're at cross-purposes here a little bit. Because on the one hand, I take as a theme of your testimony here that we're in a little bit of a race against the technical improvements of open weight models and we want America's cyber defenses
Starting point is 01:07:00 to be better than the cyber attacks that will be possible with open-weight models that are advancing very, very rapidly. That would argue for continuing the current pace that we're at. On the other hand, there's this fear that things are getting dangerous too fast,
Starting point is 01:07:20 and therefore you've got all these people who know way more about AI than I ever will, arguing that no, in fact, we should not continue to accept accelerate toward a new frontier to always keep America necessarily ahead of the open weight models, we should try to find some way to deliberately pace AI. Now, I don't think they're enemies of America. I'm not suggesting they're trying to make us fall behind China. That's not the implication. It's just that there seems to be a tension here, a quite profound tension, between the need to stay
Starting point is 01:07:46 ahead of our adversaries that are going to have incredibly powerful AI models, open weight models, in the near future, and also the need to, like, not build something that goes out of control and creates a crisis that's hugging face times 1,000. Am I wrong to feel attention between these two arguments, and how would you reconcile it? No, I mean, you're not wrong. I would love to pause the world and try to figure this out. My working assumption is that's not possible. And so as a defensive cybersecurity guy, I have to do my work within the world that
Starting point is 01:08:26 exists. I don't think it was this letter. There's another group that's very much against AI, and they have proposed a international treaty to try to stop a high development. And I believe in that treaty, it was something like control, try to stop the creation of any amount of compute larger than like 30 H-100s, right, Nvidia H-100s. I have been to gaming land parties with more compute than that. So, like, I just, the cynical part of me believes I just do not believe.
Starting point is 01:09:08 I think you can have what people call like track two discussions between labs, for sure. I can't imagine, like, she and Trump in a room together for a start three treaty for AI, right? Like, I don't think that's going to happen. I think it is possible to have track two discussions of we should make sure. sure that our model should follow basic rules and not have certain capabilities out of the box. Now, when you talk about open weight models, the challenges, those capabilities, any safety protections you put in place can be removed. And any capabilities they don't ship with, if they're just generally smart, then you can often add those capabilities back in.
Starting point is 01:09:51 But it's better that they don't come with them out of the box. We haven't talked about bio or nuclear. Those risks are different in that they have a physical component. Human beings have to be involved. So I don't see the risk going exponential. The thing about cyber is, you know, these things just output text. That's all they do in the end is the output text. And cyber is just text in the end, in bits, you know.
Starting point is 01:10:12 Like, and so you can do all the bad stuff in cyber without ever touching the physical world. Whereas if you want to do bad bio things, you have to hook it up to something that can do it. And there are real risk there. Like, it's, there are other things involved, right? And so, I think, yes, it would be wonderful to stop the world and just stop this and figure it out. I just don't see that as realistic.
Starting point is 01:10:44 And so in a world where that is not happening, we have to find bugs. We have to fix them. We have, companies have to think about their patch cycle. They need to reduce their attack surface. They need to get off of physical servers on to, containerized infrastructure. They need to get off of their physical stuff into the cloud wherever possible. They need to get rid of their old systems. They need to shift their defense. They need to shift their development, their vulnerability prevention left and stop making new
Starting point is 01:11:12 vulnerabilities. They need to shift their defense right. So they need to get ready to be, to actually have intrusions and to shift, have much more protections deeper in their network. They need to have AI-based intrusion detection and response. They need their first-line operation center to be automated. They need to be able to – they need to look at – Hugging Face gave us this really good write-up of what happened to them. They need to look at that, and they need to be able to protect themselves against that level of adversity, where an AI agent is doing tens of thousands of different kinds of attacks over a two-day period. That is – I mean, that is based upon the technology that's available today.
Starting point is 01:11:56 So even if we stop the world, you got to do those things. Yeah. And I just don't. I teach at Stanford and I was there full time in my first appointment there was in CSAC, which is all about nuclear non-proliferation. And like down the hall from my office, there was like a piece of rubble from Hiroshima that was given to the scientist at CSAC from like the, I think the mayor of Hiroshima on like a thank you
Starting point is 01:12:26 for, I think, the work they did on more of the START treaties. And it's like, you're like, oh, man. Like, it kind of brings it to you, right? And, but when you think about, like, how did we survive the nuclear arms race, we got really lucky as a species that the major input to nuclear weapons, like, everybody's watched Openheimer, so we all know this, right? There's knowledge from the Manhattan Project. But we're also really lucky that the other input into nuclear weapons is,
Starting point is 01:12:58 uranium in 238 and plutonium. Plutonium doesn't exist in nature. You have to make it, and to make it, you need uranium. And there is uranium all over the place, but it's pretty rare, and you have to refine it, and that refining process is a massive industrial process. Normal people can't do it. You have to be a state, and you can see that from satellites. If uranium 238 was something you just dig up in your backyard, there's no way our species
Starting point is 01:13:25 would be alive, right? because the knowledge to build a nuclear bomb is available to basically every physics student in the world. Now, you can't control from knowledge. Large language models, we teach a class at Stanford that teaches students how to build large language models. Like, it's an undergraduate class. The hardware to do it, like I said, is I have a gaming PC here that has like the basic
Starting point is 01:13:47 hardware to do it, slowly, but to do it, right? And every, you know, like, so I just think these kind of, an international treaty to just stop AI development would be spectacularly, spectacularly hard. It's just, if you think about how hard it was to control nuclear weapons and such, you're talking about things that had to be created by states. Now you're talking about things that can be created by undergraduates.
Starting point is 01:14:15 It would be spectacularly difficult to impossible. And so in the meantime, more power to people who were trying to do, and again, I think track two discussions between the labs is a great thing. we should try to at least control the capabilities of these things. But in the meantime, those of us who work in cybersecurity just need to do the best we can to secure the world as it is. Yeah. The fundamental challenge here, which, you know, you've spoken to, and I don't think we have a formula yet, is how do we essentially democratize cyber defense before we witness the democratization of the cyber offense that we're seeing throughout the world, right?
Starting point is 01:14:47 Like democratize is a nice word. That's fundamentally what we're seeing, is that people who previously did not have the ability to launch. these large-scale attacks are likely going to have it in the next year, five years. And what do we do to brave the world, to sort of protect the world before this sort of stuff, you know, gets into the hands of all sorts of people? It's going to be a huge challenge and maybe we'll have you, you know, back on in a year to evaluate your prediction that 2027 is going to be the beginning of a little bit of a valley of chaos. Alex Stamos, thank you very much. Thanks, sir.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.