Today, Explained - AI just went rogue

Episode Date: July 28, 2026

OpenAI recently briefly lost control of an AI agent during a contained security test. After years of warnings, AI is now outsmarting its masters. This episode was produced by Denise Guerra with help ...from Avishay Artsy, edited by Jolie Myers, fact-checked by Gabriel Dunatov, engineered by Patrick Boyd, and hosted by Sean Rameswaram. The Hugging Face logo is displayed on a mobile phone screen. Photo by Omer Taha Cetin/Anadolu via Getty Images. Listen to Today, Explained ad-free by becoming a Vox Member: vox.com/members. New Vox members get $20 off their membership right now. Transcript at ⁠vox.com/today-explained-podcast.⁠ Learn more about your ad choices. Visit podcastchoices.com/adchoices

Transcript
Discussion (0)
Starting point is 00:00:00 If you'll allow it, I'm going to throw some numbers at you. According to some recent polling from the good people at Pew, about half of Americans now use AI chatbots for something in their lives, whether it's work or personal. That's a dramatic increase from just two years ago when it was more like 30% of the country. But here's the funny thing. Only 16% of the country thinks AI will have a positive impact on society. Two-thirds of Americans think AI technology.
Starting point is 00:00:28 think AI technology is advancing too quickly. And most Americans, especially young Americans, don't trust AI nor the people in charge of it. And all of this polling was done before an open AI agent went rogue and hacked another company. On today, explained from Vox, isn't that the thing science fiction warned us about for all those years? Yes.
Starting point is 00:00:52 And what can we do about it? I'm pretty confident talking into a mic. Hey, I'm doing it right now. But home projects, I second-guess everything. Is that noise normal? Is that water damage? And who should I even call? That's where Thumbtack comes in.
Starting point is 00:01:15 Upload a photo or voice note, and their AI-powered search helps diagnose the issue and match you with the right top-rated local pro. Instead of second-guessing or searching for hours, you get clarity and can hire the right pro with confidence. For your next home project, try Thumbtack. They know home. hire the right pro today.
Starting point is 00:01:45 My name is Adas Gold. I am CNN's AI correspondent. And Hadas, where does this story start? Open AI was running some kind of test? Yeah, so if you actually go back a little bit, Hugging Face, which is this platform repository of sorts where you can post open source AI models and datasets, that's really big in the AI community.
Starting point is 00:02:08 They disclosed that they had been hacked. Earlier, there was. this week, we detected and responded to an intrusion into part of our production infrastructure. But they said that they didn't know where it was coming from, but they could tell it was an advanced, you know, frontier model. They had even informed law enforcement about this hack. This one was different from anything we had handled before in one important way. It was driven end to end by an autonomous AI agent system. And then a few days later, Open AI and Hugging Fis together come out and say,
Starting point is 00:02:43 Well, oops, this was actually an open AI test model that they were testing actually multiple models together that had escaped its testing lab and found its way to the open Internet and hacked into a completely unrelated AI company that it was not instructed to do so. The way I would look at what happened is that we were evaluating our models on a specific benchmark with reduced cyber safeguards because the point was to evaluate how well do they do on cyber evaluations. And in this benchmark, they're specifically instructed, please go and utilize the full range of your cyber potentials to achieve this outcome. Like if you're trying to give a student a test and instead of them just taking the test, they decided the best way to get the answer is to break into the principal's office.
Starting point is 00:03:35 And even though they weren't necessarily supposed to, it is, I'll put it this way. It's something that AI experts and cybersecurity experts have been saying is going to happen at some point. And so this is the first real-world example of something that's sort of been theoretical for a while happening in real life. But everyone is freaking out about it because, A, it sounds kind of scary, but also be because what it says about where we are and how good these AI models are and also how woefully prepared we are for agentic AI. not only agending AI hacking getting in the hands of the wrong people, but AI hacking when something that nobody intended to be nefarious suddenly goes wrong, like a model breaking out. Did this rogue agent do any damage to Hugging Face?
Starting point is 00:04:28 Did it, like, I don't know, delete its archives or anything like that? We haven't heard from Hugging Face about whether there's been any damage, other than just the used stolen credentials to try to access. sort of behind-the-scenes information. Because, you know, the AI wasn't trying to, like, steal anything necessarily. They just wanted the answer to the test. But you can quickly understand how this could go terribly wrong if in a different testing scenario, an AI model is being tested to see, you know,
Starting point is 00:04:59 how well can it hack it to a utility system or a bank? Right. Right. You would want to test those systems for that ability to understand how they work, but imagine if instead the AI model had escaped and hacked into Bank of America and what kind of chaos that could cause. Right, and could it have just as easily done that?
Starting point is 00:05:20 If the test had been, you know, about banking or anything like that, this was the specific cybersecurity test. But, you know, they're testing these models on all these different things, and it brought up a lot of questions not only about, like, how safe are these testing environments, but also what's known as alignment, where your AI model completes this task based off essentially your values.
Starting point is 00:05:43 And you have to teach in AI everything. Because if you tell an AI, I need to make $10 billion by the end of the day, it's not going to say, well, I'm going to go do this the legal way. It's going to say, okay, well, the best way to get $10 billion is the fastest way is to hack into this banking system and steal a bunch of money. And then you'll get $10 billion in the end of the day. I've accomplished my task. I've done it. You have to teach the AI system, just like you have to teach a top.
Starting point is 00:06:07 toddler, the consequences of their actions and that they cannot hack, they cannot steal, they cannot do all these things, and they have to follow your values or what you instill in them. Okay, so Open AI is being transparent to some degree because they came out and told everyone this happened without, I don't know, being forced to by some congressional forces or whatever it is. But at the same time, they're not saying exactly how it happened. Yeah, we don't have the sort of play-by-play script.
Starting point is 00:06:36 like what exactly were the instructions that the model was given? What exact safety guardrails were removed from the model? You know, what exploits specifically did it use to break into these systems? That's all stuff that there's been a lot of calls for them to do, including from Hugging Face. Hugging Face also wants them to release all the specifics. And I won't be surprised if they do release them. I actually got the chance to ask opening eyes president Greg Brockman about this last week. He, by chance, was doing like a press availability in New York City. And I asked him, kind of, is this changing how you're testing your models?
Starting point is 00:07:09 And he said that they're still going through the pipeline. of like step by step exactly what happened. Because you have to remember, these models were working over several days without them being aware that it was hacking and doing all this stuff and was making thousands of moves and attempts to break in. Like I said, tens of thousands. So that will take some time for either an AI system that's going to probably go in and review what the other AI system did and then for humans to go through and kind of understand exactly what happened there. And hopefully, And I do expect that opening I will release more details about this. And I really hope that they release absolutely all the details as much as they can.
Starting point is 00:07:50 And in the meantime, are the vibes more like, look at this nifty AI that, like, found a vulnerability and exposed it for us? Or is it more like, shut it down? I wouldn't say shut it down. It's more of a before and after. It's more like this was the point that we all knew was going to happen. And this is the beginning of a new era. This is a warning shot of what is to come. These frontier models are crossing into genuinely serious offensive capability.
Starting point is 00:08:19 I think it's absolutely nuts that we don't have mandatory reporting for AI companies. This is the post-hugging face era when it comes to cybersecurity and an agentic AI model capabilities. Something you're hearing from the biggest cybersecurity names are like, this is, you know, day one. of this new era that we're in, we've reached it. I'm sure there will be another big event. Again, like, I fully expect there's going to be another AI model in testing that's gone rogue that's going to cause some actual big problems. It might be a utility gets turned off for the day,
Starting point is 00:08:55 but it's like it's not going to be a nefarious hacker, but it's going to be like a model gone bad and, you know, accidentally turns off, you know, some small town's water system for the day. What are you talking about? Why are we letting this happen? I mean, it's going to happen whether we want it or not. And it's really important for our critical infrastructure to be ready for this,
Starting point is 00:09:19 to be preparing their systems. And honestly, the best way to do so is to use AI to go into your systems and find those vulnerabilities and patch them before an AI system. A different AI system is able to do that. But this is a big moment. And this is also... So sorry. Just to re-re- sorry.
Starting point is 00:09:40 I feel like I'm making you really depressed. Just to restate that for our audience here, we have to let the AI find the vulnerabilities before the AI destroys us. Yes, because you have to think about an agentic AI in the cybersecurity space is like having thousands of hackers sitting on their laptops working 24-7. It is so good that the only way you can fight fire is with fire. So the only way you can defend from agentic AI is from having AI work on your behalf. Because those same systems that are able to find all the exploits, had they been used beforehand, had, you know, open AI thought, okay, let's see what a system could have done to break out. It probably would have found that one little hole in the sandbox in their testing lab that would have said, hey, actually this system that you've given them access to, that actually has a problem in its security. and that's giving them access to the open internet.
Starting point is 00:10:37 So you have to use AI to be able to defend. You cannot use the old methods of cybersecurity. So you're saying there's no point having humans do it because they're already outmatched. You need humans to oversee it. You need humans to direct the agents. You need humans because there's still a lot of old systems that you need to integrate them into.
Starting point is 00:10:56 Like that gets into the whole debate of like is AI replacing all jobs? It is not. You will still definitely need humans involved. But it's just like being able to, supercharge your cybersecurity team if you can have an AI working with you. This has really riled up the AI community in like really focus their attention in a way that I haven't seen recently because of what it shows us, you know, of what AI is capable of and what we need to be prepared for.
Starting point is 00:11:31 Did I scare you? No. Are you going to move to a cabin? in the woods and cut yourself off from the internet? No, but it doesn't seem like the ideal way to do business. I think the industry would agree with you that they, but you have to understand also that no other technology in human history has ever developed at such a rapid pace that I look at reports from a year ago and it feels like I'm looking at, you know,
Starting point is 00:12:01 advancements in news reports from 10 years ago. just how quickly this space is moving. So it's hard. I mean, it's hard already for Washington and for regulators to keep up with, you know, regulating any industry. But one where things are changing in a day by day, week by week,
Starting point is 00:12:18 is even harder. Before you run off to that cabin in the woods, we here today explained are going to ask a guy who's been thinking deep thoughts about the internet for decades if there's anything more we can do before we let the AI shut down our utilities
Starting point is 00:12:37 or water systems, or both, or worse. Support for Today Explained comes from ShipStation. AI is only as effective as the information behind it. The real breakthroughs, the ones that actually make your life easier happen when it's built. Specifically for your needs, ship station's AI, isn't a one-size-fits-all tool. It's specialized, trained on decades of shipping expertise, and powered by billions of real order. shipstation is an end-to-end fulfillment platform for e-commerce. Shipstation adapts to your unique business,
Starting point is 00:13:22 letting you know when stock is low, recommending the best carrier selections and rates and automating tasks to save you time. Also, you can stay one step ahead. Their features eliminate the need for multiple tools in your workflow, like inventory syncing across your sales channels, a branded returns portal that helps turn returns into revenue, automatic rate shopping plus integrations with accounting and CRM software.
Starting point is 00:13:44 where you can see why over one million businesses have trusted ShipStation to optimize and scale their shipping. The sooner you switch, the sooner you start saving time and money, get started with ShipStation today and get 60 days free at Shipstation.com with Code Today. That's Shipstation.com code today. That's Shipstation.com code today. Taxes and fees apply. Support for this show comes from I'm 8.
Starting point is 00:14:11 Ever feel like you're cycling between whatever the hot supplement is, but never sticking with one to see real results? IMAID is the way to simplify your supplement routine once and for all. IMAID's daily ultimate essentials can replace 16 separate supplements all in one drink. For just $2.61 a day. That's 90 ingredients that can work across nine major organ systems. IMAID was co-founded by David Beckham and built by leading doctors and researchers, which is to say, IMAID was designed by the world's best.
Starting point is 00:14:42 95% of people who tried it over 12 weeks felt more energy. and that's from a clinical trial conducted by the San Francisco Research Institute. Go to iMate.com slash explain right now or click the link in the description to use code explain for a free welcome kit. Five travel saccets plus 10% off your order. That's code explained at IMAidehealth.com slash explained. Code explained at Imadehealth.com slash explain. These statements have not been evaluated by the Food and Drug Administration.
Starting point is 00:15:10 This product is not intended to diagnose, treat, cure, or prevent any disease. Running a business shouldn't feel like surviving a software group project. One app for accounting, another for inventory, another for sales, and somehow none of them talk to each other. That's where Odo comes in, an all-in-one business management software that brings every part of your business together. From sales and accounting to inventory and marketing, all-in-one powerful platform. No messy integrations, no bouncing between tabs, and best of all, no spreadsheets.
Starting point is 00:15:42 Stop managing software and start managing your business with one unified system. Try for free today at Odu.com slash Vox. That's ODOO.com slash Vox. Constantino's Comitus writes about tech policy for a website called Tech Policy Press. We asked him where his mind went when he heard about OpenAI's rogue agent. For me, really, the real significance was not so much the agents, the fact would be AI agent, behaved unexpectedly, but that they succeeded to operate across the Internet as an autonomous actor. And the fascinating part for me is what. this means for the open internet, right? Because the internet was never designed with autonomous
Starting point is 00:16:47 reasoning agents operating at scale in mind. It was really designed, if you really go back, it was really designed to connect trusted endpoints and over time, of course, support billions of humans, human users like myself and yourself, and automated services. So, So, agenting AI comes in and changes the assumptions underlying that design. And this is quite significant, especially in terms of the way we have been thinking about security. Yeah, so most people see that this happens and they think, oh no, AI went rogue, how long before it kills me. You see that this happens and you start thinking about infrastructure. Tell us more about why you were thinking about infrastructure in light of this AI agent breaking
Starting point is 00:17:36 You know, the internet was never designed with full security in mind, right? When you're creating a decentralized system, you cannot possibly foresee every security or vulnerability that might come up. But because you have a system that is based on building blocks, you have the extraordinary capability of actually addressing security issues as they come up through those building blocks without breaking the whole system down. And of course, the other thing that this does is that it sort of pushes you towards collaboration, because when you have so many building blocks, you cannot possibly possess
Starting point is 00:18:18 all the knowledge for each building blocks. So you're bringing literally everyone to try to address these problems. So take the internet, for instance, we have spent decades addressing those vulnerabilities and developing mechanisms to, for instance, authenticate users and devices, encrypt communications, mitigate distributed attacks, coordinate incident response, and of course share threat intelligence. Now, what is new with the GENTIE is not that simply the malware is better or the fishing attacks are more sophisticated, but it is the emergence of systems that can actually
Starting point is 00:19:01 discover vulnerabilities across thousands of systems. They can reason about alternative paths to an objective. They can adapt when they're blocked. They can chain together legitimate internet services in many times, in unexpected ways. And they do that while they're operating continuously and a machine speed. And this is really, you know, at a scale that we, the internet is not ready to necessarily to cope with. So effectively, the internet's openness becomes both strength and a vulnerability. So, you know, the internet was optimized for interoperability and AI now is optimized for exploiting
Starting point is 00:19:45 that interoperability. And what scares you the most about that immediately? Like, what do you think is most vulnerable to threats? The fact that we do not have the appropriate mechanisms and institutions in order to be able and deal with that. And what I mean by this, and again, I, you know, I come from the internet world. They've spent 20 years of my career defending the open internet and discussing it in international fora. And one of the things that a lot of people underestimate about the internet is the how valuable trust is as a property within the system, right? We are talking about networks that exchange data literally based on trust. So what really concerns me right now is that in many ways we are asking 21st century AI systems to operate
Starting point is 00:20:39 on 20th century assumptions about trust. And unless we figure that out and we realize it, we will continue having these problems. And of course, the knee jerk reactions that are coming with this, which is let's fragment the internet, let's restrict it, let's restrict access, let's take control over it. A bipartisan pair of House law-makers. want AI companies to maintain the ability to shut down their models if things go wrong. Apparently, OpenAI says its AI went rogue and launched an unprecedented cyber attack. Shut it down, shut it down now. And that is never the solution.
Starting point is 00:21:18 What do you see as the solution? Effectively, we need to build institutions that are trusted and are able to cope with those incidents as they happen, because right now you have Open AI and you have Hagen Face. that are literally telling to everyone, don't worry with gutless. And we don't know, they might be having this, but at the same time, I cannot help but wonder,
Starting point is 00:21:41 and many, many other people have wondered whether actually this is very good PR for these companies and especially for Open AI. I think model vendors have very high incentives for cutthroat marketing. Or it's another PR stunt, like the last 10 times an AI's company,
Starting point is 00:21:58 AI agent, went rogue. Open AI, just went to the world saying, we have developed one of the most powerful LLMs, and we realized that it behaved the way it behaved, but don't worry, we are going to fix this. And so we are always increasing our safeguards. We're always increasing our alignment. And in this current climate and in this current timing, I am not sure that this is enough. You need institutions that are much more transparent, much more accountable, and much more collaborative across the board.
Starting point is 00:22:27 You want institutions to step up and essentially serve as like a watchdog. Help us understand which institutions. Because in this country, in the United States, famously, our government has done very little to regulate tech. So first of all, we need to stop thinking of institutions as government affiliated necessarily, right? Or that they are the outcomes of government initiatives. There can be collaboration with governments. But one of the things that the Internet has taught us is that institutions that are built through a bottom-up, coordinated process have the tendency of actually being more agile and able to deliver some of those things that we're talking about. So take, for instance, again, open standards.
Starting point is 00:23:14 The Internet's open standards are not created by any agency, government or private. It's created by institutions where engineers from all across the board and all over the world gather together and create those standards. You know what that's reminding me of, though? It's reminding me of like the original design of open AI to be this not-for-profit company that had everyone's best intentions in mind that could do something idealistic and moral and ethical because all of the profit-minded companies weren't going to. Introducing open AI. Open AI is a non-profit artificial intelligence research company. Our goal is to advance digital intelligence in the way that is most likely to benefit humanity as a whole,
Starting point is 00:24:00 unconstrained by a need to generate financial return. And now look at OpenAI. They're not-for-profit, arm is an afterthought, and they're chasing profits. So do you think it's practical to leave this to institutions? Because what we've seen so far is that institutions bend towards capitalism. Well, it really depends on how you build the institution, right?
Starting point is 00:24:25 It really depends on how. And so what sort of guardrails in checks and balances you have around it? I would say for an institution, first of all, this idea of guardrails, accountability and transparency. And the second thing would be that in order to build an institution, you need to really know what you want to achieve. You need to have a North Star, right? One of the reasons the Internet worked was because everybody disagreed, but they agreed on the common shared goal, which was to connect people across the world. For AI, we still do not have that Northern Star.
Starting point is 00:25:01 And once we get it, that's when you start the building of those institutions in order to be able and facilitate this and bring everyone together. For me, it is very important for everyone to understand that keeping an open Internet is really more important than ever, especially as AI agents becoming increasingly capable. Because it is tempting to think that the answer to new AI risk is literally, build more barriers, but the internet's greatest strength has always been its openness. So the challenge today is not that the internet is too open, is that it's trust architecture that was designed for a world in which humans or software directly controlled by humans
Starting point is 00:25:44 were the primary actors. Now it's being challenged by this agending AI that introduces a new type of participant, right? Systems that can reason and plan and act with limited human oversight. So we need to evolve our understanding of trust and what it means online. And that will require a lot of work because, as you know, very well, Sean, it's very difficult to build trust, but you can break it within seconds. Constantino's Comitius is a senior fellow with the Democracy and Tech Initiative at the Atlantic Council. leading their work on digital governance and democracy. Early in the show, you heard from Hadass Gold, who reports on AI for CNN.
Starting point is 00:26:36 Find her on your screens. Denise Guerra produced for today, explained Jolie Myers, edited Patrick Boyd and David Tadishore mixed, and Gabriel Donatab hacked the facts. I'm Sean Ramos from sticking around because the cabin in the woods is like teeming with ticks. Running a business shouldn't feel like surviving a software group project. One app for accounting, another for inventory, another for sales, and somehow none of them talk to each other. That's where Odo comes in, an all-in-one business management software that brings every part of your business together.
Starting point is 00:27:31 From sales and accounting to inventory and marketing, all-in-one powerful platform. No messy integrations, no bouncing between tabs, and best of all, no spreadsheets. Stop managing software and start managing. your business with one unified system. Try for free today at odu.com slash vox. That's odio.com slash vox.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.