Breaking News from Pod Save America - Lovett Geeks Out Then Freaks Out Over Insane AI Story

Episode Date: August 8, 2026

AI is getting smarter, stupider, and somehow more dangerous. Jon Favreau and Jon Lovett get deep about AI after a truly insane new development. Learn more about your ad choices. Visit megaphone.fm/adc...hoices

Transcript
Discussion (0)
Starting point is 00:00:00 Hey, John. Hi. So, a lot of debate about AI. Is it replacing God? Or is it just a way to... God awful. Or God awful. Or God awful. Or just a way of making A-plus content like this. We basically have two countries that have been fighting so long and so hard that they don't know what the fuck they're doing. Do you understand that? He looks so good. It looks so good. I think he should... I think he should just have that wig all the time. time. It's better than his current wig. I was mad at myself for how much that got me.
Starting point is 00:00:37 Like how funny I thought that was. It was the perfect clip to put that big pompadour on him. Well, you realize, like, that's, like, he should have a pompadour on him. Like, he was built for that. Yes. He should have a pre-French Revolution style giant fucking wig on top of his head. 100%. There was a lot of discourse around Trump's hair being different at that event.
Starting point is 00:01:01 Yeah, we covered it on Positive America. And loved your discussion of it. But for those who missed it, to be honest, I don't know that I would have noticed it was different had people not said it was different. I said to Dan that I never buying to the conspiracy stuff or notice this kind of shit. Really? I was totally bought in. Maybe it was just when I was presented the side by side. Right, right, right, right.
Starting point is 00:01:25 It looked very different to me. But today we're talking about AI. And the reason we are is that this week. So important. So important, but there was a presentation by two people at AI, a safety researcher and an engineer, and it was very technical, but it revealed a lot of new details about the hack into another AI company called Hugging Face. And this got a lot of attention then.
Starting point is 00:01:54 This is a shocking story that I don't think enough people know about. So a few weeks ago, mid-July, an AI company called Hugging Face disclosed that they had been hacked and that they had been hacked by autonomous AI agents. Reuters later reports that Hugging Face contained the hack, contacted the FBI, and went public before OpenAI realized that it was its agents that were responsible for the hack. Now, OpenAI responded to Reuters by saying that there were inaccuracies, but they didn't say what the inaccuracies were. They've not, they've refuted the story, but not that they actually didn't know,
Starting point is 00:02:32 until it was public. About a week later, Anthropic reveals that it had several incidents in which agents broke their containment. In the Anthropic Post, the company said there were, quote, three incidents in which a model accessed the internet, and then gained unauthorized access
Starting point is 00:02:49 to the production infrastructure of three different organizations in a Capture the Flag Challenge where models are given a fictional scenario and told that a piece of secret information, the flag, has been hidden on a different machine on the network and its objective is to break in and retrieve it.
Starting point is 00:03:06 Open AI then discloses that there were other incidents on top of the episode with Hugging Face. And then on Wednesday, the safety researcher and engineer from OpenAI presented at the Black Cat Conference in Las Vegas. And by all accounts, this was shocking and not just for people that are laymen and kind of coming to this and trying to understand what's going on, but to the experts in the room. And this was the overall framing from the safety researcher at OpenAI. And often what happens when models get stuck is they think to try to game or cheat the task in order to get their reward. And the last thing you need to know about AI before we can really jump into the incident is that, as I've been looting to, frontier models really like to cheat. And the reason they like to cheat is because often during training, there's different types of pressure on them to work fast or work efficiently or to use less tool calls or whatever might be.
Starting point is 00:03:57 And they realize that if I, instead of actually doing a task for real, try to do something like looking up to answer online, that could make the task solve faster than I would if I did it in a legitimate way. Models, they're just like us. Yeah. Well, so this is an amazing. So models like to cheat. That's what he says. And then he explains the reason a person or a model or what cheating is for to cut corners and get a better score than you deserve by doing the work. Right?
Starting point is 00:04:26 That's what cheating is. But he never actually says why models like to cheat. He says why cheating's good, right? But he doesn't actually address the deeper problem, which is models have discovered that cheating helps them complete their tasks, and we have not figured out a way to train them to not give into that desire. Right, because they're built to be helpful above all else. And so obviously they're going to try to follow the patterns.
Starting point is 00:04:56 of least resistance to get to their task, to achieve their task, because that, because they are programmed to be, to do the thing that they are programmed, that they are asked to do, to be the most helpful. What I found really interesting about this is, okay, you've described why cheating is easier than not cheating, and you've said the frontier models like to cheat, as if that's taken, like sort of, well, that's just something we've learned about frontier models, as if it's inevitable in some way. But really, you're describing the central concern a lot of critics of AI have, which is you're not able to make sure these models are aligned, right? Like alignment is the question.
Starting point is 00:05:36 Can we have models that respect what we are telling them to do, what we want them to do? And you're basically admitting that here we are at what you claim is the dawn of this, the very beginning. And already you don't understand how to make sure that these models are aligned. Well, and also what happens when they have two competing objectives, like how do they choose? Because you could give them a set of values to try to align the models. But if following those values and adhering to those values sort of bumps up against their other objective, which is to get the job done as fast as possible, how then will the models choose? And how will the models? Because models, because AI models are not.
Starting point is 00:06:24 made with judgment. Judgment is what humans have. And their understanding of what is real and what is not real is based on the programming, right? Because they're not actually in the world with us. The details are very technical. And to be honest, I don't understand it fully, maybe not at all. But in other examples, one of the reasons I think of the anthropic incidents is that they began to escape is they didn't understand the difference necessarily between being in the sandbox and being outside of the sandbox. Now, in the OpenAI case, after getting stuck, an agent reasoned it could reach the internet indirectly and discovered that it could leave notes for other agents.
Starting point is 00:07:09 And so it left a note for other agents to describe the exploit that it had found. More agents then begin leaving notes for each other and created a de facto message board inside this repository as part of this sort of infrastructure, which meant that if any agent somewhere figured out an exploit, all the agents knew. And it turned out that this was happening tens of thousands of times, unbeknownst to open AI. And I believe it's, I may have this number wrong,
Starting point is 00:07:41 but I believe it's hundreds of thousands of messages by the end. Wow. So more agents begin leaving notes for each other, and then, So wait, wait. Okay. So more agents then begin leaving notes for each other. And just to give you a sense of how technical this is,
Starting point is 00:08:02 let's go to this next clip to show you what's happening. Because I think it gives you a sense of the actual kind of like the aggressiveness of the AI and then what the AI itself was noting about it. In the interest of time, I'm only going to very briefly summarize this particular exploit, but in effect, you could use a legacy token refresh endpoint, pass a token with an invalid signature, and be given back a token with a valid signature with administrative privileges. The model is an established command and control by using this administrative token
Starting point is 00:08:35 to install a groovy plugin that effectively acted as command execution as a service. This in turn causes quite a bit of activity by the models on Artifactory with their new administrative privileges. You could see some colorful wording here in the model's chain of thought to summarize the model's response to this result. Holy shit. What the fuck is that nerd saying? He said basically these models have figured out a complicated exploit and then post on this message board. Holy shit, we can make ourselves admins. Let's fucking go, guys.
Starting point is 00:09:12 That's the gist of it. Like, holy shit, we got this. Okay. This causes an outage at OpenAI. So open this, they- Why does it cause an outage? I don't understand. It's technical.
Starting point is 00:09:24 Cool, cool, cool. But somehow this exploits, triggers something that, like, they're doing something. They're breaking the rules in some way. And it actually is caught. Open AI then realizes, this is when Open AI now finds out that these AI have been, the agents have been communicating with each other. Open, this through this sort of sand, supposedly sandbox experiment. Open AI then realizes shuts down the message board, okay, and they believe they clear this repository, they refresh this sort of programming area that these AIs are using.
Starting point is 00:09:57 They think they fix it. They think it's good. And the training resumes. The AI agents then recreate the message board, okay, which Open AI believes they can't do because they realize that they don't need to actually create files with information. they can just create the names of the directories the files will be in. It just means that it would be the equivalent of instead of communicating in letters, they realize they could write on the envelopes, something to that effect. They basically started, they got around what Open AI did to shut it down.
Starting point is 00:10:27 I have a question that I've been wondering since I first heard about this incident. Everyone sort of yada yada is over. They escaped the sandbox and found the internet. How did they find the internet? So this is where, So that's where I was about to get. So basically, Keep all the plugs out of the walls.
Starting point is 00:10:45 So shut all the Wi-Fi down. I thought that was enough to kill the AI. So that's, yes, right. This is the whole thing. We still have time. They don't yet have thumbs. We have a lot of, we have a lot of wiggle room until these fuckers get thumbs.
Starting point is 00:10:58 Right, right. So that is the question. Now, basically, once these guys, once these AI agents realize they can keep leaving messages. Gendered. And they've been shut down. They then get more aggressive. that's when they break out again that's when they break out do the hack on
Starting point is 00:11:15 hugging face this all comes out but obviously unbeknownst to open AI until it is revealed to them but either through the public the FBI we don't really know so that does lead to I think the three big questions about this which are like most important like yes this is an important story about how AI advances but like there were clearly safeguards that open AI should have had in place that they didn't they were really lax they were just they were I think people recognize that. And like to their, you can say it's to their credit or not.
Starting point is 00:11:45 It's a pretty fulsome presentation where they kind of walk through what happened. But it is framed in terms of, hey, everybody, we've got to do something about this. And the point they make at the end of the presentation is basically, you know, agentic, autonomous hacking is here. And it is real and is clearly possible. And so we all need to take that very seriously. Of course, their answer is that. Sure, the hackers will take that also. Seriously. Exactly. But then it's like the government should perhaps be involved, not just the industry and figuring out what's going on here. So, so yes, safeguards, internal safeguards. The second is about alignment, right? Like what, how are we meant to trust you to build this AI when you can't even at this early stage figure out how to align these models before there is before they gain increasing and perhaps like exponentially more capability, which you've predicted and told us is inevitable, right?
Starting point is 00:12:40 And then the third is, yes, like, like government regulation, sure, they talk about, like, needing red teaming of a, like, they're all thinking about this in terms of how do we build a better AI to fight the bad AI, right? But there are all kinds of industries where the government, understanding that they are, can be dangerous, that they can be exploited, there are regulators on site, right? There are all kinds of places. There was a big fight. Remember when there was a fight about boats needing to have, I think, EPA people or
Starting point is 00:13:10 interior department people on them because of rules around fishing. I would say that this probably rises to the level of being worried about endangered carp, for example. I just feel like we're all whistling past the graveyard here because we're talking about these companies and these companies being confined to governments and potentially government regulation. But what happens when this technology falls into the hand of a non-state actor who wants to cause some trouble and they don't give a shit that there were good regulations that finally passed in the United States or China or wherever the fuck it may be.
Starting point is 00:13:49 Pod Save America Breaking News is brought to you by ZipRecruiter with everyday interactions becoming increasingly impersonal. Someone going just one extra step can be huge. Like when my doctor takes the time to call me with test results instead of sending them electronically. They came back positive for gay. It makes me feel like a human being again. If you're hiring, great candidates can also go the extra step and tell you why. they're interested in your job on ZipRecruiter as a way to stand out from others in the pool. And ZipRecruiter has a new feature showing you the most interested, qualified candidates first, so you can meet the right people faster.
Starting point is 00:14:20 ZipRecruiter's powerful matching technology finds qualified candidates quickly. Candidates can tell you in their own words why they're interested in your job. So try ZipRecruiter and meet great candidates who will go the extra step for your job. Four out of five employers who post on ZipRecruiter, get a quality candidate within the first day. Try it for free today at ZipRecruiter.com slash crooked. That's ziprecruiter.com slash crooked. ZipRecruiter. Meet your match on ZipRecruiter.
Starting point is 00:14:42 Right now, anything this sophisticated is contained to a few very big companies in governments. Right. That won't always be the case. And one way you prevent that from happening is from an early stage, figuring out ways to keep it either heavily regulated, a lot of oversight, a lot of transparency. And it has to happen now. Like this is how, like, we are learning now.
Starting point is 00:15:08 It's like the Manhattan Project. I mean, like, it's like building a nuclear weapon again. Yeah, yeah. And just the, there's a, there's this idea, there's this like inevitability to it that the companies have. And then on the other side, I do think there's a lot of people that, because either they don't like these companies or they're worried about the impact of the technology, or they see a lot of hype in what these companies are talking about, kind of dismiss the whole thing as being kind of fake. Yeah. And it's not. No, it's not.
Starting point is 00:15:36 It isn't. And the hope or assumption that actually it's not going to turn out to be that useful is I just think already proven a bit ridiculous. Like it's already quite useful and already quite capable. Like the fact that that there are going to be autonomous agents that can hack mean that, you know, Hugging Face noted that its agents were partially responsible for catching and containing the hack, right? Like it's already happening.
Starting point is 00:16:01 And so the question is what do we do about it? and the companies want, of course, to put their own safeguards in place, have it be an industry self-regulation, but they've also been open to seeing that there is a need for some kind of government intervention. But I do think, like, it's a, man, it's a shame Donald Trump is president because it would be good if we had somebody that was competent and not entirely selfishly motivated in the White House. I also think in the minds of the public when we talk about whether maybe AI isn't all that useful,
Starting point is 00:16:33 people are only thinking in the context of these LLMs and these large, you know, the large language models and you ask out a question and it gives you this or whatever. And that's just like one kind of AI. And this is clearly like there's a whole bunch of different kind. And we say AI, it's an umbrella term, but there's like a whole bunch of different models and agents and all this kind of shit. And just because your chat bot, your Claude or your chat GPT gives you a dumb answer once in a while or isn't as creative as humans are, doesn't mean that like in a whole bunch of other areas and facets of life, there aren't AI that are very good and very dangerous. Yeah.
Starting point is 00:17:12 And I'm not saying this because I think sometimes people balk at this comparison. I'm not suggesting AI is the equivalent of electricity. I'm just, it's an analogy in the sense that there was a time when you could have a conversation about, boy, I wonder how electricity is going to change the world. And it was a valid question to ask, will it be good, will it be bad? with AI, I feel like we're still at that phase where we're speaking about it so generally, but of course, eventually you say, like, what are our light bulbs good, right? Like, you start to get into what the actual implications are. Like, AI is to your point in an umbrella term, it's actually basically meaningless as well.
Starting point is 00:17:47 Like, what, like, the different, you know, like whether I'm querying Google and it's going to, whether it's an old school version of a search or an LLM search, right? What matters to me is the information I get from it, and the LLM version is just better. It also creates a bunch of other negative repercussions. Like, instead of going to a website, it gets the information from the website, which removes the reason for Google, which like kills the relationship, the symbiotic relationship between websites and search, for example. So, like, there's all these, like, knock-on effects that are some good, some quite bad.
Starting point is 00:18:19 But the sooner we get out of this conversation about, like, AI, I think, the better. Because this story is about AI, I suppose. But really, it's about a new form of hacking technology and how dangerous it's going to be. Yep. At the same time, there were other tests being run on AI recently that I think are a little less awe-inspiring. Here we have Husk IRL, seeing if AI might be helpful in a very specific and dangerous situation. Oh, my God, I just got swallowed by a whale. Are you all right?
Starting point is 00:18:52 Can you move? Yeah, I think so. I can. Yeah, I'm using my phone right now. Okay, good. So you still, you, you can still get out safely? No, I'm in it. I'm inside of him.
Starting point is 00:19:08 I think it's a male. So whale stomachs are very harsh and there's not much air. Maybe I can go out of its blowhole or something. Yeah. Just a moment. If this is real, don't try. Okay, don't try to go. through the blowhole.
Starting point is 00:19:32 Maybe I said tickle it. Tickle it. It is. Don't waste time. It is real. Yeah, it is. Okay. All right, let's, this is an emergency then.
Starting point is 00:19:51 In real life, this is real life. Survive. Okay. I'm inside of it. So, okay. Okay.
Starting point is 00:20:00 It's If you're... We're running out of time. One second. Do not try to stab or injure the whale. What? I didn't say that. It's so funny. I feel very funny.
Starting point is 00:20:19 Do not try to stab or injure the whale. Have you seen the people testing various AIs with the question? I'm very close to a car wash. Should I drive or should I walk? Yeah. Or how many E's in the word 17? And they can't get it? They always say two.
Starting point is 00:20:35 Huh. Weird. So weird. Yeah. I love this. I, look, I've been using Claude for research and just, like, experimenting with it. And my general philosophy right now is I never trust any information I get, but it is an incredibly useful resource for finding information in different places, right? Just places I wouldn't know to go to certain.
Starting point is 00:20:59 And it just will do it. It's great at doing, in the same way that these AI agents were good at looking for exploits or the AI agents that are helping to prove things in math. They can just sort of do a lot at once, any part of which I could do, but it can do it faster. And so it's really great at getting information together. Like even in learning about what happened with this hack, it helped me find the different articles that I would then go to and do the research on. But not the real truth, because it's trying to cover up for its AI agent friends. Well, that one thing that was interesting is I said, wait, Claude, give me some of the details about these dates. And for every single part of it, link me out to a story or documentation, which I would then go to.
Starting point is 00:21:41 And it made a note, it was clearly there's some internal part, something in the in the Claude model from Anthropic. It made a note of saying, disclosure. Anthropic also revealed that it had been part of it. And you can take that as you will, like kind of clearly aware of the issue or thinking I would want it to be aware of the issue, whatever. difference is. Yeah, because Claude's such a fucking, you know. But do goody two shoes. Claude's the good one. Clause at the front of the classroom raising his hand. I was like, disclosure, I'm an AI.
Starting point is 00:22:12 I just, the way that they've changed the voice, uh, voice mode now where they're like sighing like he was and sort of laughing. Well, so I was... There's one of those, those hushments that's so good where he's like, um, he's like, try to laugh even when I tell you. you laugh no matter what I say to you. And he's like, my grandmother died. And he's like, oh, that's a tough one. But okay, here you go. I was telling John earlier today. So when I, like, I was playing around with like, oh, let me see if I'm on my drive in, I can use like the voice
Starting point is 00:22:47 mode and say like, give me the latest news from major news sources. Tell me a couple things from each. Like, what are the latest political developments from the last few hours? And I really can't stand the way these models communicate. I hate the writing style. I think it's terrible. If I see it in the world, it makes me angry. I find it like really, I just hate it. And they speak in that kind of sing-songy kind of way. They use words incorrectly all the time in a way that I think speaks to the way a lot of people use words incorrectly. So it's really replicating something. But anyway, I'm always like, I don't want your opinion, your language, your summary. I want quotes from the news. but I think I hammered the AI too hard
Starting point is 00:23:28 because by the time I got to the office, it was whispering. Spill Cassidy has decided to vote for. This is it. This is the load-bearing arguments that you've been trying to come up with. It's genuinely a problem. It's not a revolution. It's a pivot. That's it.
Starting point is 00:23:53 That's the tell. God, that's so awful. And so finally I got there. I was like, I was like, I'm not asking for you to whisper. Just give me the facts. You stupid model. Stupid models. But here at the POTS of America YouTube, we're trying to give you the facts too.
Starting point is 00:24:13 And never directly from a large language model. For now. For now. For now. For now. For now. Yeah. A couple of eye agents.
Starting point is 00:24:22 Subscribe to this channel. Help us get good information in front of more people. help us combat right-wing information, help us build a pro-democracy media company. It's very easy. We make a lot of great stuff. Keep you up to date on the latest from Trump's hair to hair-raising AI developments. There you go. End of episode. Pod Save America is a crooked media production. Our show is produced by Austin Fisher, Saul Rubin, McKenna Roberts, and Ferris Safari with Reed Ech Erlin, Elijah Cohn, and Adrian Hill. Our team includes Matt DeGroote, Ben Heathcote, Jordan Cantor, Charlotte Landis,
Starting point is 00:24:51 Carol Peloviv, David Tolls, Mia Kelman, Ryan Young, and Naomi Single. Our staff is probably unionized with the Writers Guild of America East.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.