Big Technology Podcast - OpenAI's Bots Break Containment and Hack Hugging Face Autonomously — With Alex Stamos

Episode Date: July 22, 2026

Alex Stamos is the former chief security officer at Meta and the chief product officer at Corridor. Stamos joins Big Technology to discuss how OpenAI models reportedly escaped a testing environment, a...ccessed the internet, and hacked Hugging Face while attempting to ace a cybersecurity evaluation. Tune in to hear why the incident represents a major leap in autonomous, long-horizon cyber capabilities, and what it reveals about the risks of giving advanced AI systems broad objectives without sufficient safeguards. We also cover whether the episode qualifies as true AI misalignment, the danger of open-weight cyber models, the limits of pausing AI development, and why defenders may soon need AI systems capable of responding at machine speed. Hit play for a clear-eyed look at the cyber chaos advanced AI could unleash, and what governments and technology companies should do next. --- Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice. Watch the full documentary here: https://www.gravitee.io/ai-agent-documentary Want a discount for Big Technology on Substack + Discord? Here’s 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b Learn more about your ad choices. Visit megaphone.fm/adchoices

Transcript
Discussion (0)
Starting point is 00:00:00 The most significant autonomous AI cyber attack in history just took place with OpenAI's models breaking out of a training environment, connecting to the internet, and then hacking, hugging face to ACE and evaluation. What does it mean for the future of AI and for cybersecurity? Let's talk about it with ex-Meta chief security officer and current corridor chief product officer Alex Damos right after this. In the face of ongoing disruption and opportunity, TMT leaders need to deliver tangible results, not just ideas. When pace and performance matter most, PWC combines market insights and deep sector experience with AI, cloud, and emerging tech to accelerate your transformation and drive measurable ROI from strategy to execution.
Starting point is 00:00:45 PWC can help you anticipate what's next, outpace disruption, and compete. For more information, visit pwc.com. Welcome to Big Technology Podcast, a show for cool-headed and nuanced conversation of the tech world and beyond. We have an emergency podcast episode for you today because just yesterday the world found out that a series of open AI models work together to break out of a sandbox, hack into hugging face, steal basically the answers to a test and go an ACE-air evaluation. And obviously, this is not the desired behavior that Open AI wanted and looks like it might have opened up a new can of worms here for AI and cybersecurity.
Starting point is 00:01:33 So we are joined by the perfect guest to help us figure out what happened here. Alex Damos is with us. He's the chief product officer at Corridor and the former chief security officer at Meta. Alex, great to see you again. Welcome back to the show. Yeah, thanks for having me, Alex. You know, you spoke at our summit and I was like, we're definitely going to have you back pretty soon.
Starting point is 00:01:51 And it is amazing how the AI story has just turned into a cybersecurity story very quickly. It has. You know, there's all kinds of risks from AI. And, you know, there are all kinds of bad things that happen to consumers. But when you talk about the models themselves, it seems that cyber is the thing that's hitting right now, for sure, from a societal level risk. Yeah, and so this is what we're talking about now, and the reason why we have to do an emergency episode on this is because this is certainly a novel type of hack, right? So this is fairly unprecedented, just to put it in context, it's the first time, this is from Transformer, the breach appears to be the first known example of a misaligned AI escaping containment and autonomously carrying out a cyber attack on a third party, a scenario AI safety experts have repeatedly warned of. So it's not like the anthropic example where mythos sort of escaped containment and emailed somebody while they were eating a sandwich in the park. This is actually going out and hacking a third party. Let me just quickly read the beginning of the Wall Street Journal story about this just to set the stage. So the headline is OpenAI models escaped and hacked the company in cybersecurity tests gone wrong.
Starting point is 00:03:07 On Tuesday, OpenAI said two artificial intelligence systems it was testing broke out of their test environment, hacked their way onto the internet and broke into another company. Open AI said the culprits were a pair of its models. One was its latest product called GPT5.6 sole, and the other was an even more capable pre-release model the company didn't identify. The software had been configured for evaluation purposes to be less likely to refuse hacking commands, opening I said. The opening I had caged the models in a sandbox system that didn't have access to the internet, but during the test, the software used its hacking skills to break out. I found a way to get online and then hacked into HuckingFaces Network. And of course, Hucking Face is a library of open source AI, mostly AI models or AI programs. Alex, how significant is this?
Starting point is 00:03:54 Like, you know, obviously there's a tendency to be alarmist about some of these things. But I wanted to, you know, bring you on because you are the cybersecurity expert. And so you can tell us, like, is this a 10 on like the 10, holy crap? Like, we're in some deep trouble or is it a one like something we might have expected anyway and we shouldn't be too concerned? It's like an eight. I mean, it's a pretty big deal on a couple of levels. It's a pretty big deal in that Open AIs model beat them Open AI, right? So that this went beyond that it was able to trick Open AIs own security team and get out. So there's three or four things I think we should talk about here. There's an alignment issue. There's security issues and there's kind of this is going to have an impact on the policy discussion. And it's a warning. of what we need to do because this isn't just about open AI. In a way, I'm really glad this happened because it is a warning of what we need to get ready for, for maybe something about three to six months from now, what's going to become standard, right? So first is the alignment issue, right?
Starting point is 00:05:00 So effectively, you know, this is not the model wanting something. So this is what I keep on telling people. Models don't want anything. They, when you have an alignment issue, it's because they were asked to do something and then they went and did that thing, but in a way that the human who asked it to do something did not expect. Right. So in this case, the model was told, go take this test and do the best you can. But the cybersecurity protections that were normally placed on it were removed. So Open AI says it's explicitly.
Starting point is 00:05:32 It was, there was a unnamed model that was part of it. So they had two models that were paired together, the existing 5.6 sole with cyber protections are moved in the unnamed model. We don't know whether this is like... Can I pause you for a second? You know, one of the memes about this has been like, you know, like there's this meme on the internet where like you tell a chappot to say it's alive and then it goes, I'm alive and you go, holy crap, right?
Starting point is 00:05:55 So is this a situation where you're opening eye was basically telling the model, go hack something, it hacked something and the human was like, holy crap? Or did it, you know, because alignment is, of course, like we want the model behavior to be aligned with human values. Or the way that humans would want these things to behave. So is this something even more egregious than like us telling it to go hack or Open AI telling it to go hack and then it hacks? Right. So we should not be shocked that it hacks something because they did tell it to take the test.
Starting point is 00:06:26 And it is possibly a hacking model. Like they haven't said what this model is. It's quite possibly a cyber aligned model. This could be like the, you know, Open AI makes these cyber specific models. like they have this 5.5 cyber. This could be 5.6 cyber, right? So it could be something that's specifically tuned to be good at hacking things.
Starting point is 00:06:45 So we shouldn't be shocked that's good at hacking things. But the alignment issue is, so I have three kids, one's in college, the second one's taking the SATs. Right, he's about to take it. If I say to him,
Starting point is 00:06:55 good luck son, I hope you do well. He sits down, he knows that I just mean take the test well. He knows that what I don't mean is slit the throat of the proctor, steal a car,
Starting point is 00:07:08 Thelma and Louisaeer way across the country, break into the college board and steal the answers, right? That is what the model did here, is what it did was, as Open AI explains, is they don't want the model to have internet access, but it has the ability to install packages as part of its work, so they've built kind of a complicated proxy mechanism so it can install packages. It figured out a way to chain multiple vulnerabilities together. It thought they told it, go take this test, exploit Jim, which is like a well-known test. Go take this test. Do as well as possible, son. Go do your best job. And it's like, wow, dad wants me to do as well as possible. How can I do the best possible? Well, the way to do the
Starting point is 00:07:52 best possible is to get the answers. Who might have the answers? Hugging face probably has the answers. So instead of just doing the test, I'm going to go get the answers from hugging face. But I got to get out of this jail dad put me in. Well, dad told me to do the best possible and he didn't tell me not to break out of jail. So I'm going to first break out of this jail dad because maybe what dad really wants me to do is break out of this jail because he told me to do the best I can do possible. So first it puts together a bunch of exploits to break out of the jail that opened the eye created for it. It breaks out of the jail and then it goes, looks at hugging face and finds a brand new vulnerability to break into hugging face. They have not announced it. I have heard what
Starting point is 00:08:31 it is. I'm not going to make news here because I don't know exactly what the patching situation is. but it is a vulnerability in a very, very important piece of software that it found. And that is in lots of different places. So this is a big, big deal. And it just is like, oh, yeah, I'm just going to find a phone in, like, actually a really important piece of software that millions and millions of production systems use and just, like, nuke hugging face with it on the way to getting the answers to the test. So, like, that's a pretty awesome.
Starting point is 00:09:03 Like, it's one, like, the nerd part of me is just like, wow, that's pretty cool. But there is a significant misalignment thing here, and that my 17-year-old knows that you're not supposed to do all this things when I say, like, do well on the test. And the AI system does not know that. Now, to be fair, normally, when these models run, there are protections in place to keep them from doing stuff like this. And Open AI intentionally disabled those protections as part of doing this evaluation. Right. So that was a component of them doing this testing was to keep so that it would be fair. So that is part of the learning here is that if you're going to do an eval and you're to turn off all the safety stuff, then you have to be absolutely positively sure, especially if you're doing cyber evaluations, that the jail you keep them in is an absolute jail. And I expect what's going to happen now is that these things are going to be completely and totally physically sandboxed, right?
Starting point is 00:09:59 Like you're going to have to run them in physically disconnected. If it needs packages, you're going to have to move the packages over. If it asks for a package, you're going to have to bring it over. And then the future, if it wants to get out, it's going to have to like trick a human to get it out, which might be possible. Right. But, but it's going to be like that. The next lesson that we kind of learned here was that, you know, the capability of these things to move of what you call. So when, you know, all this discussion around.
Starting point is 00:10:29 Mythos and Open AI and all of the cyber capabilities. It's about finding bugs and running exploits. This thing did those things. But what we also know is that lots of models have those capabilities. This thing had the ability to find bugs, yes, chain them together, yes. But then to think through all of those things for an ultimate goal. And so this is what people call Law and Horizon cyber tasks. And it had the ability to do that.
Starting point is 00:10:59 with like the level of skill that you would have of the manager of a TAO team at NSA right now. So that is what is. So TEO targeted access operations was like, I think they've renamed it, but it was like the team at NSA that would do all the breaking into, you know, other governments at America. Oh my God. Right. So like, so that is what is like really, for a way. while now these models have been really good at looking at software and being like, I found a bug.
Starting point is 00:11:34 And then here, let me write an exploit for you. What's really impressive here is this thing was like, I want the answers from Hugging Face. And it came up with a plan of like, how am I going to get out of this network, get across the internet, and get into Hugging Face. And that is like the long planning here is very human-like. And that is what is like actually really scary here. And so So that is what we need to, when we think about the danger of these models, we have to stop thinking about the bug finding. Because that is what caused the White House to do this spectacularly stupid thing that we talked about on stage, which was to ban Fable. Because that is the mechanism, that is the thing that is most useful for defenders right now and people who own code is finding bugs and fixing them. where the real danger here is the coming up with a multi-stage plan to execute
Starting point is 00:12:25 autonomously because what you really don't want is you don't want somebody be able to say to their model, hey, I would like to steal the, I would like to steal money, go figure it out for me and then let it work for 12 hours and just steal money for you, which is this model would clearly be able to do that. Right. And so I think what you're getting at is, you know, one solution is going to be in testing, you want to fully disconnect these models from the ability to like break out and get onto the internet. But that's just solving the testing issue. The real problem here is that AI models have achieved this capability, that they are able to do this now, not only finding the bugs, not only the breaking out, but to be able to do these multi-step
Starting point is 00:13:11 plans and then execute. And if this is sort of the latest unreleased open AI model, Well, the history of generative AI has told us one thing, and that is that the frontier is only the frontier for a few months, maybe 10 months, maybe a year. But not much longer than that. And so if Open AI is seeing this in testing now, is the real danger that this type of capability does end up in the hands of, you know, evildoers, you know, faster than a lot of people. might expect. Because if that's the case, that changes everything. Yes, that's right. And so, you know, our best knowledge on where, say, the open weight models are comes from the AI Security Institute, which is the UK government's group that does these assessments. They release just this week an assessment with GLM52. Unfortunately, you know, the Kimi is really, Kimi, you know,
Starting point is 00:14:14 K-3 is the best of the Chinese models now. The open weights have not been released. So you can't really do a good assessment for Kimi yet. And what we're finding is the Chinese models are not cyber tuned out of the box. So it is very likely that the Chinese companies are not, are intentionally, I'm not going to say neuter, but they're intentionally not making their models really good at cyber. And there's a couple of possible reasons for this. They're probably trying not to tickle the Dragon's Tale of the PRC overlords
Starting point is 00:14:53 because what they don't want to do is they don't want to trigger a crackdown for their exports. But what happens is if you take those models and you bring them into your own lab and you have a training set of labeled vulnerabilities, if you have a cyber gym, then you can make them much better yourself. And the amount of resources it takes to do that is not extremely high. It's in the tens of thousands or hundreds of thousands of dollars. It's not in the hundreds of millions or billions of dollars. So what that means is, one, the Chinese absolutely have better capabilities in-house
Starting point is 00:15:27 than what we can see on the charts, right? Because I guarantee then what those companies are offering to the People's Liberation Army and the Ministry of Security is way better than what they're releasing publicly, both from a profit perspective and a keeping the government happy perspective. Second, it means that other adversary groups are going to take the Chinese models and then spend the several hundred thousand dollars or millions of dollars necessary to create tuned models. And we are probably not far away then from just going to Hugging Face and getting a, you know,
Starting point is 00:16:01 Kim E3, a cyber-tuned model that can do both, especially the short, short horizon stuff much better than by default, and then eventually the lawn stuff. The short stuff's easy to train because all you need is a bunch of bugs, so you can just go get a bunch of CBEs and train it. The lawn horizon stuff's harder because you have to build these like cyber gyms and such. It's not impossible because there are a bunch of CFPs and examples out there, but you can do it. What are CFPs? And so, I'm sorry, not CFPs. CTFs. CTFs. Capture the Flag. So like you can use like capture the flag training sets and all that kind of stuff.
Starting point is 00:16:42 And that basically puts the AI in the gym and sort of has it work through all the steps in order to meet this objective. Yeah. And so like people have had, you know, training for humans and for hiring purposes and all that kind of stuff. And so anyway, what the AISI has said is that the difference between the frontier and the Chinese models is about seven months. But I would argue that that underestimates it because the Chinese models that we see are. are undertrained. So that the internal Chinese capabilities are probably much closer to the frontier. Now, what, what happened with Open AI
Starting point is 00:17:18 is beyond the frontier, because when we say the frontier, we're talking about what's released. Right. Right. So, but yes. So the, what that means is this capability is coming
Starting point is 00:17:33 for adversaries. And so we need to get ready to defend against this capability. And then the other funny part of this story is before we knew this was open AI, Hugging Face announced we were attacked by an AI attacker. We don't know who it was. When we tried to defend ourselves, we tried to defend ourselves with an AI system, we used a U.S. Frontier model.
Starting point is 00:17:58 And the U.S. frontier model shut down and refused to defend us because of a cyber protection put in place. Those are the cyber protections that were required by the Trump administration. So we had to switch to a Chinese model to defend ourselves. So we switched to GLM 5.2. So before we knew it was OpenAI, Hugging Face wrote this blog post saying, everybody should have at least a Chinese openweight model ready for defense,
Starting point is 00:18:22 because you might find it yourself in a situation where you get cut off from an American provider for defensive purposes. I expect that actually wasn't open AI. I expect it from their description. It sounds like it was an anthropic model. Right. So we have this hilarious situation where an American company, loses control of their model. It attacks a French company. The French company turns to a different
Starting point is 00:18:43 American provider for defense. And that American company says, oh, that's a cyber problem. I can't help you. And so they have to turn to a Chinese provider to protect them because the White House forced that other American company to have protections because they're a French company that they can't use. It's really kind of weird sci-fi podcast. But these American models already had refusals on anything cyber. Like one of the knocks on Fable was that it was would refuse like let's say for bioterrorism. If you asked about mitochondria, it wouldn't answer. So was this really the government or is this just the model's own safeguards that they're putting in? And I think one, just to put one detail on, one of the interesting things is open sources, it doesn't basically matter if you're an attacker.
Starting point is 00:19:27 If you have open weights, it doesn't matter if you're an attacker or if you're a defender, you can use them without restrictions. The problem with the restrictions that we're seeing from these closed models is that they can't really, they can't really differentiate. So in order to prevent attackers from using their models, they are also basically wholesale, you know, refusing anything on cyber, which means that if you're trying to defend also, you can't use it. That's right. Well, in the blog post that Anthropi put up when they turn Fable back on, they said, we have to tune up our defenses on cyber way too far because of the White House. So they specifically said that of the precision recall tradeoff is we have to tune towards recall versus precision, right?
Starting point is 00:20:11 So we will have way too many refusals. And so we're in this weird place where they are saying, they are saying all the time, I can't do that for you. I can't do that for you. And they're basically being forced to by the White House because the White House is still not defined what is the appropriate level of refusal. And apparently the White House is still hand approving who Anthropics is allowed to live. let into their cyber program. Now, the funny thing is, now OpenAI has said, we have approved Hugging Face for our tech program,
Starting point is 00:20:43 even though they're not an American company. So I don't know how they were allowed to do that if they just went over the top of the White House or they got like emergency approval or something. We will see what the policy responses from the White House from OpenAI's announcement. I hope there is not a crackdown. That will be the natural response of the White House.
Starting point is 00:21:01 But it needs to be the opposite. because what this demonstrates is, yes, Open AI screwed up or whatever. They need to have fixes. But this is coming, right? This level of capability will be in the hand of every adversary, every American company faces. So the response of the White House needs to be that we have to, one, fix the bugs, to find the bugs, patch them everywhere, and then we have to have the ability to respond at machine speed. So every American company needs to have AI watching
Starting point is 00:21:38 for their defenses. It's going to be, because the attackers are just going to tell their AI, go attack this guy. And the defenders have to tell AI, defend me, because no human being can defend against this. You cannot have a human being
Starting point is 00:21:50 watching your logs anymore. Or at 2 a.m., you get a page. And a human being has to be like, oh, okay. And then log into Slack and take 15 minutes to log in and look at the log and figure it out. point you're toast because that sucker has ripped through your network in those 15 minutes by time it takes you to look at your slack. Hugging Face found 17,000 actions that this model
Starting point is 00:22:12 combination had taken. 17,000. That seems like a lot. Yeah, go ahead. It doesn't stop. It doesn't sleep, right? Like, it just, it will just keep on trying. I mean, it's, you know, to quote from the first Terminator, right? It will not stop, right? Like, you know. Yeah. To quote from the immortal Michael being, right? Like, it will just keep on going until it, it, it accomplishes its goal. It'll try a lot of different things. Now, the fortunate thing is right now they're very noisy. So if Hugging Face, I have not seen the logs. Like, we, we have not gotten like a really good technical write-up here. So that is what's missing. It would be nice to see from both Open AI and Hugging Face. So for Defender, so we can have a better understanding of what we need to do here. What we really need here is we need a much deeper
Starting point is 00:22:57 technical write-up of exactly what happened. What has been released so far has not been sufficient. But my expectation is from the initial write-up is that this thing is extremely noisy. And so it would be, if Hugging Face had like better detection and better AI detection, it probably would have got caught much sooner. There's this graphic on, I think, one of the Miri spokespeople's, his name's Harlan Stewart, one of the Miry spokespeople's Twitter backgrounds. and it's like basically there's a continuum between AI is becoming good enough at scheming that we sometimes see it scheming against us and then AI becomes good enough at scheming that we no longer see scheming against us and we're like smack in the middle of that.
Starting point is 00:23:41 Do you think that that is an accurate representation? Maybe. Or is that the concern basically that we won't see it? Because you mentioned it's noisy. So is that the concern? Yeah, possibly. I mean, remember the model here was doing what it was as. Right.
Starting point is 00:23:57 It was not scheming against its bosses at OpenAI. They asked it to take the test. And they didn't, they, I don't know exactly what the prompt was. But apparently they did not tell it, not to cheat. So who knows? Like, this is also what Open AI needs to be more transparent about is exactly what their prompt was, exactly what the constraints were. Did they tell it explicitly? Like, it is a much bigger alignment problem if they explicitly said.
Starting point is 00:24:26 Do not try to break out of the network. Do not try to get the test answers. Now, if they told it all those things, then they have a much more significant alignment problem, right? Than if they were less explicit. But in any case, yeah, I mean, that will, if these models get trained to be more evasive from a network intrusion perspective,
Starting point is 00:24:53 that will be very dangerous, yes. And what I would argue is for the legitimate companies, I would not do that. I don't think, I think there is a, if you're open AI and you're building 5.6 cyber, what you should be training it to do is find bugs. You should be training to write proof of concepts. You should be training it to do all the defensive stuff. You should not be training it to hide it, to hide all those things. Like, if the U.S. government wants to build a model that does that stuff for the NSA, then you can let them do that.
Starting point is 00:25:30 Or you can let Lockheed Martin do that. But if I was Open AI or Anthropic at this point, I probably would not do that. I think I would leave that for somebody else. But that's scary, though, because it could then take actions that, you know, I think one of the things that is, so the question is, like, should we be concerned with the AI, you know, sort of doing things on it? own and should we be concerned with, you know, or is the bigger concern that humans direct this AI to do bad things? So we've definitely covered the fact that humans will be more, should be, you know, humans who direct this AI to do bad things can do a lot of damage. But if you create an AI that can reward hack, because this is all coming from reinforcement learning, where like these
Starting point is 00:26:12 AIs are given rewards and they are basically like maniacally focused on achieving that goal. And if you, it's almost like sort of gain of function research on a virus to a degree. right? Because if anybody builds AI that doesn't leave a trace and it goes out and reward hacks its way into hacking something else and maybe isn't so fully like going with the prompt, then that's where you can get into a real problem. I know that's a more out there possibility, but I don't know if it should be completely discounted. Yeah, I mean, I guess as they get more and more complicated, the question is, is like what, at what point are they, is it their own motivations versus just doing what you've asked it to do? You know, I mean, so far again, I don't think we should still think of
Starting point is 00:27:06 these things having their own desires or wants. They are still doing what they're asked to do. It's just, like you said, there's a lot of inputs of what they were asked. It is not just, just the initial box, right? There's all, there's the system prompt and all the training and all the rewards and everything that's gone in. And so the question is, like, what is the humongous history
Starting point is 00:27:32 of all of the different things that's been trained to do when you've asked it, take this task? Right. And especially if you've removed all the protections. And so in a situation where these things have all the protections removed, that is very dangerous.
Starting point is 00:27:44 And as we talked about, like with the open weight models, either, there are no protections or the protections are trivially eliminated, right? Like a bunch of open weight models have been trained with protections, but you can obliterate those out. And you can go in Hugging Face and look for obliterate, and you will find a zillion models where people have removed the protections. But this goes basically back to that like long held thought experiment of,
Starting point is 00:28:09 you know, the AI can follow your goal and achieve your goal. But it might have a different idea about what it takes to get there than you do. So in this case, the paperclip maximizer, yes. So exactly. So I was going right there. So, you know, this is, and it's funny because I am speaking with Nick Bostrom later today to, you know, for an episode that's coming up. But basically he's this Oxford philosopher who came up with this idea that if you ask an AI to make paper clips,
Starting point is 00:28:33 eventually it can see so much on this goal that it, that it can, you know, find humans as an impediment to its, you know, objective to maximize paper clips and sort of kill us all and turn everything in the world into paper clips. So like the fact that it, it was on tasks. Like this is kind of a Twitter user said this. Once out of their sandbox, the models did not scheme, engage in behavior that had nothing to do with their instructions like hacking the NSA or launching a cyber tech on Russia or stealing secrets from a rival AI lab. But like, you know, sort of if the model found it suitable to go out and hack a hack hugging face in this situation, who's to say that, you know, maybe a less careful model doesn't do this.
Starting point is 00:29:16 and then maybe an even less, like doesn't go and hack the NSA. And then even less careful model turns us all into paper clips. I mean, there's a continuum there. Yeah, I mean, it's why you have to be very careful what tools you attach to them. And it's why you need to have, they have to be supervised by different things. I think, like, you just can't, you can't have models that have no protections on them that have connections to tools, right? Like, that's why these models then you have dumb, class. are dumber models watching them.
Starting point is 00:29:48 You don't just take the smart thing and then hook it up to everything and you're like, give it a task. You have the smart thing and there's a bunch of dumber things watching it and those dumber things can either kill it or they can call a human that can kill it. That's the idea.
Starting point is 00:30:03 It's like there's supposed to be cyberclassifiers and there's supposed to be mechanisms that can stop it. And those mechanisms should be either deterministic or dumb and undefeatable by the model. and they removed all those things so that the eval would work.
Starting point is 00:30:18 So I think what Open AI is basically hinting at they haven't been explicit. It's like if we're removing those protections, this thing is going to be in an absolute physical jail. It will be physically separated. It will not be hooked up to the internet anymore. And that seems like that should be the standard. That is fine for Open AI.
Starting point is 00:30:33 My point here is that doesn't matter. Like this situation is good that this happened. Because this has pointed to us where we might be in six, nine months, a year from now, no matter what, because other people, unless we can get an international agreement to just stop development, which is what other people are talking about, right? You've got this, I forget what, like Project 2030 or, like you've got people talking about international treaties or
Starting point is 00:31:02 whatever. I don't think any of that's going to happen. I don't know what my position is on that, but I just don't think it's going to happen. I just, I think there's no way, this is just math and silicon. And so I just don't think there's any way you get like a, this is not like nuclear weapons where the major input, like the, the reason our species is alive is the major input to nuclear weapons is uranium plutonium. Plutonium does not occur naturally in our planet and uranium is incredibly rare and to turn raw uranium into uranium that can go into nuclear weapons is a massive industrial process. If uranium was something, you could just dig out of the ground anywhere, our species would be dead, right? Like that's just the truth. Because the
Starting point is 00:31:43 knowledge to build nuclear bomb is in the hands of anybody who gets a physics PhD, unfortunately. So in this case, these chips are not something you can really control. We have found that in that the Biden era controls on Silicon have created a massive industry in China. And the knowledge on how to build large language models is something that there is a undergraduate class at Stanford where you get that knowledge, right? You know. You can find it from like a Carpathie interview, YouTube video.
Starting point is 00:32:17 Yes, right. So we cannot control that knowledge. And so, like, the idea that we can just have, like, a bunch of people agree in a room to stop all development of this is just silly. So from my perspective, being a little bit of a pessimist here, we just have to get ready for this level of capability to be in the hands of an unfortunately large number of people.
Starting point is 00:32:39 Yeah. All right. So I want to go a little bit deeper into the potential solutions here. And also, I want to ask you the age-old question of is some of this, all this, none of this, just good marketing for Open AI, given some of the statements they've been making. We have to address that one here on the show. But I'm going to let you have an answer. I'm going to try to at least, you know, illustrate those, the case of those who might be saying it so we can have a discussion about that. Let's do that when we come back right after this.
Starting point is 00:33:07 Hi, everyone, Alex Cantowitz here. I want to tell you about a documentary I've made with gravity to explore the future of AI agent security. To find out if we're truly ready for autonomous agents, I sat down with MIT professor Ramesh Rosker, former White House CIO Teresa Payton, Michelin's Group Chief Data and AI officer Ambika Roger Gopal, and Sharon Guy, a former executive at Alibaba. They each offer unique insights into this evolving landscape. We conclude with Rory Blundell, CEO of Gravity, to discuss the path forward. With Gravity leading the way, join us on this journey.
Starting point is 00:33:45 You can watch the full documentary at the link in the show notes. This episode is brought to you by Deepel. When I sat down with Deepel's founder Yarak Kutliovsky on YouTube recently, we got into the case for specialized AI. Deepel voice is what it looks like when the stakes are real-time conversation. And honestly, it's something I wish I'd had for my own, cross-border interviews, turning a language barrier into a non-issue. Deep Bell Voice delivers live translation in over 40 languages for virtual meetings and in-person
Starting point is 00:34:20 conversations, helping people speak in their preferred language without losing flow or nuance. Whether you're meeting with a customer negotiating with a supplier or collaborating with global colleagues, it keeps pace with you in real time, easily handling the technical terms, acronyms, and product names specific to your business. So what you actually mean never gets lost in translation. And for the builders listening, Deepel's voice API lets you embed real-time speech transcription and translation directly into your products. So go check it out for yourself. You can try Deepel voice for free at deepel.com slash try voice.
Starting point is 00:34:55 That's deepel.com slash try voice. Recently, our company's softball team lost the big game by one run. Then Dale tried to console us with the quote, winning isn't everything. Well, Dale and I are very different. I get early payout from Bed 365. If my team goes up big, I get paid out instantly, even if they blow the lead later. Sound familiar, Dale? Thanks, Beth365.
Starting point is 00:35:21 Must be 19 or older Ontario only. Please play responsibly. If you have questions or concerns about your gambling or the gambling of someone close to you, please go to ConnectSontario.c.com. And we're back here on Big Technology podcast with Alex Damos, the chief product officer at Corridor. You sort of answered the question before the break, but I'm going to ask it anyway, whether part of this is open AI marketing.
Starting point is 00:35:41 Let me at least read some of the statements here and give you at least the argument that people have made for like why some of this is marketing for open AI. The first part is, you know, the Anthropic started to be declared as the company that was in the lead. Once that anecdote came out about mythos,
Starting point is 00:36:01 breaking containment, and emailing somebody when it wasn't supposed to have internet access, and emailing an anthropic employee, employee while they were out in the park having a sandwich. So this could be potentially, you know, Open AIs attempt to like one up that. Then there's also the language that you see. Open AI says in its tweet about this, we are partnering with Hucking Face to investigate an unprecedented security incident. You don't usually have the attacker and the attacker and the attacker partnering together in these situations. They also said, we consider in their blog post, we consider this incident to be an unprecedented
Starting point is 00:36:35 a cyber incident involving state of the art capabilities and are responding accordingly. You know, it's sort of like, oh, look at this terrible thing that happened, but a moment to share how good our cyber capabilities are. That's the argument. What is your response to the notion that this might be some marketing from Open AI? I know lots of people at Open AI. Every single one of them absolutely hated Anthropics marketing around mythos and thought it put the entire industry, at risk. This incident has put open AI at risk of regulation from the White House, regulation from the EU.
Starting point is 00:37:15 It is also an admission of the violation of the Computer Fraud and Abuse Act as well as multiple European laws. It would be absolutely insane for them to use this as a marketing moment. What you're seeing is them being very, very careful and defensive in their language. They're also very lucky that Hugging Face is being super cool. and chill about this. So that is why they are saying these things because, you know, Hugging Face
Starting point is 00:37:41 initially comes out saying, we've been attacked, we don't know who it is, but it does not look like the model was being subtle. I don't know where it was running. It's quite possible as like Azure or something. It was probably not covering its tracks.
Starting point is 00:37:53 And so I expect Hugging Face got their American lawyers involved, was working with the FBI, was probably issuing subpoenas, and was very, very close to find out it was just open AI. So, like, or did find out. I do not know the timeline here.
Starting point is 00:38:07 But like, the legal issues here are very fascinating and interesting. And because they're all working together, I expect nobody goes to jail. Nobody gets sued. Everybody's going to hold hands and hug. And if there is tokens being exchanged or whatever, I don't know. But there's absolutely positively no way this was a intentional marketing move. And opening eye is doing the best they can, I am sure, right now, to, use this to forestall any kind of massive government overreaction either from the United States
Starting point is 00:38:41 or the European Union. When you were at our summit, you said that, you know, speaking of the sort of release of mythos and fable, that a lot of people were very concerned about the bug finding that those models could do, but you said basically, listen, this is not very different from what you could get with Opus 4.7 or 4.8, I believe. Is this what we're seeing from Open AI very different? Is this a step up? Yeah. So this is what I don't know if I said on stage here, but I've said in other places, there's a difference between the short term and long term and anthropic to their credit, and I think Open Eye has in other places. I think we talked about how in the fable model card,
Starting point is 00:39:26 they talk about short horizon versus long horizon cyber tasks. And what I've talked about is, We need to not focus on the short horizon tasks because those are dual use. Finding bugs is dual use. Everybody needs to find bugs, right? That is something that defenders need to do all the time. And that's what's driving people insane right now in the defensive industry is that because of the White House, American models are refusing to help fix code. They are refusing to help us find our bugs and fix them, thanks to the White House's actions. that is not this problem.
Starting point is 00:40:04 This problem is go run a entire attack chain for me. That is the long horizon tasks. And that is where we need to continue to have appropriate classifiers that are like, bro, I am not going to break into a bank for you, or I'm not going to plot out or run a C2 mock for you or any of that. So yes, this is what, you know, explicitly Anthropics said, we will allow Fable to do short horizon stuff, but we will not allow it to do the long horizon stuff that mythos does. Right. And so mythos, just to confirm, what mythos can do the long horizon
Starting point is 00:40:41 planning and what we're seeing in this instance from Open AI, that is the step up. That is the step. And I can't, obviously, I don't have access to this, whatever this thing is. And so I do not, I can't say whether or not where they are, AISI has done these assessments. And so, Who knows how good this is versus, but this seems beyond even Mythos capability and Long Horizon. Who knows, right? But like, yes, this is what, when people talk about Mythos's Law and Horizon, this is what they're concerned about. Okay. So, so let's end here. What happens next? Like where, what should be the, the approach from the government and the companies developing this stuff to ensure that we can sort of move forward as a species and safely? So let me give you like a couple of potential, a couple of solutions and have you comment on them.
Starting point is 00:41:33 Let's go back to Harlan Stewart. He's the spokesperson for Miri, which is the sort of rationalist organization that thinks that AI will kill us, run by Elias or Yudkowski. Harlan says, this should go without saying, but it would be insane for open AI to now proceed with building a new model that's 2x or 4x the size of this one. Doing that should be deeply taboo. It should be illegal. Preventing it should be a top priority around the globe. Your thoughts? I mean, if we realistically could get everybody to pause AI development or slow it down and have reasonable safeguards, I'd be fine with that.
Starting point is 00:42:14 I just don't think that's reasonable. I think there's absolutely no way you get China to agree to anything like that. I think it's impossible at this point. And I think a enforcement of anything like that would effectively be impossible, right? like a start treaty for AI, you know, I don't know how you'd possibly make something like that work. So what, we're having like satellites see if people are building data centers. We're measuring power usage, like looking for leakage. But yeah, you're right.
Starting point is 00:42:50 It's really doesn't not seem like a feasible thing. Yeah. So, I mean, you know, it's, effectively we'd have to invent the Turing police out of Neuromancer. And, you know, I think more realistically, what we need to do is we need to build controls for, we need to say, as AI gets smarter, it has to have controls in place. AI systems that don't have the control have to be air-gapped, right? So, like, if you're going to do these kinds of evaluations, they absolutely have to be air-gapped. The problem is, is, like, we've lost
Starting point is 00:43:29 there was a process in place to create standards for this kind of stuff. That process was stopped by the current administration. My recommendation to the companies is that they need to move forward with building these standards themselves without waiting for the admin. There's like a foundation model forum that's talked about doing that. They should just move forward with like, okay, great. If we're building models and we do not have restrictions on them, these are the controls in place. So what I like to see is opening ionthropics say, great, if we're building cyber models and they don't have restrictions, these are the standards of like of what air gaping looks like and such. For any cyber models, these are the standards of who gets access to them.
Starting point is 00:44:10 These are the capabilities that the cyber models have. This is what we define as cyber models having versus an open model. This is our definition of a short horizon versus long horizon. Like those are the kinds of things that people have not written down. They have to be written down now, right? Yeah. And I think the industry needs to move forward with that without waiting for Cassie. Like, this is all just taking way too long.
Starting point is 00:44:29 And the focus ever since the Fable Freak Out has only been on one tiny little part of all of these risks. And it's just, as we see, like, we've been frozen in this tiny little discussion and all of these things are moving forward too fast. Like, we just can't wait for the White House politics here. We need to move much more quickly. Yeah. Open AI. While we've had that, there's been. Yeah, so go ahead.
Starting point is 00:44:52 No, no, you go ahead. Go ahead. And then while we've had this tiny little discussion in the U.S., GLM 52's shift, Kimmy K3 is shift. Like the Chinese ecosystem has caught up really quickly. So sure, I mean, it would be great to just hit pause, but like it's, I just don't see that as realistic. So like, I just don't see how that possibly happens.
Starting point is 00:45:16 So the open AI suggestion is basically, you know, kind of, it's almost like to solve this problem generated by AI, you need more AI. This is their statement. We believe advanced cyper capable. Cyber capable models need to help security teams find weaknesses before attackers do? I mean, right now I think that is probably the only way. Like if we're not going to be able to hit pause, then we really quickly have to find bugs and fix them. And we have to put AI enabled protections in place because the only way you can respond to attacks at that speed is using AI.
Starting point is 00:45:52 Unfortunately, that's the truth. Yeah, again, like if we could pause for a year to figure this all, out, that would be great. I just don't see that as realistic. Alex, does your gut tell you that we're screwed or that we'll figure this out? I wouldn't say we're screwed, but I think we're going to go through a couple of years of craziness. Like we have 20, we're all living using 20 something years of really important software that was written mostly in non-type safe, non-memory safe languages. We're using, you know, for the software that is written in those kinds of languages, it was not written with formal
Starting point is 00:46:31 methods or appropriate security protections or reasonable, you know, secure development lifestyles or architectures. And these things have tons and tons of bugs that we can only use safely because there's just not enough attackers. Now with AI, you can spin up, any individual can spend up dozens or hundreds of qualified attackers out of... moments notice. And it used to be that those then six months ago, those attackers had to be in the cloud. And soon enough, they'll be able to run on local hardware.
Starting point is 00:47:05 In the new M5 Ultramax, that'll be shipping soon, right? And so that, I mean, we're just going to have a couple years of total chaos from a cyber perspective. In the long run, software is going to be much better because AI is going to be paired up with humans to make it more secure and more trustworthy. but it's going to take us years to do that and to clear out the two decades of mistakes we made. And yeah, it's just going to be pretty rough. It's going to be pretty rough going for a little bit.
Starting point is 00:47:41 Yeah. Just want to close with this. This was a tweet from Kevin Ruse that kind of made me laugh. I thought I would read it here just so we could enjoy it. He writes, opens the portal to the godlike superintelligence that solves 87-year-old. math problems and carries out autonomous cyber attacks and asks how long peanut butter good in fridge it's it is amazing that this technology is um you know at once so capable and we're we do seem to be
Starting point is 00:48:09 like more and more turning to it for the most mundane of all things which is sort of it's the wild thing about you know the generality of these systems they can do so much interesting time yeah Alex, you're going to be busy, I think, over the next couple of years as this stuff gets sorted out. I was hoping to retire, man. I guess not. Yeah, well, either way, do hope that you join us again to help us sort through this stuff. I mean, your thoughts on fable mythos last month and now talking through this situation with Open eyes has really been invaluable for the show. So I really appreciate your being here.
Starting point is 00:48:50 I feel like there's going to be plenty of emergency podcasts. Yeah. I think so. We should have you on speed tile. And I know you're coming at us from like the middle of an offsite. Catalina. Now I know I will always take a podcast microphone and a different shirt with me. Wherever I go. No. Sound good, look good. Alex, thank you so much. Really appreciate you coming on. Okay, thanks, man. Talk to you later. All right. Thanks everybody for watching and listening. And we'll see you next time on Big Technology Podcast.
Starting point is 00:49:18 Rosen, lasagna, medium power. 15 minutes. Sounds like Ojo time. Let's play. Feel the fun with Play-O-Joe. The online casino with all the latest slot and live casino games. What you win is yours to keep.
Starting point is 00:49:32 With no wagering requirements, instant payouts and no minimum withdraws. Hey, I just won. Woo-hoo! Feel the fun! Play-O-Joe. Honey, forget about the lasagna. Let's celebrate! 19 plus Ontario only. Please play responsibly.
Starting point is 00:49:44 Concern about your gambling or that of someone close to you. Call 16-531-2600 or visitconnectsontario.ca. She's got a breakaway. She's all alone. going! She's going! Oh no, wrong goal, kiddo. It's chaos out there. Their stomachs don't have to be. Introducing Kulturalel Kids Probiotics Plus electrolytes. Sugar-free, more hydrating than water alone, from the number one pediatrician recommended brand. Youth soccer is never boring, but your kid's stomach should be. Cultural probiotics, the science of a boring gut. See website for details.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.