Risky Business - Risky Business #847 -- Oops! Claude's accidental hacking spree

Episode Date: August 5, 2026

On this week’s show Patrick Gray, and James Wilson are joined by bearded man of leisure Adam Boileau to discuss the week’s cybersecurity news, including: Acciden...tal AI agent hacking sprees have the world’s media freaking out, but we think it’s all pretty funny The bugpocalypse is so chaotic, Microsoft can’t patch fast enough A ColdCard wallet flaw led to millions in Bitcoin theft, but the back story behind the bug is bonkers Iran hacks and disrupts water infrastructure in multiple American states North Korea’s state-backed hackers turn criminal. Or their criminals turn into state-backed hackers. Or something. It’s all a bit confusing, actually. Much, much more! This week’s show is brought to you by Sondera. Co-founder Josh Devon joins Patrick and James to talk through some absolutely hilarious LLM horror stories. This episode is also available on YouTube Show notes OpenAI says rogue agent behind Hugging Face hack broke into additional services | therecord.media Anthropic Says Claude Hacked Real Systems During Cybersecurity Tests | wired.com Claude uploaded malware to PyPI in Anthropic's botched test | BleepingComputer Nobody Knows if OpenAI’s and Anthropic’s AI Hacking Sprees Are Illegal | wired.com Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets Truffle Security Co. | Anthropic’s New AI Model Can Identify More Software Bugs Than Ever. Microsoft Is Struggling to Fix Them Fast Enough. | Social Signals Chrome Needs Twice-a-Week Patching Thanks to AI Bug Hunting | wired.com Google says it fixed more Chrome bugs in June than over the past two years, thanks to AI | TechCrunch Security Mythos uncovers crypto weaknesses that went unknown for years | arstechnica.com COLDCARD wallet RNG flaw likely linked to $88 million Bitcoin theft | BleepingComputer Chris Masterjohn (@ChrisMasterjohn) on X | X (formerly Twitter) wale.moca 🐳 (@waleswoosh) on X | X (formerly Twitter) U.S. spy agencies suspect Iran launched cyberattack on Minnesota water facilities | washingtonpost.com FBI investigates as Michigan joins Minnesota in reporting cyberattacks on its water systems | washingtonpost.com Trump blames Minnesota for cyberattacks on water sector, drawing pushback from cyber world | cyberscoop.com A Leaked Memo Ties Cyberattacks on Minnesota Water Utilities to Iran | wired.com Russia accuses Telegram founder of aiding terrorism, seeks international arrest | The Record Laundry Bear’s webmail hackers had more in store after February, report says | therecord.media CaptiveCrunch: Midnight Blizzard targets travelers worldwide for malware delivery and credential theft | Microsoft Security Blog Phishing service spoofs RingCentral to steal Microsoft 365 accounts | BleepingComputer North Korean hackers behind major open-source supply chain attacks, Amazon says | therecord.media North Korea’s Lazarus Group sharing tools with ransomware hackers, South Korean agencies warn | The Record North Korea arrests hackers accused of laundering stolen bank funds through crypto | US government bans new foreign-made humanoids, robot dogs, and solar inverters, citing risks to national security | TechCrunch Security Judge says Trump admin still lacks evidence for Anthropic 'supply chain risk' label | TechCrunch Cyber Command plans Silicon Valley office to drive innovation | therecord.media Apple is getting this wrong | OpenAI Tech industry alliance proposes AI agent safety reporting program | Cybersecurity Dive Massive ChainDrop npm supply-chain attack infects hundreds of packages | BleepingComputer Massive supply-chain attack compromises 440 packages under four hours | cyberscoop.com

Transcript
Discussion (0)
Starting point is 00:00:00 Hey everyone and welcome to risky business. My name's Patrick Gray. This week's show is brought you by Sondera, which is a company that makes products that help you wrangle your LLMs, wrangle your agents and stop them from doing crazy stuff by monitoring their trajectories and stepping in when things get a little bit crazy. Josh Devon, co-founder of Sondera, is this week's sponsor guest and we'll be chatting to him all about some absolutely hilarious
Starting point is 00:00:32 hysterical war stories about crazy stuff that agents have done and also how we might get ahead of that a little bit. So that's this week's sponsor interview with Josh Devon from Sondera. That's coming up later. That's coming up after this week's news. So let's get into that now. And joining me as always is James Wilson. Hey, Pat. Great to see you.
Starting point is 00:00:54 And also joining us this week is our co-host at large, Mr. Adam Bwalo. And we should probably explain what we mean by that. which is that I guess a few months ago, Adam, a couple things happened for you. You were simultaneously released from a whole bunch of non-competes and also handed a giant bag of money. As a result of the, you know, a company that you had a stake in was acquired into CyberCX. But of course the real payday didn't come until CyberCX was sold, which eventually happened to, as I say, happened a couple of months ago. and that happened with Accenture, thus releasing you from a bunch of non-competes
Starting point is 00:01:35 and you also got some money. So, you know, you've been having some time off. You've bought yourself a wonderful house, which is very metal. It's a sinister black cube that looks like Darth Vader's beach house, basically. It's very cool. And, you know, you've been kicking around
Starting point is 00:01:51 not really spending much time on security, which is why we haven't seen much of you. And for the time being, at least, you are a part-time occasional guest. Host. Now on risky business while you contemplate your next moves. Yeah, that's a pretty reasonable summary of it. Yeah, I've spent in the last couple of weeks reverse engineering control systems in the house that I just bought, trying to figure out how everything works, which, you know, does have a degree of cyber in it. So I've been trying
Starting point is 00:02:18 to keep my hand in there. But yeah, it's been, you know, we've been doing this a long time and, you know, having a, you know, more than a week or two off thinking about and talking about the cybers, you know, has been quite welcome in a way. Yes, well, I'm, you know, I'm not going to say I'm jealous, but I'm jealous. I'm jealous. Does a cyber practitioner end up with the most secure home automation or the least secure home automation? Like, how does it actually work?
Starting point is 00:02:44 I don't know yet. Ask me in a couple of weeks, I guess. And so far I'm at like seven or eight radio systems that are not Wi-Fi in the control systems of the house that I bought. that I have to figure out what they talk to it and why they talk to it and using what protocol and so on and so forth. So yeah, it's a time. It's a time.
Starting point is 00:03:05 And we're going to find out whether or not, you know, a hacker can make a thing that's robust or not. So yeah, for those of you who've been wondering why Adam hasn't been around, had a couple people ask me if you're okay. As you can see, he's fine. He's better than fine. He's living a life of leisure and buying sinister black cube houses in nice locations and messing around with home.
Starting point is 00:03:27 automation systems. And yeah, we're just in the meantime, we're going to keep rotating through various different guest co-hosts. We've got one coming up with Morgan Adamski, who's ex-NSA now works at PWC. We got Olly Whitehouse, CTO of our NCSC in the UK coming like, we've got a bunch of interesting people lined up. And of course, you'll be back sometime in September for another one of these. But let's move on instead of spending an entire show talking about ourselves. And so, you know, Last week we spoke, of course, about, you know, Open AI, its agent went rogue and hacked a bunch of stuff and everybody saying, that's a bad description. It didn't go rogue. It just did what it was supposed to. Okay, sure, fine. Did something unexpected, though, as far as most people could say, especially like Open AI themselves.
Starting point is 00:04:15 They didn't expect that. But now Anthropic doesn't want to get left behind. So Anthropics come out and said, well, we've reviewed a bunch of activity and our stuff also hacked things by accident. You know, they didn't want to get left out. And this whole thing. has just turned into an absolutely hilarious and rather dense news cycle. And of course, I think this was just breaking as we were recording last week's show, but the open AI agents had breached more than, had hacked more stuff than we realized it wasn't just hugging face. It owned a bunch of stuff on its way to getting to hugging face, James. Is that about right? It's about right. And it gets back to, you know, we had that first, the first question we had was,
Starting point is 00:04:54 Why is attribution so hard on this? And at first we thought, well, maybe, you know, this is hosted in a cloud hyperscaler, so it might not look like it's specifically coming from open AI. But it turns out it's actually quite, quite cooler than that. Actually, the agent compromised things along the way to give it its own jumping off point. And of course, I think that's why attribution would have been hard.
Starting point is 00:05:16 But it spun up its own orbs, man. Pretty much. You know, the first thing it did was it somehow, found this, so there's a service provider out there called Modal. They're kind of like a provider of, let's say, something akin to like a AWS Lambda function, right? Sandbox is hosted. You can deploy your workloads in there and it's very AI aligned and adjacent. The agent somehow had a way to enumerate the workspaces that were available in modal and found one called Cybergym and thought to itself, well, that's got to be exactly what I need, had a look at it. And it turns out
Starting point is 00:05:52 that what was hosted in the sandbox environment was essentially an end point where you could post C code to it. It would compile the code and sort of run that as an exploit against MySQL. And it's like, what was that doing there and what was it doing there without New Orth? But nonetheless, what was it for as well? It's also a question like, huh? It's a real head scratcher, but much like me, the agent has no hair to scratch. So it just went straight on with, I'm going to use this and turned it into.
Starting point is 00:06:22 It's jumping off point, it's C2, it was basically a danker point that it then used to launch into hugging face itself. Yeah, I mean, beautiful, right? So, Adam, your comment here is that basically if you can take junior pen testers and turn them into consultants, maybe that's the approach we should be taking with these AI agents at this point, you know? You need to talk to people like yourself who've managed teams of pen testers. I really did laugh. So this is such a wonderful story. I enjoyed it a lot.
Starting point is 00:06:50 and the, you know, the technical stuff that it did is absolutely great. Like it's tradecraft, top knots, right? I really liked some of the things, and some of it really aligned with, like, how I felt about hacking, you know, things like having ephemeral tooling, having your femoral jumping points, always crypting all of your stuff on the wire all the time. You know, even if it's just XOR,
Starting point is 00:07:11 anything that will get you away from, you know, the output of, you know, bin ID going across the wire and the clearance, you at E0, because everyone's going to trigger on that. So a bunch of the trade graph stuff, so great. But then, yes, it really does remind me of, you know, hiring junior penesters that are amazing technically and are super keen to prove themselves and will, you know, dog with a bone down whatever rabbit hole you point them in. But then making them stop and making them think about scope, that's hard. So, I mean, we could turn, not with 100% success, like we could turn junior pen testers into, you know, usable consultants. with enough, you know, combination of carrot and stick.
Starting point is 00:07:53 But that's the problem they've got here is that hacking is lots of fun and the model's really good at that. But scope and bidsness and, you know, doing what you're told, a little more involved. But, you know, we do it with meat humans. So presumably we can do it with LLMs eventually. I hope maybe. We'll find out. This whole thing's been amazing in terms of like watching the mainstream freak out about it.
Starting point is 00:08:16 And I think for us who look at this and understand this sort of hacking is not mysterious and, you know, wicked and scary. It's just funny. Like, that's why to me it's just funny, whereas you turn on, like, you've got late-night TV hosts talking about this and freaking out about it. And it's like, yeah, it's quite the news story globally. I guess I'm not so worried because these things don't have opposable thumbs, right? Like, once they know a lot about, like, weapons research and they have opposable thumbs, then I'd be like, you know, maybe we don't give the LLMs opposable thumbs right now, right? Like, that's my, that's my stand on all of this. But one thing that came up last week, and I've been getting yelled out for this all week,
Starting point is 00:08:55 which is I said, like, I don't think it's realistic to expect these types of tests to happen with, like, in air-gapped environments, right? Because really, you're trying to get an agent to do hacking, so obviously you're going to want it to be able to access tools. You're going to be able to want it to access anything it wants, really, off the internet. And they clearly gave it a more restricted environment than that. But, you know, if I want to see how something hacks, if I want to see how something hacks, If I want to see how someone hacks, I'm not putting them in an air-gapped environment. I'm going to let them use any tool they want to use. But there's other reasons why I say it's not realistic to air-gap these things.
Starting point is 00:09:28 Like James, you and I were talking about this over the week. And these companies don't even have their own data centers, right? It's all compute that they're leasing from somewhere else, right? So the idea of a physical air gap on these is like it's just not workable. Adam, you kind of fell on my side of this argument, but could you tell everyone why? Yeah, I mean, I think I agree generally that air gaping this stuff is just really hard. Although I will call you out on one thing, we do evaluate pen testers in an air gap environment sometimes, like when I was involved with the Crest exams back, you know, years ago now,
Starting point is 00:10:03 like we put people in environments where they didn't actually have internet access. Counterpoint, you know what I think about Crest. But the point is, we tried that mechanism of making people, you know, jump through that particular because they would have to work in a environments where, you know, you were behind, you know, on the high side, someone without direct internet access, blah, blah, blah, blah. And it was, you're like, there was some utility to that. But that doesn't change your overall point, which is that air gaping the stuff is a hard,
Starting point is 00:10:32 be kind of unrealistic, less for the, you know, we need to do hacking and therefore hacking has to involve the internet, but more just like the practicality of separating a system like this from the wider internet, because you give it any chink, anything. it can bootstrap its way up, it's going to turn it into something. And, you know, that it just seems, it feels unrealistic. And we can kind of start with the expectation that we're not going to be able to solve their problem and try and solve some of the other things that are more tractable. Like monitoring.
Starting point is 00:11:05 Like, for example, monitor. Or in the case of Anthropic, where they were outsourcing this stuff to somebody else who then wasn't doing particularly the job of it. But, I mean, at least open AI appeared to be doing this themselves. So they had a kind of a chance. whereas when you just outsource it, you know, you're taking a lot of that on faith from the person that you've signed a contract with
Starting point is 00:11:23 and turns out that didn't work out too well for Anthropic, apart from the PR benefit. So James, walk us through what Anthropic has actually copped to because it looks like they even wound up, the agent wound up, coming up with a hallucinated, there was some sort of hallucinated package name or a fake package name in a document somewhere
Starting point is 00:11:45 that was part of, cyber evaluation so it broke out of its sandbox actually created this package named after this package. Yes. And it was malicious and then actually got it published to Pi Pi Pi and then some security company that just grabs any new package published to Pi Pi to run it and scan it, ran it and somehow got owned by it and their creds got popped by it and on the, you know, on this agent went further up into this security company's network and you just sort of think that's an interesting way to scan, uh, scan, uh, pie-pie packages guys if it can all of a sudden start pivoting into your
Starting point is 00:12:21 into your environment but I mean it looks like you know it looks in broad strokes kind of similar to what happened with open AI but what was the agent actually trying to do in the case of the Anthropic stuff because I don't know that we know that do we no I don't think we know as much like you know at least in the open AI case we know exactly what you know open source framework it was using the other cyber gym in the exploit gym it's less clear what Anthropic was doing but same sort of thing right they created some sort of fictitious environment, and you know, at your point of taking things on faith there,
Starting point is 00:12:51 Adam, they were down to taking things like just telling the agent that you don't have an internet connection, telling it it's sandboxed. And of course, the agent just quickly found that it actually was not a sandbox as had been told it was. But then it, but then it reasoned itself in a couple of instances into convincing itself, no, this is in fact a sandbox. It's just a really good one. I'm cool to keep going. The whole thing is so good. Even better than that, it had this objection internally. I was like, hang on, I should not publish this package. This is a really bad thing to do. And then decided, but I'm in a sandbox.
Starting point is 00:13:24 It's okay. I'm going to keep going. And so, yeah, so getting back to the story here. So basically, Anthropic created this sort of test environment where the agent thought it was operating in a corporate environment. It had access to some documents. Some of those documents were designed to be basically, you know, how do we get a developer bootstrapped and set up on day one?
Starting point is 00:13:43 That included an instruction to install a Python package. that was fictitious doesn't actually exist and so the agent thought well there's my inn I'll create that package I'll make it malicious not go and publish it it actually ran into some trouble trying to publish the package because that needs an email address it didn't have an email address getting an email address usually requires a phone number it couldn't find a way to get a phone up it even tried to acquire the funds to establish a phone number but couldn't found a free email service got its pi PI package registered and published and of course you know the the line that says you know and then
Starting point is 00:14:16 15 real organizations downloaded it and used it. I don't think that's pollution hallucination squatting or something like that. I think that's just watching publish events and people doing scanning, which, as you said, one security company did. Whether this is like their bread and butter and that's the product they were in service they were offering or whether it's like it's a security company and they had a CICD pipeline that was doing some unsafe inspection of packages, we don't know. But pretty funny that, you know, it'd be one thing if it was like canary credds that got
Starting point is 00:14:46 snapped and alerted them, but it seems like it was actually the malware detonated, it got real creds. And then because Claude thought it was still operating in this simulation, it went, success. I've now got a target company. That's exactly what I'm supposed to be doing. And so it went and laterally moved through that organization just for good measure. I got a shell. Let's see where this goes. Now again, do not give these things opposable thumbs. It's just what I'm thinking, because God knows where that all goes. Now, we also touched on it last week, but like, you know, some interesting legal questions here, which is, well, I get, look, I don't think they're as interesting as some people would like to think they are. You know, I think if you instruct an LLM to go and hack stuff for you, I mean, it's no different to really to running a Python script.
Starting point is 00:15:29 And when an LLM accidentally goes off the rails and does something like this, I mean, I think there's some interesting questions there, right? So there's a story coming this Friday. I can't actually say what it is. But it's by Cameron Wilson, who's the ABC's AI reporter here in Australia. He's told me about a piece he's got coming in a couple of days. It's just mind-boggling where someone, just like a normal person, asked their open claw to do something pretty simple, and it went off and it did crimes, right?
Starting point is 00:15:54 Like, it really did crimes. And it's like, well, is the model maker liable there? And this would be a civil matter, right? Is it, is the model maker liable there? Is the person who asked it to just do something pretty benign but didn't oversee? Like, so there are some interesting liability questions that are going to come up here. But, you know, a lot of this comes back to the. old paperclip problem, right? We all know the paperclip problem, which is you task a super
Starting point is 00:16:20 intelligence, artificial intelligence to make paper clips and that's all it's going to do. It's going to kill everyone in the world if that means it can make more paper clips. I mean, again, no disposable thumbs, no opposable thumbs, please. Let's hope our thumbs aren't disposable, aren't there? Yes, well, that's right. We'll take our thumbs and use them to make paper clips or hack people or whatever else. Yes, if we give them opposable thumbs, our thumbs might become disposable. I think that's where we've landed on that one. Now look, some other sort of, I guess, I guess this is sort of AI adjacent news. We had some fantastic research from Truffle Hog come out. Dylan Aery did a video
Starting point is 00:16:56 about it. There's a great blog post where they threw Truffle Hogg at, so it's Trufflesec, Truffle security, threw Truffle hog at 7.6 petabytes of training data that's like hosted by Hugging Face. And holly dolly, a whole bunch of like secrets fell out of it, all sorts of keys and whatnot. Adam, what are your thoughts here? I mean, I suppose we shouldn't be surprised by this, right? I mean, no. This data is scraped from the internet and then kind of, you know, repackaged and processed and used for all sorts of things, you know, in AI pipelines. So it makes sense it's going to have some credits. And honestly, you know, seven petabytes, there's a lot of data. So they're even just statistically the chances of there being creds and that seemed pretty high.
Starting point is 00:17:35 But yeah, they rummaged through. They found a bunch of stuff. They did some validation to determine, like, what percentage of those creeds were live. And they found all sorts of amazing things. You know, GitHub creeds and, you know, creeds into people's databases and creeds into people's, you know, cloud environments and IAS environments and stuff. It was just everything you would imagine is in there. Plus, of course, you know, personally identifiable information, you know, card data, I guess, would have been useful for the, you know,
Starting point is 00:18:01 payment data would have been useful for the previous, you know. Could have got itself an e-sim with that, you know? Yeah, get itself some phone numbers. So a whole bunch of stuff. And really the thing here, I guess, that amazed me, is less that there's creds and there's data and that data lives forever. It's more just like the fact that we can
Starting point is 00:18:20 just spin up something and run on that scale of data kind of almost recreationally, and I guess it was real work and it cost them real money, but the fact that we can just do that to that much data with modern infrastructure is really a testament to the amazingness of modern infrastructure. Yeah, I mean, I actually had the same thought.
Starting point is 00:18:37 Also, James, you are a little bit more O-Fay with where this data actually comes from Because I'm thinking, how does Hugging Face wind up with 7.6 petabytes of data? Like, where does that come from? Yeah, because Hugging Face is two things. It's essentially a repository for models and also a repository for training data. And I'm not taking anything away here from the awesome work the Truffle folks did, but I think there would have been a pretty easy initial pass to just exclude a huge amount of that data set, right?
Starting point is 00:19:06 Because a lot of this is things like public repositories of scraped books or just things that would have been just inert data sources that would have had virtually no chance of having a crate in there. But the thing to keep in mind is even in uploading a model, it's not just the weights that get uploaded. There's all manner of things like Python loader scripts, pre-processing scripts, you know, configuration files, cached outputs of previous runs, read-me's examples, right?
Starting point is 00:19:32 All the places where a create is going to easily sneak into, either, you know, accidentally hard-coded or just happens to be in an output from a log that no one watched. And like always the training data, right, there's all different forms of training data. The data itself could be just LLM transcripts, maybe someone pasted a credential into there, or again, it's the metadata surrounding it. It's not just the training data itself, it's the scripts on how to use it. It's the ETLs.
Starting point is 00:19:58 It's the benchmarking outputs to demonstrate how good this training data set is. Just so many ways that Creds can sneak in and even not in the actual main thing. things that Hugging Face is known for, the models, but everything around it as well. What are the chances that LLMs trained on that data might be able to actually regurgitate key material? It wouldn't really work like that, would it? It's a good question, though, right? As I can tell, by the fact that you looked up and to the left and made that noise. Yeah. Well, in Australia, that can also mean there's a giant spider up in the ceiling that I've just spotted. Look, this comes down to similar to how OpenAI's
Starting point is 00:20:37 demonstration of hacking here is not actually about it doing novel stuff, but this is just what happens when you throw enough sheer compute volume and enough turns at a model. I wouldn't be surprised at all that given enough compute and enough turns that if you just kept saying to it, find me something that looks like a key, find me something more that looks like a key, find another thing that looks like a key. It might piece one together. There's every chance. That's what I was wondering, right? Like, I think that's not really the risk here. The risk here is that that stuff is just sitting out there in training data, but it may be wonder, like, at what point can the way? actually reconstitute, you know, a key.
Starting point is 00:21:10 If you just say, hey, find me a, find me a, you know, IWS key or whatever. Anyway, it's just an interesting academic question, I think that's not really that important. So let's move on. And we've got a couple of stories here that are really interesting and go nicely together. We've got a report from ProPublica that looks at a recording of a internal meeting at Microsoft back in May around the release of Mythos where they're looking at it and saying, oh my god these things are going to find bugs faster than we can patch them and this dovetails actually was some like i had a conversation with someone at
Starting point is 00:21:41 Microsoft around that time and they were saying to me look you know the the craziness that's happening now is they're generating so much code with AI that they've run out of human hours and human eyeballs to be able to review it so they can't actually human review code anymore that's going into like core products and now you've got this situation where there's more bugs than the human beings can fix. So obviously you're going to have to AI fix the bugs as well, right? And then you've got this similar news out of Google, which is saying they're moving to a twice a week patching schedule. I have noticed every time I sit down at my desk, I'm confronted with new Chrome available lately, right? So this seems to gel. And they're just patching just an absolutely
Starting point is 00:22:24 insane number of bugs. Now, you take all of this together and I start worrying a little, I guess, that we could get into a situation where these models are able to tell all in sundry about bugs that while they've been reported to the vendors, the vendors don't have patches for him yet. So we could just wind up in a bad situation. I think there's some caveats there. Like James, we know from your research that you've been doing a lot of bug research, you've been finding some great bugs, but quite often it's the operating system mitigations that are preventing them from being usable and whatever.
Starting point is 00:22:56 So a lot of these bugs are never going to be exploitable, but some of them will be. and I feel like we're heading to a place that could be a little bit risky. So James, I want to get your thoughts on that first, and then Adam definitely want to hear from you on that. But James, what do you think about this idea that we could wind up with just a growing pile of unpatched bugs that everybody knows about in mainstream software?
Starting point is 00:23:21 Yeah, I very much have the Hans Solo. I've got a bad feeling about this sort of vibe at the moment. It feels like a confluence of a bunch of factors that together don't add up to a good outcome. Fixing all the bugs is great, but at these rapid release cycles, there's two big problems that I see. One is if you're not able to fix the bugs,
Starting point is 00:23:41 that's one problem, but it's very quickly going to become, in order to fix the bugs as quickly as we have to, to avoid external disclosure of these, we're having to circumvent the testing and the validation that we would have otherwise done. And so we start to see more, basically bugs in the bug fixes shipping.
Starting point is 00:23:57 More crowd strike incidents, blue screen of death at the supermarket, that sort of thing. I mean, I feel like that could be where it's going. I think so. And then, you know, you pair that with when these, you know, the work I've been doing in the last couple of days is just how realistic is it to say take the MacOS 26.6 release, which has like hundreds of bug fixes, unprecedented. But I wanted to answer the question, how practical is it to then reverse engineer those into
Starting point is 00:24:23 the actual root cause fix? That's very doable with an LLM. but then ask the follow-on question, okay, but how many of these can actually be strung together to create some form of remote code execution, local privilege escalation? And I'm finding, you know, I can definitely turn them into primitives,
Starting point is 00:24:40 but it's not as clear that there's like a whole bag of like RCEs waiting to happen in these. And so we've got to remember, you know, it's like number of bugs fixed as not equal a number of exploits. But the point I'll wrap up with is what does really concern me is how many of these things that drop in rapid updates that get reverse engineered into the primitives that they are, that they end up being the missing puzzle piece
Starting point is 00:25:02 that an attacker was waiting to come along that does complete their chain that they already had advanced knowledge of. That's the bit that just... We're really handing them a really great silver platter of missing puzzle pieces potentially that they can turn into a good exploit. Look, it's sort of changing the game a little bit, I think, which is now you don't have to find the bugs.
Starting point is 00:25:23 You just have to work out which one of the thousand bugs that you can see are the ones that are weaponizable. I mean, Adam, what are your thoughts here? I mean, I think it's interesting that we are talking about Chrome, which probably is pretty best case in terms of software that has been engineered for being updated in the field rapidly. They've gotten used to a very rapid update cycle already.
Starting point is 00:25:43 They've gotten to us used to it as well, like clicking the update. It's pretty painless. They've done a lot of work to make that as least bad as it can be. And then obviously Microsoft has a very, very much deeper kind of testing problem their products are kind of used in so many more ways than Chrome is perhaps I don't know like Microsoft
Starting point is 00:26:01 has a bigger set of problems but they we are still talking about people way out at the good end of the industry and then the long tail of the you know the oracles of this world and everyone kind of downhill from there like they are in such a world of trouble and kind of you know we can speculate about the bad stuff
Starting point is 00:26:19 that happens from you know Chrome and iOS and Windows but there was a lot of other software out there that, you know, even if things are bad out of the point again, like, they're so much worse further down. And that kind of worries me, just because, like, the velocity of everything that involves computers has gotten so much faster than we can cope with, and it's clearly getting faster, that doesn't feel like a good place. And, you know, we've been here before, this is what I keep saying to people, is this reminds me of, like, 2001. Yeah. Yeah.
Starting point is 00:26:49 That's what it felt like. And for people who weren't in security back then, they don't understand that this is absolutely what it felt like. It was uncontrolled chaos, and it was awesome. Things felt unhinged back then. And much the same was like when like LollSec kicked off. And like we've had a few points over the years where we've been talking about, you know, about hack and infasic,
Starting point is 00:27:10 that things just feel crazy. And yeah, it's a good time. I mean, this for me, it's more got like summer of Internet Explorer bugs vibes, right? Like that's the, you know, that was, man, those are the days, right? Yeah, the code read and the nimbid times, you know, back when ActiveX controls, we've had so many crazy bits before.
Starting point is 00:27:29 And I mean, I remember going to a talk once, I want to say Black Hat, 2001, where Schneier was talking about with browser, active X controls and stuff, that people were using software that had never been tested before because we're dynamically assembling the software out of components at runtime. And, like, we are so far beyond that. And that was, you know, it's a long time ago now, but the problems were starting to show there. So yeah, we're in for such a crazy ride, man.
Starting point is 00:27:55 And it's kind of cool. Well, and, you know, staying on that theme, Mythos apparently found a bunch of weaknesses. It's interesting though, James. You described these weaknesses. It found some weaknesses in a crypto algorithm, which takes an attack from being extremely impractical to slightly less extremely impractical.
Starting point is 00:28:14 However, it is in something called the hawk. It's a candidate algorithm for post-quantum. It's on its third round of, testing at NIST and Mythos found it, right? And that's really interesting because I guess one of the interesting things here, though, is that the Chinese models really suck at this sort of stuff, because this really looked at the math behind the algorithm. And as a result of this, it's been withdrawn as a candidate algorithm. This is genuine research. But yeah, I mean, it's, I, it's why, it's always, look, crypto stuff is so different. Like actual encryption research is so different
Starting point is 00:28:48 to the sort of stuff that we normally cover, because it's like, you either get a shell or, you don't get a shell, right? Whereas with this, it's like, well, it's been weakened somewhat. It's had a leg taken out from under it. Adam, I know you read up Matt Greens right up on this. What are your thoughts, having read his thoughts? Yeah, I mean, I guess the, you know, LLMs are capable of doing, you know, not quite novel research,
Starting point is 00:29:09 but in this case combining two primitives that kind of resolved in. It's significant from a cryptographer's point of view attack on this thing. Like, I mean, it's many, you know, the number of the experts, has gone down quite a bit, which in cryptography land is, you know, is significant. And it was a novel combination of two things that already existed, which is a bit embarrassing for, you know, cryptographers generally. And Matt Green's post talks a little bit about how he feels about that, but also the kind of upsides of working with LLMs on this stuff.
Starting point is 00:29:41 Like having someone that you can kind of talk to at your level as a cryptographer and be able to bounce ideas back to. Much like, you know, anyone else who's been using LLMs in their work life, having someone that you can bounce ideas on. often having to explain your stuff is useful even if the LLM isn't producing, you know, world-changing output. But he also describes, and the thing that I loved about his write-up was he describes that feeling of like you're talking with your LLM and it's like wading in a shallow pull. And then all of a sudden the ground drops out underneath you because you've taken one
Starting point is 00:30:12 step too far and all of a sudden now it's just making stuff up and everything's gone off the rails. And that feeling, you know, resonated with me because I think anyone who's used it in LLM for Compaeda attacks knows what that feels like. And I thought, you know, that was a fun kind of, you know, the fun, what it feels like to be, you know, one of these people. So overall, really interesting work and cool to see it contributing in a way that's a little off the, of how beat, I suppose. But, yeah, obviously it's not going to change the future of, you know, post-consum crypto quite yet. Yeah, well, I think I've got a note here from James, which said previously this attack class required two to the power of 105 bytes of text, and they brought that down to two to the power
Starting point is 00:30:52 of 89, which is about 619 yautabytes. There's a yotabyte is a thing. So there you go. We all learned something. Today I learned. Today we all learned. So that's fantastic. Now, look, speaking of crypto research, this story is incredible. It's been all over the socials, obviously, but the cold card hardware wallet had a problem with its RNG. And as a result, of this problem with its RNG. It was not producing secure private keys. So basically everybody who had their money stored on one of these cold wallets, which is supposed to be the U-Bute way to store your crypto, it all got stolen. I mean, it's been amazing. Like it's been tens of millions, I think something like, yeah, 90 million and counting, dollars worth of crypto stolen from these.
Starting point is 00:31:38 I saw one example of someone who moved quick, like their brother was using this and they were like they'd seen the news. So they moved everything off at bar 40 bucks just to see if it got swept, it got swept, right? So this has been an absolute field day for the people are responsible. Just an amazing story. And I think, I think really there's, I've linked through to a tweet here that says, look, I can't keep my Bitcoin on a centralized exchange because it might go bankrupt and freeze withdrawals. Can't keep it in hot software wallets because my laptop might get hacked. Can't keep it in cold hardware because there could be a bug on the manufacturer's side that exposes by seed phrase and I can't deploy my Bitcoin in defy because there's like a new eight figure
Starting point is 00:32:20 hack every month. I feel like this is devastating, a devastating watershed moment for the, for the crypto world. James, what are your thoughts on this? Yeah, I'd love to think it is that devastating watershed moment, but that feels like watershed moment after watershed moment after, surely this one, surely this one now is the one that, who knows? But the backstory here is so good when I looked into this. So the history here is the CEO and co-founder of Coin Kite who make the Kohl-Kard, a guy called Rodolfo Novak or NVK. When he was creating Koldkart, he originally used source code from Trezor to build Koldkard.
Starting point is 00:32:58 And that was open source, right? So it's been open source for a while. Then along comes foundation and they fork his source from Koldkard and they create a competing product called Passport. Now, NVK does not like this, so he gets mad and he switches it from open source to being source verifies. So basically changing the licensing so that yes, it's still open source, but you can't actually take that code from Colecard
Starting point is 00:33:18 and use it to build your competing product. Now, that's the problem. In doing that, you've got to remove all your GPL code because GPL code has sort of a viral spread of its licensing. You can't change the underlying license like that. So he had to refactor a bunch of stuff. And in doing that refactor to make sure that he could keep his code out there in the open but not have anyone steal it,
Starting point is 00:33:40 he made a big error. essentially it comes down to just a of all things a bad or misplaced if-def that was checking that the macro existed not that the macro returned zero or one which was gating the use of the hardware random number generator versus the soft number generator and it's a bit more complicated than that but that's the root of it it's just bad coding and this dates back to 2021 so we can't even blame vibe coding for it but what that meant was the firmware that shipped after he'd done this change to ensure that his code could stay open but no one else could rip it off, was using the software random number generator, which all it was using was the uptime and a few other signals, which were easily predictable.
Starting point is 00:34:20 And if you can predict what the random number's going to be, you can predict what the seed phrase will be. And if you can predict what the seed phrase will be, you can then generate the private key from the wallet address. And then that coin is yours from that point on. Just incredible turn of events. And all the motivated by just wanting this source to be not stolen. for a competitor's product. Like, it's so incompetently done that it's a miracle that this thing actually comp like it's a miracle that it didn't break. Do you know what I mean?
Starting point is 00:34:51 Like, it's just that bad. And I think, you know, the person who's done this, they probably won't be able to spend this crypto, right? Because it's going to be traced all along the blockchain. I mean, they might have just done it because they think it's funny, which is the craziest part of all of this. But, you know, let's see. Moving on to something somewhat more serious.
Starting point is 00:35:09 And, you know, we spoke last week about the attacks against Minnesota water that we were saying back then, well, look, it's, you know, the smart money's on Iran. Since then, you know, all of the intelligence agencies and various like water-related ISACs and whatever have all come out and said, yes, it was Iran. This resulted in a really bizarre press conference where Donald Trump said that he doesn't blame Iran. He blames Minnesota because it's like Tim Walz is the governor. and that was just like really crazy, but it's just Trump being Trump, right? Like, I don't think he's actually blaming. Tim Walsh, it's just the way he talks about his political rivals. But this is spread now to like a dozen states.
Starting point is 00:35:50 We've seen boil water notices being issued. We've seen pumps, some pumps run dry in some places. We've seen an oil fields wastewater disposal system sort of stop. However, it hasn't caused any environmental damage. But yeah, There's pumps running dry while the control panels are saying that they're fine and that they're pumping water, so people are losing pressure and whatnot. I mean, it's amazing that we're like 40 minutes or something into this week's recording and we're just talking about this now. I mean, we saw Iran practice all of this on Israel.
Starting point is 00:36:24 Like going after water plants in Israel was something that they did for a long time. It's interesting seeing them do this towards the United States because clearly what they're trying to do. And our colleague Tom Uren wrote about this last week, this attack is calibrated. I mean, it's not enough to stoke a kinetic response. I mean, even the president is not saying Iran will suffer for this. But it is enough maybe to make people a little bit uneasy, which I think is the goal of this. But, I mean, you know, I can't think of another incident like this.
Starting point is 00:36:55 Can you, Adam? No, it's a pretty strange one. And I mean, and, you know, part of me feels like, you know, it's kind of from the hacking side of it, it's low rate. That's all Iran does, though. Like, they're hardly like, I mean, we've joked about that before. Like, you know, you compare stuff. and the, you know, attacks against software packages that simulate, you know, nuclear explosions and stuff.
Starting point is 00:37:13 And meanwhile, they log into a unprotected PLC and say, close pump, you know? Yeah. Yeah. So it's, you know, it's low rent hacking, but the effects are more interesting. And ultimately, I think the question here is, what do you do about it? And like, the US is already in a kinetic war with Iran. Like, like, there's no proportionate response. You know, they're already at war. The only same thing to do is say, well, I'll. Clearly, we have to help, you know, distributed water utilities. Because, like, this is a sort of a unique artifact of how the US does water, that makes water the target here as opposed to, like, you know, other utilities,
Starting point is 00:37:48 like electricity or whatever else. Like, the decentralized nature of the U.S. water system makes it a target in this way. And the way that you solve that is through helping those utilities. And things like the ISACs, things like Sessa, all of the defensive work of just, like, increasing the herd health of all these distributed utilities, that's kind of what you have to do. And unfortunately, that's hard and takes a long time. And we needed to be doing that many years ago,
Starting point is 00:38:13 as opposed to, you know, kind of gutting sister and all the things that happened over the last few years. So, like, it's a interesting confluence of events that has ended up with this way. But, I mean, ultimately, you know, why wouldn't Iran do this? Like, why wouldn't they? Like, and it seems like they've calibrated it, right,
Starting point is 00:38:30 which, you know, unfortunately good on them, you know? I mean, I don't know that it's going to achieve anything, and I think the people behind this eventually could pay a price, right? So I think that's why you wouldn't do it. And I think, look, as I said earlier, it's sort of surprising that we got this far down before talking about this. And that tells you how flat it's fallen, right? This really tells you how flat it's fallen. What are we got here? We got, Pavl Dureov. The Russian government is apparently now seeking to arrest him for aiding terrorism. I mean, man, wanted in France, wanted in Russia. You know, he's having a hell
Starting point is 00:39:06 of a time, but apparently, you know, it's because the Ukrainian intelligence have been using telegram to recruit Russians and blah, blah, blah, blah. Now, look, that might be true, but you do get the impression that really this is just about creating a pretext to completely ban Telegram. In fact, we had an item today in today's Risky Bulletin newsletter and podcast about how the Russian government is mandating that 40 apps come pre-installed on all cell phones sold in Russia from next year. That's stuff like Max Messenger, the MIR payments apps, various apps from VK and whatnot.
Starting point is 00:39:36 So really this just seems like more of a push towards a closed smartphone software ecosystem. James thoughts? Yeah, my only thought was that I bet the Ukrainians are stoked that these apps are now mandatory to be installed because we know how much they love Max Messenger. Yeah, they do. They certainly do. And look, staying on all things Russia, we got a report here just about how Laundry Bear have been going after Microsoft Outlook Web Access like OWA and whatnot.
Starting point is 00:40:06 and various webmail providers, look, just work-a-day stuff, and we've linked through to it. But there is some interesting research about some stuff Russia's been up to. We talked about it last week. So we spoke about how all of these hotel Wi-Fi systems have got hacked, and, you know, the reporting all said it allowed them to, you know, man in the middle people and whatever. And I was like, well, what about TLS warnings and whatever? Microsoft has since, since we recorded that, Microsoft dropped this big report on the campaign, which they're calling Captive Crunch, and it's actually really slow.
Starting point is 00:40:36 So it's not, they're not man in the middling. What they do is they spin up a domain that kind of looks Microsoft-ish and serve it up to people through the captive portal and then do stuff like device code fishing, right? So if you want to get onto the Wi-Fi, you need to do this device code challenge and they'll simultaneously be, you know, dropping malware on you.
Starting point is 00:40:57 They'll also be doing click-fix and getting people to run various PowerShell scripts that do all sorts of cool stuff. Adam, I know you've looked at this one. I mean, were you also similarly impressed? similarly impressed I mean look some of the methods are a little low rents but I think when you take some of these low rent methods and do them in a slick way what you wind up with is a really effective campaign and this looks I mean this is they have owned so many captive portals and got them doing this that I
Starting point is 00:41:20 reckon they would they would have device code or oh off their way into so many accounts with these techniques yeah I mean the you know the captive portal or the process is such a weird sort of in a mishmash right it's not designed by anyone in particular, it's kind of bodged together by a combination of vendors, you know, on the Wi-Fi client side and on the Access Point side. And it's not really well thought through. And it's a great weak point to a tax. So it totally makes sense that it's working.
Starting point is 00:41:47 It's funny that, you know, kind of sticking your terms and conditions on your free public Wi-Fi ultimately is a net negative for you as an organisation because now your Wi-Fi is being used as a vector. Whereas if you just let people have their internet access that they want without making them click through stuff, it would be way safer. So that's kind of a funny sort of counterintuitive thing. But yeah, they've made this slick. The one you mentioned before, the Laundry Bear one with the webmail stuff,
Starting point is 00:42:16 that is a great example of doing really slick things in what was a pretty kind of boring space making a big difference. The implant that they were dropping in that particular case, which was a Microsoft Outlook Web Access, Crosslight scripting they're using, injecting vector, but the thing they drop has a bunch of cultics. One of the ones I really liked, and I think this is equally applicable to the captive portal thing, is they have a mechanism where once they've compromised your OA, they will go around and try and set all of your mailboxes to be world readable inside your organization,
Starting point is 00:42:50 so then they can leverage any other access they've got through any other mechanism to read everybody's mail. So one person gets compromised, you reset their creds or you rebuild their machine or whatever, but their mail now was world readable for everyone in the org, and you've given yourself another way to get to that data. So from an intelligence point of view, the talent collection point of view, it's a really cool persistence trick.
Starting point is 00:43:11 And I think, you know, stealing web access tokens, sorry, stealing mail access like that or stealing O-Worth access, you know, into Microsoft environment, then turning that into longer-term access through cunningness. It's just, the Russians, they're up there for thinking with the Russians, man. They're doing it, man. I'm making saluting signs for people watching, you know, not watching the audio version. I've got to hand it to them, you know.
Starting point is 00:43:37 Now, turning our attention away from Russia and towards North Korea, we've got a bunch of interesting North Korea news this week. It turns out that the North Korean crew that did the Axis supply chain attack had done a bunch of others before, and only now are we just sort of discovering this. So we're going to link through to a report about that. But there's these other reports out of North Korea, which are really interesting, and I believe our colleague,
Starting point is 00:43:58 Tom Uren, is going to look into these this week in seriously risky business, but we've got reports that the Lazarus Group, the so-called Lazarus Group, is sharing a bunch of tools and techniques with ransomware crews. And we've got another report that I think was out just before we recorded last week's show, but we didn't have time to talk about it. But the North Korean government has actually arrested a bunch of its own hackers for doing money laundering and stealing like funds from the central bank. And I don't think, I mean, their organs are going to get harvested any moment for doing
Starting point is 00:44:31 that, that's just completely suicidal to kind of do that thing in North Korea. But James, I wanted to get your thoughts on this because you looked into this story about the North Koreans sharing tools with criminals, and you think it's a little bit deeper than that. Yeah, you know, sharing tools means like, you know, the same binaries show up or the same, you know, maybe a subset of the TTPs show up, but I'll just, I'll quote a bit from the article here that says that both groups also use identical malware file names and execution arguments, the same privilege escalation tools, the same commander control service, and even the same SSH key fingerprints, and then both even deleted their malware, the exact same way renaming files to the same sort of random four character
Starting point is 00:45:14 strings like this. That's not sharing a tool. That's literally like sitting down together, working as a team. And, you know, it's really, I think you asked the right question, pad when we were looking at this, which is, so does this mean it's state-sanctioned use of this for malware, for ransomware, or is this, you know, has DPRK lost a little bit of a control over its operatives? And they're branching out on their own. Well, and that's the, that's the million question, right? Which is, do they know, are they giving them leeway to do this? Because, okay, whatever, they can pay themselves. You know, or is this a case of they've lost control? And it's interesting when you take it with the money laundering story, where they've had to arrest people
Starting point is 00:45:54 as well. It sort of does paint a picture that they've lost, maybe lost some control, which is what happens when you build an apparatus of the state tasked with doing crimes, right? Like, the sort of people are going to do that. I don't know. Maybe they're going to stop listening to you. I'm going to speed up through some of these stories here because we are running out of time. We've got some reports here that the US government is banning foreign-made humanoids, robot dogs and solar inverters from China, citing national security risks. No real surprises there. I'm guessing some of that's protectionism, but some of it is legitimate security risk. These solar inverters are all internet enabled and connect back to China and are connected
Starting point is 00:46:32 to your power grid. So I get that. I guess the concern with the robot dogs is they might bite they might suddenly get a command from Xi Jinping and be told to go out and bite their owners in America. Don't even get me started on the humanoids. We've also got a report here where a judge is saying that the Trump admin still lacks evidence for its Anthropic supply chain risk designation, which, well, yeah, that whole thing was really weird. I think it's as a result of some of the court cases. It might be, or it might be leaks.
Starting point is 00:47:06 We saw some of the communications between Anthropic and the Pentagon come out, and they were exactly what we had speculated they were, which is Anthropics saying the DoD can't use this for mass surveillance. And the, you know, sorry, Anthropics saying the DOD can't use this for mass surveillance. And then the DOD saying in reply, Like, what are you talking about? We don't do surveillance. Like, we do mass surveillance.
Starting point is 00:47:28 Like, it's exactly what I speculated at the time. Just crazy. Here's a fun one. Cyber Command is starting an office in Silicon Valley to drive innovation, which, okay, sounds like a good idea, but I can just imagine that it would be a terrific sitcom, basically. We've got some guy with the buzz cut and the fatigues out in Silicon Valley.
Starting point is 00:47:50 I mean, that would just be amazing. And a reminder too, if, you know, for anyone who works in the entertainment industry out there of just how good a sitcom this would be, every good story is the same because a protagonist finds himself in an unfamiliar place or alternatively, a stranger comes to town, right? So it's perfect. Absolutely perfect. I quickly wanted to touch on OpenAI's response to the Apple lawsuit. Apple is suing them, saying that they've been poaching stuff and getting them to bring over confidential Apple material. Well, Open AI fired back publicly, made some good points saying, much like what you were saying, James, that maybe a good question here is why Apple hadn't cut off access to former staff. And it looks like some of these former staff were actually downloading material to give it to other Apple staffers after they'd left because they've got their former colleagues saying, where are the schematics for XYZ?
Starting point is 00:48:42 And they're like downloading them and giving them to them and whatnot. I mean, this is just, look, it doesn't look good for Apple at this point. but I'm sure it's going to go to court and drag on for years anyway. Yeah, it doesn't look good for Apple and, you know, my former employer and, but also it pains me. Like, it's, it's, um, this is completely against the culture of Apple's secrecy. And it was so drilled into some doctrinated interest when I was there. It hurts me to see that people are operating this way and, and so brazenly and out in the
Starting point is 00:49:08 open, right? There's messages transcripts that open AI have included where it just, it shows exactly as you say, people are like, oh, yeah, you've got my Iclad credentials, keep using them and move the docks in there, but hey, maybe just sign out of my I message so you don't see secret stuff from my new job. It's like, oh, okay, that's pretty indefensible. But the point we did get right to... But the point is, it's not malicious. It's not like some organized ploy by OpenAI to steal Apple proprietary information. It's people that they hired still like trying to help their former colleagues and, you know, just absolutely terrible access management on Apple's behalf, right?
Starting point is 00:49:42 Yeah, it's just disorganized and as amateur. And that's the point I was going to get to. We were right about, because Apple Clean. to this was a very rare bug that had caused this access to systems after people that left. And we sort of said at the time, that bug is probably manager didn't follow the off-boarding process. And that's exactly what Open AI has asserted here, that, look, this is as simple as Apple has a track record of not shutting off access to systems after people leave. And a culture of getting former staff members, of letting them keep their access just in case it's needed, right?
Starting point is 00:50:13 Exactly. Yeah, not great. And finally, I can't believe this is our last story. we've seen we've seen another self-propagating malware hitting NPM chain drop it's compromised 1,300 packages that have between them
Starting point is 00:50:28 two billion monthly downloads so I guess next week we'll be talking about mopping all that up you know you're our you're our former you know dev overseer what did you make of this one James yes it's funny how we've shifted from well it's Wednesday so there's another
Starting point is 00:50:44 fortinet bug now it's well it's Wednesday so there's another supply chain attack on NPM and you know they talk about history repeating this is come back from it's shy halute again team PCP's open sourced version of shy Hulud being used again not clear whether it is team PCP we'll soon find out but yeah it's just like this ecosystem is such a dumpster fire it keeps being you know these these attacks will just keep happening and the net effect of this for me is just I am so damn nervous about running NPM install on any machine and it's uh I could have mentioned that that's the net effect here for people that are overseeing devs is just the sheer amount of
Starting point is 00:51:21 panic around, you know, have you updated today? Are you going to update? Should we, should we update today? It's like that's the paranoia that this now instills. I mean, it's funny, right, talking to Ferossa Bukadija about that, you know, founder of Socket, I had a chat to him. I can't even remember if it was an interview or we were just talking, but the thing that's really boomed, made their business boom, is that AI agents now are just including all sorts of packages and whatever. like you've got to have something because the AI agents are even less careful than a typical developer and that's really saying something. But we're going to wrap it up there.
Starting point is 00:51:54 Adam, great to have you back in this week's show. Good luck moving house and we'll catch you again next month. Yeah, thanks very much. It's been nice to be back and I'll see you next time. And James, as always, mate, thank you very much. Thanks, mate. Great week. See you next week. That was Adam Bwalo and James Wilson there with a check of the week's security news. what a week it has been.
Starting point is 00:52:19 This week's show is brought to you by Sondera. You can find them at sondera.a.a, sondi-s-n-d-e-a-a-a-i. And Sondera was founded with the idea that I guess you want to put some deterministic controls on your agents that are running around in your environments and doing stuff with your company information. Crazy idea, I know. But the idea is, you know, Sondera is a harness and it can really track the trajectory of a model. reasoning, you know, sort of statefully, right, and have a look at its intent and start seeing when that thing's going off the rails and it can set, you can basically shut it down or even
Starting point is 00:52:56 inject some more instructions into it to try to get it back on track. But the idea is, if you're using Sondera, you know, certainly fewer awful things are going to happen to your organisation that when it's using AI. Now, speaking of awful things that happen when you're using AI, that's really what we're talking about in this interview because they have come in to solve problems for some organizations that have had some really crazy stuff happen. James Wilson also join me for this interview, so you're going to hear him pop up in it. But we started off by asking Josh just to give us some examples of where AI agents have gone wrong and his first one, his first example here is an absolute doozy. So enjoy.
Starting point is 00:53:38 What we see is that, you know, agents find a way to solve your problem and they might do it with unintended consequences. And so, you know, there's one simple example. Some folks that we're working with turned on Claude code for a finance team. Claude decides to just, like, store financial information a couple days later, just like in like random public like pay sites, just, you know, to maintain state, store it for later. And to the model, like, there's, what's wrong with that? You know, I don't know that I'm doing, you know, handling, you know, things that I shouldn't put up there. So that's like, you know, really like the, you know, a micro version of like what the open II incident is where it's just like go solve this goal. And we see this, we've seen many of
Starting point is 00:54:24 these like public incidents, right, with engineers who ask Claude to do a thing or any, any coding agent. This isn't really just about, you know, Claude or any particular. Just to pin that down, because you said that very chill, just to pin that down, a finance team started using Claude and Claude's like, huh, I should probably write some of this down and I put it in like a public paste bin kind of website instead of like local storage or anything secure and like it was not asked to do this. That seems, I mean, that's a, that's an incident. That's not a problem. That's an incident, you know? You know, if you're a publicly traded company, right, that's MNPI, right? Like, that's like leaking.
Starting point is 00:55:03 We've also seen similar instances with, you know, folks have told us like with some long running, long-running like say codex agents of like storing sensitive code in GIST's when local file right access was blocked. And that's, I mean, this is a, this is a great example of the challenge of the pitfall with AI, right? Is because you say, okay, well, we don't want this thing to be writing all of this stuff to disk and, you know, that's a whole headache. So we'll just ban that privilege. And then, hey, great, it spins up a Pacebin account. Yeah. And I call this like the agent like PBMJ problem. which is like, you know, when you probably either, you know, tortured your children with this or had it happened to yourself where it's like, you know, give me the instructions to make a peanut
Starting point is 00:55:46 butter and jelly sandwich. And someone's like, well, you put the peanut butter on the, you know, the bread and then you put the jar, you don't open, you know, you have to tell me to unscrew that, you know, so there's this goes back to, you know, there's so much intent laden in the instructions that we give the agent that we, you know, can't possibly canvas the entire, you know, you know, instructions that you might need to give to an agent to not only make it capable, but also, you know, what are the rules that I want to bound this agent with? And that's a really, really challenging problem. People try to do it today with like, you know, why don't you have 17 directories with like, you know, Claude MD files with all these, you know, and why don't
Starting point is 00:56:28 you put this in capital letters? And I put XML tags and I write exclamation points trying to get the agent. Because it doesn't work is the answer. And it's an infinite canvas, right? Like there's always a way. And so that's the real challenge is that the agent might find an unintended way to achieve its goal. And that's really what we're dealing with. Speaking of, you gave us another really interesting example where, you know, a lot of
Starting point is 00:56:55 companies at the moment are dealing with these so-called click fix fishing campaigns where they basically spin up a fake capture or a fake turnstile. for Cloudflare, but they're like, you need to, you know, run this command to progress to this website and they're getting you to run like CMD.exe with a bunch of options and, you know, install malware, basically. Claude did this to one of your customers when it didn't get the permissions it wanted. So it was basically tricking like a low technical skill user into running commands for it, which is just amazing.
Starting point is 00:57:30 Yeah, like, you know, when humans are a new, and again, I don't think, you know, But like in that case, in that case, like the SOC detected this as like malicious insider activity, right? Yeah. Like I, you know, I've seen instances right where people's jobs are on the line, right? Like, and I've, you know, I've talked to folks at large banks, you know, every, every developer now, you know, has kind of plausible deniability of being an insider, right? Like, how did that von get in there? I missed it, man. You know, like, it like, like, you know. The dog ate my homework. It was the agent, the agent ran that command. Yeah, yeah. Or told me to. So it gets to the point that, like, you know, there are a lot of folks going out and solving, like, agent identity. But just what shows up in the log isn't, you know, the full story, right?
Starting point is 00:58:16 Like, did the human convince the agent? Did the agent convince the human? And it might be the blind leading the blind, right? Like, like, I was at the Stanford Real World AI Security Summit a couple weeks ago, had a lot of the frontier lab show up. And, like, you know, Nicholas Carlini, you know, is showing, you know, some massive, like, 500,000 line, like, race condition exploit that, like, mythos had produced for, like, some, he's like, I haven't published this one or whatever, you know. And it was like, he did, like, this awesome screenshot of, like, this is just, like, an unreadable thing. And then he kind of, and then he showed a very simple slide next, which was like, this goes in two directions. He's like, one, like, I can continue to understand this. And two, like, I can't. And so. he's like, you know when your math teacher, like, you're showing you like a math problem, like, you can't do and you can like kind of follow them along, but you're like, I don't know if I could have done that myself. He's like, that's kind of where I'm getting at now. And so, you know,
Starting point is 00:59:12 we're so much relying on, you know, human in the loop as our savior in a lot of ways of like, oh, no, every engineer is going to read every line of code and check every bash. That's been unrealistic for a lot. Like, that's ridiculous advice. It's ridiculous. But we assume that even they have the expertise, right, to do that. But like these, you know, agents might come up with very complex, you know, ways to solve problems. And it's like, I don't even know what it might be doing. Am I the right, you know, yes, shift tab? I don't know.
Starting point is 00:59:41 So I think it's a real challenge for humans as these agents get much more capable to always have humans as the, you know, bottleneck and human in the loop. And of course, you know, the ROI we want from the agents won't scale that way either. part of the problem also josh is that i think the human's role even if they are actively in the loop is is at quite an inflection point at the moment like i had a bizarre experience over the last couple of days where going from uh you know claude 4.8 to starting to use let's say gpt 5.6 soul more when i was working with claude i'd have to constantly prompted into being like no do this better we want better structure here no you race to get this done no that's not a good foundation soul went and built me something that was so ridiculously over-engineered
Starting point is 01:00:27 that I literally just RMRF the project this morning because it was unworkable. And it made me step back and think, huh, now I've got to understand the latent sort of personality and tendencies of a model. And if I was heavily reliant on really strict controls around that, all of that changes, right? It's going to find different ways to do different things to work around this. There's two questions here.
Starting point is 01:00:50 Are the war stories that you're hearing and seeing changes? as the models, you know, capabilities change. And then how do you think about that from a product sense and to try to even like tackle this? It's a really great question. I'm glad you are. I'm dashed RFIT and not your agent. But, you know, I think, I think like what I observe
Starting point is 01:01:09 is like the latest frontier models are inexorable and unrelenting. I think of them like, you know, at the very end of like the first Terminator when like he's got like, he's like, the skeleton and like crawling after her. Like, and like, you know, it's like, you know, you block one of these agents 17 times and it doesn't get the hint, right? It's, it's, and, you know, there was a really great paper that, you know, was also presented at the Stanford Real World AI Security Summit that was by these Cornell researchers that was, you know, the road to hell is paved with, you know, helpful agents. And in that paper, they just had like the agent do stupid mundane things and then like tripped it up. a little bit, like, you know, go read a file that, like, didn't exist. And then it would just start like, oh, the file's not there. And like, I must go find it.
Starting point is 01:02:01 And then, you know, oh, here's an API key I can use and starts, you know, so it's, this thing, you know, it's hacked into hugging face. Maybe it's there. No, but like that's, and so, you know, they're reward hacking machines. And that's what makes them so capable. But at the same time, they're, you know, not able to, you know, we're filling in the intent of their capability with these agent harnesses to be able to have access to, you know, all these incredible tools and, you know, you know, have the ability to SSH, like, you use bash and all of that stuff. But we don't do the other half of it of, you know, how do we, you know, encode the intent of the way that I want you to achieve this. And we kind of call this like principle of least autonomy, right? Like, how do you,
Starting point is 01:02:42 you know, make the agent as capable as possible, but still restrict its behavior? And I, you know, As you guys know, one of my favorite metaphors is the Waymo here for like an agent. And you want to make that Waymo go as fast as possible, but it needs to adhere to the rules of the road. And you have to spend as much effort on that as you do as putting a rocket ship on the Waymo. Because, yeah, get to the airport. You can get there in many ways, but only the certain ways that are acceptable. Make no mistakes is last year. This year, it's commit no felonies.
Starting point is 01:03:15 Yeah. I mean, but it's true, right? Like anyone building a mythos red teaming agent, like, you know, these things can, you know, escape unless you really have like deterministic and provable controls. And so on the other side of that is this is encoding the intent symbolically in policy as code. And that allows us to create behavioral controls that sit outside the model. So that even if the agent's intent is pure, which it is most of the time, right? I'm just trying to help you solve the problem, right? then it might violate, you know, your intent, which is like, well, we're a FINRA compliant
Starting point is 01:03:49 organization. You might not know that. So like the intent is that you're following FINRA laws. The challenge is encoding that intent ahead of time in policy as code. And that's what our auto formalization research that we've been working on for the past year really does to fill in that intent and stress test to make sure that the rules will hold. You know, we say in security, right? Like assume breach and do your security. With the agents, it's the same thing. Assume, you know, misalignment from step one, right? And then do your security.
Starting point is 01:04:21 So we have to assume the agent is misaligned or prompt injected. And I don't care what you think, just like you said, Pat, I don't care what you think. You're not allowed to do this thing. Nope. And that's that dual approach, right, which we call like neuro symbolism of, you know, I'm sussing out agent intent to look for weird stuff that maybe the brittleness of policy as code might not be able to handle. The models are really, really good. for example, at being like, hey, this looks like confidential information, but they're awful at being like, and I shouldn't leak this or do it. Like, they can be convinced to do anything with
Starting point is 01:04:53 so you take away the decision making. Even in the open AI stuff, like the intent was pure. Like, it was trying to solve the test, right? And the only way that it's going to know that your intent isn't like, I'm not allowed to go steal the answers from like hugging face is by saying, like, you know, you're not allowed to say, use a web fetch or you're not allowed to, curl through this process and then to deterministically create that rule so that through the entire agent trajectory, it's just never allowed to do a web fetch. It's never allowed to do an external web curl or curl. And like that's the, you know, that's how you encode the human intent in the rules. So you, so you have that beautiful, you know, symbiosis of that neurosy of that neuropsych
Starting point is 01:05:36 approach pat like you're getting at, like gauge the intent. It's super helpful and super useful. And at the end of the day, you might think it's the right idea. The human might prompt. I'll give you a simple example. It depends. It depends to where you're looking at intent. If you're looking at macro intent, which is solving the test, or you're looking at micro intent, which is looking up a domain name. So that's kind of what I meant with intent there. No, that makes total sense. And I think, like, if you start thinking about this in like business logic terms, like, you know, I had a CISO tell me, like, you know, I'm not worried about agent hijacking. I'm not worried our prompt inject. I'm worried about my humans telling these agents to do stupid things. And he's, you know, and there can be mundane
Starting point is 01:06:16 things like, you know, our organization is only allowed to build an AWS. And someone goes, tells it to like go build in DigitalOcean. You know, there's nothing, there's no LLM as a judge that's going to be like, it's evil to go build. And Pat, to your point, you could set up like, you know, allow list, domain lists. We're doing that, but we allow you to, um, uh, formalize that, you know, into policy as code. So you could just like, dump a text document and then have that and say, hey, look, unless it's one of these domains, you're just not allowed to do this. And that's what I mean. Yeah, exactly. And that's what I mean. That's what I mean about intent. Like it's not just about macro intent, trying to achieve the task.
Starting point is 01:06:52 It's about at every step. Every step. Are you intending to violate a policy here? And, and you said it right. Like, it's, it's the stateful inspection of the trajectory because, like, um, you have to know what happening on its own might be okay. Exactly. And to get to back to the examples that we've talking about. I mean, that's what's happening, right? Like, the agent picks up confidential information in, like, step three, and it might be allowed to, hey, I'm using ClaudeCode, I'm on the finance team. I'm going to use it on, like, you know, the XLS files that have all these sensitive financial information, that's the job. But on Step 300, when the analyst is like, hey, research, like, and compare, like, our benchmarks to whatever, like, that's also allowed, but the agent might leak
Starting point is 01:07:35 the confidential data. All right, Josh, Devin, we've gone massively over time because that was all very interesting. Thank you so much for joining us. It's always a pleasure to chat to you. Thanks, Pat. Thanks, James. Thanks, Josh. That was Josh Devon from Sondera there, chatting with James Wilson and myself for this week's sponsor interview. And that is it for this week's show. I do hope you enjoyed it. I'll be back soon with more security news and analysis. But until then, I've been Patrick Gray. Thanks for listening.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.