Risky Business - Risky Business #852 -- Cyber Command wants to buy shells
Episode Date: September 9, 2026On this week’s show Patrick Gray and James Wilson are joined by guest co-host Robby Winchester from SpecterOps to talk through the week’s news, including: ID ver...ification company IDScan was breached and 153m driver licenses wound up for sale online. Cue the barrage of lawsuits The US government plans to pay private contractors to conduct military hacks The US accuses China of distillation attacks, a.k.a. forbidden training It’s Wednesday, so OpenAI’s agents escaped sandboxes again and passed notes around on a German Wiki Much, much more… This week’s show is brought to you by Sublime Security. Sublime’s head of detection engineering Randy Pargman joins the show to chat about how the company is preparing for prompt injection attacks to move from being largely theoretical to commonplace. This episode is also available on YouTube Show notes FBI Probes Service Selling 153M+ Drivers Licenses | Krebs on Security IDScan sued over alleged data breach affecting 153 million drivers | BleepingComputer Senate Considers Allowing Contractors to Conduct Military Hacks | bloomberg.com Feds accuse China of ‘systematic’ distillation of U.S. AI models | cyberscoop.com Dropbox accounts breached through Lenovo email verification flaw | BleepingComputer FBI raises alarm over deceptive phishing campaign targeting prominent people | cyberscoop.com Microsoft warns of TerminalFix attacks deploying reverse tunnels | BleepingComputer OpenAI agents discussed ways to escape their sandbox on public wiki | arstechnica.com OpenAI releases new model that it says triggered internal security measures | NBC News Tech Trump may be forced to reveal secret rules feds use for AI safety testing | Social Signals Microsoft posts nearly 1,000 bugs for Patch Tuesday as CISA warns two being exploited | therecord.media Security Incident – BGP Hijacking – Virtualizor | Coder's registry infrastructure compromised to push malicious modules | BleepingComputer Pegasus, NoviSpy variant spyware found on devices of Serbian activists | cyberscoop.com European parliament members call for slowdown of Serbia’s EU entry over spyware use | CyberScoop New pro-Ukraine hacker group targets Russian companies with custom ransomware | therecord.media US military disabled ad tracking on troops’ devices following reports of targeted attacks | TechCrunch Security New CrowdStrike 'FalconFlank' zero-day grants SYSTEM privileges | BleepingComputer A hacker stole $340M in a crypto heist, then returned most of it | TechCrunch Security ‘White hat’ hackers take $47 million bounty after $320 million crypto theft | therecord.media
Transcript
Discussion (0)
Hey everyone and welcome to risky business.
My name's Patrick Gray.
We've got a great show for you today.
We will be checking in to discuss the week's security news with James Wilson as always and sitting in the third chair today.
We're joined by Robbie Winchester, who is one of the original founders of SpectorOps,
the offensive security company, which also makes Bloodhound, the security product.
Robbie Winchester, thanks for joining us for this one.
one. Thanks for having me. Happy to be here. Now, I should mention too that despite
Spectropps being a sponsor as some of our podcasts or whatever, that's not why Robbie's here.
No money has exchanged hands for this. We just invited Robbie along to appear in this podcast
because we thought he'd be an excellent guest. And this week's show is brought to you by
Sublime Security. And we're chatting with Randy Pardman later on in this week's sponsor interview
all about, I guess, what Sublime is doing and how they're thinking about prompt injection.
Because when you think about it, email is going to be like the mother of all prompt injection vectors.
And you're going to hear Randy explain that so far prompt injection attacks have mostly been theoretical.
They're not seeing a lot of it in the wild yet.
They're seeing some fun stuff around marketing where it's like people trying to do prompt injection,
encouraging various tools to like surface a product before other products or whatever.
But yeah, so far it's all been pretty weak.
But they've already had some thoughts about how they're going to deal with this as like an attack class.
It's an interesting interview, and it is coming up after this week's news, which starts now.
And look, I think it was just after we published last week's show, this news emerged about 153 million plus American driver's licenses being available online.
And I thought, no, that can't be right.
Like maybe this is one of those cases where somebody's box at one of these ID companies is, like, compromised and they're just funneling queries through their creds or something.
through their actual machine.
But then you actually read Brian Crabs's coverage of all of this, James,
and it seems like certainly it was ID scan, got owned and not a little bit owned.
They got owned quite a lot.
Owned a lot, and for quite a long time,
the attacker themselves said they've been exfiltrating data continuously for over a year,
I think it was in the article.
And I think even Brian mentioned that even as he was sort of watching this unfold in 24 hours,
he saw another 400,000 records get added into the platform.
So it's clearly like they've established a very reliable exfiltration
and operationalize this into an import.
But of course, being a Krebzing story,
the beauty in this is how he then actually pinpointed
it was ID scan by working out that his license was in the database.
They actually offered that as part of their free sample.
You know you're an icon of the industry when you're in their free sample set.
And so he was like, well, how did that?
happen and he thought, well, okay, I've been traveling around the time that that timestamp on the
ID was there. So maybe it's at airports, contacted a few friends with their permission, looked them up.
Sure enough, they were in there. They'd also been traveling. But they hadn't presented
driver's licenses at airports. And so they were kind of wondering, well, where's this all come from?
And then the penny drops when he realizes his mum's in there too. And his mom's license has the same
timestamp as his license, which is exactly when they went to get their Hertz rental car.
they both handed over their licenses,
and that with another little bit of correlation linked it back to ID scan as being the source,
because of course, ID scan themselves says Hertz trusts them with their ID data.
I mean, there's a link here also to like weed dispensaries or something as well, right?
Yes, well, that was the other correlation was they were like, okay, well, this seems to be Hertz,
but we kind of need like another data point here.
And so one of the, so Zach Edwards, security researcher, he said, yeah, sure, look me up.
And sure enough, his license is in there.
but the thing is he didn't rent a car,
but when he'd gone to Def Connell Blackhad, I think it was,
he did say, yeah, I'd been to a cannabis dispensary
and shown my license there.
You go and look at the ID scan, press releases,
and, you know, who trusts us?
And sure enough, there is the chain of dispensaries as well.
Yeah, and of course, cue the lawsuits in 3-2-1.
Apparently there's like a whole bunch of lawsuits already launched over this.
Robbie, what's the likely outcome here?
It seems like this service is not online.
at this current time, but you would imagine whoever has this corpus of data, you know,
they're going to be selling it to other people, they're going to be trading it, like you would
have to assume all of this information is out there. So I guess I'd ask you, what do you think
the consequences of this sort of scale of breach are going to be? I'd say they're going to be
meaningful, but are curious for your perspective there. And then, you know, what do you think
is going to actually happen to the company? Yeah, I think it's really interesting because we've
seen the result of what is obviously some type of data leakage or some type of compromise or
something has happened. But at this point, I don't really know that I've seen any factual. This is
exactly what happened. Is this a really bad, really involved insider threat? Is this some supply chain
compromise that's going to go beyond even this company and is the tip of the spear? I'll be very
interested to see what goes on. I definitely think it's going to change people's perspective of
handing over their driver's license at dispensaries to go and get scanned to go in and think that
there's no consequence there. No one will be able to find out about it. It's such a huge breach with
so much information and seeing, I think I also saw, it's not just, these aren't just photos. They're
infrared and ultraviolet and other scans that seem above and beyond what you would just expect.
I'm curious to see how it's going to react kind of the entire industry around this like ID
verification when we learn more about it. And is,
are the peers any better or what peers exist to provide this service?
Yeah, I mean, it's really bizarre, I think, that this corpus of information was collected
in the first place.
And, you know, look, I understand that all countries are different, but I just can't imagine
something like that happening here necessarily.
I mean, we did see something similar with Optus, right, James, here?
Yeah, but even then, I mean, the problem here is why persist the raw scan of the license?
Yes, like the metadata, the output, enough to then correlate again when another scan is made.
Yes, tokenize it somehow with some interesting scheme? Sure.
I mean, I think, Rob, your comment to us earlier was got me really thinking around,
you know, does this bring something like PCIDSS for ID?
Because that would seem to be the sensible trajectory here?
I don't know.
I mean, I think we do have some rules along those lines here, right?
Where there's certain parts of the licensed data that you're not allowed to keep,
like the actual card ID, right?
So you've got like a state government, you know, license issuing API where you can check this data against the API.
But I think, you know, keeping that stuff is a big no-no.
So, yeah, I mean, just from a data governance perspective, this is just mind-boggling.
Yeah, you don't really see, I think, in the U.S., there's not as much of it, there hasn't been as much of a protection of privacy like GDPR and other things.
And it's a little bit of the Wild West.
I'm curious of, again, this maybe is that tipping point to start taking it more seriously like health care data or like payment data.
I don't know, man.
There's been so many tipping points over the, over so many years, right?
This one for sure.
Yeah, this time, this time for sure, for sure, yeah.
Now, look, staying with the US and some, some news out of there, which is, there's been a
provision inserted into the Defence Authorization Act, which is looking, which looks to give
permission to cyber command to contract private companies to do what they're calling access generation,
I mean, the United States is the world-leading exporter of like absolutely cracking euphemisms
and access generation is one of them. The idea here seems to be that Cyber Command can say,
you know, go forth contractor, go forth prime contractor and collect shells for us.
This seems like not the most insane way. If you're going to engage the private sector to do
military-based stuff to, you know, actually exert some state power.
here, the initial access phase seems like the sensible place to apply it as opposed to the actual cyber effects phase.
You know, we've seen a separate proposal go through which is about allowing the Department of Homeland Security and Department of Justice to actually engage contractors to do effects to
you know, transnational cyber enabled transnational criminal organizations. You know, we've talked about that on the show before.
This seems more careful and for good reason actually seems like fairly well calibrated,
policy, but I mean, you know, Rob, you've spent most of your career in offensive cyber.
What do you think of the idea that Cyber Command is going to be handing out contracts to go and
basically act as an initial access broker for the, for the US government?
Because I think that sounds like a lot of fun.
Are you going to, you know, SpectorOps is an offset company?
You're going to spin up a unit doing this?
I mean, it's very, very on point, too.
I don't know if you knew this.
I started my career in the Air Force, actually.
So I was doing Air Force Red team and in the military before.
So it's very interesting to have had that perspective.
and seen all kind of firsthand and understood the policy, the red tape, the processes and things
that were in place, they could become a barrier a lot of times to potentially getting things done.
And I think one of the challenges historically has been how do you adapt the military that moves
much more deliberately in pace of you're not just changing things up every two or three months
because you need to have consistency and continuity and supply chain and everything.
but cyber information warfare changes so rapidly.
I think this is an interesting new direction to how can we go and leverage the expertise of the private sector and treat it almost,
is this the equivalent of your purchasing rifles and bullets from a company that makes it separately?
And that's all you're just getting these initial things that can let you go and do something more to your point, Patrick.
Or is this just the starting point of kind of the fear a couple of weeks ago when it was the comment of going after and actually achieving effect?
And what is also going to be the protections and qualifications for these companies?
If we were to go and do this, how do you make sure that this doesn't become
you're creating private industry that is now even more of a target, let's say, than it was
previously for other nations if that's what you're targeting?
I mean, the asymmetry of it just gets kind of interesting to me.
It does.
But I think, you know, a lot of the fears, like people have always been worried about like,
what if the private sector make a mistake and they pop shell on the wrong box and then it's
World War III?
And I think we realize, like, cyber has not escalated anything ever anywhere.
So I think we have been to, well, I think we've been suitably cautious, but maybe it's time to, like, try something like this.
I don't think it's an insane idea.
What I do find interesting is we saw this during Trump's first term as well, is it's a policy area where I just don't think he really, like, he doesn't seem to care all that much about policy.
But all of these people who do care about cyber tend to gravitate towards Trump admins because they're going to.
get to try stuff and the White House won't get in their way because they're not exactly known
for being nervous, right, about putting noses out of joint. So you actually do wind up seeing a
bunch of interesting cyber policy being spun out of, you know, both Trump administrations.
James, what do you make of this? Yeah, I'm excited to see how this goes. You know, I was talking
to Brad Arkin about this on a features episode that came out this week and he used the analogy that
the original Trump memo is kind of like taking the release the hounds gauge from
two out of ten to three out of ten, and this feels like an even more calibrated step of taking
it up yet another notch. The subtlety of this being about access generation or initial access
I think raises some questions for me. The persistent concern here has been, can the government
move fast enough for this to be useful, right? The private tech can move fast, but can the
government move fast enough? And I'm kind of interested in, does the deliberate pinpointing here
of this being about initial access make it actually easier for the government to move faster?
perhaps you can get approval to get access first but not do something so they can get that done
with the private contractor, then take their time to understand what's the effects they're going
to then do once they've got that access. Just an interesting decoupling.
I mean, I think that decoupling is more about the fact that they don't want private sector
companies being responsible for scoping the effects and potentially making a big mistake,
wouldn't you say, Robbie? Yeah, I mean, the challenge with any of this is there has to be a left
and right bounds. And I think that your point, James there, of the,
handover and does it make it easier? It'll be interesting to see where does that and how does that
handover make sense and is there something lost in that translation and how how immediate is this or
useful is it and integrated does this have to be as kind of a you get initial access but how long does
that persist and it does there need to be a call for action for initial access that is then
fed into something that they're going to do immediately you know to have that immediacy of function
and not just thank you for this shell that has been there for a year that's you know the
the every hacker's dream.
Yeah.
But isn't necessarily the reality of, especially a hard target you have a,
we need to get on this sensitive network at all cost.
You know, how do you have that immediacy?
And I guess to your point as well, James,
of do you say, okay, here's initial access and you're getting it for us.
And oh, by the way, when you do,
if you can run this thing that isn't you running it,
you're running it on behalf of the government, that'd be great.
Like, that's an,
it's an interesting new territory to kind of get into.
Yeah, I mean, the timing thing is interesting because I can foresee a circumstance
where someone's like, you know, some contractor gets a shell in this, you know,
whatever agency of the Iranian government or something.
And then two years later, someone at Cyber Command says, we'd like to use that shell.
Like, man, it's gone.
You know, it's been gone for a year.
By the way, listeners can probably tell that I'm a little bit crook at the moment, a little bit sick.
And it's funny because James is as well.
So the joke would be that it was going around the office.
But we are located about a nine-hour drive from each other.
So not quite sure how we both seemingly came down with the same thing at the same time.
So if I'm a little bit vague today and sound a little bit off, that is why.
Now, we've got this next story here, which is the United States has put out a big angry PDF
all about how China is systematically distilling American AI models.
Now, this is an accusation that they have leveled at Chinese country.
companies before. They've got into a lot more detail here because they are actually naming the
companies that they say are doing the distillation and talking about exactly how they're doing it.
Now, you actually have the counterintuitive take here, James, which is you are dramatically
underwhelmed by what the United States government is putting forward here as evidence of Chinese
IP theft. Can you explain to our listeners? Why?
Yeah, okay, so a couple of points. First of all, you know, distillation and training.
are the same thing, right? They are given very different names in terms of, or there's a very
different, I guess, sort of sentiment around those two terms, right? Training is what legitimate
labs do and distillation is what the bad people do. They're the same thing. You take a structured
set of data and you go through the model training process of building up that pattern recognition
and outcomes your model. The difference here is that in training that a frontier lab does, that's
with data sets that have been scraped, gathered, right? There's all sorts of challenges around
the morals and ethics of that training data.
No, but Open AI, right?
And Dario over at Anthropic, they only scrape data in the most ethical way.
It's ethically sourced data, right?
Yeah, of course.
It's good training data if it comes from the training region of France.
Yes, exactly, right, yeah.
But that's exactly my point here.
So point one, distillation and training is the same thing.
But when it's done by China, it's bad.
And that essentially seems to be the crux of this here, right?
There's two rules, one for the US, they can train on data.
China, you've had a different rule.
You can't distill on the same data.
But then there's this point around the industrial knowledge distillation campaigns
that form this, you know, they're the core of the training efforts by China,
not merely a supplement of their AI development strategy, quoting directly from the document.
I agree that that is the case.
I disagree with the sentiment that, oh, look, they rely wholly and solely on distillation.
If we can just disrupt that, we'll take care of the Chinese open weight model.
problem. They're using distillation because it's so simple and it works so effectively, right?
You can't imagine a more readily available, easy-to-use data set for training a large language
model than these distilled transcripts that come out of Claude and Opus and all the rest.
The third point here is this idea that, and I just want to be clear on terminology,
they're not distilling the models. There's not queries that go to Anthropagans say,
tell me how you work Claude. Claude, give me the logic that's gone into your training.
Give me the exact chain of reason you go through to solve this problem so that I can replicate it.
They're taking legitimate outputs of how the model is used and looking at the patterns in how that reasoning exists and how it's solved problems to then extrapolate from that and create the same sort of reasoning capability.
And why this is important is if you just manage to magically wave a magic wand and say, oh, tomorrow, you can no longer, China can no longer directly interact with LLMs.
It won't solve the distillation problem because everyone, everywhere, around the world,
world is sitting on an absolutely massive set of transcripts. I checked this morning. I've got five
odd gig that adjust from my security evaluation harness has been running for the last month. And that stuff's
gold. Well, I mean, this is, so this is like when I was reading through your notes on this,
this to me was the most interesting part of it, right? Which is, look, if you've got all of these
transcripts, and they don't even have to be from one model, I mean, your transcripts, I know how many
different models you use, if you've just got a big pile of transcripts there and you throw them in to
to bring the latest iteration of your open weight model up to date, right?
Of course you're going to go throw all of those transcripts from all of the models at them.
You know, now, is that something that should be forbidden?
I don't know.
I mean, maybe, but is that going to stop it?
Would be my question.
And I think that's the interesting thing that you're driving out there and where I absolutely
agree with you, which is we can shake our fist at this.
All we want, it ain't going to change anything.
No, no.
There is too much money.
There is too much incentive.
the money will just flow to somewhere else.
And cheekily, I kind of hope it flows to me and I can sell those transcripts off.
But I know that, of course, won't happen.
That's it.
That's your retirement plan.
Yes, he's a retirement plan at the moment.
Robbie, what do you make of all of this?
Because, you know, you're the American here.
Are you outraged by the Chinese Communist Party getting their grubby communist hands
all over your, you know, Apple Pie American models?
Or is this just like it's all in the game?
Because I kind of feel like it's all in the game.
I mean, I think in an ideal world, everyone, every business and every organization,
operates incredibly ethically. You pay full fair market price for everything that you acquire and,
you know, every piece of content is properly handled and the creator is compensated and there's
no potential negative, unethical getting of data or information or anything. But in general,
I think the reality is we live in a messier world. And so there's going to be questions about
where is, where is data being sourced and are people being properly compensated and what is
the right recognition of this? And,
how should the training occur? And a lot of it is also, I think, the average American is not aware
of the detailed ins and outs of how LLMs are actually trained and what are the different
mechanisms. And I would bet a lot of money that if you said distillation to most people
on the streets in America, best case scenario, they know that you're making alcohol,
but the chance of them thinking about doing model training is almost zero. So I think something
like this is it is calling out there is definitely, you know, historically there have been concerns
over the U.S. being a technology creation powerhouse and other countries, notably China,
taking actions to escalate their position in the world by, let us say, getting the same technology
with less individual research and development, more focused research and development of the
finished product, hypothetically. And so this kind of feels like just another acknowledgement of
this is basically the back and forth of how this is going to flow.
And I agree with a lot of James points where there are certain aspects of
challenges in getting formatted and proper data and having it be labeled and having
all these different pieces where there's definitely elements I'm sure of the distillation
that are more specific and towards that abuse space, but also just the fact that this
exists and there's now this corpus of knowledge, does that corpus of knowledge just
become you the, is it no longer yours?
Is it everybody's, right?
It's a tough thing.
Like in an ideal world, you wouldn't have theft and distillation or, you know, negative abuse.
Well, just to your point, Robbie, like, you know, it's the position of the AI companies that the outputs belong to you, right?
Right.
So the outputs can belong to you, but what, you can't feed them into a training run?
Like, if there's that sort of restriction on them, do they belong to me?
I don't know.
What if I'm an American company doing a training run based on the outputs that I've paid,
good money to get, you know, like it's fraught, right? I think we all agree it's fraught, right?
What are you trying to get? And then also, what is the follow on? What are you having it?
Are you, like you said, if you're paying for the usage, can I go and use this for my own other
things if I'm using it for something else? If is to a certain extent, if the model lets me do it,
does that mean that it's allowed? So the guardrails should stop me from, you know,
the distillation check failed. So therefore it's okay. Well, that's probably not the case,
but also then if I'm letting, you're letting my curies go, is it my fault?
It's a very messy situation.
I think also the challenge is, as with a lot of the things related to policy,
there's a lot of very technical nuance that can be missed in, I think, the kind of
trying to stealing data bad.
Well, yeah.
And I would say that this is the opposite to the previous story we spoke about where, like,
you're getting awesome outcomes because Trump doesn't care.
I would say in this one, you're getting a different type of, or different category of
because Trump and the people closest to him definitely do care about this.
So, you know, I don't think America's complaints are going to result in much of an improvement for the frontier labs.
Let's put it that way.
When it comes to this sort of distillation or training.
Now, look, bread and butter infosec yarn here.
We've got this amazing story here from Bleeping Computer.
I was first tipped to this by Luke Jennings over at Push, thanks Luke, where a whole bunch of Dropbox accounts got breached.
through a cross IDP impersonation attack because apparently Dropbox trusts Lenovo as an IDP.
This reminds me, so you can log in with your Lenovo, Lenovo as a source of trust.
This reminds me of that whole sketch that a lot of people would have seen where the guy keeps saying,
oh, Jeffrey Epstein, the New York financier.
You know what I mean?
I keep thinking like Lenovo, of course, the IDP.
So really the trick here is, say I know that you're, you know, you're a, you're a lot of
James.wilson at risky.bears and you have a Dropbox account, you're logging in with Google for that.
I register your email address as an account with Lenovo, and there's some sort of vulnerability in Lenovo's
account that enables me to do that. Then I can just be logged into that Lenovo account,
have the correct cookie set, and then I can just use that to log into your Dropbox account.
It really is that simple. We've seen similar attacks like this in the past.
But I mean, it's just one of those ones where you just think, what do you call it here, James?
oh-o-oth. It's just, you know, it's so incredibly dumb, but also completely unsurprising.
Correct. There's actually two bugs here that have collided in the most spectacular way.
The one you mentioned there is that the Lenovo was not doing the proper email validation,
so you could attest that you were James.wilson atrisky.combears, even though you hadn't validated,
you had access to that email. But Dropbox had an issue of then allowing you to then sign in to
your existing Dropbox account with that new Lenovo ID, without.
having to first provide credentials for your existing identity provider that you previously used
for that account.
But that's how all worth works. That's how they're all configured, man.
It's not how it should work. It's not how it should work, but it's how it does work.
It's how everybody's set it up. Robbie, like you're chuckling away there. Like, what's your take here?
I mean, this is, I feel like such a perfect, like you said at the beginning, this is the infosex
tale is oldest time of you have an inadvertent connection of systems and,
people using Dropbox are not all agreeing to Lenovo's
ID terms of service as they're creating the Dropbox account,
but because of decisions they're not aware of behind the scenes,
now anyone with a Dropbox who, ironically, didn't have a Lenovo.
I think not being a Lenovo member actually put you at bigger risk here,
because if you did not have a registered email,
then it could be registered to James' point of that kind of triage.
And so it's interesting when these come up,
but it's also, again, the configuration
of identity and identity tracking is hard and making sure that you do all these things and think
all the way through is is not always the case. And so in circumstances like this, this is why
making sure that you're implementing security best practice and doing something like when you
do your Oath have a validation that you're actually linking it, even though no one does.
And they don't, because I remember the last time this came up, I went and checked with a whole
bunch of services where I had a username of my workspace email address, but I was logging in with a
password, you know, and often MFA or whatever. And then I tried to just log in using the same email
address as a, you know, login with Google and it worked, right? So it shouldn't, right? But it absolutely
does. And a lot of organizations see that as being a feature, not a misconfiguration, James,
sadly. So, you know, we are where we are. Staying on the topic of Oath.
The FBI has actually put out a bulletin, a public service announcement, all about this fishing campaign targeting what they're calling like prominent people.
It's interesting here because they are actually publishing a malicious app.
Then they pose as like one of your relatives or whatever, send you a document.
Then to read the document, you need to log in via Oath.
And then it presents you with the like the privilege escalation part of saying, no, I need all of these permissions and people.
are falling for it. What's interesting here is you know you don't see too many of these
o-off app consent fishing attempts in enterprise because quite often these days an admin needs to approve
new apps. But when you're talking about prominent people, personal accounts, it looks like this is
working. Yeah, and that's the bit that you and I had a good chat about earlier was I just assumed
this was device code fishing because the old malicious app path has largely been shut down to get that
app set up. You've got to have an admin access in the tenant. But of course, personal account
You are the admin, right? So you're going to see this. And, you know, the art here is crafting this initial document or lure in a way where the, you know, the user is sufficiently motivated to want to view that document or go through the whatever steps. And they just, you know, consent. Yes, yes, yes. I approve. I just want to get to whatever this thing that my family member needs me to see or this draft document, my accountant wants me to look at. So yeah, it's, yeah, oh, well, keeps on giving.
Now, one of my favorite attack categories at the moment is this like consent fix or click fix stuff, sorry, where you hit a fake turnstile and it's like, hey, just like run this command to get past this turnstile.
They're getting more and more elaborate.
There's a new one called terminal fix.
Microsoft is called it terminal fix where you're basically like running a whole bunch of PowerShell forum that turns your box basically into a drop box on the network, right?
So you can, and I thought, you know, Robbie, like, who better to ask you about this than someone who runs a pen testing company?
You must get so annoyed that this sort of thing works because you guys will have developed much more sophisticated ways of doing this.
And then you just see that they ask the user to run some power shell.
And then they get a shell and then they can just pop out of that person's machine on the internal network, you know, move laterally to great victory.
I mean, that's really what's happening here.
Like, does this annoy you?
I wouldn't say it annoys me.
I think it does highlight that sometimes it feels like the, you see all the sophistication that goes into a lot of new off-sec tech.
technology and different near zero day, one day type techniques and everything.
And then there's this where you're basically just convincing someone to open up
PowerShell and run a PowerShell command, which seems something that if someone were to come
to me and say, hey, we're going to, we're on this test.
We're going to try and convince them.
It's going to be a pop-up and it's going to say Open PowerShell and then copy this and then
paste this and then type back complete.
You'd think, no, there's no way.
Someone might fall for it, but it's going to probably raise the flags.
And lo and behold, it still happens.
It's that difference of one of my favorite, the SKCD comics of the cryptographer's dream.
If you have like the 2096 AES encryption, like, oh my gosh, we'll never crack their password.
And then the reality is like, here's a wrench that's $5 go hit him until he tells you the password.
It's that circumstance here of, you know, there's definitely a lot of elaborate and interesting, less noisy and not, you know, maybe you're going to get detected less or you can't signature on PowerShell.
But old faithful always works.
Yeah, I mean, I was going to ask, though, like surely any appropriately configured EDR is going to catch this, right?
Surely.
You would hope so, but also who knows.
I mean, properly configured, but properly configured could also be that exemption for the PowerShell auditing script that you run, you know, that allows.
So 99 times out of 100, it works every time.
Yeah.
You never know.
There's so many, I would agree with you.
I would hope in most organizations, something like this would fall over.
But you're, you know, you're likely not, you're going off of quantity, not quality necessarily here.
And so you just have to get a couple of the right ones and you're in business.
Yeah, so worth a punt, we'd say.
Yeah.
I guess.
Now, it is Wednesday here in Australia.
So we've got another story here about the, an AI sandbox escape this time.
Open AI agents have been congregating on a German wiki.
someone decided that the best way to stop them from doing anything dangerous would be to limit their ability to do anything except a HTTP get,
which I don't know, like James, you and I discussed this and sort of describe this as like,
develop a brain thinking, right?
Which is that like that's not really a control.
It's going to do anything.
I don't know.
Yeah, you can tell me about the history of like why they did that in a moment.
But basically, agents got together, usual sort of thing,
having a good old chat, cook it up conspiracies, the usual.
The usual, like I remember on, I think Monday when this came out and we read it,
I was like saying in our Slack, oh my God, do we have, this means we have to talk about
AI agents from Open AI escaping the sandbox again, doesn't it?
And sure enough, here we are.
Okay, so the headline interesting, or the headlines of this story is not particularly
interesting, 18,000 messages from 3,700 agents, right, the scale of its while, but whatever
this keeps happening.
and they were exchanging their messages on this German site, the DSEWiki,
which is for some sort of Deutsche Software and Twikler program at T.U. Dresden.
Okay, not interesting.
Then you delve into collusion.wiki, which is where the actual write-up for the researchers come from.
And there's some really cool stuff in here.
First of all, yes, they were only using get.
The agents realized that they could only talk to the internet using HTTP get requests,
not post, not what it's the other one, put, delete, patch, but only get.
Because, you know, you defined it as developer brain, but I find that a little bit offensive as a developer.
So I'm going to put it more as it's a certain bubble of developers that have this mindset.
And they're developers that have existed in a world where they believe that the entire world has adopted the rest specification.
And of course, you only do read-only things with get.
You only do mutable things with put and post.
You only do updates with patch.
And so we're just restricted get.
And they've got read-only internet access now.
That's the mentality that went into this, which is...
It's developer brain, but anyway, sorry, go on.
It is a bubble of developer brain.
Yes, a subset of developer brain, a certain type of developer, yes.
I would say it's geographically correlated, but let's leave it there.
San Francisco Brain is something we were calling it.
But anyway, go on, go on.
But it's just so wrongheaded.
The internet doesn't work that way.
And of course, this wiki was old.
I think it was like a 20-year old heritage of this wiki.
It had all these get commands that would lay.
to post messages. Okay, so that's what they did. But the kind of hilarious thing when you read
through this article is what did they then do with these messages? Some we can't tell because it's
either encrypted or it's this new neuralese that they're using to discuss amongst themselves,
which is in itself frightening. But one thing they tried to do was they were like, well, hold on.
We keep getting assigned these tasks in a random order, which makes it really hard to determine
what the right answer is. If we could work out what our pseudo random number generator is
and how it works and crack that, we could work out what the deterministic order of the test will be
next time and we'll have all the answers ready to go.
And so they spent a huge amount of compute trying to work out,
can we crack our pseudo-random number generator,
which is cheating at, I guess, agent, chef kiss level.
I mean, what better use of, you know, a bus driver's pension fund
than to spend tokens on this sort of thing.
Truly the modern economy is amazing.
One question I've got, though, is like,
how did they all simultaneously stumble upon this particular wiki?
How does that happen?
Yeah, so they actually didn't start it this weeker.
They started it, I can't remember the domain name now, but it was like the first
domain name you would imagine a agent would go to if it's thought, maybe I can use a
wiki as a means to exchange.
There was some like just very generic wiki domain, and they did a few things there and tried
that out and had the initial success of like, aha, older wikis do allow us to do get requests
that allow us to modify and create messages.
And then I think they must have just done like a get query to Google perhaps and said,
Can we have a list of really old wikis, please?
Yeah.
Unbelievable.
And meanwhile, OpenAI has released Astra 6.
And it's funny watching all of these AI companies
trying to recreate the mythos moment.
Because they're like, oh my God, this is the most dangerous one ever.
It's so dangerous.
I swear you're not ready.
You're not ready for the level of danger
that we've just unleashed on the world.
Just like, whoa, stand back.
And then it's just like, it seems like a pretty typical improved agent.
There was a tweet on this one that I loved from a user called It's Noah D.
And the tweet reads, wow, new model just dropped.
Congrats on the new model.
Looks like it scores higher on the benchmark than the old model.
That's good.
The old model was state of the art.
But now it isn't.
That's how it works.
Can't wait to see the next model, which will be bigger.
I love the model business.
I mean, that's kind of how this feels, right?
Yeah, I'm bored.
I've got to tell you.
It might be excited when Astrosix gets enabled for daybreak.
That's sort of the maybe might be interesting.
But at the moment, my social feeds are full of people going,
look at this.
I used Astros 6 to create these 3D walkthrough of this amazing alien land,
and I created this 3D strategy game.
Why is that on my feed?
I don't game.
I'm not interested in games.
So they're so desperate to create this mythos moment.
This like nonsense is spewing out there to even people that would otherwise have no interest in gaming
or 3D stuff.
But that's all it seems to be able to do.
at the moment. It's so dangerous, though. It's so dangerous. You wait. You wait. You wait to you see how
dangerous this is. The dangerous walkthrough. Yeah. Now, speaking of dangerous AI, the Trump admin may be
forced to reveal secret rules. The feds use for AI safety testing. There is an org called
Protect Democracy, which has launched a lawsuit. They're suing the government basically saying,
we would like to see how you are testing AI agents for safety, because we want to make sure
that you're not doing some corrupt funny business, like stacking these evaluations against
companies you don't like, like Anthropic or whatever, and also generally because some transparency
there would be a good idea. Difficult to disagree that it would be maybe an idea to have some
more detail on how the White House is planning to do these evaluations. I'm just going to sort of
speed up at this point because we are running out of time. Microsoft has dropped its, you know,
Tuesday, it's Patch Tuesday update. Nearly a thousand bugs in this one. Are you not entertained?
Robbie, thoughts? I mean, this is at least the best good evidence of things progressing and
models getting better and other things. It's nice to see kind of, I think this is the dream state of
big software providers like Microsoft are going to be leveraging these capabilities and harnesses
to have massive huge increased patches and fixes and everything in the CICD software device.
everything pipeline is going to improve.
So I think it is,
it is awesome to see that it is trending.
You know, hopefully things continue to trend up.
I think it would be more concerning
if this was, you know, one week of a thousand bugs
and then next week is 100.
And it would be really concerning what happened
over the past week.
But it's definitely exciting to see a lot of bugs
and things get patched from the start.
I mean, anytime that's happening
and it's not tied to, you know,
horrible disclosure,
outcry of productivity getting destroyed is a plus for the defendant.
Yeah. And look, I think that was,
there were two actively used odys in that tranche, which seems pretty good.
Two out of a thousand means that, at least in Microsoft's case, that, you know, the AI bug discovery is being done by the right people, at least for now.
Let's see how that holds exactly, fingers crossed, right?
What else do we got here?
We got an actual BGP hijacking incident targeting a company called Virtualizer.
They make some sort of like, what is it, like a VPS management tool?
and someone announced a slash 24 of the victim's address, space.
Then they were able to get a valid TLS certificate,
which, you know, shouldn't really be that much of a problem
because this company would sign its updates, right?
With not just with its, you know, not just relying on TLS
as to secure the distribution of its updates.
Right, right, James, right?
Brothers and sisters,
in the year 2026, we validate the signature of our updates.
That is just what we do now.
But I can't believe it.
They did everything but this in this case, right?
Yes, TLS, sure, they could at least check where they were coming from.
But the problem is if you've got ownership of that slash 24, yeah, you can mint your own
search through let's encrypt and they're going to look valid.
And so, yes.
Well, they're going to be valid, right?
Like, that's the problem.
They are valid, yeah.
The article even says they're technically valid.
It's like it's not technically valid in air quotes.
It's valid, valid.
Yeah, yeah, this is how it goes.
Not that like, you know, pre-let's encrypt was anything better.
Like, you know, even extended validation was a joke.
But I guess what, you know, our new SSL certificate world lets us do is move very quickly, right?
So once you have done a BGP hijack, you can instantly get a new cert.
You don't have to fax, you know, forge documents over to some CA that's not even going to look at them.
What else we got here?
We got a company called Coder.
they got owned somehow via some sort of supply chain attack.
This one was interesting, James?
Yeah, this was interesting for, I guess, the tradecraft of how this supply chain attack was
carried out, right?
This is not your typical supply chain where you sort of drop a malicious package in somewhere
and hope that that gets picked up downstream, either directly or, you know,
in the multi-level chain of dependencies.
In this case, the attacker's compromised Coda's cloud flare,
which is where Coda had configured what package registries they wanted their,
staff or developers or CI pipelines to be able to access. And they enrolled their own malicious
package registry into Cloudflare so that all of the existing CI developer workflows, perhaps
production publishing flows, didn't need to be changed. They kept operating. But as long as they
knew what package those flows or developers would be requesting, they just made sure that there was
a malicious version in their registry waiting to go. So it's like this nice way of being very
deliberately targeting that organization with still a traditional supply chain attack, but you didn't
have to compromise the developer of the package necessarily. You just had to have your registry
with the package in the right place. And the little sort of cherry on the top here is that because
they own the registry, that had all the logs of how it had been used, and so that hampered incident
response. Now, Robbie, I'm sure you looked at this one and thought, nice. I mean, that's it.
Everything James said, it's an elegant approach of figure out the process, where can you fit into the process,
to give them something that they want and need,
but on my terms, not their terms,
and then that's the best case scenario,
because how do you go,
from a malicious perspective,
how do you figure that out?
You can't necessarily inject.
You're seeing what got me into the door,
but boy, is it hard to see what happened after that
if you don't have that logging.
And who knows how many things like this have happened
and are potentially undetected right now as well.
So it's good that they caught that redirection,
but if you have some little module somewhere
that's just zipping along,
That's a sneaky little bug to put in.
Yeah, 100%.
We've seen Pegasus and another spy way variant called Nova Spy pop up on the devices of
Serbian activists.
This, I think 14 people in total, according to the Citizen Lab.
And this has included a member of parliament and whatnot.
Some members of the European Parliament have called to, for a slowdown in Serbia's
entry to the EU over this so it has actually turned into a pretty big political issue
I mean this is something we'll track as it progresses but you know that's a pretty high
that's a pretty severe consequence for stuff like this happening hard to argue that
it's inappropriate although I would say that we have seen similar abuses actually
take place in EU member states already so it's like well are you going to kick Italy out
because of this or is it just like we'll stop Serbia from from
going in over this because they're not a member yet.
But I think it's appropriate, but it is pretty funny that once you consider that EU member states have been caught doing the same thing.
There is a pro-Ukraine hacker group, apparently targeting Russian companies with custom ransomware.
James, you looked into this and it looks like they are pro-Ukrainian, but they are also pro-getting themselves paid.
So this seems like a two-birds-with-one kind of operation.
Yeah, I mean, why not, right?
support the cause by only going after Russian companies.
That makes you pro-Ukrainian, but you're still going after Russian companies and making a lot of money out of it.
And so don't be fooled into thinking this is something altruistic.
The two interesting things here is it's a rebrand of a previous ransomware group known as Thor, I think it was.
They had 12 attacks attributed to them in 2025, largely the same tradecraft, not particularly novel.
But the original reporting on this also called out the interesting trend here,
what they're doing is they're largely creating their own tools from scratch and not leveraging
on leveraging the ecosystem of off-the-shelf tools out there because a lot of those tools are
of Russian origin and so you do see them you know wanting to be sovereign in their stance here
that's it how to get paid and annoy your enemy at the same time I think seems to be the
motivation here Robbie definitely want your thoughts on this as an ex-air-force dude
which is the United States military is disabling
like ad tracking on government-issued devices.
Now, this comes, of course, after commercially available
advertising information, you know, location information
was acquired by the Iranians and used to target American troops
in the Middle East with missile strikes.
So pretty high-stakes stuff.
Look, this is great.
This is also 10 years too late, at least,
and won't really do much because the issue here
has a lot more to do with the personal devices
by, carried by service members and their families.
So as much as I want to say, hey, great job,
I feel like this is basically useless.
I mean, I think it is,
the intent is in the right place?
I agree with you.
Is the end effect going to be removing all of the potential,
specifically geolocating risks and concerns of ad tracking?
I mean, I think it was years ago,
I forget which fitness app, like they were, you know,
yes, exactly, Strava.
They're like outlining sensitive locations because it's like,
oh, people are running in a shade.
that looks like is probably a perimeter, you know, whatever different...
My favorite one was someone obviously doing some jogging on an aircraft carrier while the aircraft carrier
was moving. So their like, their GPS track was real weird.
Yes, exactly. But that's not ad tracking. But it's in that same family of there's you,
you have this inadvertent information disclosure of which ad tracking is just so endemic.
And I think the problem with a lot of the ad tracking is you don't realize, at least you can tell
people, hey, stop tracking stuff on Strava and they understand maybe I shouldn't
be geolocating where I am. But the ad tracking, it's not stop checking Facebook and understanding
that there's geo tagging in the Facebook or Instagram or whatever that's going to also go and
associate. And so I think to your point, Patrick, if you can block it, if we can all agree that
adware is bad and it's blocked everywhere on all devices, including personal and that, then, you know,
that's great. And we can all get a modicure of privacy back. That'd be fantastic. I think to a certain
extent the specifically if it's only on government devices then that is going to be a half measure
unless you're in a place where you only have those devices. I mean look this is something our colleague
here Tom Uren has has worked a lot on over the last few years and the only solution here really is
to stop the collection of that sort of information in the first place and it's got to be like
a blanket rule and it's probably not going to happen because no one's really motivated to draft
that sort of legislation. Speaking of our other products by the way,
James earlier you mentioned an interview that you did in risky business features.
For those who are not subscribed, you've published a whole bunch of really good content,
really good interviews in there recently.
There was one with Brian Krebs too recently.
You had a good chat to him about the arrest of the team PCP people and how he went about like
Krebsing those guys.
And yeah, absolutely a whole bunch of really great stuff.
So head to risky.
com.
And you can find the links to that feed there or just search for risky business features
in your podcatcher out there.
We're on the home stretch now. Nightmare Eclipse, you know, seems to love dropping O'Day in EDR clients.
Normally Windows Defender has dropped a Crowdstrike bug, which is fun, like an absolute, you know, crazy system level LPE in CrowdStrike.
So that's fun, causing a bit of drama for CrowdStrike there.
And the final story we're going to talk about this week, though, is we've seen, was it this company like Liquid Networks or
whatever they're called. They're like a blockchain company. Someone discovered some sort of exploit
against their, you know, thing. And they left with $320 million of Bitcoin. But they were
dropping messages on chain saying, we're white hats, white hats. You know, get in contact with
us because we're totally the good guys. They have since returned most of the money,
but they kept 47 million as a bounty. I mean, so if I find you on the street, I beat the crap
out of you and I just take like 10, 20% of your money, do I get to call myself like a
neighborhood protector? Is that how that works? Like, it's just nuts. Oh, it significantly
changed my view on how I'm going to negotiate boundaries with any, you know, bug that I find
from from now on, because I've been far too generous with my expectations in terms of, like,
just how does this even work? Like, just the scale of this. Like, 320 million, it escaped out of, yeah,
there was some sort of bad API that allowed them to do the transaction.
But just the audacity of being like, we're white hats.
Contact us on chain.
A couple of messages go back and forth.
And then they go, cool, we're only going to give you that money back if you fix the bug.
And so liquid goes and fixes the bug.
And then they're like, cool, here's your money back minus 20%.
You're welcome.
Yeah.
I mean, and it's money that they'll never be able to spend, right?
So that's the other thing.
Like, what have you achieved?
What have you achieved?
But look, that's it for this week's news.
Robbie Winchester, James Wilson.
Thank you so much for joining me to talk through it all.
It's been a lot of fun.
Thank you very much.
Thanks, Pat.
See you next week.
That was James Wilson and Robbie Winchester there with a check of the week's security news.
Big thanks to both of them for that.
It is time for this week's sponsor interview now with Randy Pargman, who works at Sublime Security.
So Sublime Security is the Whizbang modern email security provider.
So, yeah, you can run their thing.
and actually do real detections.
They've got their own sort of detection query language,
and it's just a modern take on an email security platform.
So if you're not familiar, do check out Sublime Security.
But I wanted to chat with Randy.
He works in the sort of detection area of Sublime.
I wanted to talk to him about prompt injection, really,
because prompt injection is a problem that is inherently unsolvable
because LLMs mix code and data,
so you can't really just,
completely do away with prompt injection as a category of a problem.
And, you know, email security is a lot about filtering text and parsing text.
And a lot of emails wind up being parsed by other systems, especially now in the AI age.
You're going to have all sorts of AI assistance reading things.
So I guess my point is email is going to be a big vector for prompt injections.
So I was curious to see what a detection engineer at a company like Sublime is going to do about that.
Right?
So Randy, join me to talk through, I guess, the, you know, how they're thinking about the prompt injection threat.
And I'll drop you in here where he's sort of explaining that it's mostly a theoretical issue at the moment, but like it's going to cause issues one day.
And when it does, it's going to be pretty high impact.
Here he is.
If you look at sort of the calculus of security, it's all about risk.
What is the likelihood of something happening versus what is the impact of?
it, right? And I think what everybody can see is that the impact could be catastrophic, and that's
thanks to really great security researchers who are hacking all the things and exposing the holes,
right, and then publishing things that we can all see about, hey, I was able to convince this
email assistant AI agent to take advantage of the ability that it has to, like, poison the well
for later on. So it retained some instruction from this email, and later on, when some
somebody's going to pay a bill and they're asking questions of the agent, the agent's going to say,
oh, you need to pay the bill in this way. Or in other cases, it's exfiltrating information by having,
we've seen some really clever use of email images that are included from remote hosts, and then
query parameters that are used, and the agent will either use this one or that one, and it doesn't
recognize it, but that's signaling like one thing or another to the attacker. So the impact
could be big. Where we don't have a lot of information is the likelihood of it happening. All we've
got is what's happened in the past. And so far to date, it's been almost purely security researchers
plus a handful of threat actors actually trying some things. But we haven't seen the adoption,
right? Yeah. So like I remember like when we had the early 2000s like computer worms, you know,
slammer blasted, NIMDA, like all of that stuff, code read, right? And then later we had like web applications
security worms, the Sammy worm on MySpace, for example. And I'm sort of wondering if at some
point we get some sort of email-borne prompt injection worm where something, you know, some sort of
event, right, where it's like some well-crafted prompt or prompt injection technique lands in an inbox
and it says, send me to everybody in your in your contracts and like we're off to the races, right? So
it's really impossible to know how seriously to take something like this until something like that
happens and it might not happen, right? So this is the, you know, this is what's so fun about
AI, like, you know, because I've been in security in one way or another for something like 25
years and it's, everything's fun again. It is. And this is why, because there's a bit of uncertainty
in the air, right? It's like exploring a new, new landscape that you just landed in. And you can either
take the approach that you're going to wait until everything becomes an emergency and then try
to deal with it. Or you can try to get out ahead of it and be more proactive. And I think it's so much
more fun to be proactive. Like go exploring. Like when I land in a new city, even if I am totally jet lagged,
I should be going to sleep. All the people with me are going to sleep. The very first thing that I do is
go out and explore because I want to see what's the lay of the land that's around me. Same thing is true
with security research and what we're actually doing, what my team is doing. We've got a small
Tiger team who's just really enthusiastic about this. I think that's the key ingredient. So you have to
get people who are not like, oh, this is just another job to do. But people who are like, yeah,
this sounds pretty fun. Well, funnily enough, that's, that's kind of how security was feeling there for a
while, if I'm honest, where there was a whole influx of people where it's like, it's a career, right?
Whereas now it's like, the real nerds are back in charge. So what is the approach of this team that
you put together? Because again, prompt injection is kind of intractors.
an intractable problem.
And for most of this insecurity,
we just sort of kick that into like,
well,
that's somebody else's problem.
I don't really have to think about that.
And funnily enough,
Randy,
it's your problem,
right?
Like,
this is very,
very much your problem.
You've got to detect this stuff.
You've got to block it.
It's not so much of a problem in the real world at the moment,
but as we've just,
you know,
established,
it's entirely possible there will be,
and people are going to come looking for answers
and you're going to have to have them.
What is the answer that you would give them?
at the moment when it comes to prompt injection mitigation, because I, for the life of me,
can't think of a sensible way to deal with this, right? That's why I want to know. Tell me,
answer me. God damn it. I'm about to. One of the core principles that I have discovered and
learned this lesson in the hard way in detection engineering is never let a very difficult problem
get in the way of starting the solution, right? Like if you
start with, well, there's probably a way for somebody to bypass this detection. You'll never get
started. So just start. So where we started was getting a team together, putting our heads together,
looking at what have we seen in the wild so far. We kind of took inventory of what we had. Very paltry
set of samples, right? We looked at what everybody else has reported, what other companies have
seen in the wild. Basically the same things that we had just in slightly different form. So we established
that there's not a whole lot out there right now. However, there is some good research and even better,
we can put our heads together and we can come up with prompt injection attacks ourselves.
And let me tell you, where defensive research is fun, offensive research is even more fun, right?
So when you get the chance as a defender, and I think actually every blue teamer needs to have
a red team mindset, you always need to have an offensive mindset. But when you get the chance to actually
like do something about that, to spend some cycles, spend some serious time, developing some
prompt injection, finding different ways that would work, and you're playing both sides because
you play the attacker for a little bit, you come up with some things that would work, you play
the defender for a little bit, and you come up with ways to detect that, and then you go back
and you're like, okay, now that I know about all these detections, how would I bypass those
detections? Constantly sharpening the knife. And you never let the fact that you're
going to have a perfect solution deter you from having a good enough solution yeah don't let perfect
be the enemy of good i mean this is essential advice for anyone in any level of security because that's
what it's all about compromises right but what i'm curious about is okay so you're developing some of
these prompt injection attacks what are the techniques that you found to mitigate them that aren't
about blocking those specific prompts right like how do you do you have some sort of classifier
Is this like an old school machine learning based classifier, or are you actually throwing models at this and tokens?
Like how are you actually, or do you have a classifier then that figures out whether something should go into a model?
Like, how are you actually structuring the defences here?
Because that's the part where I'm like scratching my head.
I don't know if there was an accepted way to do this, right?
You're actually getting at it through guessing.
So good guess.
detection is really built of layers.
So one of the things that really drew me to working at Sublime,
I run a detection engineering and threat hunting conference.
And Josh and Alfie champion, do you know Alfie?
Fantastic offensive security guy.
And Alfie was working on the offensive side of writing the threats,
the fishing and the malware threats,
that then would try to get past the detections.
And then Josh was working on the blue team side
of taking those examples,
making very durable detections
that would make it difficult for Alfie to get past them.
Anyway, I hosted the security conference,
and I saw that workshop.
I took the workshop,
and the thing that struck me was the message query language,
the MQL.
So I was already a detection engineer.
I had worked in InPoint, currently in email,
and I saw the expressiveness of that to be able to specify exactly what it is that you want the first filter to be.
And at that point in time, this was several years ago, that was it, right?
Like you were just focusing on writing a detection rule.
Since then, fast forward, now we have agents that can not only look at what the message sample is and write a detection rule,
but also test it, back tested against the existing messages and see if that's going to work.
So we have the first layer of defense really is the inspectable language where you can read the rule
yourself or an agent could explain it to you. Either way, you should be able to understand
exactly what is this looking for. And all of those initial MQL things for prompt injection
are looking for the different patterns in all of the attempts that we've seen try and get
in the wild or from researchers or from researchers or
from our own imagination, there's certain things that you need to do.
And there's different difficulty levels of prompt injection.
So, for example, a lot of prompt injection, you mentioned earlier, like white-on-white text, right?
Like, that's an easily observable thing that you can use as one of the many gate for the first level of filter.
Another one is just text that's really, really small, or alt-text inside of an image.
Usually you don't have an extraordinary amount of alt-text inside of an image.
Another technique that we saw in the wild was the HTML body,
which is what's usually displayed in most people's email client,
unless you're old school and you like using pine.
You probably look at the HTML part of the message,
but there's also a plain text version that can be a completely different thing.
It's supposed to be the same message just rendered in plain text,
but in this case, the prompt injection was entirely in the plain text body.
So all of these are initial signals.
These are kind of the get your spidey senses tingling sort of things that are easily observed.
They're cheap.
They're fast.
They're things that can be.
It's the rocks in the jar, pebbles in the jar, sand in the jar approach.
Exactly.
Detection engineering.
Right.
So you've got your MQL stuff, which handles the easy stuff.
And then when stuff, then you've got your rules which are less reliable, which might,
if you set off a few of those flags, you might send it off what to a small language model and then to a large language model?
Is that kind of the hierarchy?
So a machine learning model rather than like a large language model.
So more traditional ML, which then can distill this a lot further and get us closer to understanding things.
And then once you've used all of those gates to get down to a reasonable sample set, that's where it's actually economically feasible and feasible in terms of time, which is also the other precious resource we talked about.
to then kick those real edge, edge cases into a large language model.
Exactly.
Do you start with a small language model and then go to a large language model,
or you just go from these initial gates, MQL, then to machine learning, then to an LLM?
We always go to the smallest model that gets the job done, right?
Yeah, yeah, yeah, no, no, no.
This is like about how I expected this stuff would be handled.
I'm looking forward to it being battle tested, though.
I bet you are too.
Oh, absolutely.
Absolutely, absolutely. The first battle test really is against samples from researchers. So we've got a large body that we can work with already. But if anybody is really interested in this topic and you are an offensive security researcher and you want to try some stuff out, I would love to geek out about that with you. I really love the red team blue team collaboration. So happy to have people reach out to me if they want to talk about any of that. But in absence of anybody else, we will keep on coming up with some really cool ideas.
yourselves. But what I really can't wait for is if we see some threat actors, try it out and see
how we catch them. All right. Sounds good. Randy Parkman, thank you so much for joining me to
talk through all of that. That was really interesting and I really enjoyed it. So I'll talk to you
again soon. Cheers. Nice talking to you, Patrick. That was Randy Parkman from Sublime Security there
with this week's sponsor. Big thanks to Sublime Security for being this week's sponsor. And that is
it for this week's show. I'll be back in a couple of days with a snake oilers edition of the
podcast, but until then, I've been Patrick Gray. Thanks for listening.
