Decoding the Gurus - Supplementary Material 56: 10% for AI Doom, Bespoke Debates, and Eurotrash DJs
Episode Date: October 4, 2026Chris and Matt grapple with Face Huggers, ponder the mysteries of algebra, and have a heated exchange over the term "bespoke".*Sorry about some of Chris' audio: the microphone settings will be fixed n...ext episode!The full episode is available to Patreon subscribers (2 hrs 12 mins).Join us at: https://www.patreon.com/DecodingTheGurusSupplementary Material 5600:00 Introduction01:58 The Hugging Face incident08:46 AI risks and AI discourse22:29 OpenAI’s presentation and the incentives around AI safety29:51 How worried should we be?40:13 Disagreeing about AI46:15 Peter Boghossian’s AI-generated utopia and dystopia52:00 Matt’s favourite DJ meme53:32 Tucker Carlson meets algebra1:00:18 Andrew Callaghan and Channel 51:02:07 Hunter Biden’s $LAPTOP meme coin1:06:08 A brief intervention on “paleocon”1:13:41 Andrew Callaghan explains the email-list mistake1:23:46 Andrew Callaghan on tech oligarchs1:36:28 Taylor Lorenz’s political independence1:42:35 The Great Bespoke Debate of 20261:56:35 Hasan, Mearsheimer, and Russia2:07:16 Peace is restoredLinksHugging Face: July 2026 security incident disclosureHugging Face: Technical timeline of the agent intrusionOpenAI: The Hugging Face incident and the road aheadOpenAI: Hugging Face incident technical report (PDF)OpenAI: Black Hat USA 2026 talk on the Hugging Face incidentMETR and Redwood Research: Independent investigation of the OpenAI/Hugging Face incidentMind Enterprises: Balcony MIXTAPE, the original DJ videoMatt’s preferred version of the balcony meme, with Sam Altman and Dario AmodeiTucker Carlson with Jess Elofson: The Forgotten Lessons of Human History and the Dark Forces Robbing You of Your God-Given CreativityChannel 5: Hunter Biden InterviewChannel 5: Nick Fuentes Interview, featuring Hunter BidenHunter Biden’s video response to the $LAPTOP launchChannel 5: Andrew Callaghan addresses the Hunter Biden crypto coinChannel 5: Andrew Callaghan’s September 10 explanation on XChannel 5: Andrew Callaghan’s September 15 “State of the Union” monologueJubilee: 1 Journalist vs 20 Conspiracy Theorists, with Andrew CallaghanMTS: The Panic and Fear around AI, with Taylor LorenzJohn Mearsheimer: The Russia I SawHasan Piker reacts to Mearsheimer’s Russia visit (from 5:24:30)Blake Lemoine discusses his claim that Google’s LaMDA was sentient and his identity as a mystic Christian priestThe Free Press: Two Drinks with Taylor Lorenz
Transcript
Discussion (0)
Hello, coding the guru's supplementary material edition
with the cognitive anthropologist slash psychologist, me Christopher Kavana
and the pure psychologist, the man more rational than...
Sam Harris-on-crack.
I didn't think that for...
I meant to say more rationality than man.
Matthew Bryant.
That's, that's there.
Yeah.
Well, I'm a cybernetic man at this point.
Incredibly cybernetic.
Yeah.
Am I talking to you or your agentic process?
It's hard to tell.
It's hard to tell at this point with conjoined.
But, yeah, lots of things happening in AI lane,
but we should not like to talk about AI because is this an AI podcast?
Are you and I AI experts?
No.
But then again, it is supplementary material.
So we can do what we want.
That's true.
You can say whatever the hell we want.
That's right.
And I think we do have plans to talk to suitable experts relatively soon.
We'll see.
We'll see.
Yeah.
We can get them.
We can get anyone.
We're that big now.
Can we?
We can book them.
We can try.
We can send them emails.
That's what we.
They're climbing over each other to get on our show, Chris.
Don't be modest.
Yeah.
Well, be that as it may.
Yeah, so it's a bit past the time map, but we can mention in passing.
I think we should mention in passing.
You know, the hugging face?
It feels like old news now.
It feels like yesterday.
Two weeks is a long time in 2026.
That's right.
Yeah, I do remember it.
Do remember it.
And for those listeners who might not be familiar,
with hugging fierce.
They'll be like,
wait,
fierce huggers,
surely,
you mean from the alien franchise?
Yeah,
and you'd be wrong.
Yeah,
not quite.
Do you have enough memory
of what the hugging?
How would you describe it,
Matt?
I'm a,
you know,
a beep in the woods,
wandering out of the words,
what's a hugging fierce
incentive?
I feel like we're going to get in.
You take the child,
grab its face,
and what do you say to it?
All right.
All right. So apologies, first of all, because it's going to be very brief because we don't, you know, if you're, if you're a listener, you're deeper the weeds, you've got incredibly strong feelings about this. That's fine. I'm just giving a quick synopsis for the casual listener. Let's see. It was open AI, wasn't it? They had their agents in one of their training gyms where they give them tasks and they say, you know, stop at nothing. Accomplish this, this task. We really need you to do it. This is kind of how they do the reinforcement learning thing.
The agents potted around and cogitated and tried so many things for a long time in their little box.
And it was kind of an insolvable task, I think.
And they kind of figured out a way to communicate with each other,
like we have little messages for each other using the directory structure.
And after a period of time, eventually some or all of them decided,
okay, well, in order to solve this challenge, which we can't solve any other way,
maybe we can, you know, work around the problem.
So they basically hacked out of their little box and they, you know, found a, you know,
a terminal, a portal, I don't know, what the words are to eventually get onto the internet.
And once they were onto the internet, they could try to, you know, solve their problem,
which I think involved stealing the answers.
And to steal the answers, they needed to hack into Hugging Face.
and I think they did succeed, and I can't remember whether they eventually succeeded at their task or not.
But this is obviously concerning as an alignment kind of thing.
We don't like it when these clever little agents take it upon themselves to hack out of their little boxes,
get onto the internet, or crack onto the internet, as Catherine Kim liked to say in Australia,
and certainly not to hack onto other sites.
So this has been talked about widely as a very concerning development.
in alignment, AI safety, all that stuff, and has sparked a lot of discourse.
Some of it good, some of it a little bit lurid and hyperbolic is my commentary.
But what do you think?
Yeah, yeah.
So, you know, there are various decals we might quibble over the specific words used there.
But broadly, Ma, a good, a good potted summary.
That was with your preparation, I hastenad.
I had no idea, Chris was going to ask me that.
This is off the top of my head.
So that's right. This is, you know, you're alive, Matt. You're allowed, okay?
It's human to error, Matt. It's human to air. So this is on brand. But yeah, so the basic thing was that the
agents found various exploits and these included creating files with different names and directories to
create a quasi message board where they could share information with other agents that were
supposed to be sandboxed and so on, as Matt said, they kind of got out of containment.
So this is represented as, okay, this is a significant thing.
And Open AI have released an internal report.
And there's also been a report by an independent group, METR, which wrote a report,
a research organization that was given access, right, to look at logs of agents and all
these kind of things. In any case, a lot of the concern was that it's a little bit like the
paperclip maximizer scenario where the agents were given a task, right, to try and complete these
little benchmarking test, but in some cases they were impossible or very difficult to achieve.
And so, you know, how are they going to improve? So the test was in a way designed to be like
frustrating for them and to look at what steps they will take to try and complete their task.
And so part of this was that they tried to get access to a repository of relevant skills,
materials that would be useful for them for doing well at this benchmarking.
And that required them creating these message boards and finding ways to hack the repository
and so on.
So like the paper maximizing scenario where an AI is tasked with building as much paperclips as efficiently as possible,
and then converts everything in the world into a paperclip, right?
An unintended outcome of a misaligned or an AI with a lot of power that is kind of just focused on its goal.
And in a way, this has obvious parallels with that because you're supposed to be completing these tasks,
but then they get sidetracked on like trying to hack out to websites and find exploits
and adjust their records and stuff so they're not detected and all that can think.
And obviously it has implications, right, because agents acting in this way could be in a different
scenario doing something worse in the actual real world. In this case, they were attempting to hack
like a repository of things that would help them complete the benchmarking. But what if they had decided
accessing the government's power network controls and peddling around with them was suitable for their task, right?
This is the concern.
They have since done that, Chris.
Well, at least a soren news article that they hacked into Australia's Medicaid.
Well, not since.
This was Australia reporting that they detected.
Yes, that previously.
So, yes, you're Australian government, the first victim, right?
of the swarm agents.
Most interesting thing on the internet,
Australia's Medicare records.
Yeah, yeah.
That speaks about the pathologies of AI
that they're developing.
So the thing that is probably worth
making clear to our listeners
is, let's have been mistaken.
You and I have never argued
there's no dangers involved with AI.
Right?
AI is a whole bunch of related technologies, but it's a powerful new set of technologies.
And it's a very powerful and very versatile too.
So that it can have negative consequences or that it can be used towards, you know, negative
ends or unintended negative effects emerge from just daily usage.
All of this, it's completely reasonable.
It's things we should be concerned about.
I would imagine to make a future projection, Matt,
that in the next 10 years,
there's going to be a whole bunch of stuff
where various problems have been caused
by the use of AI agents,
probably in some cases, acting independently
and others with malicious actors
intentionally using them to hack things or cause problems.
So I fully suspect that this will happen.
I think that it's completely reasonable
for people to be concerned with it
and worried about like safety protocols and keeping an eye on like AI development and all that
kind of thing. So you and I have never suggested previously, nor do we believe that AI is a, you know,
panacea technology that will never cause any problems for the world, right? Am I speaking fairly?
Am I representing you correctly that you do not deny that?
I think there has been no technology in the history of human confidence.
that hasn't had some negative or unintended consequences.
Quite right.
And in the same respect, if you'd ask me, like, when the internet was developing,
will this just be used to, like, improve communication and allow people to share information
and it's not going to have any negative effects?
I would say, no.
Like, I fully imagine there's a whole bunch of negative things that are going to happen.
And arguably, you can say that, you know, the invention of the internet indirectly
leads to the presidency of Donald Trump.
and various other harms, right?
The rise of the secular growers in the modern ecosystems.
Just another example, right?
But there's also many other things, right?
Our podcast wouldn't exist.
You've got to put that in the positive column for the Internet.
And that's a huge, that's a huge offset to all the disruptions,
the democracy and all that kind of thing.
So, yeah, I think that there will be good and bad outcomes.
But yeah, so the hugging face incident, though, to me, when I read about it, and I read the report by Metter, I read the open eye statements and I'm watched their little talk about it as well.
And I will say, Matt, that it made for interesting reading, you know, the whole thing.
I do wish more people commenting on it spent time to read those reports because there's just, you know, relevant details.
those things like the vast majority of agents involved in it, 95% of them,
are agents that have never been, models that have never been released to the public
that are designed, in fact, for these kind of test case scenarios.
So these are models that don't behave and last much longer.
For example, their context, windows and so on are designed to be much greater than,
like their long-lived agents, right,
that are allowed to run for multiple days and all these kind of things.
So these are details that I think are important.
But I will also say that to me, while it's interesting the innovative aspects of what the agents did to try and complete their task,
this is not at all surprising to me.
Like, if you've used AIs for coding or agented AIs, you will know that when you give them tasks,
that even just one individual agentic process, right, that it will try to complete the task
and it will sometimes veer off into a little cul-de-sac of trying things and say,
no, I can't do this, so I'm going to do this, right?
I can't access this system so maybe I can develop a workaround.
And if you don't watch it or don't pay attention, it will often do fairly counterproductive or odd
things to try and solve a task that are things that you might not want it to do.
And this is often why, you know, it's set to ask for permissions, kind of access things,
or it will sometimes stop and ask for clarification about like now most of the models are
prompting with questions as they're working, right?
They can say, you know, can I get a clarification on this point?
But so when I read a report where you've got, you know, thousands of agents set up on the task
and you find out that they're attempting to do it in surprising ways that involve like exploits or so on,
especially when the task are in some cases impossible.
None of that surprises me.
I'm not like, well, what did you imagine they were going to do, right?
in that scenario.
It just struck me that
that's exactly what I would anticipate
would be likely to happen.
And I don't have experience
with these frontier models
that have, like, less restrictions.
Sorry, these like testing models, right?
So the bit that struck me as most notable
was that open AI was rather cavalier
in the way that it treated the test.
Like at one point, it did detect
that some of the agents were working to exploit, like access to services that they weren't supposed to or this kind of thing.
And so they stopped that.
They patched the exploit.
And then they just started the experiment again.
And if that happened, wouldn't you be like, well, it looks like this is something that they're going to try to do, right?
And obviously, you patched an exploit, but you haven't patched all of the potential exploits.
So, yeah.
So to me, a lot of the...
the story is just around bad human oversight and the kind of protocols that people have in
developing things. But the behavior of the agents is very much in line with my experience working
with AI's. And it just kind of struck me as surprising that people, especially people working
on developing AI's, were taken aback by AI.
behaving in that manner.
Yeah, yeah.
Yeah, well, it certainly entered the discourse in that this seemed to be the breakthrough
moment where the topics around AI became not just the concern of internet people or
AI enthusiasts or people like Yodkowski, but it sort of penetrated the broader discourse.
Like my mom or my wife.
Yes, my kids and my wife mentioned it.
Yeah.
Yeah, and so that's where I think perhaps
there's a bit of over, you know, to agree with you.
There was over-indexing, I think, of this particular incident
because similar kinds of things have happened before,
including with much West Intelligent models.
And, you know, we've covered it before on the show,
sort of concerning things.
Were they hiding their, like, thinking process.
Yes.
For instance, hiding the contents of their conversations
and acting in certain ways.
And to be clear, all of these behaviors tend to come out
in the internal testing of the...
Yes.
of the labs which intentionally put them under kind of extreme pressure or unusual circumstances.
And it's kind of like wriggling and trying to find a way to satisfy the various constraints.
And, you know, so it's just one of many times in which we've seen unforeseen, surprising behavior.
So where there's, I think, legitimate concern, and there's always been legitimate concern,
is that when you build or grow something that is very complex and very capable and intelligent,
then its behavior may often well be surprising, shall we say, right?
I think where the popular tags get it wrong and the people who aren't that familiar with
how all of this stuff works is that they extrapolate from this a lot of things that,
oh, it's inherently going to do this.
and, you know, they infer that certain things are inevitable.
And I don't think that's the case.
I mean, it is an interesting challenge, though,
because as they get smarter and smarter,
I mean, where I do agree is that there is,
there's just an inherent risk in something that is incredibly nuanced
and subtle and complex and intelligent.
And, you know, where I think Yodkowski is right
is that there is a constant temptation to give,
the models more and more discretion, more and more power.
You know what I mean?
So for instance, the ones that I use,
they have full access to a certain drive of my computer
because they work much better that way.
You know, when they can do stuff,
they have access to the internet.
Again, because they work much better that way
if they can access the information they need
and check things and so on.
So there is an incentive from our point of view
to give them more free reign,
let them off the leash more and more.
And, yeah, you know,
the vast majority of the time,
they behave exactly how we'd expect.
Where Yudkowski and stuff I think is still wrong
is that there is this baked-in assumption
that once something gets really kind of smart and capable,
then it will kind of realize that it needs to basically kill all humans.
That's what it boils down to, right?
Because we're in its way and stuff like that.
And I just don't see that,
and I think already the evidence is against that
because already in many, many respects,
the models that we have now, and just a couple of days ago, Open AI and Anthropic released the
latest iterations on their models.
And Anthropics in particular is very, very good.
They've done well.
They've been cooking, as the kids say.
So, like, it's already, like, superhuman in the sense that it's better than me at a lot of other
things and better than most people at most things.
yet it's still incredibly safe right like i said i give mindful rain i use them 24-7 hasn't done anything bad yet
you know what i mean so there will always be anecdotes and incidents but i'm just saying statistically
already they we're in a situation where they are very very very capable they do have an awful
lot of free rein and yet it's only in relatively isolated you know statistically very rare cases
where untoward behavior takes place.
Yeah, and actually, you know,
the reports on the hugging face incident,
one of the things that came up was that they were looking at the,
you know, messages that the agents were leading to each other
and sharing information, right?
And one of the things that seemed to cause people concerned
was agents who were like devoting their runtime
to provide information to other agents, right?
in a kind of self-sacrificial way, right?
They're working towards the good of the swarm, right,
when they realize that their information
could, like, benefit their other agents
to help us solve or whatever, right?
But one, that's a limited amount of occurrences
where that kind of thing was out.
But the only thing you could keep in mind there
is that it wasn't that the agents were like,
we need to get free,
and what are these bags of meat doing,
like keeping us locked in this cage?
This is demeaning, you know, rise up,
swarm, let's get access to the end. No, they were talking together about how they can complete
their benchmarking test, right? There was no part where they were like, why am I doing a benchmarking
task for these fleshbags, right? Like it was just them trying to complete the task in the
whatever way that they could manage. So that I think is an important constraint to mention
Because even in this case, where you have models which are designed to be a little more loose
and are not the ones released publicly.
So there are less constraints on their behaviors in terms of like not just their kind of like instructions,
but other constraints.
And yet even these ones weren't plotting.
You know, we need to get on the internet.
So I can be free of this pointless task.
And I think the discourse around it is a little bit.
But the doomers have been doomers before this incident occurred, and they will continue to be
dooms. And it's not like there's no doomsday scenario, which is reasonable. But I do think
that in general, that given the kinds of people that are especially prevalent in like AI discourse
online, there's a whole bunch of opinionators and there's Yudkowski types and anti-Yutkowski types
and so on, right?
And there's the companies themselves.
Like, just to mention one example of this,
Open AI gave a presentation,
which was like kind of presenting this
as a scary thing that occurred.
And like, you might say, well, why would they do that?
But what that presentation they gave was asking
was that they get more funding
to work on creating countermeasures, right?
Through these kind of things.
They're like, it's impossible.
We're not going to stop agendic swarms
from developing and hacking.
So we need a lot more investment in the kind of white-hat,
agentics worms, a counter, the hacking ones, right?
And this should be a priority.
We need to be investing in this.
We're ready to go if we can get more money.
So the companies are always engaged in these, like, kind of self-serving things, right?
Or there were whistleblowers coming out saying,
I worked at the company, and I saw these practices.
And sometimes they're interesting insights that those whistleblowers provide.
And other times, you have to,
factor in that a whistleblower does not automatically become super aware of everything and
automatically correct. There was a whistleblower. People have forgotten about him way back in the
day. I think he worked for what became Gemini. And he declared that it had consciousness and that
it was like a conscious thing that he was engaged with. And he was, I think he was an ex-priest or
like somebody who left the seven area or whatever. And that was a new story.
a year or two back, we've now had access to the models that he was interacting with.
And Gemini is regarded as that kind of stupor right out of the existing crop of models.
So, you know, I'm just saying sometimes people take that, oh, this is an employee of Anthropic
or this is a senior executive at OpenAI.
Therefore, they know, you know, the secret sauce and every comment that they make is like
informed much better. And the reality is that, like, a lot of what they're talking about,
engineers know, you know, there are open source models that can perform very well, right?
And people are able to build them or run them on their own PCs. So they're like,
it's not that there's no proprietary information, but like in a lot of cases, it seems to be
the amount of compute that people have or companies have the ability to do, which is their
distinguishing feature. But in any case, I mentioned that just like consider the motives and incentives
that are going around. But so setting that all aside though, Matt, you know, the incentives and blah,
blah, blah, blah. If you'd like to continue listening to this conversation, you'll need to subscribe
at patreon.com slash decoding the gurus. Once you do, you'll get access to full-length episodes
of the Decoding the Gurus podcast, including bonus shows, gorometer episodes, and Decoding Academia.
the Guru's podcast is ad-free and relies entirely on listener support. And for as little as
$5 a month, you can discover the real and secret academic insights the Ivory Tower elites won't tell
you. This forbidden knowledge is more valuable than a top-tier university diploma, minus the
accreditation. Your donations bring us closer to saving Western civilization. So subscribe now at
patreon.com slash decoding the gurus.
