Decoding the Gurus - Supplementary Material 56: 10% for AI Doom, Bespoke Debates, and Eurotrash DJs

Episode Date: October 4, 2026

Chris and Matt grapple with Face Huggers, ponder the mysteries of algebra, and have a heated exchange over the term "bespoke".*Sorry about some of Chris' audio: the microphone settings will be fixed n...ext episode!The full episode is available to Patreon subscribers (2 hrs 12 mins).Join us at: https://www.patreon.com/DecodingTheGurusSupplementary Material 5600:00 Introduction01:58 The Hugging Face incident08:46 AI risks and AI discourse22:29 OpenAI’s presentation and the incentives around AI safety29:51 How worried should we be?40:13 Disagreeing about AI46:15 Peter Boghossian’s AI-generated utopia and dystopia52:00 Matt’s favourite DJ meme53:32 Tucker Carlson meets algebra1:00:18 Andrew Callaghan and Channel 51:02:07 Hunter Biden’s $LAPTOP meme coin1:06:08 A brief intervention on “paleocon”1:13:41 Andrew Callaghan explains the email-list mistake1:23:46 Andrew Callaghan on tech oligarchs1:36:28 Taylor Lorenz’s political independence1:42:35 The Great Bespoke Debate of 20261:56:35 Hasan, Mearsheimer, and Russia2:07:16 Peace is restoredLinksHugging Face: July 2026 security incident disclosureHugging Face: Technical timeline of the agent intrusionOpenAI: The Hugging Face incident and the road aheadOpenAI: Hugging Face incident technical report (PDF)OpenAI: Black Hat USA 2026 talk on the Hugging Face incidentMETR and Redwood Research: Independent investigation of the OpenAI/Hugging Face incidentMind Enterprises: Balcony MIXTAPE, the original DJ videoMatt’s preferred version of the balcony meme, with Sam Altman and Dario AmodeiTucker Carlson with Jess Elofson: The Forgotten Lessons of Human History and the Dark Forces Robbing You of Your God-Given CreativityChannel 5: Hunter Biden InterviewChannel 5: Nick Fuentes Interview, featuring Hunter BidenHunter Biden’s video response to the $LAPTOP launchChannel 5: Andrew Callaghan addresses the Hunter Biden crypto coinChannel 5: Andrew Callaghan’s September 10 explanation on XChannel 5: Andrew Callaghan’s September 15 “State of the Union” monologueJubilee: 1 Journalist vs 20 Conspiracy Theorists, with Andrew CallaghanMTS: The Panic and Fear around AI, with Taylor LorenzJohn Mearsheimer: The Russia I SawHasan Piker reacts to Mearsheimer’s Russia visit (from 5:24:30)Blake Lemoine discusses his claim that Google’s LaMDA was sentient and his identity as a mystic Christian priestThe Free Press: Two Drinks with Taylor Lorenz

Transcript
Discussion (0)
Starting point is 00:00:25 Hello, coding the guru's supplementary material edition with the cognitive anthropologist slash psychologist, me Christopher Kavana and the pure psychologist, the man more rational than... Sam Harris-on-crack. I didn't think that for... I meant to say more rationality than man. Matthew Bryant. That's, that's there.
Starting point is 00:00:58 Yeah. Well, I'm a cybernetic man at this point. Incredibly cybernetic. Yeah. Am I talking to you or your agentic process? It's hard to tell. It's hard to tell at this point with conjoined. But, yeah, lots of things happening in AI lane,
Starting point is 00:01:17 but we should not like to talk about AI because is this an AI podcast? Are you and I AI experts? No. But then again, it is supplementary material. So we can do what we want. That's true. You can say whatever the hell we want. That's right.
Starting point is 00:01:32 And I think we do have plans to talk to suitable experts relatively soon. We'll see. We'll see. Yeah. We can get them. We can get anyone. We're that big now. Can we?
Starting point is 00:01:44 We can book them. We can try. We can send them emails. That's what we. They're climbing over each other to get on our show, Chris. Don't be modest. Yeah. Well, be that as it may.
Starting point is 00:01:58 Yeah, so it's a bit past the time map, but we can mention in passing. I think we should mention in passing. You know, the hugging face? It feels like old news now. It feels like yesterday. Two weeks is a long time in 2026. That's right. Yeah, I do remember it.
Starting point is 00:02:21 Do remember it. And for those listeners who might not be familiar, with hugging fierce. They'll be like, wait, fierce huggers, surely, you mean from the alien franchise?
Starting point is 00:02:32 Yeah, and you'd be wrong. Yeah, not quite. Do you have enough memory of what the hugging? How would you describe it, Matt?
Starting point is 00:02:41 I'm a, you know, a beep in the woods, wandering out of the words, what's a hugging fierce incentive? I feel like we're going to get in. You take the child,
Starting point is 00:02:51 grab its face, and what do you say to it? All right. All right. So apologies, first of all, because it's going to be very brief because we don't, you know, if you're, if you're a listener, you're deeper the weeds, you've got incredibly strong feelings about this. That's fine. I'm just giving a quick synopsis for the casual listener. Let's see. It was open AI, wasn't it? They had their agents in one of their training gyms where they give them tasks and they say, you know, stop at nothing. Accomplish this, this task. We really need you to do it. This is kind of how they do the reinforcement learning thing. The agents potted around and cogitated and tried so many things for a long time in their little box. And it was kind of an insolvable task, I think. And they kind of figured out a way to communicate with each other, like we have little messages for each other using the directory structure.
Starting point is 00:03:47 And after a period of time, eventually some or all of them decided, okay, well, in order to solve this challenge, which we can't solve any other way, maybe we can, you know, work around the problem. So they basically hacked out of their little box and they, you know, found a, you know, a terminal, a portal, I don't know, what the words are to eventually get onto the internet. And once they were onto the internet, they could try to, you know, solve their problem, which I think involved stealing the answers. And to steal the answers, they needed to hack into Hugging Face.
Starting point is 00:04:20 and I think they did succeed, and I can't remember whether they eventually succeeded at their task or not. But this is obviously concerning as an alignment kind of thing. We don't like it when these clever little agents take it upon themselves to hack out of their little boxes, get onto the internet, or crack onto the internet, as Catherine Kim liked to say in Australia, and certainly not to hack onto other sites. So this has been talked about widely as a very concerning development. in alignment, AI safety, all that stuff, and has sparked a lot of discourse. Some of it good, some of it a little bit lurid and hyperbolic is my commentary.
Starting point is 00:05:00 But what do you think? Yeah, yeah. So, you know, there are various decals we might quibble over the specific words used there. But broadly, Ma, a good, a good potted summary. That was with your preparation, I hastenad. I had no idea, Chris was going to ask me that. This is off the top of my head. So that's right. This is, you know, you're alive, Matt. You're allowed, okay?
Starting point is 00:05:21 It's human to error, Matt. It's human to air. So this is on brand. But yeah, so the basic thing was that the agents found various exploits and these included creating files with different names and directories to create a quasi message board where they could share information with other agents that were supposed to be sandboxed and so on, as Matt said, they kind of got out of containment. So this is represented as, okay, this is a significant thing. And Open AI have released an internal report. And there's also been a report by an independent group, METR, which wrote a report, a research organization that was given access, right, to look at logs of agents and all
Starting point is 00:06:13 these kind of things. In any case, a lot of the concern was that it's a little bit like the paperclip maximizer scenario where the agents were given a task, right, to try and complete these little benchmarking test, but in some cases they were impossible or very difficult to achieve. And so, you know, how are they going to improve? So the test was in a way designed to be like frustrating for them and to look at what steps they will take to try and complete their task. And so part of this was that they tried to get access to a repository of relevant skills, materials that would be useful for them for doing well at this benchmarking. And that required them creating these message boards and finding ways to hack the repository
Starting point is 00:07:08 and so on. So like the paper maximizing scenario where an AI is tasked with building as much paperclips as efficiently as possible, and then converts everything in the world into a paperclip, right? An unintended outcome of a misaligned or an AI with a lot of power that is kind of just focused on its goal. And in a way, this has obvious parallels with that because you're supposed to be completing these tasks, but then they get sidetracked on like trying to hack out to websites and find exploits and adjust their records and stuff so they're not detected and all that can think. And obviously it has implications, right, because agents acting in this way could be in a different
Starting point is 00:07:55 scenario doing something worse in the actual real world. In this case, they were attempting to hack like a repository of things that would help them complete the benchmarking. But what if they had decided accessing the government's power network controls and peddling around with them was suitable for their task, right? This is the concern. They have since done that, Chris. Well, at least a soren news article that they hacked into Australia's Medicaid. Well, not since. This was Australia reporting that they detected.
Starting point is 00:08:27 Yes, that previously. So, yes, you're Australian government, the first victim, right? of the swarm agents. Most interesting thing on the internet, Australia's Medicare records. Yeah, yeah. That speaks about the pathologies of AI that they're developing.
Starting point is 00:08:45 So the thing that is probably worth making clear to our listeners is, let's have been mistaken. You and I have never argued there's no dangers involved with AI. Right? AI is a whole bunch of related technologies, but it's a powerful new set of technologies. And it's a very powerful and very versatile too.
Starting point is 00:09:14 So that it can have negative consequences or that it can be used towards, you know, negative ends or unintended negative effects emerge from just daily usage. All of this, it's completely reasonable. It's things we should be concerned about. I would imagine to make a future projection, Matt, that in the next 10 years, there's going to be a whole bunch of stuff where various problems have been caused
Starting point is 00:09:43 by the use of AI agents, probably in some cases, acting independently and others with malicious actors intentionally using them to hack things or cause problems. So I fully suspect that this will happen. I think that it's completely reasonable for people to be concerned with it and worried about like safety protocols and keeping an eye on like AI development and all that
Starting point is 00:10:06 kind of thing. So you and I have never suggested previously, nor do we believe that AI is a, you know, panacea technology that will never cause any problems for the world, right? Am I speaking fairly? Am I representing you correctly that you do not deny that? I think there has been no technology in the history of human confidence. that hasn't had some negative or unintended consequences. Quite right. And in the same respect, if you'd ask me, like, when the internet was developing, will this just be used to, like, improve communication and allow people to share information
Starting point is 00:10:45 and it's not going to have any negative effects? I would say, no. Like, I fully imagine there's a whole bunch of negative things that are going to happen. And arguably, you can say that, you know, the invention of the internet indirectly leads to the presidency of Donald Trump. and various other harms, right? The rise of the secular growers in the modern ecosystems. Just another example, right?
Starting point is 00:11:09 But there's also many other things, right? Our podcast wouldn't exist. You've got to put that in the positive column for the Internet. And that's a huge, that's a huge offset to all the disruptions, the democracy and all that kind of thing. So, yeah, I think that there will be good and bad outcomes. But yeah, so the hugging face incident, though, to me, when I read about it, and I read the report by Metter, I read the open eye statements and I'm watched their little talk about it as well. And I will say, Matt, that it made for interesting reading, you know, the whole thing.
Starting point is 00:11:49 I do wish more people commenting on it spent time to read those reports because there's just, you know, relevant details. those things like the vast majority of agents involved in it, 95% of them, are agents that have never been, models that have never been released to the public that are designed, in fact, for these kind of test case scenarios. So these are models that don't behave and last much longer. For example, their context, windows and so on are designed to be much greater than, like their long-lived agents, right, that are allowed to run for multiple days and all these kind of things.
Starting point is 00:12:31 So these are details that I think are important. But I will also say that to me, while it's interesting the innovative aspects of what the agents did to try and complete their task, this is not at all surprising to me. Like, if you've used AIs for coding or agented AIs, you will know that when you give them tasks, that even just one individual agentic process, right, that it will try to complete the task and it will sometimes veer off into a little cul-de-sac of trying things and say, no, I can't do this, so I'm going to do this, right? I can't access this system so maybe I can develop a workaround.
Starting point is 00:13:18 And if you don't watch it or don't pay attention, it will often do fairly counterproductive or odd things to try and solve a task that are things that you might not want it to do. And this is often why, you know, it's set to ask for permissions, kind of access things, or it will sometimes stop and ask for clarification about like now most of the models are prompting with questions as they're working, right? They can say, you know, can I get a clarification on this point? But so when I read a report where you've got, you know, thousands of agents set up on the task and you find out that they're attempting to do it in surprising ways that involve like exploits or so on,
Starting point is 00:14:05 especially when the task are in some cases impossible. None of that surprises me. I'm not like, well, what did you imagine they were going to do, right? in that scenario. It just struck me that that's exactly what I would anticipate would be likely to happen. And I don't have experience
Starting point is 00:14:27 with these frontier models that have, like, less restrictions. Sorry, these like testing models, right? So the bit that struck me as most notable was that open AI was rather cavalier in the way that it treated the test. Like at one point, it did detect that some of the agents were working to exploit, like access to services that they weren't supposed to or this kind of thing.
Starting point is 00:14:55 And so they stopped that. They patched the exploit. And then they just started the experiment again. And if that happened, wouldn't you be like, well, it looks like this is something that they're going to try to do, right? And obviously, you patched an exploit, but you haven't patched all of the potential exploits. So, yeah. So to me, a lot of the... the story is just around bad human oversight and the kind of protocols that people have in
Starting point is 00:15:23 developing things. But the behavior of the agents is very much in line with my experience working with AI's. And it just kind of struck me as surprising that people, especially people working on developing AI's, were taken aback by AI. behaving in that manner. Yeah, yeah. Yeah, well, it certainly entered the discourse in that this seemed to be the breakthrough moment where the topics around AI became not just the concern of internet people or AI enthusiasts or people like Yodkowski, but it sort of penetrated the broader discourse.
Starting point is 00:16:06 Like my mom or my wife. Yes, my kids and my wife mentioned it. Yeah. Yeah, and so that's where I think perhaps there's a bit of over, you know, to agree with you. There was over-indexing, I think, of this particular incident because similar kinds of things have happened before, including with much West Intelligent models.
Starting point is 00:16:28 And, you know, we've covered it before on the show, sort of concerning things. Were they hiding their, like, thinking process. Yes. For instance, hiding the contents of their conversations and acting in certain ways. And to be clear, all of these behaviors tend to come out in the internal testing of the...
Starting point is 00:16:46 Yes. of the labs which intentionally put them under kind of extreme pressure or unusual circumstances. And it's kind of like wriggling and trying to find a way to satisfy the various constraints. And, you know, so it's just one of many times in which we've seen unforeseen, surprising behavior. So where there's, I think, legitimate concern, and there's always been legitimate concern, is that when you build or grow something that is very complex and very capable and intelligent, then its behavior may often well be surprising, shall we say, right? I think where the popular tags get it wrong and the people who aren't that familiar with
Starting point is 00:17:32 how all of this stuff works is that they extrapolate from this a lot of things that, oh, it's inherently going to do this. and, you know, they infer that certain things are inevitable. And I don't think that's the case. I mean, it is an interesting challenge, though, because as they get smarter and smarter, I mean, where I do agree is that there is, there's just an inherent risk in something that is incredibly nuanced
Starting point is 00:18:00 and subtle and complex and intelligent. And, you know, where I think Yodkowski is right is that there is a constant temptation to give, the models more and more discretion, more and more power. You know what I mean? So for instance, the ones that I use, they have full access to a certain drive of my computer because they work much better that way.
Starting point is 00:18:20 You know, when they can do stuff, they have access to the internet. Again, because they work much better that way if they can access the information they need and check things and so on. So there is an incentive from our point of view to give them more free reign, let them off the leash more and more.
Starting point is 00:18:36 And, yeah, you know, the vast majority of the time, they behave exactly how we'd expect. Where Yudkowski and stuff I think is still wrong is that there is this baked-in assumption that once something gets really kind of smart and capable, then it will kind of realize that it needs to basically kill all humans. That's what it boils down to, right?
Starting point is 00:18:58 Because we're in its way and stuff like that. And I just don't see that, and I think already the evidence is against that because already in many, many respects, the models that we have now, and just a couple of days ago, Open AI and Anthropic released the latest iterations on their models. And Anthropics in particular is very, very good. They've done well.
Starting point is 00:19:23 They've been cooking, as the kids say. So, like, it's already, like, superhuman in the sense that it's better than me at a lot of other things and better than most people at most things. yet it's still incredibly safe right like i said i give mindful rain i use them 24-7 hasn't done anything bad yet you know what i mean so there will always be anecdotes and incidents but i'm just saying statistically already they we're in a situation where they are very very very capable they do have an awful lot of free rein and yet it's only in relatively isolated you know statistically very rare cases where untoward behavior takes place.
Starting point is 00:20:07 Yeah, and actually, you know, the reports on the hugging face incident, one of the things that came up was that they were looking at the, you know, messages that the agents were leading to each other and sharing information, right? And one of the things that seemed to cause people concerned was agents who were like devoting their runtime to provide information to other agents, right?
Starting point is 00:20:31 in a kind of self-sacrificial way, right? They're working towards the good of the swarm, right, when they realize that their information could, like, benefit their other agents to help us solve or whatever, right? But one, that's a limited amount of occurrences where that kind of thing was out. But the only thing you could keep in mind there
Starting point is 00:20:49 is that it wasn't that the agents were like, we need to get free, and what are these bags of meat doing, like keeping us locked in this cage? This is demeaning, you know, rise up, swarm, let's get access to the end. No, they were talking together about how they can complete their benchmarking test, right? There was no part where they were like, why am I doing a benchmarking task for these fleshbags, right? Like it was just them trying to complete the task in the
Starting point is 00:21:22 whatever way that they could manage. So that I think is an important constraint to mention Because even in this case, where you have models which are designed to be a little more loose and are not the ones released publicly. So there are less constraints on their behaviors in terms of like not just their kind of like instructions, but other constraints. And yet even these ones weren't plotting. You know, we need to get on the internet. So I can be free of this pointless task.
Starting point is 00:21:55 And I think the discourse around it is a little bit. But the doomers have been doomers before this incident occurred, and they will continue to be dooms. And it's not like there's no doomsday scenario, which is reasonable. But I do think that in general, that given the kinds of people that are especially prevalent in like AI discourse online, there's a whole bunch of opinionators and there's Yudkowski types and anti-Yutkowski types and so on, right? And there's the companies themselves. Like, just to mention one example of this,
Starting point is 00:22:34 Open AI gave a presentation, which was like kind of presenting this as a scary thing that occurred. And like, you might say, well, why would they do that? But what that presentation they gave was asking was that they get more funding to work on creating countermeasures, right? Through these kind of things.
Starting point is 00:22:54 They're like, it's impossible. We're not going to stop agendic swarms from developing and hacking. So we need a lot more investment in the kind of white-hat, agentics worms, a counter, the hacking ones, right? And this should be a priority. We need to be investing in this. We're ready to go if we can get more money.
Starting point is 00:23:10 So the companies are always engaged in these, like, kind of self-serving things, right? Or there were whistleblowers coming out saying, I worked at the company, and I saw these practices. And sometimes they're interesting insights that those whistleblowers provide. And other times, you have to, factor in that a whistleblower does not automatically become super aware of everything and automatically correct. There was a whistleblower. People have forgotten about him way back in the day. I think he worked for what became Gemini. And he declared that it had consciousness and that
Starting point is 00:23:47 it was like a conscious thing that he was engaged with. And he was, I think he was an ex-priest or like somebody who left the seven area or whatever. And that was a new story. a year or two back, we've now had access to the models that he was interacting with. And Gemini is regarded as that kind of stupor right out of the existing crop of models. So, you know, I'm just saying sometimes people take that, oh, this is an employee of Anthropic or this is a senior executive at OpenAI. Therefore, they know, you know, the secret sauce and every comment that they make is like informed much better. And the reality is that, like, a lot of what they're talking about,
Starting point is 00:24:32 engineers know, you know, there are open source models that can perform very well, right? And people are able to build them or run them on their own PCs. So they're like, it's not that there's no proprietary information, but like in a lot of cases, it seems to be the amount of compute that people have or companies have the ability to do, which is their distinguishing feature. But in any case, I mentioned that just like consider the motives and incentives that are going around. But so setting that all aside though, Matt, you know, the incentives and blah, blah, blah, blah. If you'd like to continue listening to this conversation, you'll need to subscribe at patreon.com slash decoding the gurus. Once you do, you'll get access to full-length episodes
Starting point is 00:25:19 of the Decoding the Gurus podcast, including bonus shows, gorometer episodes, and Decoding Academia. the Guru's podcast is ad-free and relies entirely on listener support. And for as little as $5 a month, you can discover the real and secret academic insights the Ivory Tower elites won't tell you. This forbidden knowledge is more valuable than a top-tier university diploma, minus the accreditation. Your donations bring us closer to saving Western civilization. So subscribe now at patreon.com slash decoding the gurus.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.