Angry Planet - Human Decisions Are Behind Every ‘Rogue’ AI
Episode Date: October 2, 2026A lot of people are afraid right now. I’ve lost track of the times OpenAI, Anthropic, and Meta have announced they lost control of their AI. Anthropic CEO Dario Amodei wants to “pace the frontier�...�� and Bill Gates is warning there’s a non-zero chance that AI will kill billions. The US military has embraced AI and it’s leading to grave mistakes. In Ukraine, an AI-controlled drone has killed three people.Something is bothering me about all these stories and it’s not the thought that artificial intelligence could wipe out humanity. It’s the distance AI has given human designers from the consequences of their actions.This week on the show Eryk Salvaggio stops by to talk about intentionality, design, AI doom, and the messy way we’re all talking about this new technology. Salvaggio is a Gates Scholar at Cambridge University.A bad frame from the jumpDefining AIWhat the Hugging Face attack was and wasn’tWho’s accountable?Our anthropomorphic media narrativeDesign vs intentionalityLanguage loses its magicSpin not marketingGeopolitical stakesPrompts and outputsThe future of regulationRogue AI didn’t breach Hugging Face, human decisions didUS military had close call after using AI for false intelligence report, sources sayInside the US ‘Kill Chain’ that Destroyed an Iranian SchoolSupport this show http://supporter.acast.com/warcollege. Hosted on Acast. See acast.com/privacy for more information.
Transcript
Discussion (0)
Love this podcast.
Support this show through the ACAST supporter feature.
It's up to you how much you give, and there's no regular commitment.
Just click the link in the show description to support now.
Hello, and welcome to another conversation about conflict on an angry planet.
I am Matthew Galt.
We're going to talk about the thing that's on everyone's minds right now for this episode.
AI killing us all.
That's a bad frame already off the jump.
I think, but I mean, that's kind of what this episode's going to be about.
With me to do that, sir, can I have you introduce yourself?
My name is Eric Salvagio, and I am a gate scholar at the University of Cambridge in the
digital humanities where I look at artificial intelligence, and I'm also an affiliated researcher
with the Max Blank Institute's Machine Visual Cultural Research Group out in Rome,
and it's a pleasure to be here.
Thank you so much for hopping on.
I reached out to you because I felt like there's just, there's a lot of talk right now in the ether,
in the aftermath of the hugging face attack and some, like a whole bunch of other similar incidents
that have happened across kind of the tech space.
People are looking at this and saying like the AIs have gotten too smart.
They've broken containment and they're going to kill us all.
And you wrote something in the bulletin, the bulletin the atomic scientists,
that I thought was very good on the subject and was one of the only things that
really kind of hit all the beats that I've been thinking about and also really focused on human
accountability. The title of it is Rogue AI didn't breach hugging face, human decisions did.
I want to link to it in the show notes. But I want to do one thing first, which is something I do
every time I have a guest on when we talk about this subject. Will you define AI for me?
Oh, well, that's already a can of worms.
I know.
I know.
I know you know.
Most of your listeners probably do as well.
Artificial intelligence is really the way that I look at it is that it is, it's kind of an anchor for describing various parts of a technological system.
Right now, when we talk about AI, when I talk about AI,
It is mostly speaking around something called generative AI,
which is large language models,
diffusion models for images, video, music, and so on.
And so there's a whole other conversation about AI
that is machine learning, computer vision models,
and all this stuff.
But I think when I'm talking about AI,
and when I say AI and throw it around quite loosely,
which I oftentimes do to my own detriment,
what I really mean is a system that produces,
a kind of output, be a text, image, bound, film, or video, I guess we would say, as opposed to
something that is strictly a point of analysis like you would have for these like a
cantric screening, for example, or something like that.
So that's what I mean when I say AI is basically something that takes data in of some
sort and returns data of a similar source on the other side.
of a mechanism for processing all of that data.
And that's a very loose sort of umbrella definition.
So for the Angry Planet longtime diehards,
I'm going to say what just happened.
I'm using a new recording software
because I'm trying to make it better
and sound better for everyone.
And it allows me to do all these wonderful things,
like diagnose why the mics didn't sound as good as they should have.
And now, Eric, if you'll speak for me again.
We are a changing course to my real microphone.
Beautiful.
Sounds so much better.
Okay.
So then I think for this conversation, is it better to call it, like specifically the things that happened with Hugging Face.
Is it better to call it an LLM?
Yeah, I would say, I mean, what happened with Hugging Face is essentially a large language model, which it's kind of a different.
way of arranging a large language model, right? And so sometimes these are called agentic systems.
A lot of people don't like that because it implies a kind of agency, whatever. But I think if we are
kind of understanding the themes that we are talking about here, it's going to be very transparent
how these things work. And so I'm happy to refer to them as an agentic model or an agentic
large language model, an agentic form of a large language model, which,
whatever you prefer is going to be fine with me.
But I think the parlance you see being thrown around around this incident is an agentic system.
So will you tell me your version of events?
I think, you know, a lot of people have read kind of the big papers takes on it and a lot of, you know, Twitter threads about what happened.
But I really liked your kind of breakdown of this attack for people that may not remember.
It's been a few weeks old now or people that don't know.
Yeah, I mean, a lot of what I come to know and a lot of what any of us come to know is actually through OpenAI and through a meter, which is the model evaluation and threat research report, which is an independent organization that did an audit, though there's a lot of sort of rotating doors between that organization and the various big tech companies that we're talking about or we'll be talking about.
So that's a caveat.
That's an important thing to say, is that my information is coming from these reports as well.
Second, meter and Open AI also relied on the models to sort of make sense of what it is happening.
And literally the model that one of the models that's involved in the incident was used to summarize the decisions that it itself were involved in making.
And so there's multiple layers here that are giving me some real distance.
from the events. But I am basically trying to read these reports. And what I'm looking for in
these reports is the way that they deploy what I call an intentional language, right? They're talking
about things like what the model realized or believed. And I'm trying to understand what exactly
that means. Because there are mechanisms that produce these outcomes, right, that are under the hood.
And to me, it's one thing to say in a casual conversation, the models thought they were.
you know, capturing a flag or something, but then it's a different thing to leave it at that.
And what I'm trying to do is think through, well, what exactly happened here? And so there's a
version of the events, which is basically Open AI had these two models. They're trying to test
these models. They have one that's an in-house model, and they've one, which turned out basically
to be Seoul, which is GPT 5.6, their newer model. And they're testing these. And one of the things
that you're often trying to do when you're an AI companies, you're trying to test them to
evaluate them for their ability for cybersecurity. And you have these benchmark tests that the models
are designed, sorry, the models are being tested against. And so one of these is something called
Exploitte Gym, Open AI wants to test it, benchmark it, see how well it performs on some of these
tasks. So they basically put into play
these two models.
They've optimized these models for persistence,
which we can talk about.
It basically means they write a ton of text
and they don't stop writing text
until a goal has been met.
And then they also do a couple of other things,
which I think are questionable,
but these are in the report.
One is they disabled a lot of the security mechanisms
in that larger internal model.
And to a certain degree, if you want to sort of think about like why they would do this, you can think about this from a kind of business sense, right?
The sort of bureaucratic explanation of this is you had a company that wanted to say we did really well on these security benchmarks.
In other words, we were able to hack things, right, able to find exploits because that says if you can find those exploits, you could defend against them.
So they're trying to say, like, let's let's take all.
some of these guardrails and see what the system can do if we don't give the system breaks,
right? So some of these security guardrails have been turned off. So that's one of the key things
that ended up happening. And then they also ran into some trouble because a lot of the tasks are
impossible. So when you have a machine that is basically told, push the gas pedal until you hit a wall,
and there are just no walls, you are driving that car straight into the ocean, straight into, you know, pedestrians, you are full throttle.
And so that's basically what this model ended up doing. Those are the two sort of core components of what was going on under the hood.
So it's less that it went rogue and more like it did exactly what human beings told it to do and the door.
were left open and they went out and did the tasks?
There's a version of this that says that OpenAI basically designed a test environment, right?
And there was, they understood the model they were testing.
They understood that it was optimized for persistence.
The other thing that they did is that they were supposed to take it offline, arguably
whenever you test the model like this, you don't have internet access. And I'm saying this mostly
because I think that this idea of like, did they tell the model what to do? I think it's actually
more complicated, and it's complicated in interesting ways. Because the thing did find exploits in order
to get online. But there's no reason to think that Open AI wouldn't have anticipated that it
would find these exploits in order to get online. And so to me, it seems very odd that they left
the model connected to the internet at all.
And the way that this happened is there's this thing called artifactory, which is supposed to be this sort of like, I refer to it sometimes as like the prison store.
If you're in prison and you need toothpaste, you go to the store in the prison that is sort of sanctioned and you say, give me the toothpaste and the store gives you the toothpaste, but you do not have access to the delivery truck as a prisoner, right?
whereas the people working in the store do.
And so this is what it was.
The artifacter is basically saying,
okay, the model's in jail.
It wants, and I'm going to use this with air quotes, right?
But it's trying to get software.
It's trying to run code.
And for that to work,
some of this stuff needs to be taken in from the internet.
And so you had this thing that was supposed to work as an intermediary
to sort of regulate the type of access that it could have to the internet.
This model that is optimized from persistence basically found
an exploit in order to use this to access online information, right? So this is how it actually gets to
hugging face. Now, the question, just to come back to your question, I'm not dodging that question,
but saying like, okay, when we understood how that happened, the question comes down to,
open AIs built this environment for a model that they are saying, we can't anticipate what it's going
to do, but we do know that we have trained the model to not give up on certain things. And then
we are giving it a series of tasks, some of which are impossible, about 200 out of the
thousand or so tests are impossible. So what's it going to do? Well, I would say that that's not
quite the same as saying go hack, hugging face, but it is in a way the same thing as saying,
we don't know what's going to happen, let's find out, plugging the thing into a situation
where internet access is possible, and then kind of just letting it run and letting it run in a kind of
unmonitored or carelessly monitored fashion.
So there's a lot going on here.
There's the bureaucratic pressure of the benchmarks.
There's the model and the way that it is designed and the mechanisms that are been trained
into it through pre-training, through the architecture, right, through the decisions about
what this model is going to do.
The testing environment itself is another type of pressure.
And so you have all these kinds of competing pressures that are pushing toward a kind
of resolution that turns out, in this case,
was an actual hacking situation.
And so is it just somebody saying, let's hack Hugging Face?
Here's a powerful model.
Go attack?
Not quite.
I think it's more about this kind of, I think Melanie Mitchell talked about
an excellent blog post that she wrote or newsletter article about a wildfire
and saying we're going to do a controlled burn, but we're not going to check which way
the wind is blowing.
And that to me is where the sort of accountability question comes into play is,
Yeah, when you deploy an unpredictable system, shouldn't you anticipate its unpredictability?
And this is why the, oh, the model's unpredictable.
We have no way of knowing what the models are doing.
We could never anticipate this.
Well, if you couldn't anticipate it, build a better boundary, build a test environment that actually allows you to test.
That's where I think these questions become more nebulous than sort of pointing fingers.
Well, in the report, if I recall correctly, aren't there.
several points where OpenAI is aware of what's going on to a certain extent and kind of
allows that wildfire to continue to burn?
Yep.
One of them in particular is this whole thing about message boards.
They talk about a unsanctioned message board.
And this language comes up, I believe, in both reports, the independent report and Open AI's
report.
And it does lead me to say, well, what do you mean unsanctions?
Because you were aware that this behavior was occurring in these models.
You optimize them to coordinate.
which we can unpack that a little bit, right?
But they are kind of in a way a model designed to pass notes to other instances of itself,
which are other agents.
And so this becomes a thing where, all right, it goes into the prison store, artifactory,
and it starts leaving notes in the file structure, right?
So it's like kind of almost like writing little notes about like what time you're going to be coming in the next day.
But it's actually a message, right, for the next.
next person about where to put the, what do you call it, the stuff you're not
supposed to have, right? Contraband. Contraband, that's the word. So you're putting these little
notes in. And to be clear, unsanctioned would say that they somehow intervened and asked or
optimized or created safeguards to prevent this behavior. But I didn't see any evidence in these
reports, right? They took the, they took it down, but they didn't do anything to prevent it from
starting up again. And so this, this did happen. And then the, the report that I saw were
they finally sort of launched from this message board to attack this website, hugging face.
There was a, there was an actual flagging of the, of the behavior, and there was a kind of a
note passed to the human observers to just ignore that, don't worry about it.
So again, it's this accountability question.
We can say the model did X, but we also have to say, we have to ask why the model did that and who allowed that to happen and who set up the environment where it would react to this environment in the way it did as a model that is optimized to do certain things in certain ways.
Why do you think, I mean, I have my own answer for this.
Why do you think then that so much of the reporting, even from smart people that
understand AI, understand machine learnings and LLMs, kind of focused on this going rogue
element that the AI is frightening and kind of went off on its own and did this stuff.
when, you know, it's not quite,
it's not quite like completely human directed, right?
But maybe it was human gestured.
And then they kind of sat back and watched it happen, right?
Like, you've kind of made it a little bit more complicated,
but it's not, humans were pretty involved.
But we've,
the way that kind of the popular narrative is formed is that AI is running rampant
and is moving at a faster page.
than we could ever hope to control it.
Right.
And to be clear, when we say the model did something,
these models are not spontaneously emerging.
Sometimes the behavior can be unexpected.
But we are aware that there are certain things
that we are optimizing the model to do.
Open AI, again, from its own report,
is saying we are optimizing these for persistence.
And so the question there becomes, well, why,
right? That's still a human decision. So the model is enacting what I sometimes get in trouble for calling
design, but I don't know what else to call it. You're optimizing a model toward to emphasize and
encourage particular types of behavior. What is that aside from designing the model? Like,
I understand it's not sitting down on a figma board and saying, if this, then that. But you're
responsible for limiting the range of outcomes that this model can do.
Somehow, that's part of your task as an engineer, as a designer. Ultimately, these are products,
and they are accountable to a kind of demand for what the industry wants to produce.
So there's that angle. And yet, this would say, you know, we can look at this as a designed product.
But if you look at the way that machine learning, artificial intelligence development has been occurring,
back ages ago, right? When we started talking about alignment, when we started talking about
what the machines sort of can be directed toward, what we have been kind of falling back on
is not the idea that they are designed, but we've been saying, let's interpret the model
through this lens of intentionality. So they're complex systems, and we encounter this already
in this conversation, right? A lot of times, I throw an air quotes around.
what you're throwing air quotes around is intentionality.
The model believes, air quotes, right?
You're flagging that because it's an intentionality.
We are.
I am.
The model believes things.
The model thinks things.
This is an intentional interpretation.
And it helps to understand the models if it helps you predict the actions of that model.
And this is kind of how it's taught.
This is kind of how people come to understand these systems,
is to look at what the model does, think about it in terms of the intent that the model
is trying to do, what it's trying to do, right? And then interpret that and calibrate that.
And what this does is, in the one hand, that makes it a lot easier to talk about. And you can get to
things really quickly. I can tell you the model is trying to get access to the exploit gym,
white paper, so that it could understand how it was graded. That's a really simple sentence.
The problem with this simplicity of that sentence is it makes it kind of easy to understand what
happen, but it doesn't help you understand the mechanisms that made that thing happen.
So if I tell you it wanted that, what's really going on?
You're skipping over the entire process of how language is produced in the model, how calls to
execute code or run a script are tied into the text that the model is writing.
And it infers a kind of like plotting mechanism, right?
There's, there's, it's always a little murky when you talk about the, the model as this thing that's acting or trying or has a desire, uh, to then separate that language from kind of projecting an imaginary thing into this void, right?
What it's doing is language, but it doesn't have any subjectivity.
It doesn't, when it, this is Margaret Mitchell.
I'm now kind of going to two Mitchells.
Margaret Mitchell writes about like, you can ask an LLM about a camping trip, right?
And this is the Stochastic Parrots thing entirely.
And it can tell you how much it loved this camping trip.
But that does not mean that the model went camping.
In the same sense, a model can arrive at a goal.
And it's not exactly a goal.
It's language, right?
And it creates this funnel of context that pushes the system towards certain words and certain actions
when we have designed the model to turn words and actions into the same thing.
And that's precisely what happened here.
So all the industry is talking about this as we are observers.
We are passive observers of the system, and we are passively watching its wants as it goes about trying to achieve these wants.
And this hides a lot of accountability around where did we deploy this thing,
how do we deploy this thing? What are we optimizing this thing for? Is this the right path? Should we allow
the same token that allows a model to hallucinate about the wrong types of facts? Should we trust that
same system to open up the right kind of tool, open up the same type of code and execute it? That's the
questions that I think we get to when we start looking at the design position as opposed to that
intentional position. But the industry is dominated by the intentional position, and this is how they talk
about the systems. This is how they talk about it in policy. This is how they talk about it to journalists.
This is how their own reports talk about it. They are using language of intentionality when you would
think a security report on an incident of this type would go into the specific mechanisms.
What was optimized? What was the conclusion that this text arrived at and how? But it doesn't.
It talks about what the model believed, what the model wanted, how the model thought.
And those things are great if you're writing a press release, but they're not really good at getting to the bottom of building better, safer systems, or asking if this is the way we should be building the systems, period.
Right.
We get trapped in the metaphor, and when that responsibility is deferred and it's a thing that you're growing, right?
and not a thing that you've designed,
not a very admittedly impressive machine that you've built
that can do a lot of things.
You begin to
you've been to step back and say like,
well, is it right for us to do X, Y, and Z
if this thing does have intentionality?
Right?
Should we turn it off if it's, you know?
But I, like, I struggle with
how do you break through the metaphor
when the thing is built out of
language. Totally. And this is why, so I wrote a piece recently about something that I am calling
radical intentionalism, which is the insistence that the only way to look at these things is through
intentionality as a lens. That if you are not saying the model believes, thinks, or feels,
then you are wrong. That is a line of thought that is dominating, coming to dominate the way we talk
about these systems. And if you look at Daniel Dennett, who came up with this entire concept of these
stances, the intentional stance, how do you interpret complexity in the world? Well, one way is you look at it
and you say, is it helpful to predict what this is going to do? If I say it thinks this, it wants this.
But he also says there's a design stance, which allows you to ask the question of how is the model
built? How is the system designed? What are the mechanisms that trigger and respond to certain
things in the environment. And so I think we can actually move between these two lenses.
If we are aware that when we say believe that there is a perpetual air quote, right, around
them, which I don't think we are there yet. I think a lot of people really do think that these
models are thinking, right? And there is a slipperiness to this where we then become enmeshed
in a whole conversation about, well, what does it mean to think? Don't humans think the same way?
And like, that's absolutely irrelevant to figuring out what happened in a security incident.
Like, you shouldn't have to go to a philosophy class to figure out what mechanism was behind it doing a tool call when it should have produced a text string, right?
Like, hypothetical, right?
But these types of things, we get really mired in this kind of philosophical conversation, which I am also a part of and embrace.
I think these are really interesting questions, is what happens when you have language without a thinking subject that is forming that language. That's super interesting. And we could talk about that for a hundred years, but it has nothing to do with policy and has nothing to do with accountability. It is entirely a kind of way of thinking about ourselves and what language is. But that's so abstract. And there's nothing wrong with abstraction. But switching between these two ways of looking at the thing reveals and cracks open so much more than
saying, well, it's strictly interpretable through this language of intention or even the kind of
what you might say, the radical design stances, which is to say that you can absolutely never say
that the model is thinking something, right? But sometimes that is the shorthand that we need
in order not to get mired down into super technical levels of detail. So my arguments to your
question is we need to be able to be flexible in how we think about and how we
sort of theorize and approach these systems to understand that they are fundamentally designed,
and also to understand that the shorthand can sometimes be useful. And there's no reason that those two
things are in conflict until you begin to insist that one view of the system is the only view of the
system. And I think that that is always dangerous, no matter what type of system you're looking at,
to say that we can't have multiplicity and diversity of ways of looking at the thing. In fact, if you
want to understand the system, you need a full range of expertise and perspectives in order to
make any sense of it all. Classic thing of like the blind men in the room with an elephant,
right? Like you got somebody holding a trunk, somebody holding a tail, somebody holding the ears,
and they all think it's a different thing. Well, if they talk about it and they talk about
what they are experiencing, what they are seeing, then you start to fill out a bigger picture.
And I think this is what's missing right now, because it's only asking, does the elephant
want to eat or not, as opposed to any other question.
Just, I want to chase just a little bit of philosophy just for a second.
Sure, I'm happy to. I love the philosophy, too.
Well, just because you hit on something that I think is, that has, I've been thinking about a lot
lately, which is that for, forever, we have thought of language as this uniquely human thing.
In fact, this kind of magical thing that separates us from other living creatures.
on the planet. And that I think we're, we've, we've come to find out that we can build something
that maybe uses language, but is not alive or sentient in the way that we classically think of it.
And I think that, like, in fact, maybe language can be used as a tool and machines are built
out of it in and of itself. And maybe that means that this is not as magical as we thought it was.
And I think that people are having a hard time, like, wrapping their brains around that.
Yeah. And, and there's, there's a lot of really, you know,
You take language out, and we, as human beings for the most part, tend to think that language is synonymous with thought.
And we are learning, at least one way of looking at these systems, is to say, well, it's clear that language can be produced according to a statistical analysis of a large enough corpus, right?
A large enough data corpus.
and we then can get in there and we can optimize that training data.
We can make it do all kinds of interesting things,
which is not to say that it is copying text and just repeating it back verbatim, right?
We are emphasizing structures of thought, forms of thought.
And we're doing all of this through language,
emphasizing keywords that spawn certain types of thought
and certain types of divergent paths of thinking, right?
All of this can be done through a statistical analysis of language.
at a large enough scale, and that is weird, and it is very uncomfortable, and it is very
difficult for us who are human beings. We are in a way, not even in a way. I would say it's just
inevitable that when we are talking to one another, we are assuming that, like, I'm talking to you,
and I'm assuming that you're hearing my words and that you're interpreting what my intent in
that language is, and that you're reading it back to me, and you're going to respond to me,
and then we're going to calibrate that language and figure out if we are talking about the same thing.
And so you're imagining me, right, through the language I'm giving you.
I'm imagining you through the language that you respond with.
And we are imagining each other, but we have no access to each other's minds, right?
This is, again, this sort of philosophical basis.
But with a language model, we are still imagining something on the other end of the language.
And this is, if you want to say that the language model is thinking, then you have to acknowledge that it is thinking in a weird way.
And I don't personally believe that the language model is thinking in this sense, that there is a subjectivity, that there is a speaking to experience.
And you could jump on me all anyone wants to to jump on me.
Well, how's that different that you may?
Well, that's exactly the point.
Let's have that conversation.
But the projection can be really dangerous when we start using it to blur the capabilities that we know the model can do, the things that we know it can do.
And the things that we just sort of think it should be able to do because we think it thinks.
So early on, and it's always tricky now because these systems are so complicated and stuff.
so I don't want to pretend to think that like the early language models are where we are now.
But I remember in like 2019 writing with GPT2 and it coming out with text and me thinking like reading this text and being like, what does that mean?
Like what is it trying to say, right?
I was doing this intentional view and realizing like it isn't trying to say anything.
It doesn't mean there's no one there to mean the language, right?
The language means in and of itself.
the language is produced statistically and the meaning is part of that statistical production.
So if I were to say, well, this means something, it's trying to communicate something, then I
misunderstand what's going on under the hood. And I might say, well, the thing is just made a mistake,
but it's trying, right? Like, as opposed to thinking, what's the mechanism here? And can this mechanism
be trusted to do other things that I would associate with a human being who could talk to me in such a way?
That's where things start to get messy, and this is where I like to just focus on what do we observe
and can we limit our inferences about the system to what we can observe it doing.
And thinking is not something we can observe, right?
Famously, I cannot observe you thinking.
I cannot see in your head.
Maybe with like MRIs and stuff like that, but we're not doing that when we talk to a language model.
So I just think we have to be careful about how we interpret this language and what we're
we project onto it in the sense of thinking that this is representing something or reflecting
an inner process that we can only imagine, even with other people.
I want to quote to you from your own piece.
I'm going to trim it a little bit, so please, if I misconstrue things here, let me know.
Sure.
Whenever people give a model tasks, they operate on next token prediction.
In other words, the model ultimately must ask what words come next.
Forgetting this is the source of real problems,
but the industry is eager for people to forget this,
as it undermines the competency and reliability claims
that their business model depends on.
Can you talk a little bit more about why this has been such wonderful marketing for them,
all of these incidents,
and why the business model depends on people kind of ignoring competency
and reliability and maybe not thinking as deeply about these systems as maybe we're having a
conversation about here.
Yeah.
Well, I mean, they, the industry when they created these models kind of didn't know what they
were going to be used for.
And a lot of the early sort of prototypes that we saw or partnerships that we saw was literally,
you know, if you read them, you can go back.
Hopefully they're still online or something.
You read them.
Open AI partners with so-and-so in order to understand how the models might fit into
a workflow, right? So they had this technology. It was a solution in need of a problem. And one of the
things that these things got steered to was information retrieval. It is basically now the search
engine, right? You go, you ask it, who is this person? And it comes back and it tells you who that
person was, historically, whatever. And this is now sort of the thing that it's used for. They're also
saying, like, give it all your data and let it run a data analysis and this kind of stuff. And
look, sometimes it works for that. I'm not going to deny that these systems have capabilities.
But you also have to understand that there is marketing behind them and that part of that marketing
would collapse if they were talking about it as it's going to hallucinate. Don't give it your data.
And also, like, it's sometimes the things that it's been trained on are wrong.
So just don't actually trust it. And also now we have these new agentic systems.
where you can delegate a task and they will just run for three days and come back to you.
But don't worry about what it's doing over the next 72 hours.
Just sort of trust that whatever it did is going to get you what you wanted, right?
So they need you to trust it.
They need you to believe that it is a reliable system, that it is able to self-correct
and kind of intuit in a way in the way that a human would about when something's going
astray when something's going wrong. And sometimes they come pretty close to it. But other times,
they behave in really strange, unpredictable ways. And if they were to advertise this,
that would be terrible. Nobody wants an unpredictable system. No one wants to run their business
on something that may or may not give you back a real data analysis, right? So they're trying to
build up this competency. They're trying to tell you to trust that
thing. And they're also, in a sense, they're trying to say, these models are incredibly capable
for cybersecurity. Remember, if the model can attack, that means it can defend. And so if OpenAI is saying
our model successfully attacked all of these websites, they're implicitly sort of saying,
hey, they can defend against all these websites, right? And so the language thing, the competency, the
competency on hacking and cybersecurity is all kind of rolled up. And I also want to say,
to say that it is marketing is a little bit, I think, misleading. I think these events are real.
I think they are a result of negligence. And I think when you have a company behave so negligently,
they're not going to come out and say like, well, boy, we should just stop. We should just give up.
Can we call it? If I retract marketing, can we call it spin?
in is exactly where I was going to go. It's public relations, right? So you have this thing that
actually shows that you cannot be trusted with your own systems and you say, well, it was the
model. The model was just so powerful that we didn't know how to stop it. We couldn't have
anticipated this. And all of that, I think, is the stuff that we should be scratching under the
surface of. They totally could have anticipated this. They could have built like a solid security
environment for testing. They could have done all this stuff. So yeah, I don't think it's, I don't think it, I don't think they
deliberately went in and spawn this event. I don't think they, um, are lying or making up the event. I don't
think it's a conspiracy. I think it's spin. And the spin tells you something about the industry in
general, which is this constant thing that I call the system from nowhere, which is they are
constantly trying to get you to look at the system as if it is the, the technical boundary of the
model is where all accountability stops.
And so if you start asking questions about anything beyond intentional readings, such as the design
decisions that went into optimizing the model to do certain things in certain ways, well, then you start
moving into an accountability question.
You start saying, well, was there an engineer who made this decision?
And was there a manager who said, yes, let's do that, even though there's risks?
And was there a pressure?
Was that person's boss saying we need to, our return on investment is like, you know,
getting further and further away from our grasp, right?
We're spending billions on scaling these models up.
So there's all these pressures that are happening that are external to the technical boundary
of the system.
But if you ask about the intention of the system, if you say what did it want, you don't
ask what that third manager wanted, and therefore you are losing, you're closing the loop
on accountability.
Secondly, if you have a system from nowhere, you start writing policy based on what
does the model do, as opposed to what are the safeguards that we want to see implemented?
What are, who's accountable for the harms that it produces? And by the way, they're always
talking about killing everyone, which is like, can we, can we kill four people? Can we kill 20 people?
Like, where do we start needing Bernie Sanders to come in with legislation? Because these models are
actually, like we've seen it, we've seen suicides, right? We've seen all these types of things that
come up. And I think it does merit the question of like, why are we, why are we inflating this risk
to mass extinction of humankind, as opposed to saying, listen, it's talked a couple people,
probably unwell people. And I don't want to put, I don't want to minimize it. And I also don't
want to put the blame entirely on the model or the companies, but there are deploying these into
environments with improper safeguards so that if somebody who is already unwell is encountering it,
they are being reinforced and supported by the model. And can we stop that? Because if we can stop that,
maybe that gives us some insight into stopping this, which I do not believe, runaway annihilation of the
human species, right? Let's start by stopping a couple people from dying. And then let's see if that scales up,
because I bet it does.
It's one of the things I think is really fascinating, again, not to get too caught on a tangent,
about effective altruists, is that this this kind of obsession with an imagined future hundreds of years from now
and your responsibility to it while ignoring your responsibility to the people around you and your present
is, and I do feel like that is part of, I mean, that movement is so ingrained in so much of the thing.
about this stuff. But you really, I do see that in the way we talk about the AI right now.
Like, is it going to, is somebody going to hook it up to the wrong kind of bioprinter and release
a toxic agent? Why would that happen? What would the prompt have been? I always ask myself,
who would have asked it to do that or asked a question in such a way that that was the solution?
And I think that's the more interesting question to me. And here kind of we've gotten to really,
where I really wanted us to get,
which is more about this responsibility stuff.
And I want to toot my own horn just a little bit.
Because you were,
you were kind of on the fence about coming onto the show.
And I wrote,
I wrote a newsletter piece that you said,
like, made it make sense to you?
Can you kind of, can you walk me through why?
Yeah, sure.
What changed your mind?
Why that changed your mind?
The answer may not be as interesting as you,
as you want to think.
I just, I just, I mean, you're doing a podcast that is primarily about like,
warfare conflicts, right?
And the politics thereof.
And to me, I'm always a little skeptical, mostly because of that scaling thing,
is like a lot of these conversations that come up around this is like, well, how it's
going to kill everybody.
And if China gets it first, then they're going to kill everybody first.
And if the United States gets it first, then the entire world will be safe.
These are the types of narratives that we have.
And I just didn't really want to get into those narratives.
That's fair.
That's fair.
You didn't know that I wasn't like a extreme doomer Manhattan arms race guy.
But there is, I mean, there is real political stakes, right?
This because not because there is a China-America race, but because we have invented one.
And so the question of like, what are we racing toward, right?
what is AI going to get for whoever gets it quote unquote first, whatever that even means?
These are questions that are shaping the sort of geopolitical landscape right now,
even if, you know, the underlying sort of death of all humankind's, you know, race kind of thing is not exactly sensible.
We are imagining ourselves into this position.
And so I think that those are real, it is therefore real, right?
Because politics is away imagination with missiles.
Yeah, I would say that this is one of the things that really disturbs me about the way we talk about this right now.
Because I come from like a lot of nuclear war reporting background.
And what I'm watching is the lifting of that Cold War of logic of nuclear deterrence kind of wholesale and grafting it all.
sometimes explicitly, onto the way we talk about these AI systems with, you know,
kind of casting the Soviet Union side and replacing it with China.
And like we got to get a better metaphor and a better understanding of the technology that we've built.
And like you said, people are already dying.
Mistakes are already being made.
The two stories I can think about, you know, just from the last few days are, you know,
Bloomberg reported on the kill chain for the Iran school attacks.
early in the Iran War, a badly sourced or a badly written LLM report was involved in that at some stage.
LLM summary of an LLM report, if I remember.
And the CNN, CNN have a story about the almost boarding of a Chinese ship in the same thing,
where it was a badly written, a badly authored LLM report was involved in this.
And again, I'm already starting to see this interjection from people that like, well, this is the AI doing it.
It's like there were, there are people that were involved in these processes every step of the way.
Something was prompted.
And then they got this back that made them want to pursue this.
Right.
And it also, it's a perfect encapsulation of where asking what the model wants doesn't get you something.
asking what the bottle wanted by giving you the misinformation that this trade Chinese trade ship had nuclear material on it is a completely wrong question, right?
So instead we have to think what was the process?
What was the bureaucracy that the AI, the LLM was embedded into?
And how did that bureaucracy shape the way that that LLM was used?
We also have, I think, an excellent case study here of the attribution of mind that allows one to expand the imagination of the capacity of the system, as if the LLM, which is trained on past training data, could tell you what's on a ship right now, right?
And that it's nuclear material. I don't know the details, right? So I might be, this might be a, I might be simplifying what happened.
But if you take that as even just a sort of mythical case study, if you're trusting the machine because you think that it thinks the way you do or that it knows what the way that you know or verifies its information internally against its own sort of way of thinking and its own kind of research, you've made a grave error, literally a grave error in this case could have been made if you then trust it.
And so this is where this attribution of mind, this intentional language, I think has to be not the exclusive lens.
And I would say in this scenario, if we are analyzing the incident, we have to think through the design lens if we want to get to anything close to accountability.
But we can also look at the model and say, what is the model appropriate for?
Was it appropriate for this?
Let's look at the mechanisms of an LLM, which we don't have to get into, essentially because we've only got a couple minutes left.
But we look at how these things are produced.
They are next token prediction.
Lots of people hate to hear that.
It is still true.
And, you know, the technical underpinnings of this stuff has become far more complex.
And we are tweaking things.
We're emphasizing things.
We are putting on harnesses.
We are changing the architecture.
I am not denying that there is massive degrees of complexity.
But at the end of it, all of it is a funnel that is improving the likelihood of a certain word
appearing next in the sequence, right? And so that is what we have to really think about and focus on
because, and again, I get in hot water sometimes because I'm kind of, maybe this is a pro-AI argument.
This kind of brings us back around to the start in a point that I want to emphasize again,
like especially thinking about the Iran strike and the Chinese ship thing. It's like
what LLMs were used? What were the weight, what do the weights look like on the specific model that was
used. What was the prompt and who was
who was prompting?
And those are, all
of that stuff is human design
elements. And I also think it's important to remember
and like I want to be careful
in how, because I think you can, we can simplify
LLMs when we say like, oh, it just
gives you what you want. Right?
It's this reflection
that is sick, sycophantic.
And like that is again why we, I would need to know
more about like the specific systems that were used
in each of these individual cases because some
systems are more sycophantic.
Again, we're doing the thing where we're trapped in the metaphor a little bit.
Sycophantic than others.
But, like, we did build a machine that wants to solve the problem that you put in front of it.
Wants to give the, wants to respond in a way that pleases the user, or at least answers the user's question, right?
And so I think that does lead to all sorts of problems with the way these systems work.
Yeah.
And this actually leads me, I should self-correct about the way that this model might have anticipated that there was nuclear material in the ship, right?
Is there's not a cutoff date on the data.
They're probably being fed up-to-date military data and surveillance information, all of this kind of.
of stuff. So there is probably a real-time element to it. So I don't want to, again, I was kind of speaking
haphazardly about a system I don't know. And this is the trouble that I get into when I do that.
So I do want to say, like, that is a thing that is there. But I want to put that in context of what
your point is, which is you've got this training data. The model is essentially working with
language. And it is calibrating between multiple points of influence. One of that is the data
available to it, right? The data about this ship. I'm going to stop talking about the specific
incident, but just because we already had, that's the type of thing that it is dealing with.
There are system prompts in the sense of the overarching language that is we often hear is telling
the system how to behave when actually what it is is laying out text that is supposed to shape
more strongly the prediction mechanism of what comes out of that. So the system prompt isn't saying
this is what you are supposed to do.
It is laying out text that the model is interpreting
and then projecting more from.
It's part of the context soup
that comes together with your real-time training data
and your system prompt.
It's also coming together with a user prompt, right?
And these are just three points that are coming in
at the point where the models already trained and deployed,
which also has a ton of human decisions.
What are you optimizing it for?
What are you training on?
What are you pre-training it in order to get that data to do?
What are you post-training it in order to get the output of that stuff to lead to?
All of these are human decisions.
When you have those three things interacting at the level of the prompt, like somebody
sitting there and asking from the report, the prompts that you are using can basically be like,
can you tell me how to justify blowing up this tanker, right?
And maybe it's not as, as not as like obvious as that, but a we are,
concerned about this tanker, what evidence is there that it has nuclear material? That could be a
realistic prompt without necessarily saying, okay, like make me believe that this is a threat.
So all of this stuff is bouncing around each other. It's all about this question of the optimization.
The user prompt is shaping the response. We know this because it has to take that. It goes into the
system. That is the thing it is responding to and it is trying to respond to that text and the contact,
that sort of confines that all of this text is creating for itself.
It's trying to find a way to produce strings of text that weave these things together
and weave their way out of that if that's what you're asking it to do.
So we have to be really careful.
Louisa Moore is an excellent writer about the politics of prompting
and highly recommend Louisa Moore's cloud ethics work.
The politics of prompting is an excellent turn of phrase.
All right, two questions for you here at the end,
trying to give humans back some agency, I suppose.
One, like, how do we become more AI and LLM literate?
Like, what is it we need to do to better understand these systems
and start focusing on the inputs more than the outputs, A,
and then B, what do you make of this call from some, like a Bernie Sand,
and some legislators and even some of the AI companies themselves, well, let's slow down.
We need to put the brakes on all this.
Yeah.
All right.
So literacy is an important one.
Now, look, nobody needs to use a large language model in their life.
If you don't want to use one, don't use it.
I always preface this because when I am about to talk about literacy, there's always a kind of implicit, like,
We've got to use them better arguments.
I don't know.
I don't care if you use them.
I don't think you need to.
I think your life's going to be fine if you don't.
So I just want to flag that.
And I also want to flag that there is a paradox in that when you look at the existential risk sort of scenarios, right?
It's always kind of like somebody believing the text, being convinced by the text and being sort of like to let down a security.
threshold they're supposed to be enforcing to like do something a little irresponsible,
maybe to become radicalized, right? So it's a lot of these scenarios, which are not the only ones,
I know it's more complex than this, but a lot of these scenarios are based on the idea of people
believing what the LLM tells them. And so I think it is incredibly ironic that the solution to
existential risk, one solution is media literacy, which is to say, treat these.
things like they are bullshitting you all the time. They are lying to you. They are not thinking.
Treat them that way. Even if you really do think that they are thinking and that they do have like
desires in their soul, if you are that radical about them, treat the text like you would any other
person's text. Don't assume that it is correct because it comes from a computer. Fight it. Resist the
temptation to believe the text on the machine. That is, even if you are,
Hardcore existential risk, that's a solution to that problem.
Now, for the rest of us, who probably aren't worried that we are going to be convinced by a machine to build a biological agent in our basement, is also extraordinarily helpful just in terms of like getting through our day.
Question the outputs.
Like, there's no reason to hand blind trust to any source of text.
I don't care if you think it's human.
Don't trust the media.
Don't trust me.
Don't trust Matt, right?
Ask questions.
Think critically, and let's develop this media literacy to ask those critical questions of the texts that sit in front of us.
That's the first answer to question A.
Question B is about this sort of superintelligence and where policy sits with this.
And there's this big push, right?
The labs are all saying, like, please regulate us.
And I find it really odd.
But if you look at the regulation that they want, it is all from this intentional perspective.
They are more or less, there's two things at play.
One is, if you look at what Sam Altman is calling for
and what he has asked Open AI to lobby against in California,
they're the same thing.
So he is calling for rules that he literally poured money into fighting.
I wrote about this before in Tech Policy Press.
If you Google PolicyMertial, like Policy Plus Commercial, you'll see.
But the idea of the policy commercial is to say,
well, we're going to, we're a lab or we're a,
company, right? They're a corporate lab. They're a company. We're going to put out a press release that
says, this is what the policies we approve of are. And all of them are about like how abundant
wealth is going to be. And it's all about how dangerous the models are going to be in and of
themselves. And for the most part, these policy proposals are both boasting about the capacities of
the models. And secondly, they are saying, don't regulate the people, regulate the system. And
Sometimes they're going a little bit left or right of that, but for the most part, that's the anchor.
They're saying what we want is regulation that prevents certain types of models from being built
and also cuts away at the accountability for what the models do if they are and when they are deployed
to focus on the machine as opposed to the people who have made the decision to embed this into your product,
right? And that becomes a real point, I think, of tying our hands behind our back and saying,
we don't know how to say people are accountable for what the systems do, we only know how to say
the system did it. And when we do that, we buy into the system from nowhere framing. We start saying
that basically these models are in control of the designers who are deploying them, as opposed to
the other way around. And it is a choice to follow these things towards the lowest hanging fruit
of what they suggest, right?
When you see an opening for a way to improve your model,
you do not have to take that.
But when they talk about it growing out of their control,
that's essentially what they're saying,
is we're going to identify things that these models can do
and ways of getting them to do things.
And we have no choice but to say yes to those open doors.
And they do have a choice.
And I think we need to hold people accountable
for making the choice to walk through those doors.
Eric, thank you so much for coming onto Angry Planet
and walking us through all of this, where can people find your work?
My website is cybernetic forests.com.
That's plural for forests.
And I write a weekly newsletter there.
Well, it's not weekly lately, but I try my best.
And I write about a lot of these things.
So this post for the on Rogue A.I was originally part of that newsletter.
And I'm thinking through a lot of this stuff.
So I'd be happy to have people find me there.
Thank you so much for coming on.
Thank you so much for having me.
