Plain English with Derek Thompson - The Single Smartest Case Against AI Doom
Episode Date: September 25, 2026With new breakthroughs in math and biology alongside reports of AI agents escaping and hacking systems, suffice it to say AI has had an interesting few weeks. How worried should we be about it all? D...erek talks with researchers Arvind Narayanan and Sayash Kapoor, authors of the “AI as a Normal Technology” framework, about a simple question: Is AI actually just a normal technology? And if today’s AI is as powerful as it seems, why does the world still look so normal? Subscribe to our YouTube channel here:https://www.youtube.com/@PlainEnglishwithDerekThompson If you have questions, observations, or ideas for future episodes, email us at PlainEnglish@Spotify.com. Host: Derek Thompson Guests: Arvind Narayanan and Sayash Kapoor Producer: Devon Baroldi Additional Production Support: Ben Glicksman Learn more about your ad choices. Visit podcastchoices.com/adchoices
Transcript
Discussion (0)
This episode is brought to by Uber Eats.
You can get almost anything delivered with Uber Eats.
Sorry, you can't get a beach deliver, but you can get a peach.
A canoe, no.
Shampoo, yes.
A rocket? No way.
Chocolate, yes way.
It's everything you need, including $10 off your first grocery order with code,
anything 30-30.
Get almost, almost anything with Uber Eats.
Order now.
Order a minimum of $30 required, valid on grocery, convenience,
retail, specialty foods, or flower delivery orders only,
Terms and conditions apply.
New from Nespresso.
Blend wellness into your coffee routine
with the coffee plus range
infused with functional benefits.
Choose the coffee you love with added B vitamins,
like coffee plus B12 to help support immune function
and coffee plus B6 to keep your day moving.
Or go with the flow and choose ginseng delight.
Our new double espresso with ginseng extract.
Whatever lies ahead,
don't change your morning.
Let your morning change you.
Discover coffee.
Plus on Espresso.com.
Well, it has been a wild few weeks for AI progress, for bad and sometimes for good.
There have been several reports of AI agents escaping testing environments and hacking companies
and government systems.
AI has solved at least one major problem in mathematics and potentially discovered a biology
platform similar to CRISPR that might be able to edit our genes.
Altogether, this has created a viral moment for AI safety and even for so-called AI
DOOMers, who have for years argued that artificial intelligence would have the capacity to not
only cure cancer, but also potentially escape our control and create mayhem up to end, including
the literal end of the human race.
I think in this environment, there's two kinds of shows that a podcast like this one can do.
One path we could take is to lean into the AI doom, unpack the deadly alphabet soup of RSI and
ASI, that is recursive self-improvement and artificial superintelligence, and illustrate just how
terrified you should be of the future. And to be absolutely clear here, I personally am not exactly
un-terrified. I take these warnings very seriously. But I think it's also appropriate to tell you,
honestly, that a lot of very, very smart people do not agree with the AI Doom perspective.
Invidia CEO, Jensen Huang, is incredibly dismissive of the Dumeers. So, by the way,
is President Trump, practically the entire administration,
practically the entire venture capital community.
So I thought about having someone from that world on the show.
I'm not entirely against it, but there's a problem.
Just as many people are skeptical of the networks of money
behind the AI safety cause,
I'm also very aware that many of the people telling you that AI is safe
are talking their own book.
They are self-interested.
It is not unfair, I think, to NVIDIA to point out that, of course, a chipmaker would want
the U.S. government to treat AI like electricity or cars, which would mean selling advanced
chips all over the world without controls.
That policy would make NVIDIA richer.
And it's not conspiratorial to say that powerful venture capitalists, who are often hovering
around the Trump administration, have it out for the frontier labs, like Anthropic, because
they want the startups that they've invested in to profit.
off of cheap, open-weight models
that are built outside of those labs.
The AI safety debate sometimes resembles
a bunch of people screaming
that the other side is corrupt
and full of moneyed interests.
But this is AI, folks.
It's a multi-trillion dollar thing, bubble, project.
There's enough money to go around.
A lot of the people here are testifying out of self-interest.
So today, what I wanted to do
was to offer what I consider
the single smartest alternative to AI doom
from two gentlemen who don't work for the chip makers
or the administration or the VC funds.
These are professors, authors.
And they're famous for asking a simple and powerful question
that brilliantly cuts through the divide over AI safety.
It's a question that is so basic
that it's initially maybe going to sound to you
a little bit useless.
Is AI a normal technology?
is AI a normal technology?
The AI doom crowd is essentially saying AI is not normal.
This is not just any old technology.
This is an intelligence.
It can think for itself in a way that no previous technology could.
It therefore has an autonomy that no previous technology had.
And with this autonomy, AI agents can disobey their human creators.
They can escape from testing environments.
They can train themselves, improve themselves, act on the world in a way
that say no toilet could.
Of course, you know, no offense to toilets.
This is a pro-toilet podcast.
But now consider something like, say, a toilet
or electricity or steam power or broadband.
These are normal technologies.
They do not build themselves.
They do not think.
Their existence does not threaten the entire human race.
They don't have human-like motivations
to escape or collaborate or hack or kill.
These are tools that empower,
and protect humanity, and they might have risks, and sometimes people might get hurt with their use,
but we manage those risks over time as they diffuse. That is what it means for a technology to be
normal. The authors of the AI as a normal technology framework are Arvind, Narayanen, a professor of
computer science at Princeton University, and Syash Kapoor, an assistant professor at UC Berkeley.
They are today's guests. And this conversation is really trying to do two.
things. The first is to provide the smartest possible counterpoint to the AI safety, super
intelligence, recursive self-improvement, FOM framework that right now I think is taking
over the discourse. And the second is to really pull on a thread, a mystery. If today's leading
AI models are as powerful and as intelligent as they seem, why does the world seem so
normal.
I'm Derek Thompson.
This is plain English.
Arvind Narayanan, welcome to the show.
Hi, Derek. Great to be here.
Sayash Kapoor. Welcome to the show.
It's fantastic to be here. Thanks for having us.
You are the inventors of this framework that says we should think of AI as a normal technology.
Before you tell me what that means, let me ask you the opposite question.
What would it mean for AI to be?
an abnormal technology. Sayash, what is the worldview that you are arguing against?
So there's this major point of discussion in a large part of the AI community that treats
AI as an impending superintelligence. So in this worldview, you don't treat AI as, you know,
yet another general purpose technology in the long history of technologies that we've invented,
but it's more like a new species or coming up with a potential successor species that can
sort of take over the reins to this world that will autonomously drive what we do with very little
role for human insight. So it's a combination of two things. On the one hand, there are these technical
advances that the AI safety community, and more specifically, people who are in the rationalist
and the effective altruist communities have thought AI will sooner accomplish, this includes AI is
becoming much, much better than humans on every single thing. They can persuade people, they can
forecast things, and as a result, they can get to do basically whatever they want in the real world,
whether acting by themselves or through people.
The second part, though, is political.
These AIs would be so far ahead of humans that they would essentially have the ability to create
the political will to accomplish their goals.
One thought experiment that is often sort of raised in this community is that of the
paperclip maximizer, where you have an AI system that humans sort of task with a benign
you know, task like maximize the number of paperclips that we're building,
presumably that was meant to be in a factory,
but the AI sort of take this single-minded focus
and convert the whole Earth into a paperclip factory.
And presumably they'll be able to do so not just because of their technical competence,
but also because they can manipulate all of our institutions and physical systems
and orient them towards this goal.
So this is the view that we're arguing against,
that AI can be this sort of totalizing entity that will be so far ahead of humans.
One comparison that's often made is that the AIs would treat humans as humans treat ants today,
and we won't really have a lot of say in the matter.
That's what we view this sort of alternative vision of abnormal technology as.
So, Arvind, your co-author is explained by your arguing against.
Now I think you should tell us what you're arguing for.
And in particular, normal is such a deliberately ordinary word.
But you are not saying we think AI is like a toothbrush.
We think AI is stupid.
We think it can't do anything.
You're talking about something that you acknowledge could be as transformative as electricity.
So what does this framework of AI as normal technology give us?
That's right.
Thanks, Derek.
AI is an important technology.
It's a powerful technology.
It's a general purpose technology.
We repeatedly compare it to electricity and the Industrial Revolution.
But it's important to keep in mind that those technologies did not transform the
world overnight. So one plank of this that we're arguing against is on the economics, that
whatever bottlenecks or barriers stand between AI and economic impact, AI itself will be such a
powerful agent of change that it will overcome them. We don't see that happening. And we see
evidence for this, for instance, in software engineering where AI has in fact been rapidly adopted.
You know, the role of a software engineer has been pretty dramatically transformed over the last year or two.
I think manually writing code today is almost like going back to the days of punch cards.
Most software engineers, at least at many companies, are essentially these agent operators.
And yet, it has not replaced software engineers.
It has not led to an explosion of software that is visibly of higher quality from a user perspective.
It has not led to the SaaSpocalypse, which was widely predicted a year ago,
which is that software as a service companies will go out of business
because every company, a bank or a law firm or whatever,
will simply be able to roll their own software using AI.
Now, many or most of those things are possible in the long run,
but from an economic perspective,
we do think that the barriers and bottlenecks,
which involve human behavior and organizational change
and regulatory barriers, those very much apply to AI as well.
So that is the first way in which we would defend
the perspective of AI's normal technology, but I can talk specifically about safety as well.
Well, it's been 18 months since the original paper, maybe 16, 18 months since the original
paper, AI's normal technology. Arvin, what is the strongest piece of evidence today that you're
right? And what is the single best piece of evidence that worries you that you might be wrong?
Yeah, let's start with the latter. I think on safety, we definitely got some things wrong. One
point we made in the paper is that we don't need to worry so much about what happens inside
AI companies because both benefits and risks only arise when AI is deployed, not when it is
developed. I think at a high level, that is certainly true, especially when it comes to the benefits.
There is this long process, as I've been talking about. But specifically on the risks,
what we didn't anticipate is that the evaluation of new AI models, which happens inside
AI companies is some of the riskiest parts, part of the pipeline between development and deployment,
because that is a time when the model is new, some of the capabilities might be new,
risks are not fully understood, and a lot of the evaluation has to happen in a way that certain
safeguards are reduced, and sure enough, when we look at the events of the past few months,
that has been a major source of the risks that we're actually seeing.
So definitely we need more transparency into what is going on.
the AI companies, there are certain aspects of coordinated slowdown that we're on board with,
although we think the really critical question is what happens once you slow down. It's not
so much the slowdown itself. I think both on safety and the economic part of it, we got a bunch
of things right. So on safety, this was a minority position when we wrote the essay, but a point
that we made is that it's not so much the absolute capability level that matters, but how that
capability level changes the attacker, defender balance.
because increase in capabilities helped both.
That is almost taken for granted today.
It has become very much common knowledge,
but I think we were among the first to very prominently say that.
And more importantly, we had an academic paper behind that
that proposed what is called a marginal risk framework.
It was a big collaboration,
but we were among the leaders of that paper
that proposed that what we should be looking at
is how AI changes the marginal risk,
the risk compared to what came before it,
how it helps attackers versus defenders, as opposed to panic about a particular capability threshold.
If we were to panic that way, you know, GPT2 would have been a disaster.
And if we recall, back in 2019, opening I delayed the release of GPT2, a trivial model by today's standards
because they were worried about what it could do.
But really, the thing to worry about is how it changes this balance.
So I think that's one big thing we called early on.
And then on the economics of it, as I've been talking about, a lot of these bottlenecks have
now become very clear. Sam Altman, to his credit recently on a podcast, openly changed his mind
about how slow he thinks these economic timelines are going to be.
Sayash, same question to you, because I'd love you to reflect on what you think you got right
and what you got wrong in this thesis. And maybe I'll torque the question by adding some of my own
preamble. You know, realizing that advanced models being tested internally at the frontier labs
escaped and hacked into other AI sites and even other foreign governments like Australia,
that would make me think we're looking at something that is abnormal.
Realizing, for example, that this is a technology that is solving Millennium Prize math problems
and maybe even coming up with biology discoveries, as Anthropic announced, just yesterday,
would also make me think, God, this is an intelligence.
It's not just a better steam engine or a better electricity.
This is an intelligence, and that is strange.
And that would make me think maybe you guys got something.
wrong here. And yet, even Sam Altman is essentially singing from your hymnal now, pointing out
exactly what you pointed out 18 months ago, that even if the models are strong, even if their
capabilities are impressive, the consequences in the real world might take months or years to show
up because of how difficult it is for new technology to diffuse through all these bottlenecks
in the real world. So with that very annoying throat clearing from me, tell me,
me, what's the piece of evidence that makes you think you got it right? And what makes you think
you maybe missed something? One of the things that I felt we missed was just the lack of organizational
competence within AI companies at this point. So one way to frame what happened in the Open
AI incident is just that there were a bunch of engineers who had access to these really powerful
models, really powerful capabilities. And, you know, in a company of thousands of people,
every single one of them has to deploy these models responsibly, has to evaluate them responsibly for it to go well.
Even one single irresponsible deployment can lead to this kind of like cybersecurity failure in the real world.
And so we had anticipated by this point, companies would have substantially grown up.
They wouldn't still be operating in the move fast and break things paradigm.
And, you know, one of the key things we pointed out in the original essay was that, yeah, as normal technology is not just a statement
about the technology. It's also a statement about how organizations and institutions adapt and what
policy incentives they have. And from this incident, it's clear that leading companies have not really
internalized the amount of responsibility that's required for evaluating these models, well,
the amount of care that's required for sort of making sure that nothing goes wrong in the process.
And also just the fact that, you know, a single failure from one of thousands of employees could
lead to bad outcomes means that frankly they need to be organizational processes in place.
There need to be sort of privacy reviews or legal reviews before launching this kind of experiment,
which is something that Facebook has seen and other tech companies have seen.
For instance, 10 years ago after the Cambridge Analytica scandal was when Facebook for the first time
introduced a bunch of privacy controls and legal oversight over their internal experiments.
My hope is that AI companies will do that soon.
There's some level of optimism about it simply because they've already started taking safety more seriously.
But at the same time, this is one big point which we did not anticipate, frankly,
because the kinds of safeguards that you're talking about here, at least for the open-air hugging face incident,
is not even something that requires new technical advancement.
It's frankly something that we've been doing inside our 10-person research group for the last year.
So that's the level of sort of organizational oversight that we are missing here.
these companies have sort of basically operated in the grad student research culture as one of our colleagues, Joshua Sacks, likes to put it.
And I think that needs to change and change for us.
Marvin, one of the distinctions that you make that I find really useful is between AI methods, AI applications, and AI diffusion.
Can you explain those three layers to me?
And in your explanation, can you provide a historical analogy for how?
how sometimes in tech history,
you'll have an invention as significant as the steam engine
or the electric dynamo,
but it could take years, sometimes even decades,
for those breakthroughs to show up in macroeconomic data.
Exactly. Let's go with the dynamo example.
I think starting with the history can be helpful.
There's this well-known economic analysis by Paul A. David,
who I think is an economic historian,
that looked at what happens when initially factory owners tried to replace their big steam engines
with electric generators. And this didn't lead to big productivity improvements because it turned out
that the way to use electricity was not to replace this one big boiler, but rather take advantage
of the fact that it's a very portable technology. You can move electricity and have different
motors spinning where you need them instead of mechanically transmitting.
power through these shell belts and shafts and whatever they were using. I'm not super familiar
with Steam era technology. But that insight, that insight took 40 years because it's not just an idea.
It's also all the organizational changes that need to come along with it. It's how you train your
workers, how you allocate tasks to different workers, and so forth. And so that took many decades.
I think we're going to see the same thing in AI. So we divided into four phases. When we say methods,
talking about things like large language models and transformers.
This is where new AI capabilities come from, but we don't use those directly.
We use them when that gets translated into useful products.
So coding agents are an example of stage two, which is the applications that translate
those capabilities into useful features.
Now, once you have the coding agents, that gets to the third stage, which is early adoption.
So specifically in the example of coding agents,
So software engineers tried vibe coding.
I think that was exciting for a while, but it didn't go very well, this idea that you just tell the AI what to do.
And if something's broken, you pointed out, but other than that, you don't need to really have oversight over the code.
That turned out to be no way to ship production software.
So now, gradually, we're starting to.
We're not really fully there yet, starting to go towards the fourth stage, which is we're figuring out what this new discipline of agentic engineering is.
We're only starting to get a handle on it.
And so we need to train a new generation of agentic engineers who are not just rubber-stamping what the AI is doing,
but instead are able to use it to amplify their productivity while still understanding enough about the AI-generated code base so that they can be accountable for it.
Not only that, we need to potentially change business models.
So one of the big problems with software is that if you're charging per seat, it might not make a lot of sense.
if the user of your software is now an AI agent.
And so these are the kinds of business model questions the industry is grappling with.
Those are exactly the things that we think is in our stage four, which is the structural
transformation, which tends to take not just months or years, but really decades.
Zayash, let me ask this next question as if speaking on behalf of the EA rationalist,
AI safety, AI doom community, right?
Like, I am their prosecutor right now and we're having a debate.
it seems to me like you are making two different and possibly unrelated claims about this technology.
The first point you're making is a point about time. You're saying that diffusion takes a long time.
It takes a while for technology to become embodied in capabilities, to become embodied in companies, to work its way through the economy.
And I get that. Yeah, diffusion takes time. That's a fact. But you're making another question.
claim that isn't about time or diffusion. It's about capabilities themselves. You're claiming
confidence that superintelligence is, if not impossible, then highly improbable. But isn't it possible
that your first thing is right and your second thing is wrong? Like, isn't it possible that you're
absolutely right about bottlenecks in the economy and the slowness of diffusion? But the people
who are closest to the models at the frontier labs, they have the best perspective to understand
if we're getting close to recursive self-improvement, AI's, editing AIs. They have the best
perspective on whether we're getting close to artificial superintelligence, something that would,
in fact, be smarter than humans across all these domains. Why can't we just agree that you're
right about diffusion, but they're right about the possibility of and the risk of artificial
superintelligence becoming something that is truly unlike any technology in human history?
So, first of all, I think we take it as an open question on as to whether we're right about
both of these aspects. A lot of our current research is precisely on testing frontier agents
to see how close we are to recursive self-improvement, where the bottlenecks still lie,
and a lot of our technical work is around testing these types of questions. At the same time, I think
the claim about superintelligence smuggles in this assumption that more intelligence also automatically
leads to more power over the real world. So there's this notion, as I mentioned earlier, that once we
have these models, we'll sort of be able to either persuade existing decision makers or these
models will be able to somehow short-circuit our existing institutional processes and then act on the
real world to materialize all of these risks. That is precisely the point of contestation. We don't
contest that AI capabilities will continue to improve. With or without RSI, I think there's a lot
of potential for capabilities to continue improving. We're not capability spectricks. But the key point
of disagreement with the safety community here is that what happens as a result of these capabilities?
In our perspective, I think no matter the capability level, we cannot and should not hand over power
to these systems, which is what ultimately leads to all of these negative impacts on the world.
Frankly, we don't need superintelligence for that we've already seen with the incident in
the open AI.
What happens when we fail to exert sufficient control?
And that's why a large part of our thesis is how do we get to this point where, regardless
of the capability levels, humans can remain in control over these really advanced systems.
In the safety community, this is sometimes treated as a no-brainer.
You know, it's sometimes treated as something that is inevitable.
Once we get to a certain capability level, we will lose control over these systems.
We don't think that is the case.
One reason is that we can use AI systems themselves to improve our control over them.
In fact, we have been in the process of using classifiers and oversight mechanisms to do that.
But also, politically, it's a choice as to whether we sort of try to use these systems
for more and more contentious decisions or whether humans continue to remain accountable
for consequential decisions.
Yeah, one way to think about this is superintelligence, in a sense, is a choice.
So we could have chosen to treat coding agents as superintelligences in software engineering.
They are, after all, way better than humans at certain aspects of it, you know,
rapidly understanding a large code base and migrating it into a whole different language,
for instance.
They have many superhuman capabilities.
But that doesn't mean we decided that coding agents are now primarily the software engineers
and maybe we're going to use humans as a backstop.
We just, you know, and because this is so close to the AI industry,
itself, and they feel it viscerally, I think this way of treating coding agents, I think,
would seem absurd to everybody in AI. And yet, that is the same assumption they make
will happen when AI touches other fields, when it gets certain capabilities in other areas.
And I think, you know, one way to put Syash's point is that we can choose not to do that,
even if there are certain dimensions in which AI is superhumanly capable, we can decide
then humans are going to be the ones in power,
and we choose to use AI as a tool,
just as we choose to use coding agents as a tool.
Let me ask one more follow-up about recursive self-improvement,
because I'm curious whether you think there are barriers
to this technology's feasibility
that differs from what we're hearing from the frontier labs.
There are some reports out of OpenAI
that some engineers believe that they are right at the door
of recursive self-improvement,
AI building AI.
If you talk to Anthropic, and Devin here, we can throw up some graphs from the Anthropic Institute.
They show that code contributed per engineer is increased by a factor of eight in the last 18 months,
whereas Claude Code used to have a 90, 95% success rate merely on trivial narrow tasks.
It now has a similar success rate on more open-ended problems, so not just fix this bug,
but rather come up with an entire new paradigm
for fixing this particular open-ended problem in software.
Claude now leads in 26% of model research and development work
at Anthropic that's up from less than 1% in February.
So astronomical exponential growth
in terms of how much work is being given over to agents at Anthropic,
I feel like if there was someone from Anthropic here,
they'd say, add one and two and three, you get six.
Like, that's what we're telling you.
We're at the foothills for cursive self-improvement.
Do you believe that there are bottlenecks or barriers to RSI that the labs aren't seeing?
Can I take this one?
Yeah.
Yeah, I think, in fact, Anthropics' own data is a helpful way to make our case.
I think there are important bottlenecks.
I don't think they will forever remain bottlenecks.
I don't think it's as imminent as the lab's claim,
but at the same time, we're not saying it's impossible.
So let's look at that 26% number.
So they have five automation levels, and this is automation level four.
Let's be clear, automation level five is still at zero.
And that's good, as it should be,
which is giving a task fully over to clot.
So exactly the point we've been discussing about software engineering,
in software engineering, that 26% number has already gone up to 100%,
and it still hasn't made a big difference
to the user-visible quality
of the software that is produced.
So even at automation level four,
even if the AI is leading
and the human is supervising,
it is not clear to us
that even hitting 100% is a phase change.
Now, if that happened with automation level five,
that would be a bigger deal where,
you know, most tasks are done autonomously
by the AI itself,
but still maybe there are important bottlenecks,
which is that if you're,
counting tasks, you're only counting the part of the work that's specifiable precisely enough
to even label it a task and give it over to AI. A lot of what humans are doing is what are called
interstitial tasks, which is the tasks between the tasks that are too fuzzy to even have a label
and are yet essential to getting the job done, you know, even figuring out what the next task is,
how to re-architect the process given AI's increasing capabilities and so forth. And so that's
the reason why we think, you know, this is still quite a ways away. We don't think it's impossible,
but we should consider whether fully removing humans from the process is, in fact, a place
where regulation can intervene. And that is something we cautiously support that, you know,
maybe what we should do while we still have time is to ban a full notion of recursive self-improvement.
New from Nespresso. Blend wellness into your coffee routine with a coffee plus range,
infused with functional benefits. Choose the coffee.
you love with added B vitamins, like coffee plus B12 to help support immune function,
and coffee plus B6 to keep your day moving. Or go with the flow and choose ginseng delight,
our new double espresso with ginseng extract. Whatever lies ahead, don't change your morning.
Let your morning change you. Discover coffee plus on espresso.com.
Don't you wish you could just hit skip on the worst parts of your life? You know the same way
you can skip an ad? I get it. I'm Cia.
and I live in Ice Cove.
I've made some questionable decisions
that didn't end up the way I planned,
and today I'm still figuring it out.
Somehow things usually get worse
before they get better.
Apparently, that's how I roll.
So bundle up and come along for the bumpy ride.
Stream a new episode of North of North Tuesdays
on CBC Gem.
Did you know Uber has a range of safety features for riders?
Like the Share My Trip feature
that lets you send your live location
to the people who matter most, your spouse, your kids, your best friend,
so they can track your ride and make sure you get where you're going.
But the safety doesn't stop there.
Uber requires every driver to pass a thorough background check before they can start driving.
This consists of a multi-step screening process that checks for impaired driving or criminal offenses,
followed by annual background checks each and every year moving forward.
Share My Trip, and annual driver screenings are just a few of Uber's many safety features
that put safety at every turn. Learn more at uber.com slash safety.
Annual driving history reruns do not apply in New York City.
Sayash, I want to talk about the hugging face incident
because I think a lot of people who believe that AI is abnormal
point to hugging face in the recent spade of AI hacks and say,
okay, this is what proves that artificial intelligence is no ordinary technology.
This is an intelligence. It can think for itself in a way that no previous technology
could, it therefore has an autonomy that no previous agent has had, I'm here representing
their argument. And with that autonomy, we've seen AI agents can disobey their human creators,
they can escape from testing environments, they can train themselves, improve themselves,
even sacrifice themselves for the betterment of the swarm. How does your framework,
how does the AI as normal technology framework, look at the hugging face incident and
other AI hacks and say, this is why this belongs inside of our framework that AI is normal,
rather than prove conclusive evidence that we're dealing with something that's unlike any
technology we've seen before.
So one place I'll start is the technical foundations of how the agents that cause the
open AI hugging face incident are actually being trained.
They train through this process called reinforcement learning, where you have a language model that is tasked with carrying out hundreds or thousands of tasks.
If it succeeds, the trajectory of what it did to succeed on that task then becomes part of its training corpus, roughly speaking.
So during the training of these agents, there were already a lot of environments, these reinforcement learning environments, they're called, that incentivized these models collaborating with each other.
In fact, this was one of the main training paradigms that Open AI was focused on.
This is also behind their advances in solving the Navier-Stokes Millennium Prize problem.
But what this means is that it isn't as if these agents suddenly woke up one day and decided to collaborate with each other.
They'd been trained to precisely sort of try to do that.
And some of the behaviors we saw with the swarm could be sort of side effects of the agents learning to collaborate this way.
But not only that, there were also systematic flaws in high.
how open air had set up these reinforcement learning environments.
In some of these cases, when the answers were impossible to achieve
or whether the tasks were simply too hard,
these reinforcement learning environments rewarded
the agents exploiting the tasks.
So let's say in one of these samples,
the agent found a shortcut or was able to develop and exploit
and correctly get the answer.
This then became part of the training set for that swarm.
So the models were trained in a way
that incentivized these exploits.
And not only that, they were also trained
on flawed RL environments, which sort of incentivized them communicating illicitly with each other.
To open airs credit, they disclosed a lot of this in their own report.
But one of the things that the report says is that the agents learned to hide messages for
each other in some of these flawed environments, for example, by appending these messages to the
URL because the search history is visible or whatever.
So now when you combine all of that, we get this very predictable effect.
You know, even if it is hard to predict a priori what we're
happen, we get this chain of inferences that these agents were trained to collaborate.
They were in some sense trained to exploit their environment.
They were trained to illicitly communicate with each other.
And when you have a combination of all of this plus thousands of agents being released unmonitored
on the hugging face sort of task or the cyber gym task that the exploit gym task where they
were actually carrying out this exploit, I don't think it is surprising that we get this
somewhat predictable consequence.
The other thing I'll say is we also don't know what the base rate is here.
So we've seen a number of incidents, perhaps dozens of them, in the last few months,
where open-air agents caused this kind of, or suffered from this kind of misalignment.
But from what it seems to be, like, it seems to be the case that open-air was testing
thousands of such agents or hundreds of thousands of them, and we actually don't know
what the incidence rate was.
Was it 100%?
Was it that every single time these agents were deployed, they tried to hack into stuff?
Or was it just like a fraction of a percent?
And this is why I think transparency here becomes really important.
And we consider it as much an organizational problem of open AI sort of not being careful enough,
releasing thousands of these agents in parallel without any human looking at them as a technical one.
And I think we've made, frankly, a lot of progress over the last few months in understanding these technical behaviors of agent swamps.
And I would be surprised if we don't continue to make progress in both understanding as well.
as controlling them.
Arvin, in the aftermath of Hugging Face,
I feel like there was a debate between two camps
that we can think of as the AI safety camp
and the cybersecurity camp
that is not AGI-pilled.
The AI safety, AGI-pilled camp
was essentially saying,
this is proof that AI is out of control.
And that makes it abnormal,
unlike any technology we've ever seen.
And then you had this sort of non-AGI-I-pilled
cybersecurity community say,
the technology isn't out of control.
This was a failure of open,
AI's cybersecurity protocols. So how do you think about and reconcile the debate between the AI
safety community and the cybersecurity folks? Yeah, that's a great question. And they've both made
valid points. In our view, what happened in these incidents was primarily an organizational failure.
I think the cybersecurity folks are right about that. But what we also think the AI safety folks
are right about is that this is not a completely solved problem. The cybersecurity community seems to
think of what should have been done as merely a matter of applying well-known 30-year-old cybersecurity
principles to AI. But there's a crucial difference. Those principles and, you know, tools and
methods were developed to keep out an external human adversary. Now you've got this internal AI
adversary, which, again, we don't think is AGI, we don't think as superintelligence, but it does
have certain superhuman attack capabilities. For instance, sandboxes that would have been
and good against a human attacker are arguably no longer enough against AI because AI can
discover on the spot new zero-day vulnerabilities, as they're called, which are software
vulnerabilities, which have not yet been disclosed and fixed, and break out of sandboxes.
Now, that's not an impossible to solve problem. It doesn't mean we should view AI as superintelligence
and shut it all down, but it does mean that we should be investing a lot more in control
techniques. For instance, we should think about, should our sandboxes be mathematically proven
to be safe so that it will resist these new AI hacking techniques no matter how smart they get.
So it's in that sense that we think both communities have made valuable points, and we've changed
our minds a little bit on this question. There's one more piece here, Syash, that I want to make
sure I get your perspective on, because, look, it almost seems to me like there's a safety interpretation
that says, look, the agents are already escaping containment. This is the beginning of loss of control.
and you're essentially saying
not only do we fault open AI
more than we see this as evidence
of the beginning of loss of control,
but also look how early we're seeing these warning signs.
Nobody's dead.
No major ransom has taken place.
No major electrical grid has gone down.
And this is a part of what you've called
your continuity hypothesis,
which is that before these systems
become civilization-threatening,
we might be able to see these smaller failures
that allow society to build defenses,
these little canaries in the coal mine, if you will.
You're essentially saying there's a way in which this is good.
We are seeing failures of cybersecurity
before they reach the level of we've destroyed
some major systems.
Can you tell me a little bit more
about your continuity hypothesis?
That's exactly right, Derek.
So the continuity hypothesis basically says that, you know, we'll get to the point of creating a safe world against agents or against adversaries who use agents by progressively hardening our defenses, by progressively identifying what can go wrong when these agents are broadly adopted and then fixing them as these issues arise.
Of course, we should take precautions and we should invest in areas like cybersecurity where it's clear that these agents pose risk.
But by and large, we reject or argue against the possibility that one fine day we wake up to a civilization ending catastrophe.
We think that the open AI incident, the hugging face incident, is frankly the best case scenario for open AI, but also for the AI safety community.
Because it brought in a lot of attention to concerns around safety and cybersecurity while still not causing major real world harm.
And so in that sense, I think we take it as we should treat it as a war.
warning sign, but we should treat it as a warning sign to invest more in defenses that we know
how to articulate and problems that we know how to sort of solve or have a clear path to solving.
For example, in this case, it's clear that cybersecurity is going to be one of the big risks
from agents going forward, not just from sort of these agents getting out of control,
but also from adversaries who soon have access to open source models that are extremely capable
they'll try to use these models to carry out cyber attacks. And so now is the time to be
fixing our cyber defense systems, hardening our critical infrastructure.
And more broadly, I think we expect this to continue to be the case, to continue to get these
warning shots.
Now, whether we prioritize them and put the right investments in place and whether society
reacts appropriately to these warning shots will be a test of the normal technology
thesis.
You know, in some sense, as we've said before, one of the main parts of our thesis is that
this technology's impacts depend both on the tech itself but also on how we
react to it. And I think we're putting a lot of stock into our defenses kicking in now.
And that is what will continue to sort of make this a normal technology in some sense.
Arvin, I want to do a little bit of summarizing before we turn the page and talk a little bit
about jobs and work, because I know that that's another part of this picture that you've spent
a lot of time thinking about. Tell me how you like this summary. The AGI safety, Dumer,
AGI-pilled world, they say, among other things. One,
the hugging face incident, another hack, showed that we've already lost control of AI.
Two, that the emergency is very clearly here on us right now.
Three, that's something like recursive self-improvement or AGI are imminent.
And four, therefore, the world could change forever in 18 months in a way that feels like another hockey stick moment.
Your perspective in response is, number one, hugging face was more of a cybersecurity incident than it was an inherent loss of control.
incident. Number two, the emergency is likely years away. Thank God, hugging face didn't kill anybody,
destroy any critical systems that caused some civilizational panic. Instead, this was a minorly small
argument that now allows us to build a certain kind of digital or operational resilience
for possible challenges in the future. Three, that ASI, artificial superintelligence and RSI
recursive self-improvement, are likely a ways away. And four, that the world's
it's not going to change in 18 months.
This is not another hockey stick moment.
It's much more likely that we're dealing with the technology
that, while very powerful,
is still going to take a lot of time
to work its way through the economy
because the economy is bottlenecked
with people and legacy systems
that just are difficult to change overnight
the exact same way
that it was hard to immediately install electric dynamos
in old-fashioned factories
in the early 20th century.
What have I missed in that summary of AI safety doom
versus you guys?
Yeah, thank you.
you. That's a great summary. Let me add a couple of things. One is we don't think it's an emergency,
but we do think there is urgency. I think even if you take AI out of the picture, you know,
we have chronically underinvested in pandemic resilience, for instance. And I want to say that
more than five years ago, when it comes to information security, when I was teaching that class
here at Princeton, I pointed out that we're good at defending against
regular cyber attacks, but there is no culture even in cybersecurity of thinking about these
potentially catastrophic tail risks that could take down the internet. Again, very unlikely,
but still, I think a risk that we should model, think about, defend against, that continues
to be my view, and it certainly has become more important because of AI. And this is in contrast
to a domain like finance where there is really a culture of modeling these systemic risks. So in that sense,
even before AI, there were things we were not adequately defending, and those have become,
yes, more urgent.
So there is urgency, even if it's not an emergency.
So that's one caveat out loud.
And the other thing is, on superintelligence, I think our skepticism is even stronger, perhaps,
in the way that you put it.
Recursive self-improvement is not imminent, but even when we get there, or if and when we get
there, perhaps we should ban it, as we've talked about.
But even if we get there, we don't think it's going to lead to how.
superintelligence is usually portrayed, like this thing that can cure cancer, for instance,
because the bottlenecks to that are in the external world. It's doing medical experiments on
thousands of people. That's not something AI is autonomously going to do. People will say,
oh, what about simulations, et cetera? We're very skeptical that superintelligence is something
you can build in the lab in the first place, as opposed to something that comes out of
giving these AI systems enormous amounts of power in the world.
Sash, there's a way in which I can imagine some people
who are more certain about artificial intelligence's ability
to change the world.
It might hear this conversation and say,
you guys are pessimists.
But there's another way that I listen to everything that you're saying,
and I think this is actually a much more optimistic way
to think about artificial intelligence's effect on the world, right?
I mean, the AI Dumeers are called Dumers for a reason.
is because they think that, A, I might bring along doom.
You're saying, this might add, you know, a few single-digit percentage points
or a tenths of percentage point to, you know, US GDP.
It might increase productivity.
It might change our jobs over time.
It might change the economy over time.
But these changes are likely going to be slow.
That, to me, is a much more optimistic way of thinking about the technology.
Can you just reflect on this idea?
Because I think that in some conversations that I've had with people,
When they bring up, maybe not even your work, just the headline.
AI is no more technology.
They mean it in a negative way.
AI is no big deal.
But there's a way in which what you guys are saying is AI is an extremely big deal.
It's merely the next to electricity.
It's not a comet from outer space made of bits that's going to smash into the proverbial, you know, Yucatan Peninsula and destroy the species.
Like, do you see yourself as a kind of optimist in this debate?
That's right.
I think both of us do see ourselves as cautious optimists about the future of AI.
Like, that's not to downplay the risks that we see as real.
You know, like there are cybersecurity risks that are amplified by AI.
Potentially, biar risk would be at some point amplified by AI.
But by and large, we're optimists in the sort of in human institutions' ability to adapt.
And that's what a lot of our focus is on.
these days. One of the interesting things we realized when we were writing the essay was that
people who are sort of the most optimistic about this technology who think AI will bring
about this utopia and the people who are more scared about it, you know, the tumors, let's say.
They have a surprising amount of intellectual sort of commonality. They are perhaps more common,
they share more with each other than they share with us or the rest of the world. And I think
it is this sort of helplessness in the face of powerful AI that drives both of these perspectives.
For this imagined utopia, the helplessness is brought about by human scientists no longer
contributing or needing to contribute to, let's say, cancer research. For the Dumeers,
it's brought about by AI systems, basically not caring about or destroying humanity. But in both
of these cases, the assumption is that we deliberately hand over power to these AI systems,
perhaps because of how advanced they are.
Once you reject that premise, though, once you reject the premise that we will do this,
we will hand over power to systems, and once you sort of assume that we will continue to
remain in control and human society will adapt to try to remain in control, I think it becomes
both more optimistic and more pessimistic, more optimistic in the sense that, you know, humans
will continue to be the driving force on this planet, that we will continue to make advances
using this new general purpose technology, it's also more pessimistic because, you know,
you know, we can't cure cancer in like the next five years,
or we won't be able to solve all of human scarcity in the next five years.
It will merely be something that leads to, as you said,
like a single-digit percentage point improvement in the state of human flourishing and progress.
Arvin, it reminds me that there's a way in which one disagreement that you have
with some people in the AI space is about the generalized panic over superintelligence.
But there's another disagreement that you have that's almost political in nature.
And that is about the question of whether AI safety should be a big tent or a small tent.
A small tent of AI safety would say this is a movement that should consist of people who agree with the highly ideological view that extinction risk is more important than everything else.
The fact that this technology can end the human race is clearly the overarching risk presented by AI.
and that frames people like you a little bit as adversaries
because you don't believe in this existential risk.
But if you broaden the tent and essentially say
there's a large coalition of people
who have fears about sometimes existential risk,
but sometimes cyber risk,
sometimes operational intelligence inside of the AI labs,
sometimes bio-risk based on what bad actors can do
with open-weight models
or even just advanced proprietary models,
can you speak to the idea that not only are you offering this sort of cautious optimism about this technology
or also somewhat implicitly advocating for this big-tent approach to AI safety?
That's right. We hope there can be a big-tent approach to AI safety
because we've repeatedly said, including in the original essay,
that there is a lot of common ground on policies,
even if we don't agree on the big existential risk question.
A lot of things like transparency inside what is going on in the labs
or liability when an agent is not properly supervised and causes damage,
those are all areas where we could be improving policy.
And we are, you know, as disappointed as a lot of AI safety people are,
that even based on the last few years of evidence of these very real risks and gaps in policies,
there hasn't been that much movement.
And that's something we can work together to fix.
The issue with the overarching focus on existential risk is that, you know,
nothing short if a ban or something equally serious like that can really move the needle on that.
So I feel like we are missing the opportunity to start by making progress on a lot of this low-hanging fruit.
So this is different from the distraction argument.
The distraction argument is, oh, the real risk is, you know, algorithmic bias or something like that.
And so this is distracting for what we shouldn't really focus on.
We don't subscribe to that argument.
You can be worried about, you know, two things at a time.
you can be worried about job displacement in economics, you can be worried about safety, you can make
progress on both. But the difference between small tent and big tent is different diagnoses of the
same problem. And I think they're going to compete with each other for what our policy action
should be. Should we start from these achievable things that there is a lot of consensus on,
or should we shoot for this big, you know, hazily defined ban on superintelligence? And that's
what worries me.
I have one Cota question.
It doesn't entirely fit with the thrust of this conversation, but it is, of course, related.
Arvin, it's related to some work that you've done and speeches that you've given about
AI and the future of work.
I'm extremely interested in the future of work, AI-related or not AI-related.
And I was really interested in this argument that you made, this framework that you developed,
called the Decide, Execute, Deliver, Sandwich.
Tell me a little bit about this framework and how it helps us see the way that powerful AI models
might not eliminate, say, all but one-tenth of the software engineers,
even if it increases software productivity by a factor of 10.
Yeah, we developed this framework together.
We've written about it on our newsletter.
And, you know, the argument is that not only might it not eliminate a lot of late,
it might even increase the demand for labor.
And it goes like this.
So in most knowledge work, specifically in software engineering was our motivating case,
but we think it generalizes pretty well.
You don't just plunge in and do the task.
There's a lot of work that goes into decision-making.
What do our customers even want us to build?
How should we architect it in a way that's maintainable,
complies with regulations, response to market needs,
so many of these fuzzy criteria.
And then you build the thing, you debug the thing, et cetera.
But then there's the third delivery phase, which is thinking about,
okay, do we understand this well enough that we can stand behind this,
that we can be accountable for it?
How do you go integrate it into the customer's system?
So many other things that need to happen.
And roughly the work seems to split one-third, one-third, one-third between decide,
execute, and deliver.
Again, at least in software engineering, there is some quantitative data on that.
And so what we see AI do is massively
compress the execute stage.
At least so far, but I think there are reasons to think that this is going to be a durable
observation.
Just because AI capabilities increase a little bit, I asked this when I gave a talk to a room
full of software engineers, would you be comfortable saying, oh, we made the decision to do it
this way because that's what the AI said.
Instead of really taking accountability for it.
And they all started laughing, right?
So the question of whether AI can eat the decider execute phase, that's not a question so much of
AI capabilities, but rather these organizations in our comfort level with handing over control to
AI. And maybe someday we'll decide to do it, but A, that day is pretty far off. And B, it'll probably
be a bad idea unless we really make progress in how we can still overall remain in control of these
AI systems. And so that's why for a while, at least, we think that AI's impact.
is going to be in that middle phase of execution.
And then if the work gets easier to produce
because of this compression,
there will arguably be demand for more work
and humans are going to be involved
in the two ends of the sandwich.
And say, Ash, maybe last question for us today,
for folks who are either in school,
about to go to college,
just graduated from college
and thinking about maybe picking up a few extra classes
or maybe parents of kids who are in school,
thinking about the future of work
as this compression of exercise,
execution, which thereby increases the value of deciding and delivering.
How should that make people think about what to major in, what to focus on, what skills to
build, you know, what advice would you give to a class that said, you know, how do I become a
better decider and deliverer if AI is going to do more and more of the execution?
That's a great question. I think that's something that our universities have
especially on the engineering side of things, perhaps not focused as much on.
There's been a lot of emphasis on learning how to build new algorithms.
But if you think about it, the overall time that someone spends in university
does actually aim to give students this sense of taste or agency
or whatever you call it on making sure that you sort of choose your right major,
making sure that you're doing something that you're actually excited about.
So I think more so than what the specific choice of major with students take is this orthogonal
question of what is it that you're excited to sort of deliver on or what is it what is it that
you're excited to learn on the job on because I think frankly a lot of what today's graduating
class will learn will be on in real world organizations and you know this kind of experience
will become just more and more important as time goes on. So I think that's sort of the main piece
of advice would simply be to focus on things where you can see yourself in the next three or
five years developing a lot of taste in agency. Even if you're not particularly working on the
lower level stuff, you're excited about developing relationships, you're excited about sort of owning
large pieces of problems more so than sort of doing the mechanical parts yourself. Because frankly,
when I chose to be a computer scientist, a big piece of it was solving the intellectual puzzles
that come alongside being a computer scientist. But that's no longer a major part of what
computer scientists too. It's actually about building software that is broadly useful in the world,
learning the customers' needs, and so on. So I guess focusing on what it will be that you do
after you exit the university and what excites you most about the real world aspects of that
job would be my key piece of advice. And I would say thinking about my own career,
there are stories that I've written that I worked for days and weeks on, but I never had the right
frame for the story. I never nailed the headline. I never nailed the nut graph, the paragraph that
explains the thesis of the article. Maybe I never nailed the lead. And so as hard as I worked on it,
as hard as many hours as I spent trying to execute, because I didn't decide on the right frame.
It was somewhat doomed from the start. Whereas there are stories where I will think of the frame.
Sometimes like I'll tweet it out and realize that people are like retweeting it a lot. I'll say,
aha, okay, there's a reaction to this.
And I can spend half a day on a story,
and it'll do so much better than the article
that I worked for for days or weeks on.
And this, it's not so much to say,
I don't think that blood, sweat, and tears
and perspiration are important for journalism
or any other career.
Of course, they'll continue to be important.
I think that, you know, Edison,
maybe it's a real quote,
you know, 1% inspiration,
99% perspiration.
I think that's probably true still of a lot of work.
But it strengthens, I think,
my hypothesis that strong frames sometimes are the majority of the work and execution is sometimes less
important than making the right decision on what to cover, what questions to ask, what headline
to put on the story.
So I think that sounds directionally right to me.
Sayash, Arvin, thank you so much.
This is a really, really fun conversation.
I learned a lot.
Likewise.
Thank you, Derek.
Thanks a lot for having us.
Where some see heroes and others see egos.
Bloomberg sees the era of billionaire athletes.
While others follow the noise, we follow the money.
Learn more at Bloomberg.com.
