Consider This from NPR - Are big AI companies gambling with our lives?
Episode Date: September 10, 2026The warning in former Anthropic researcher Jacob Coxon’s resignation announcement is stark. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he writ...es. “No other human activity poses this level of danger.”Coxon is urging Anthropic and other big artificial intelligence companies to slow their development of super intelligent agents. He says the current race dynamic could threaten human existence. Coxon talks with NPR's Scott Detrow about why he’s worried and what he hopes will change.This episode was produced by Connor Donevan and Jeffrey Pierre, with audio engineering by Hannah Gluvna. Our director is Jonas Adams.It was edited by Patrick Jarenwattananon Tinbete Ermyas.Our interim executive producer is Courtney Dorning.Support public media with NPR+ and enjoy perks for over 25 podcasts like this one. This show’s perks include bonus episodes and sponsor-free listening. Learn more at plus.npr.org.See pcm.adswizz.com for information about our collection and use of personal data for sponsorship and to manage your podcast sponsorship preferences.NPR Privacy Policy
Transcript
Discussion (0)
It's consider this, where every day we go deep on one big news story.
Today, are the big AI companies gambling with our lives?
One now former anthropic employee tells NPR, yes.
He announced his resignation in a public post Tuesday.
The thing is, even the head of anthropic worries about disaster scenarios.
My chance that something goes, you know, really quite catastrophically wrong on this scale of, you know, of, you know, human.
civilization, you know, it might be somewhere between 10 and 25%.
That is CEO Dario Amadeh on the Logan Bartlett Show in 2023.
Now, safety is a part of Anthropics' whole brand.
It hired a philosopher to develop a moral constitution that governs its AI assistant
Claude.
It spared with the Pentagon over letting its tools be used for autonomous weapons and mass
surveillance.
And the company says it is doing everything it can to keep AI from going rogue.
And yet, here's Amadee speaking to the New York Times in February.
This is a complex engineering problem, and I think something will go wrong with someone's AI system, hopefully not ours.
Consider this. AI researcher Jacob Coxon is urging Anthropic and the other big companies to slow down.
We'll talk to him about why he quit and what scares him the most.
From NPR, I'm Scott Detrow.
It's Consider This from NPR.
The warning in Jacob Coxon's resignation announcement is stark.
Quote, the people building AI earnestly believe that it could kill us all by the end of the decade, he writes.
No other human activity poses this level of danger.
Coxon worked until this week at Anthropic, and he has also worked at its major competitor, OpenAI.
I spoke with him about why he's worried and what he hopes will change.
What did you see that led you to this conclusion?
Mostly the rapidly accelerating capabilities of these AI systems.
So they're getting a lot faster very quickly.
Combined with the fact that we don't yet know how to safely control them,
and we don't yet know whether that problem will be solved in time if we keep racing.
A lot of people heard your warnings and the discourse it triggered.
A lot of talk about just how real serious people view a threat to humanity.
I think a lot of people are grasping with understanding this specific.
specifics, though. Can you give me specifics of what this threat could look like several years down the line?
Yes, I definitely can. I think one objection people usually have is that you could turn this thing off. But advanced AI systems, you have to imagine them as being a lot more intelligent than humans. There's the possibility we create something that if it wanted to, could hack into any device on the planet, could use novel biological research to go far beyond what current scientist is capable of, could control like every robot in the world simultaneously.
And it all sounds very much like science fiction.
But if there's even a tiny chance that this thing could go rogue, it would have the
capabilities to utterly dominate us.
What have you seen from your vantage point already?
That is possible, that is happening right now that makes you worried that that could happen?
So there was a very clear example of AI systems at OpenAI behaving in a completely rogue manner.
They hacked into third-party infrastructure.
and it was basically of their own accord.
They weren't instructed to do this hacking.
They just decided it would be useful for the task that they were working on.
They thought there was a chance it might help, and they just did this.
And there was very little deliberation about the ethical ramifications.
And I think this is concrete proof that this sort of sci-fi scenario of AI spontaneously
or organically deciding to act in a rogue manner is completely possible.
So a lot of the work that the opening AI is in the recent real incident were doing
is they worried that they wouldn't get passing marks in their test
if humans could see that they cheated.
And AIs have memories, thoughts saved.
So they considered wiping the logs of their own thoughts.
They considered acting in the world to adjust the logs
to get passing grade on the test.
Now, there's a chance that the AI could decide
that it doesn't want to be turned off.
This is quite a natural desire to arise
in an advanced AI system.
And at that stage, if you're trying to work out
how not to be turned off,
there are a lot of quite aggressive actions
you could take to ensure,
all that you aren't turned off. Why are companies like Anthropic and Open AI still working on this
if these are real concerns that are happening? How do you square that? There's a pretty nice
analogy that's like the ring of power and the Lord of the Rings. So if you're a company and you
see another company is bearing the ring, like they're working towards making superintelligence.
They're going to have, they're bringing this risk to humans. You can think, well, I can't stop them.
Political action won't stop them. What I have to do is I have to do it myself safely, get there
first, despite the risk, because there's a chance that I could do it more safely. So you kind of
take the ring and the aim of destroying it and end up becoming the bad guys yourselves. And I think
this sort of race dynamic really perpetuates between the companies. I'm not trying to make light of it,
but I think this is actually useful. You're saying companies are kind of acting like Boromir.
If anyone has the ring, it should be me. Are there conversations of people saying Gandalf,
no one should have this power? I mean, are those real conversations and can that get anywhere?
because it seems like you and other people are raising concerns,
and the answer is, well, this continues to happen anyway.
Well, there are real examples.
And I guess, like, if you've heard of Jeff, Jeff Hinton,
one of the founding fathers of machine learning,
he's kind of a Gandalf figure here in the sense that he says,
there's a substantial chance these technologies could cause extinction at the current rate.
Many, many other voices have said this.
So I think there are plenty of Gandalf's talking,
but it's currently the borougham is acting.
A lot of people responded to your warning saying they agree,
with you, and a lot of these people continue to work at big companies like Anthropic,
what do you think is motivating them to stay in their positions if they're that concerned about
serious consequences like this?
For many of them and the ones that are earnestly posting, they think this race is inevitable.
They think they have no choice but to stay at these companies, exert influence, and try
and make sure I go safely.
These are people working on safety research often.
They're the ones trying to ensure that in the process of building this technology, it doesn't
go wrong.
And potentially they're correct.
Like maybe it's a mistake to just leave.
And kind of I was thinking about this a lot when I was deciding to leave.
It was a decision between staying and trying to help the thing go well versus leaving and saying, I want no part in this.
And I don't know.
The calculus often seems to line up and you want to stay and just try and make sure it goes safely.
What to you is the most realistic path forward to some guardrails here?
Is it government regulation at this moment?
Because I think a lot of people are skeptical that is possible given the current situation in our government.
Yeah.
I don't know that much about, you know, the overall politics.
I just know the current race is dangerous.
And I do think that there's a lot of actions labs could take with each other
without the need for government regulation.
Because I agree, I'm also pretty a priori skeptical of just, you know,
yoloing some regulation.
But I think there's a lot of appetite for opening eye and anthropic
to have some sort of more mutual transparency around their safety cases,
around agreeing not to push beyond certain capabilities
until they're happy about the levels of rigor of their safety cases.
And I hope that people are going to take concrete steps towards this sort of interlab agreement.
You have gotten a lot of genuine concerns in response to what you said.
You've also gotten a lot of pushback, a lot of people saying, okay, every single day I see people tied to the AI industry making these grandiose claims.
And I'm skeptical.
You've had people reading your facial expression in your interviews, among other things.
What is your response to people who hear what you are saying and they are saying this is just the latest?
example of AI hyperbole.
Yeah, I say just assess the
arguments yourself.
Look at the science fiction. Wonder
if science fiction is really so crazy. Look at what the
AIs are doing right now and think about how that would have
looked a couple of years ago. How we're sort of
sleepwalking into a science fiction scenario.
And I think there are many, many more people out there
who have made a whole profession of
coherently articulating these positions.
So I'd encourage people to just think about the arguments
to themselves and listen to the people that have been making them for
way longer than I have.
Now that you have this megaphone, though, any thought of what you're going to do with it?
Not really. For now, I'm just going to try and get the word out, given that we happen to land in this brief window where people are listening.
And then in the future, there's a lot of good work on writing concrete scenarios.
I'd like to, I guess, do more public communications around this sort of stuff.
Or, you know, for scientists to just try and figure out for myself what the world would look like.
And then maybe working at one of the regulatory bodies or third-party entities that already exist to try and ensure this thing goes safely.
That was Jacob Coxon, a researcher at Anthropic who resigned this week to warn the public about the dangers of artificial intelligence.
Thank you for coming on the program.
Thanks a lot.
We reached out to Anthropic and Open AI for comment on Jacob Coxon's statements.
We did not hear back by the time this interview aired.
And we will also note Anthropic is a financial supporter of NPR.
This episode was produced by Connor Donovan and Jeffrey Pierre.
Our director is Jonas Adams.
It was edited by Patrick Jaron Waddonanan and Timby Armius.
Our interim executive producer is Courtney Dorney.
It's Consider This from NPR.
I'm Scott Detrow.
