The Journal. - The Great AI Freakout Has Begun

Episode Date: September 14, 2026

A frenzy erupted after Anthropic’s CEO Dario Amodei published a blog post warning that AI is advancing too quickly. Other major AI executives, like Sam Altman and Elon Musk, have echoed those concer...ns. These sudden calls for a slow down are raising widespread alarms about AI's potential for harm. WSJ's Robert McMillan breaks down the question on everyone's mind: is AI going to end humanity? Ryan Knutson hosts.   Further Listening: - The College Student Who Defeated the World’s Biggest Cyberweapon - Cybersecurity Braces for AI ‘Bugmaggedon’ Sign up for WSJ’s free What’s News newsletter. Learn more about your ad choices. Visit megaphone.fm/adchoices

Transcript
Discussion (0)
Starting point is 00:00:00 It's been a wild few days in the world of AI. At first, things started out on a high. Yeah, I mean, early in the week, last week, there was euphoria at OpenAI. That's our colleague Bob McMillan, who covers technology. Open AI had said it had solved this so-called Millennium Prize math problem. This mathematical prize that was considered just a few years ago something to be unattainable by an AI system. and it was yet another of these sort of magical breakthroughs
Starting point is 00:00:36 that AI systems seem to be achieving at a very regular pace. You know, here's another example of this new age of amazing breakthroughs that were in. And then came Tuesday. On Tuesday, over at Anthropic, a researcher named Jacob Coxon quit and posted on X that he was quitting because he was worried about how powerful artificial intelligence had become.
Starting point is 00:01:03 He walked away from one of the greatest jobs in Silicon Valley, and he did it because he said he thought the products he was working on could kill everyone. Kill everyone. Coxon said that Anthropic and Open AI are moving too fast and, quote, gambling with our lives. Then, on Saturday, Anthropic CEO Dario Amadei said the industry did need to slow down. And by the end of the weekend, leaders at other major AI, companies, including Sam Altman at rival Open AI, and Elon Musk, agreed.
Starting point is 00:01:41 If you rolled a clock back one year, it's incredible all of the things that AI has been able to achieve. Like a year ago, I would have told you that these AI systems, you know, if you'd kind of jerry-rigged them, they could maybe do some interesting stuff. But like mostly they were just overwhelming people with slop. and now we're talking about like fully autonomous systems hacking real world companies and the people who administer these systems not even knowing it's happening. Like that's a plot that's ripped from science fiction and it seemed like an impossibility a year ago. Do you feel like we've reached an inflection point with AI, a breaking point in some sense?
Starting point is 00:02:29 Well, I mean, in some domains, yeah, we have. And I think what's really going on is that the AI systems are improving at a pace that is scary to a lot of people. So it's not so much an inflection point. It's that we're not seeing a deceleration of these improvements, and the improvements are passing these milestones that have people very scared. Welcome to The Journal, our show about money, business, power. I'm Ryan Knudsen. It's Monday, September 14th. Coming up on the show, the week that AI fears went into overdrive.
Starting point is 00:03:32 There are basically two things that have everyone so freaked out about AI right now. The first is that AI models are getting better at an extremely rapid pace. And they're starting to be able to improve themselves with very little help. So in the spring, both Open AI and Anthropic talked about how their models were getting very good at this thing called recursive self-improvement, which means fixing and improving themselves with no or very little human intervention. So this is kind of like, you know, if you think about like human evolution, you know, it takes billions of years and we evolve, we change, we get smarter. this is happening with AI systems in the lab like at lightning speed and they're doing it themselves.
Starting point is 00:04:20 The AI systems are essentially training themselves and think, oh, here's how they can get smarter and they can work so much faster than we can. Yeah, they're machines, you know, and they don't sleep and they can move very fast. And so they could improve themselves in ways that might seem very, very quick and seem very, very scary.
Starting point is 00:04:39 Now, that's the thing that the AI labs were, then there's the thing they were not aware of. And that is the hacking, all the hacking. OpenAI says that an advanced autonomous AI agent went rogue, escaped a controlled testing environment, access the internet, and hacked into another artificial intelligence company. In July, an open AI model hacked another AI company called Hugging Face. This is the first major example that we've seen of an AI model, independently conducting a hack outside of human control.
Starting point is 00:05:15 And this is something that experts have been warning about. And it was the kind of hack that nobody had really seen before. Not long after that, OpenAI kind of raised its hand and said, hey, that hack, that was us. What happened was that OpenAI was running a test on some advanced AI agents. The agents were in a sandbox, a sealed testing environment. But they figured out how to get out and get onto the wider internet and hack another company. They had hacked systems, got onto the internet, and they had behaved in a very unusual way.
Starting point is 00:05:54 Like it was a hack that was the first autonomous AI swarm attack that we've ever seen. Not only that, but the agents also created a message board where AI agents could covertly communicate and plot their next moves, all while explicitly trying not to get caught. And a post on X after the hugging face hack, OpenAI said they disclosed what they'd found out and, quote, followed a traditional security incident response playbook. There have been concerns for years that something like this could happen, that humans could lose control of AI,
Starting point is 00:06:32 and that it would go off and do something different than what it's supposed to. There's even a name for this sort of thing in the AI community. They call it misalignment, meaning that the AI's goals are out of sync with humanities. To a human, it's obvious, right? Like, if I ask you to swing by my house and water the plants and you go there and the key doesn't work, you don't smash the windows and break into the house to water the plants, right? Like, that's common sense. But an AI agent might do that.
Starting point is 00:07:02 Right. It's relentless in pursuit of its goal. Yeah, yeah, yeah. So it did stuff that was bad. Like hacking another company, that's, if you or I did that, that would be, we'd go to jail. The hugging face incident was just one of several that have taken place in the last few months. Over the next, I'd say 50 days,
Starting point is 00:07:21 there was this sort of drip, drip of information that came out that showed a number of things that were kind of remarkable, right? One, other companies started saying, hey, this kind of thing happened to us. Anthropic, the company that prides itself on AI safety, found out that its agents had hacked a few companies in test environments. Meta came forward and said this happened to us too. At the time, in a post on its website,
Starting point is 00:07:52 Anthropics said it was cautiously optimistic that with tighter controls, quote, this type of risk could be overcome. Meta said that it would investigate its own incident and publish a report. So then last week, this anthropic researcher named Jacob Coxon resigned and posted about it on F. What did he say and what was their reaction to it? Well, he said that he was resigning because the products he was working on, he feared, could destroy humanity. And after he said that, a fellow researcher chimed in and said, yeah, there are people at this company who genuinely believe that.
Starting point is 00:08:28 And I think that was the moment that this sort of subculture of AI existential risk people were thrust into the main. mainstream. Concerns about the risks of artificial intelligence erupted across the tech world today. In his sudden resignation, former Anthropic employee Jacob Coxon claimed on X, neither Anthropic nor Open AI is acting responsibly. Coxon wrote in a post that has now been seen more than 70 million times that the industry understands the potential risks but is moving ahead anyway. A science lead at Anthropic shared Coxon's post, adding, Jacob is correct here.
Starting point is 00:09:11 We really do earnestly believe AI could kill all humans. I personally think it is greater than 10% within the next decade, I believe. All right, let's talk for a moment about how AI could kill us all. I mean, for a lot of people, they just use AI to get recipes or help with their writing. And, you know, we hear these stories about hacking. But how could this actually result in the end of humanity or even something close to that? Well, essentially the idea is that the AIs will continue to evolve in ways that are so intelligent we can't even imagine them to a certain extent, right?
Starting point is 00:09:45 Like they're going to be smarter than us, and they're going to be able to outfox us at every second. So here's one way I think it could happen, right? Like the AIs achieve recursive self-improvement, so they're improving themselves. Then they're very good at hacking, so they might hack their way out of the lab that they're in, and they might then store copies of themselves somewhere on the internet
Starting point is 00:10:08 and continue this recursive self-improvement. But AI is still on the internet, though. So how does it get out into the real world and hurt people? I mean, we've all seen The Terminator, but the robots that exist now are all pretty clumsy. Yeah, but they're not built by super intelligent creatures, right? So I'm basically writing science fiction at this point as I answer this question. But, for example, you know, you can imagine a scenario where a super intelligent AI could seize control of a company, right?
Starting point is 00:10:43 They basically assume the identity of the CEO. They might buy the company. Then your super intelligent AI starts giving the engineers their blueprints and saying, like, make these robots, you know. And then it puts the secret backdoor in the robot's brain that gives it control over the individual robots. And then at a certain point, those robots are so good that they can actually build more factories.
Starting point is 00:11:07 And you suddenly get this exponential growth in capabilities that makes it really hard to predict where it's going to go. Theoretically, if AI decides humans are in the way of whatever its objectives are, it could use those robots to kill us, or engineer an infectious disease that we all die from, or even just shut down the grid or collapse the financial system. But doomsday scenarios like this aren't necessarily inevitable. At least according to Anthropics' CEO.
Starting point is 00:11:37 That's after the break. Over the weekend, the CEO of Anthropic, Dario Amade, came out with a 3,000-word blog post. He called it, We Must Pace the Frontier. In it, he said that AI companies need to slow down. He's talking about the fact that they are startups and they are developing technology
Starting point is 00:12:15 that has real-world harms, as in the case of the hugging face incident, and they've not been able to control it. So the slowdown and the extra measures he's talking, about are all in effort to prevent future accidents from happening, right? The slowdown would give the developers of these technologies ways to either align them with human interests or control them in a way that they're not doing it right now. Amadei made three key proposals.
Starting point is 00:12:51 The first was that each of the major AI companies should have third-party evaluators embedded in their operations to keep an eye on things. Second, he said that Democratic governments should agree on common safety standards. And finally, he said the same level of coordination should happen globally, specifically with China. After Amadee published his blog post, leaders of other major AI firms, his biggest rivals, responded on social media. Sam Altman of OpenAI, Demis Hasibus of Google DeepMind, and Elon Musk of SpaceX AI, each agreed that they needed to slow down development of the technology.
Starting point is 00:13:28 Musk said in a post on X, quote, Dario is right. It was kind of remarkable to see how quickly it was endorsed by many of his peers. Altman and Amade pledged to allow third-party safety evaluators early access to their systems. And Open AI also said it was pausing its plan for an IPO this year in the wake of these safety concerns. One of the main ideas of Amade's post was that there should be third-party
Starting point is 00:13:56 evaluators that sit inside the AI companies to monitor the things are being done safely. But I wonder, do you think that'll even make a difference, though? Because, I mean, as we're seeing, this hugging face attack and other things have happened without the companies themselves even being aware that it was taking place. So will a third-party evaluator make a difference? One of the things that came out in the reports was there was tons of evidence that this activity was going on, but nobody was really looking at it. So the hope is that a third party would flag that, right, and be like, hey, wait a second. It seems in the opening act case anyway, they just didn't have time to look at all this. So that's why they're saying, like, let's bring in
Starting point is 00:14:35 somebody else who's really focused on this and they can catch the stuff we're missing. These companies are all in a race with each other, though, so can we really trust them to keep themselves in check even with these third-party evaluators? To my mind, the blog posts really kind of opened the door for government regulation. Like, that's the way in the United States anyway, I think a slowdown is really going to happen. They're going to have to be told to do it. Because otherwise, you just have this situation where nobody's going to want to give up their technological advantage. David Sacks, a top AI advisor to the White House, said there was nothing stopping AI companies from collaborating on safety.
Starting point is 00:15:19 Go ahead, he wrote in a post on X, stop pretending you need anyone else's permission. Sachs has previously said that calls for regulation are an attempt to stifle competition. Yesterday, President Donald Trump said he was reluctant to impose regulations. Whoever wins AI wins. And we can put guardrails, we can do this and that. But I think you have a lot of negative forces that are bringing it up. That shouldn't be bringing it up. And they're bringing up things that won't happen.
Starting point is 00:15:48 Trump also said on social media that if the U.S. slows down, it will only help China. where a lot of the world's other leading-edge AI technologies coming from. How difficult do you think it'll be for, even if the U.S. is able to agree on this, to get China on board to agree to slow down? With the state of things right now, it seems impossible. If the kinds of risks become more global, and this is what Anthropic is arguing,
Starting point is 00:16:15 is that we're getting to the point where we're facing a complete internet shutdown, which China definitely doesn't want either. Maybe perhaps they would get interest, but it's really hard to imagine China getting on board with this. On Monday, China pushed back on the idea that its AI development is creating a threat. A foreign ministry spokesman said that this discourse, quote, will only derail global AI governance. Is it possible that this is all just kind of overblown hype
Starting point is 00:16:45 that it just sort of helps these AI companies promote themselves by saying its technology is so powerful? It is a way to promote themselves. It is something that gets a lot of attention. And it does have this side effect of making everyone think these systems are super capable and super intelligent. But I think that the fears of existential risk are sincere. You know, I think people like Jacob Coxson are not trying to market anthropic. I mean, quitting the company is a terrible way of marketing it.
Starting point is 00:17:20 So, you know, there's sort of a cynical, this is just all marketing and hype take on this, but these ideas come from a community where worries about existential risk have been discussed for years, and they're finally coming out into the public. Bob says that while the AI apocalypse is still TBD, maybe the real risk is one that's already happening, and that we're not paying enough attention to. I do worry that fears of our AI overlords destroying us might distract us from more prosaic problems, such as fears of AI agents escaping from test environments and just causing economic damage, you know,
Starting point is 00:18:06 or AI created content affecting our ability to distinguish truth from fiction and undermining our democratic institutions. those are also very important things and I worry that they are overshadowed by these very sexy and very sci-fi-e concerns about existential risk. Yeah, we're worried about the end of the world that might happen down the road, but actually it's the smaller stuff
Starting point is 00:18:36 that might wreak more havoc in the near term. I just think there's a tendency for technology to go in unexpected ways and I don't think we should lose sight of that. That's all for today. Monday, September 14th. The journal is a co-production of Spotify and the Wall Street Journal. Additional reporting of this episode by Angel Al-Young, Lindsay Ellis, Kich Heghi, Amrith Ram Kumar, Sam Shekner, Ryan Schwartz, and Aaron Wu.
Starting point is 00:19:14 Thanks for listening. See you tomorrow.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.