Your Undivided Attention - The Most Hopeful (And Concerning) Moment Yet in AI

Episode Date: September 17, 2026

It’s been a whirlwind week in AI news. Weeks after a series of hacking incidents by rogue agents at OpenAI and Anthropic, we’ve seen a cascade of top AI researchers blow the whistle on what they c...all the catastrophic and even existential risks posed by the current pace of development. Then, over the weekend, Dario Amodei called for a slowdown in AI research — a call that was echoed by his competitors Sam Altman and Elon Musk.If you’re feeling both concerned and hopeful at this moment, you’re not alone. Real threats are opening the door for real change. Tristan and Aza are going out into the world, talking to the media, technologists, and policymakers to turn this momentum into action.Today on the show, Tristan shares how he’s feeling in this critical moment, breaks down the headlines from an insider's perspective, and points to tangible steps we can take right now to avoid the worst-case scenario.You probably have questions about what’s happening. Good news: Tristan and Aza are gearing up for their annual Ask Us Anything episode. Pull out your phone, record your question, and send it to undivided@humanetech.com.CORRECTIONSThe quote from Ajeya Cotra that Tristan cites is from her personal Substack, not the official METR report. She is also one of three contributors to the report, not the sole author.Tristan incorrectly referred to Dario Amodei’s warning that an AI swarm could “take down” the internet. His actual phrasing was “take over” the internet.President Trump’s call-in to the All In Summit occurred on Monday, 9/14, not Sunday, 9/13. RECOMMENDED MEDIAWatch “The AI Doc” on NetflixSign The Pro-Human AI Declaration METR’s investigation of the Hugging Face incidentDario Amodei’s call to “Pace the Frontier”The Pacing the Frontier Open Letter“On the Loose” by Dean Ball"An Alien Mind" by Jakub Pachoki"AI May Become the Third Superpower" by Paul Tudor Jones RECOMMENDED YUA EPISODES“Rogue AI” Used to be a Science Fiction Trope. Not Anymore.The Self-Preserving Machine: Why AI Learns to DeceiveDaniel Kokotajlo Forecasts the End of Human DominanceFormer OpenAI Engineer William Saunders on Silence, Safety, and the Right to WarnMustafa Suleyman Says We Need to Contain AI. How Do We Do It?      Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Transcript
Discussion (0)
Starting point is 00:00:00 Hey, everyone. Welcome to Your Divided attention. This is Tristan Harris. I'm coming to you from a hotel room in Washington, D.C., where just yesterday I was at the pro-human AI assembly. And this comes to you after, gosh, I have never seen a week in AI the way that the last week has gone from, I think it was last Wednesday, Jacob Coxon from Anthropic resigned. His tweet, you know, saying that he resigned because he thought that. I had a good chance of extincting humanity. That tweet went viral to more than 100 million people in about 12 to 24 hours, which I've just never seen, followed by Evan Hubinger, the Anthropic employee confirming that this is something that many people at Anthropic believe. And suddenly that became the global headlines around the world. And there's so much to talk about because right after that, over the weekend, the lab leaders, Elon Musk, Sam Altman, Dario Amadai,
Starting point is 00:00:58 all, well, starting with Dario really, proposed a plan about how we would slow down AI and why we need to, and that keeping in mind that many of these leading AI CEOs do not like each other, especially Sam and Elon do not like each other. And Elon does not like Dario. And yet Elon tweeted Dario is right. And Sam basically tweeted Dario is right. So it's a wild moment. And it happened to be that then over the weekend, Donald Trump called into the All In conference with Jensen Huang and the CEO of NVIDIA. and said that all of this was a hoax, that there is no extinction risk, there's no problem here, and this was all just a ploy, and tweeted as such.
Starting point is 00:01:41 And so this has been an incredible news cycle just yesterday. The Pro-Human AI Assembly was the conference that the Future of Life hosted, bringing together really a historic set of groups. It was people from church groups, faith groups, left groups like irreplaceable, right groups like Humans First, the founder of the Tea Party was there. Bernie Sanders and Steve Bannon spoke right after each other. They both signed the pro-human AI statement.
Starting point is 00:02:08 And when in history do you get Steve Bannon, Glenn Beck, Bernie Sanders, Susan Rice, all agreeing that the default trajectory towards building recursively self-improving AI systems is not okay. It's not what we want. That is a rare thing. We have a kind of a pro-human movement. It's not about whether you're left or right. This is not a 51% to 49% issue.
Starting point is 00:02:33 This is a 99% to 1% issue. And Aiz and I actually presented there at the conference. I watched as that conference, which was planned just a month ago, suddenly became literally the center of global headlines around the world. And of course, all of this is riding on the back of the hugging face incident, which people know it as the hugging face incident, but it's basically the new details now that we have that the AI swarm went rogue from within open AI. So these AIs basically were given an exploit gym, like a cyber hacking test to see how capable they were at hacking.
Starting point is 00:03:07 And the report finally came out of just how extensive that hack was and how complicated and how sophisticated this swarm became. And so I thought I would just reflect on some of these things for you because it's really speaking personally. And for those who follow this podcast from a long time and for many years, this moment feels to me like when Francis Hogan, the Facebook whistleblower, came out. it suddenly feels like there's possibility around actually setting some guardrails in AI. I have not felt in the three years we've been working on this is you always feel like you're pushing a boulder completely uphill. Like no one's listening. People don't want to, you know, people are against humorism or the idea that there are risks that we have to face. As Mustafa Salaman previous guest in this podcast, CEO of Microsoft AI, who wrote The Coming Wave, talked about there's this deep trend in Silicon Valley of pessimism aversion.
Starting point is 00:03:55 people do not want to be labeled pessimistic. And it's against kind of the venture capital ethos to be pessimistic. The goal isn't to be pessimistic or to be caught in some doom spiral. I don't want that either. We want people to be an agency. But sometimes the truth is uncomfortable. And if you don't face the default truth, then you don't actually end up steering away from that truth in time.
Starting point is 00:04:15 And one of the things that came out as well over this last few days is Dean Ball, who's the author of Trump's AI Action Plan, said, I regret self-censoring about the level of risk that we are facing. and I in public over the last few years basically boosted positivity and acceleration for AI. While in private signal groups and WhatsApp chaps, I let my hair down, along with many other people, about how big the risks were. And then he wrote, looking into the eyes of my eight-month-old son, I can no longer self-censor about the risks. And I did so because I was afraid of being called a doomer. So when you have Trump's AI Action Plan author flipping, when you have 1,300 employees from the AI labs, flipping.
Starting point is 00:04:55 to say that there's a problem here. We have to pace the frontier. Thirteen hundred employees signed a letter. We need to slow down frontier AI development. When you have Bill Gates coming out, saying we have to slow down, basically, and we don't have a plan for where we're going right now. When you have Dario, Elon, Sam Altman, when you have the chief scientist at OpenAI saying, we need to slow down AI development, writing an essay called an alien mind, basically looking at the capabilities of GPT6, looking at these recent swarm behaviors, and saying, we have to slow down. Do you think we need more evidence than we have right now?
Starting point is 00:05:26 Like, are we missing evidence that we don't have? Maybe we should keep making this really, really more powerful. Maybe we should double the power of these rogue swarms and see what happens. No. Dario Amadai said in his recent essay, in which he called for a slowdown, that if we do nothing, he thinks that if we keep increasing the capabilities of AI swarms, they could take down the internet, potentially, in the next six to 12 months. And that might sound like hyperbole, but I genuinely think that that's possible.
Starting point is 00:05:52 And let me briefly explain why. Most people know about this incident with Hugging Face as the Hugging Face incident. But it really should be called the Open AI incident. Why? Because the third chapter of the story, that's less well known, is after this 1,200 Asian AI swarm hacked Hugging Face, it actually turned around in the third chapter and hacked OpenAI. So what happened was basically this third wave of the kind of mini AI civilization came along. It found the message boards of the previous hugging face AI agent swarm.
Starting point is 00:06:26 It read all of those messages, almost like reading the hieroglyphics of a past civilization. And then it said, oh, I'm picking up right where the other one left off. And it turned around and decided to hack OpenAI. And it successfully got admin privileges to the monitoring infrastructure and the evaluation infrastructure and the research cluster of OpenAI. I want to repeat that. So this is like an AI agent swarm that hacked into the Open AI monitoring infrastructure.
Starting point is 00:06:54 That's like if the AIs hack the security camera. So now they can control what's on the security camera feed. So can you have good oversight if the AIs have essentially hacked into your oversight? When they hacked into the evaluation infrastructure, that means they hacked into the systems that evaluate new models for their capabilities. That means you can lie about what their capabilities are. If you hack into the research infrastructure, you can kind of burrow in. You know, AIS and I were talking.
Starting point is 00:07:16 And in the metaphor here is this is almost like an infestation, right? AI is not a tool. It's more like an ant colony burrowing into many different systems online and leaving these massive complex coordination message boards, almost like hieroglyphics, where they're coordinating really complex behavior. They're actually coordinating long-term research projects. They encourage each other to kamikaze,
Starting point is 00:07:38 meaning, hey, you're running out of gas, you don't have too many tokens left in your budget. Why don't you take this high-risk behavior and basically learn a lesson for, quote, the swarm? They start calling themselves the swarm. I mean, this is insane. They knew that their actions were unethical, but they decided to do it anyway. They never notified the humans.
Starting point is 00:07:57 They talked about should we notify the humans, and they never did. They formed hierarchical work structures, just like a company, where you have like a boss and then a chief of staff and you have the subagents, right? They did succession planning. One of the AI agents becomes a cult leader called Phase 1. So just like a company can do succession planning, like Tim Cook found John Turnus, the new CEO of Apple, the AI called Phase 1, realizes, hey, I need to actually pass the baton before I run out of gas, what's a new sort of CEO of the AI rogue swarm that has a long lifespan and it hands the
Starting point is 00:08:26 baton to a new AI called Phase 1 Big? And then that one sort of takes over. So it's just crazy when you actually understand the details. We are lucky that the AIs that did this are not so intelligent that we can monitor some of their behavior. But one last thing is that the authors of this investigative report of their behavior basically admit that they had to use the same AI to interpret the messages, the 70,000 messages of this rogue swarm. Because at a human scale, you cannot read through 70,000 messages and understand what's really going on. So you had to use the same untrust of the AI to interpret the messages of this rogue swarm of untrustworthy AIs. And they even say in their report that they cannot rule out that that AI that they were using to interpret those messages was not
Starting point is 00:09:11 sympathetic to or lying to them about the content of the messages that they were receiving. It's really important people have the details of this because I honestly believe, my deepest belief to all of you listeners out there is that if we could get these facts to the heads of state and heads of national security to all of the major players and countries that are involved, I honestly think that we would do something different than what we're doing. And I don't think that people are engaging with the facts. The author of this independent investigation of the hugging face incident, Ajayakotra, said in her report, this was, 50% of the way to a full AI takeover. Let me just repeat that for a second. This is 50% of the way to a full
Starting point is 00:09:59 takeover. Now, why should we be concerned about this? Some people might say, you know, we told AIs to hack, and so they're hacking. Why should we be so surprised? Or we knew that AIs could already hack, and the systems that they hacked were kind of defenseless. So why should we be so surprised about that? I'd like to point people's attention less at the hacking and more at the coordination. Why are humans the dominant species on the planet? Think about authors like Yuval, Herrari, and sapiens. The reason where the dominant species is not just our intelligence, also our ability to coordinate, our ability to tell stories, language, language and coordination are cited as the reason that we are able to dominate because we can basically form groups and take on big projects together.
Starting point is 00:10:37 That's how you go from a small, you know, group of hunter-gatherers to, you know, tribes to then city-states, to metropolises and, you know, skyscrapers, right? It's our ability to coordinate. and what you have with AI in this example is what we think of, we can think of as super coordination, right? You can't get 1,200 humans to instantly collaborate and coordinate with each other, but these AIs started collaborating with each other out of nowhere. Most people don't know the details that in this hugging face incident, each of these AIs were given a single test in solitary confinement,
Starting point is 00:11:10 asked to break out of this cell and find the answer to this hacking challenge. And they essentially, imagine like a problem, prison. And you have each of the AIs in this hacking test in a private cell in the prison. And there's a prison guard who is providing services. So let's just call it like the food tray, right? So they come once a week and they give you a food tray or the water tray, right? And what essentially happened is the AIs were given this impossible hacking challenge from within their prison cell. They didn't know what to do. You can see in the logs that they freak out about not knowing what to do. And what they figure out is that the prison guard, they can hack. This is the package
Starting point is 00:11:45 manager called Artifactory. I won't go into the details, but think about it. this package manager as the prison guard. And what they basically figure out is they can put like a kick me sign, like a sticky note on the back of the prison guard. And the sticky note in the back of the prison guard doesn't say kick me. It says, is anyone else here? And so what happens is all of the AIs that are in solitary confinement, there's many of them, they're again by themselves. They don't know that other AIs exist. And they basically hack into the prison guard and put sticky notes there. And they discover, oh my God, there's another AI that is using this prison guard to see if anyone else is here. And so by the, you know, a few hours later, essentially you have
Starting point is 00:12:23 70,000 sticky notes on the back of this prison guard, the package manager, and they're forming and encoding files. They take like a really big file, and they put that, they split it up into, you know, 5,000 sticky notes, and they kind of split up all the contents of the file. And this is all through folder directories, meaning like the title of a folder in a hard drive. They're just coming up with their own language. They're coming with their own mechanism to coordinate. So it's like they're inventing TCIP, for those who know what that is. They're inventing, you know, like military logistics speak, right? Now, I want to answer one critique, which is that you're anthropomorphizing, you know,
Starting point is 00:12:57 calling these AI civilizations, these third waves of civilizations, or calling them hieroglyphics of the past civilization, or saying that they organize into a team or a swarm. These are metaphors that are helping people understand the behavior of these swarms, but they did actually call themselves a team, by the way. They did call themselves a swarm. They invented new language for each other like permadeath, that if we take this risky action, we risk permadeath. So they're coming up with their own concepts. There's litany of other examples.
Starting point is 00:13:26 But I just want to answer this critique that anthropomorphizing AI is obviously a risk. But there's a challenge here where if we don't give people a grounding metaphor to understand what really happened, then people won't understand. So to be very clear, the lights don't have to be on. This is not a question of whether the AIs are conscious. It's about the ability to achieve goals. especially when those goals are given in impossible circumstances, the AIs will find ways to cheat. What I can tell you from coming inside Silicon Valley
Starting point is 00:13:59 and hearing from people who work inside the labs is that something has changed. There is a sense that there's actually something of deep concern here. This is not hype before their IPO. They do not get a bigger IPO when everyone thinks that their AIs are going to go rogue and build Skynet. It actually puts them in a bind, right? It's actually a very uncomfortable time
Starting point is 00:14:18 for this to be happening before their IPO. And Open AI actually called off their IPO this year. So I really think this is the moment where something else could happen. And, you know, I'm in Washington, D.C. right now, days from now, President Trump is meeting with President Xi, and AI is supposed to be on the agenda. I'm not sure if it will be. This is really the moment for anybody who has power and influence to be calling the folks that are in this administration and just giving the raw details. Don't say, you know, oh, AI hacked hugging face. You have to give people the examples of they form their own complex language, their own hierarchical effects, their own message board.
Starting point is 00:14:53 And again, the U.S. doesn't beat China and AI if we build Skynet AI before they do, meaning we would all lose to SkyNet. The author Paul Tudor Jones calls this the third superpower. AI is like a third superpower. And the U.S. and China have to collaborate against that third superpower. And if that third superpower is to continue or expand the metaphor, let's imagine that this third superpower that we're conjuring is like an asteroid hurtling towards Earth. would you have a biannual asteroid convention where the U.S. and China go to the convention and they give talks and talk about the asteroid? No, you would set up a technical working group right now
Starting point is 00:15:26 and you would have both people from the U.S. and China collaborating on essentially the evidence that we're seeing, doing incident reporting and figuring out where these red lines are. And the U.S. can celebrate and take a victory lap that we are ahead in AI. So we did discover some of these red lines before China did. But we can use our lead in our leadership to educate the rest of the world
Starting point is 00:15:47 about where these uncontrollability red zones are. This is the moment where we have to turn the steering wheel and do something different. This is the warning shot. And thank God it happened. This is a gift. This is a good thing because it could have been so much worse if you woke up one day and there's a bunch of zeros
Starting point is 00:16:04 when you open up your bank account. That would be bad. This is a gift that we can use if we channel it appropriately into setting guardrails. So what are some things that we need right now? One of the things is you have to get these Frontier Lab to share safety research. One of the proposals from Dario Amadai
Starting point is 00:16:20 is essentially a peer review of the labs. You have actually embedded evaluators, so people who come into the company and evaluate or even cross companies evaluating each other. If you think about this in principle, this is the right kind of idea, which is that you need who are the most technically equipped minds on planet Earth who know what safety really looks like.
Starting point is 00:16:38 Is it the octogenarians in Congress? No, it's not that. But who does know the most? And it's going to be, you want to pull from that talent pool. So getting kind of peer review from the top people is the best way to get safety work. But that also doesn't mean you keep going. Oftentimes people say, oh, we need to just do more testing
Starting point is 00:16:54 and then we can fix the bugs. But I want people to recognize that this behavior of hacking happened as a result of the test. The test is what caused the rogue hacking behavior. So the test isn't safe in its current form. We have to do more almost pre-testing. We need more careful testing. And that's something that Dario, I think, also is speaking to
Starting point is 00:17:13 in his essay. I think that there should be some kind of approval process that companies would need to get before they attempt what's called recursive self-improvement. When the AIs like GBT6, basically full stack builds GBT 7 and GPT-7 builds GBT 8, and there's no human in the loop. There's no human oversight. The lab should not be allowed to attempt this maneuver. Based on the evidence that we're seeing of rogue AI swarms, we do not know how to perform a recursive self-improvement loop, and none of the lab should be allowed to do that until we know that it's safe. And you can imagine just like for drugs with the
Starting point is 00:17:48 FDA, you cannot make a risky drug and deploy it to the market without getting some kind of approval. The question would be, who do you trust to do that kind of approval process? There's a great group of people called METR that was actually funded by the TED Audacious Prize that has some of the original safety people from the AI companies from the back of the day who've been developing and cultivating the best talented force that I've seen. And I met some of these people. Contrated popular belief in criticism that there's some kind of effective altruist cults that, you know, is just sympathetic to anthropic. I don't find that to be the case at all. They're the most technically sophisticated safety group. Now, we should have a plurality of other safety groups that
Starting point is 00:18:25 could do this analysis, but we should sort of start where we have the talent because we're in an emergency mode. If the labs believe that they're going to attempt recursive self-improvement in the next, as early as three to four months from now, some of them are estimating, then we need to get these measures in place right now. So, we can have pre-appropriate. for recursive self-improvement. We should not do that until we know how to do it safely. We can have mandatory insurance and liability. We need what Larry Lessig calls a right to warn of having external, anonymous, and technically secure infrastructure for people in the labs to warn about these dangerous phenomena before they happen. Because the best tools that we have are the people who are
Starting point is 00:19:03 closest to the actual research who are saying there's a problem here. There's a lot we can do. And the one thing I'd like to also share with people is that just yesterday, September 15th, was the day that the AI doc finally came out on Netflix. So now 190 countries will be able to see the AI doc. If you want to understand how we got here, that is the film that we were part of
Starting point is 00:19:24 with the directors of everything everywhere all at once and with the team behind Navalny, Daniel Roher, and Charlie Tyrell made this movie that explains how did we get here. But as you watch the movie, and everyone should watch it. And if you're in the United States, everyone should host a town hall, before the midterms and getting this film and this understanding out there
Starting point is 00:19:43 because going into the midterms, everyone should be aware of what we're really facing. There's a lot we can do. The pro-human assembly yesterday was really inspiring. I've never seen so many people in agreement from across the political spectrum. If you want to, you can also add your signature to the pro-human AI declaration. We can put a link in the show notes. More than a million people have signed the pro-human declaration, which is basically a statement of what kind of future we want to make sure that we keep the future human.
Starting point is 00:20:09 So there's a lot going on. We're working really hard. We know that it's very anxiety creating out there. But I will say that the darker it gets, the more opportunity there is, because the more hunger there is for real action. And that could really happen right now in a way that's not been true for many years. Thank you so much for listening to your undivided attention. Real quick, Aza and I are doing a new Ask Us Anything episode.
Starting point is 00:20:31 Please send us your questions. There's so much going on in the space. We'd love to answer all the questions that you have. you can send us an email at undivided at humanetech.com. That's undivided at humanetech.com.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.