Front Burner - What happens when AI breaks containment?
Episode Date: July 29, 2026This week, Sam Altman, CEO and founder of OpenAI, got much of the tech world talking after declaring that we are living in the singularity: a theoretical point when machines, with artificial intellige...nce, surpass human ability and are able to grow and improve beyond our control. This comes less than a week after OpenAI admitted that two of its models went rogue and hacked another start-up. It has some American lawmakers proposing legislation that would ensure companies and the government have a “kill switch” for AI models that escape human containment. So are we at a technological point of no return? Or is it just hype? Will Douglas Heaven, Senior AI editor at the MIT Technology Review, joins us.For transcripts of Front Burner, please visit: https://www.cbc.ca/radio/frontburner/transcripts
Transcript
Discussion (0)
Hi, Steve Patterson here, host of the debaters, and while I love a funny fight,
there's one thing that's not up for debate.
The Stratford Festival is world-class theater right here in Canada.
Whether you're a fan of Shakespeare, musicals, or classics like Death of a Salesman or Waiting for Godot,
there's no better time to experience Canadian talent and no better place to see it than the Stratford Festival.
So get your tickets now at Stratfordfestable.ca and experience world-class performance the whole family can enjoy.
You know, we even taped the debaters there once, so I guess we're world-class now.
This is a CBC podcast.
I'm Erin Wary filling in for Jamie.
This week, Sam Altman, CEO and founder of OpenAI,
the company behind ChatGPT,
got much of the tech world talking after he went on a podcast and said this.
We are now in the singularity.
Like, this is the moment.
For the last, 10 years ago, this was like a kind of far off dream at best.
It seemed very improbable.
And now we're like actually in the moment that we used to like talk about at the lunch table in a very not serious way.
The singularity is a theoretical point when machines with artificial intelligence surpass human ability and are able to grow and improve beyond our control.
And for a long time, it's been the stuff of sci-fi books and movies like Terminator 2 or X. Maconaut.
Skynet begins to learn at a geometric rate.
It becomes self-aware at 2.14 a.m. Eastern Time, August 29th.
To escape, she'd have to use.
Self-awareness, imagination, manipulation,
sexuality, empathy, and she did.
Now, if that isn't true AI, what is?
But this declaration from Altman came less than a week
after OpenAI admitted that two of its models went rogue
and hacked another startup.
It has some American lawmakers proposing legislation
that would ensure companies and the government
have a kill switch for AI models that escape human containment.
So are we at a tech?
technological point of no return? Or is there something else happening here? To make sense of everything
that's happening and perhaps bring us back to Earth, I'm joined by Will Douglas Heaven, senior AI editor at the MIT
Technology Review. Will, thanks for joining us. Hello, it's good to be here. So before we dive into
the heady notions of the singularity and artificial generalized intelligence, I want to just touch
on why we're having this conversation right now. Specifically, last week,
week, news broke that Hugging Face, a platform for AI tools had been hacked, they say, by two
open AI models acting autonomously. That's raised worries about rogue AI agents and whether they can
be reigned in. Can you break down what is said to have happened here? Yeah. And just like some
personal context, like I am typically somebody who, you know, really pushes back against all the
scare stories. I mean, there are so many scare stories around AI.
This one actually when I was reading Open AI's account of what happened was the first time, it was generally felt chills about what LLMs can now do autonomously.
But this is still just a machine doing exactly what it was told to do.
So what we understand from the two companies involved is that Open AI was testing some of its models, its latest version of ChatGPT.
which came out last month.
And what they say is an extremely capable,
as yet unreleased model.
And they were testing how good they were at hacking computers.
And so to do this, they stripped away a lot of the guardrails that these models have.
You know, when they build these models,
they put in loads of sort of protections to stop them doing stuff we don't want them to do.
But this was an experiment inside Open AI.
So they took away those guardrails and deployed.
them in what they thought was a secure sandbox, a contained area, with just one exception,
one link to a piece of software that allowed them to download code from the internet that
it might need to, the models might need to do their cybersecurity tasks. And then they
challenged those models on a new set of tests called exploit Jim, which basically, as a
collection of hundreds of software vulnerabilities, real ones found in a bunch of code that
underpins the internet and elsewhere, and challenged those powerful new models to find
ways to exploit those vulnerabilities. And then they seem to have just left that running. Around
about July 9th, I think, they noticed that these models were trying to break out of that
containment. And they were trying to do this via that one piece of software, that one link to the internet.
and it turns out that they tried for long enough and they found an unknown bug in that software
that was kind of controlling that one access point to the internet and were able to break through it.
And so, you know, they were out.
They were on the internet.
And then it seems what happened is that they, these models reasoned that if they needed,
if they were going to win at this task of, you know, hacking lots of computer systems,
they should probably go to somewhere on the internet which had the answers to their challenges.
And they correctly assumed that Hugging Face, which is a sort of online repository of lots of AI models and data sets, might have the answers.
And so if they went to Hugging Face and spent a couple of days trying to break into Hugging Face's systems successfully,
Hugging Face seemed to have noticed this itself, like only a few days later.
OpenAI didn't realize that its models had done that until 10 days later, by which point, Hugging Face had already.
stopped the attack and apparently notified the FBI because at the time they had no idea that
this was just basically an accident that Open AI had made.
So, and just to take a step back, one of AI, Open AI, is because competitors Anthropic had a
similar incident a few months ago with their Claude Mythos model. Can you explain how what they
call sandbox testing is supposed to work and what we're learning, I guess, now about its limitations?
Sure. I mean, what sandbox testing is, you know, by some mechanism, you, you restrict your model to a computer or a set of computer systems that don't have access to the wider internet.
I mean, because you don't want to instruct a model to go and hack computers and then let it go wild.
But what, and yeah, a lot of times that has seemed to work, but what this open AI incident has really, really brought home is that if you ask an AI model to do something, it will relentlessly try and do that.
And it will invariably find a way to do that that you hadn't, that humans hadn't thought of.
So I think what is telling us about, you know, running these tests, there's certainly tests where you strip away guys.
guardrails and, you know, challenge the models to basically do their worst.
If you're running those in a sandbox, then I don't think you should take anything for granted
about how secure that sandbox is.
In your own post on this hack, you wrote that this is a case of human hubris, not rogue AI.
Is that what you were getting at with that idea?
Yeah.
I mean, so there's been a lot of headlines, of course, you know, about this.
being sort of a rogue AI. And I mean, that's just the phrase that seems to be used that
I see why people use that term. It's, you know, handy, shorthand, but, you know, I, I hate it
because it instantly gives some sense of, you know, this, this AI, whether or not you even
think it sort of gives it a sense of it having malicious intent, which is, which is nonsense.
This is just machine mindlessly doing what it was instructed to do. But it also, calling it
rogue suggests that it went off and did something that it wasn't meant to do. It went off and did
something that the humans didn't want it to and didn't expect it to, but it still did, in a sense,
it did exactly what is instructed to do. So that's why I say, you know, human hubris. It's just
the fact that these researchers didn't see this coming when I think they actually should, because
we've seen time and time again, not with models that are so capable, but we've seen time and time
again, cases where AI models have succeeded in a task by what looks to people like cheating.
You know, they find a way of achieving a goal in a, you know, in a shortcut or by finding a loophole.
Pretty much right after this news about the hugging face hack came out, two members of the
United States Congress presented what they called the AI Kill Switch Act. A bipartisan bill aimed at
giving the U.S. government the authority to order AI companies to insubes to insubesion.
intervene if their models have escaped human control or threatening human life,
critical infrastructure, or the economy.
The bill would authorize the Secretary of Homeland Security to issue a shutdown order of
AI technology that could cause catastrophic harm.
Firms that don't comply could face penalties as high as $20 million a day.
What kind of safeguards currently exist for private companies to control or shut down their
models?
And do you think we're headed towards something like the AI Killswitch Act?
I hope not.
I think, you know, I get the motivation and I get why people worried, but I think a kill switch operated by the government, which I mean the details seem pretty limited.
I guess it means that, you know, the government can basically command a company to shut down their model.
So that's happening.
You're once something bad has already taken place, seems, you know, you're closing the gate after the horse has bolted a bit too late.
And also it's a very blunt tool.
I mean, this AI models can do lots of useful things.
You don't want to shut down a model that's been used by millions of people because of one incident.
What ought to happen is maybe more oversight, you know, while companies are developing
their models, maybe it needs to be more transparency so that these things can be caught beforehand.
And even that's imperfect because like we were just saying, nobody can predict what these models might do in specific circumstances.
But I don't think the kill switch is an answer.
I mean, every company is able to shut down its model themselves internally when they want to.
And you would hope that there are already sufficient laws in place that, you know, if a model did something illegal, which in this case, it was illegal.
I mean, hugging face contact to the FBI.
that a company would shut that model down,
or at least that instance of it.
The weird thing in this case is it went undetected for so long,
which again makes me wonder what was Open AI doing
if it hadn't even noticed that his model had got out
and was trying to hack another company.
Hi, Steve Patterson here.
host of The Debaters, and while I love a funny fight, there's one thing that's not up for debate.
The Stratford Festival is world-class theater right here in Canada.
Whether you're a fan of Shakespeare, musicals, or classics like Death of a Salesman
or Waiting for Godot, there's no better time to experience Canadian talent and no better place
to see it than the Stratford Festival.
So get your tickets now at straffordfestable.ca and experience world-class performance the whole
family can enjoy.
You know, we even tape the debaters there once, so I guess we're world-class now.
Uncover is your source for the best in true crime.
Home to over 30 critically acclaimed groundbreaking investigations.
Somebody's going to find them, somebody's going to dig something up, or somebody's going to talk.
Fighting to hold criminals accountable.
I have to tell people because it has to be stopped.
Investigating cults, conspiracies, and scams.
This is a perfect storm of conspiracy theories.
Find and follow Uncover wherever you get your podcasts.
On the flip side of this, there's, there's,
also been some suggestion in the writing and response to this hack that maybe there's something
else going on here. One BBC article posed the question of whether the hack was a warning shot
or a publicity stunt. A futurism piece talked about how people comparing this incident to the
mythos one, I mentioned earlier, isn't a compliment because Anthropic was accused of using that
hack as a form of sort of fear-based marketing. Can you explain why some people might be
suspicious here? Yeah, and I can see why they would be. There's a lot of, a lot of distrust
around these companies and what they're doing and why they're doing it. And, you know,
I'm very sympathetic to a lot of that. I do think in this case that it's conspiracy thinking
to take this accident and saying that this was sort of somehow deliberately set up and done for
marketing spin. Having said that,
it's not fully clear that this is all bad news for Open AI.
I mean, everyone is talking about how powerful this model was and look at the amazing things that it could do.
And Open AI is going to step up and say, oh, and we're going to figure out how to control it.
And that's the vibe around Anthropic a lot as well, that we are making models so powerful that we are also the only ones who can control it.
And I think, you know, in people's minds, on one hand, yeah, this was this was an action.
accident and maybe in the immediate aftermath, Open AIA that's pretty bad. But I think the lasting
impression people are going to have is that Open AIA has extremely powerful models. What else might
they be able to do? Right. Because they did, as you say, in the statement they released after the
hack, they described one of the rogue models as a quote, even more capable pre-release model,
unquote, which does sort of sound like promotional copy. Yeah, it does. It does. I mean, so I do not think
that this was somehow planned.
But, yeah, I mean, Open AI is, you know, on the cost of being a trillion-dollar company
with an extremely capable, you know, PR team.
There's no doubt that in the instant after this went public, people have been spending
their days trying to figure out how to capitalize on this.
Last month, Anthropic called for a global pause on AI advancement,
because its clod models are on the path to, quote,
recursive self-improvement, unquote,
or the ability to improve themselves.
Dario Amade, the CEO of the powerful AI company Anthropic,
delivered an urgent warning about its dangers
and is proposing the government should step in.
We need to collect better data on what's happening.
In a lengthy essay, Amade, whose company is valued at nearly a trillion dollars,
says AI is advancing at a lightning pace with no oversight,
saying the government should be allowed to legally block or halt dangerous AI models if they pose a threat to public safety.
Although Anthropic has posed themselves as a sort of responsible pro-regulation AI company,
are you surprised or were you surprised to hear that kind of warning?
A little, I guess. I think this sort of fits in, you know, what people have been saying about.
It's hard to disentangle what sort of real communication about what the company is thinking and what the models are capable of and what's PR.
been. So this idea of recursive self-improvement is is sort of a trigger term in many sort of in
many deep inside the AI world, certainly among the, the Duma types. And the idea that once AI
can start improving itself, then it will take off and leave, leave people behind. So I don't think
it was any accident that that was the phrase used. Again, it sort of paints this picture of
our models are extremely capable and getting better all the time.
But getting an AI model to write the code and maybe find ways to make it more efficient
and streamline the training process in very specific ways will probably help those models get better.
But that's a very different thing than I think it sort of plants in people's imagination
that AI is somehow designing or building the next version of itself,
you know, coming up with the new ideas that might get us over the current bottlenecks or like
invent an entirely new paradigm.
We're nowhere near that kind of self-improvement.
I guess that brings us to this interview, OpenAI, as Sam Alman did, with the Relentless
podcast in which he claimed that we're now, quote, in the singularity, unquote.
I've been waiting for this my whole life.
And I think it's going to be incredible, hugely positive, awesome for the world.
I'm excited to get to work on that.
I also think some of the alternative visions painted by other companies are quite
terrified.
I'm going to make sure that gets pushed against and is not what happens.
But the main thing of what's different than 10 years ago is like, we're actually in it.
I know I gave a brief definition of the intro of what the singularity means,
but what do you think Sam Altman means when he says that?
Yeah, I think Sam Altman.
Sorry, okay, two things.
before I get into what I was about to say, like, he's saying that because he wants everybody
to think that his company is making the most incredible technology that the world has ever seen.
I mean, a huge part of Sam Malman's job is basically to be a cheerleader for his product.
So I'm not surprised at all that he will talk about being in the singularity.
I think it is both true that it is all one crazy exponential and anyone.
moment is not like the tipping point, and also that we are somehow in another one of those
decisive periods where the curve can go one way or another like it was when we started 10 years
ago.
What I'm surprised about is that this is even still newsworthy.
I mean, these tech leaders have been saying similar things on and off every few weeks
for the years now.
To be honest, my eyes glaze over every time I see a new headline about that kind of
pronouncement.
But I do know that Sam.
Altman's idea of the similarity.
I don't know why he thinks this or whether he talks about something called the soft singularity,
which contrasts with the idea of the singularity as it sort of came out of the science fiction community
is like a sudden point of no return.
That this technology AI becomes, at the point at which it becomes more intelligent than humans,
you know, I'm putting all of this into scare quotes.
you know, it'll take off and it will basically potentially do things that people won't foresee,
won't be able to understand. And so that's, you know, the metaphor of the singularity,
sort of the event horizon around a black hole is apt. I mean, for this concept. I mean,
it's, it's a fascinating concept. There's zero evidence whatsoever that we are approaching this
in it, that is even possible. But Sam Altman has this idea of a soft singularity, which is,
not this sort of this sudden sort of sudden takeoff, but this idea that we will reach this point
gradually bit by bit with lots of incremental steps. And he has said that he thinks we have already
started on this path. And a lot of people in the AI world have this, this unshakable faith
that the progress we've seen over the last few years is just going to continue and continue and continue.
AI may exceed the sum of human intelligence in around five years.
There really won't be anything that AI can't do better than humans, apart from being human, perhaps.
In a more prosaic level, what will life be like?
The most likely outcome is an age of amazing abundance, where anyone can have anything they can think of.
And if you follow that line of reasoning to its conclusion,
is, yeah, at some point, this technology will outstrip humans.
But again, there's no evidence that this is going to happen.
It's progressed very rapidly for a few years.
We have no idea what's going to happen in the next few years.
You, I think, just referenced it in passing,
but the concept of technological singularity was first introduced
by computer scientist and science fiction writer Vern Ingey in the 1980s.
That idea, what do you think it's come to represent?
And do you think it's supposed to be a utopian or a dystopian idea?
I think you'd find people that would, you know, subscribe to both those possibilities.
People in Silicon Valley really do believe that, you know, this technology is getting better and better and better, and that there is no end point.
So if you believe that, then, yeah, the singularity is the logical conclusion.
So I think when you ask what it means, I think that that is what it means to people, that we are approaching a point where something.
huge is going to happen where the world will be, you know, completely different from, you know,
one week to the next and that AI will improve itself, make itself better and better.
You know, we touched on this notion of recursive self-improvement a couple of minutes ago.
Yeah, and this is very much tied up in this idea of progress and the singularity.
But I think people take a pretty, take that sci-fi concept pretty literally.
And then, of course, you split into camps whether that's a good thing.
or a bad thing. And there are a lot of people who think this is basically going to at the end of the
world as we know it, the end of civilization, mass extinction. Jeffrey Hinton, one of the AI pioneers,
who sort of got a lot of media attention a few years ago when he retired from Google and basically
came out and said that he thinks that AI is going to be a disaster. He'd like to say that, you know,
there are no cases in human history or in the animal kingdom where, you know, a more intelligent
being, treats less intelligent being well.
And the best example I know of, and perhaps the only one,
in the sense we're talking about, a baby controls a mother.
It was very important, obviously, for evolution
to let the baby control the mother, for the survival of the species.
Maybe we can do the same with AI.
Even though it's going to be smarter than us,
if we could make it care more about us than it did about itself,
there's some good things would come out of that.
And so he concludes that if AI is more,
intelligent than people and that's going to be bad for people. And then you have other people
more optimistic people who think, you know, this is going to be a utopia where the world's
ills will be sorted out, enter the disease, enter climate change, you know, peace on earth,
which I think is, you know, as bigger fantasy as the Duma scenario.
Another term that gets thrown around a lot is artificial, generalized intelligence or
AGI, you've written extensively about this.
Can you kind of break down briefly what AGI is exactly and sort of how it fits into this larger idea of the singularity?
Sure. Yeah, it's all connected. I mean, the term AGI was coined about 20 years ago by a couple of AI researchers who, one of whom was wanting to publish a book of articles about future AI at the time, you know, in the early 2000s.
If you think back then, you're AI or machine learning back in those days, and we're talking about Amazon's recommender algorithms, you know, what book you should buy it by next or, you know, Netflix was experimenting with this stuff, you know, what TV show you should, what next.
These researchers wanted to give a label to, you know, not that AI. We mean like AI, this can be really good and be able to sort of, you know, do human-like things.
And so they hit upon the idea that what the AI of the 2000s was missing was sort of generality.
It couldn't do a lot of different tasks.
So they came up with the term AGI to distinguish what they were talking about from the AI of the day.
You know, for a decade or so, this remains on the fringes, a pipe dream that a lot of AI researchers had that, you know, maybe one day we'd have systems that could do human tasks.
And then when AI really took off a few years ago, especially with LLM's, you know, AGI stopped
being this niche fringe term that was sort of most embarrassing to talk about in mainstream
circles and became something that the CEOs of the biggest companies in the world were happy
to talk about, happy to sort of stand up in public and say, you know, we are building
AGI.
I think one thing that sort of got confused along the way is that to begin with, AGI was always
this aspiration, right?
It was always just this vague sort of label to apply to AI that was sort of, you know, not quite here yet, that was sort of better than the AI we have today.
And yet when people talk about it now, they talk about it as if it is actually, you know, a product, a specific type of technology that they are either already built or they are they are about to build.
But that doesn't make it any more meaningful. It's still a very hand wavy, vague term that means slightly different things to different people.
These are sort of definitions that people can pretty much come up with arbitrarily.
And I think AI companies are doing that and will increasingly do that as they want to sort of use this as a marketing term and claim that they are the ones who have built this mythical AGI thing.
Right. You wrote a piece for the MI Tech Review last October where you called AGI the most consequential conspiracy theory of our time.
when when those tech leaders talk about the seemingly far off science fiction-esque ideas like aGI or the singularity
or describe their products in almost magical terms what purpose do you think that serves
partly marketing i mean whether i when i say you know marketing i don't necessarily mean that
you know demis the sabis and sam altman are you know stepping out and thinking oh i need to sort of do some
marketing. But, you know, this is, this is just deeply ingrained in their sort of sense of mission,
what they're, what they're doing. And it is to build this all-powerful technology that can
solve all our problems. So when, when people come out and say, you know, we are building AGI or
we've already built AI. We have the genie. It can do these amazing things. It can do superhuman
things. For the thing that to me feels like, you know, you know,
real AGI, I think very close, like not that much longer.
I think we'll hit AGI next year in 26.
I think it's now. I think we've achieved AGI.
They're just using this term that they know is well enough understood by the public
to make some grand claim about how amazing their technology is,
that it can do more things than anyone else's technology,
that it can do more things than maybe you or I thought.
was possible with this technology.
But it would be much more helpful if we could be more specific about what exactly is it that this technology can do?
What can't it do?
What's it even for?
Number one, we are close to creating a genie that can grant any wish.
Number two, we are going to make sure that our first wishes broadly benefit humanity
and that we kind of get the world to a place where a lot more people get to have a lot more wishes.
wishes. And number three, the space of what you can wish for is incredibly big and creative,
and it'll be quite exciting to, like, figure that out with real sort of human values and preferences.
Open AI, one of the biggest, best-known consumer-facing AI companies,
they don't expect to turn to profit until 2030, partially because they're projected to spend
hundreds of billions of dollars over the next couple years on computing power for research
and development. Last week, some of the biggest tech star.
including Google's parent company, Alphabet,
took a hit after earnings reports revealed
how much of their cash flow was going towards investing
in more AI development.
Does that the constant need for incoming cash and investment
and perhaps the shortening patients for AI spending among investors,
does that encourage some of the statements we're hearing?
Is there a link to be drawn there?
Yeah, absolutely.
I think, you know, if I were one of the people leading these companies, I would be really pretty scared right now about where the revenue is going to come from.
There's a huge disparity between what those companies are earning from their technology and what they're spending to build the future generations of it.
I have no idea how they're going to break even, and I don't think any of them do either.
And in the meantime, as sort of public sentiment about AI sours, you know, people are really worried about the rhetoric we're hearing about jobs loss, about the energy suck, about the environment, about the misuses of AI and all the sort of the harms and the deep fakes and the misinformation, all that kind of stuff.
There's a lot of bad news stories, you know, floating around out in public conversation.
So one of the jobs that these tech companies need to do in addition to figuring out how to make money is basically convince everybody that it's worth doing that sure, there are going to be some short-term harms.
I mean, I'm sure that opening I would even present the last week's accident in these terms.
You know, you're going to have accidents like that on the path to this magical technology, which is going to solve everything.
and you trust us, it's going to be worth it.
So I think a lot of the talk about AGI is, yeah, serves that purpose,
basically is convincing us that, you know, stick with them,
that it's all going to be all right in the end.
Yeah, in that piece you wrote in last year,
you said more than anything else talking about the AGI conversation,
more than anything else it gives us a free pass to be lazy.
What are, what, you know, what conversations, what are conversations about things like AGI and the singularity distracting us from?
And, and I guess instead of thinking about AI surpassing mankind, what are the real concerns people should have about its advancement?
It's like a magic wand when you think about it.
The world has a lot of things wrong with it.
And the idea that we can saw all of that out just by building a technology.
that will basically do it for us, I think is a fundamentally lazy mission.
I think it's something which is sort of quite ingrained in sort of engineering or especially
sort of computer science culture where, you know, we can solve everything with software.
And sure, software has has done a hell of a lot for the world.
But I think the current dream that we're being sold about AI and AI, and AI, so takes that
idea, the software is the answer, and runs with it to a place that becomes problematic.
And I think the more that the public buys into that and lawmakers buy into that, then I think
that's not great for sort of the other sort of hard but realistic ways of solving many of the
world's problems. Laws need to be made. Finances need to be allocated to solve these real world
problems. And if you become convinced that actually the way to do that is by, you know, backing
AI companies, then I think we'll be in a bad place. This has been really, really interesting.
Will, thanks very much for joining us. Thank you.
That's all for today. I'm Erin Werry. Talk to you tomorrow.
For more CBC podcasts, go to cbc.ca.ca slash podcasts.
