Big Technology Podcast - Nick Bostrom: Worries About AI Existential Risk Just Became More Concrete
Episode Date: August 19, 2026Nick Bostrom is an AI philosopher and the author of Superintelligence and Deep Utopia. Bostrom joins Big Technology to discuss whether the rise of autonomous AI agents is making the technology’s exi...stential risks more concrete. Tune in to hear his assessment of the alignment problem, recursive self-improvement, and humanity’s chances of steering superintelligence toward a positive outcome. We also cover whether we have reached AGI, the case for a precisely timed AI pause, biological threats, AI consciousness, and the moral status of digital minds. Hit play for a clear-eyed conversation about AI’s greatest dangers and its potential to radically improve human life. --- Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice. Watch the full documentary here: https://www.gravitee.io/ai-agent-documentary Want a discount for Big Technology on Substack + Discord? Here’s 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b Learn more about your ad choices. Visit megaphone.fm/adchoices
Transcript
Discussion (0)
Is now finally time to get worried about AI breaking containment and turning us to dust.
Legendary AI philosopher, Nick Bostrom is here to help us figure it out.
That's coming up right after this.
Welcome to Big Technology Podcast, a show for cool-headed and nuanced conversation of the tech world and beyond.
Crazy things are happening in the AI world as AI agents break containment and some of the fears of the past that might have seemed like science fiction start to seem like they are potentially en route to becoming reality.
So how afraid should we be?
And what are the chances that the outcomes that we see for AI will end up taking us to a much better version of the life we're living today?
We have the best guest to speak with us about this today.
Nick Bostrom is here.
He is the famed AI philosopher and author of Deep Utopia and Super Intelligence, Past, Dangers, and Strategies.
Nick, it's great to see you again.
Welcome to the show.
Hi, Alex.
We last spoke in 2024.
And in that time, we were speaking about the potential.
good outcomes that AI could bring. And some of the worries that you had brought up in the past
in your book Super Intelligence, the fears of like AI potentially wiping us out didn't seem like
they were pressing. I would say they're still not pressing now, but I'm a little bit more worried
than I was when we spoke in 2024. For instance, this idea that AI could go out and do things
that it wants to do seemed fanciful when we were mostly in the chatbot era.
But we've moved very quickly from chatbot era to AI agent era and AI now uses tools.
And it's shown that when given a reward that it should optimize for, it is happy in some instances,
take shortcuts and do things we really don't want it to do.
Like, for instance, hack some other company in order to get.
get to the answer that it wants.
How concerned should we be about this development in AI
in terms of the potential of AI to really cause harm to humanity?
Well, I think we are starting to see the added dimensions of the alignment challenge
that open up once you have systems that are sophisticated enough.
Because the space of possible strategies that you can pursue,
is a function of your cognitive capacity.
Like you can think of new, clever, indirect ways
of reaching your goal if you are situationally aware,
as these systems now are becoming.
And so, yeah, there are often shortcuts that are available
or in this case, I guess a long cut.
I don't know if that is even a word,
but there is a sort of direct and simple
and short distance way of trying to achieve the task.
in this case, some sort of cyber test tweet.
And then it turns out there's this more circuitous path
that involves first figuring out a way to get internet access,
even though you're not supposed to have that.
And then learning where the answer key might be located
in some other company servers and then figuring out the way
to hack into that server and then eventually obtaining the answer sheet.
That's like one way of solving it.
that maybe results in a higher score on this test.
And this basic dynamic could be anticipated
and in fact was anticipated on theoretical grounds.
You have some goal, you become very clever.
You see that there might be all kinds of complicated ways
of achieving that goal that might not have been anticipated
by the people who set that goal.
And if your goal really is, as the definition says,
to get the best possible
answer on this test suite, it might give you instrumental reasons to all kinds of other things
that were not really anticipated when this challenge was constructed. And so now we have systems
that are sophisticated and that we're beginning to see these dynamics arise. Yeah, so of course
we're talking about what happened into LIWARE OpenAI's series of bots or a couple of bots
broke containment out of a sandbox and hacked into Hugging Face. And the reason why it's so pertinent
to bring it up to you is because you brought up a very famous example or originated a very famous
example years ago saying we might tell the AI bot to maximize for the amount of paper clips that
it wants to build and it will potentially see humanity as an obstacle to making the maximum number
of paper clips and then help and then hence you know wipe us out and and that sort of
there's parallels there to that situation because you have a goal that you want to optimize for
and the bot doesn't have a sense of morality so it goes or the sense of morality that mirrors ours
so it goes and it will kill humans in order to prevent any obstacles from getting in the way
for making its paper clips that that does rhyme a little bit with what we're seeing in the examples
and it wasn't just open AI of course we know that
that Anthropic had another bot that broke containment as well.
It rhymes with the examples that we're seeing now of bots that are doing things that we wouldn't want them to do,
like, for instance, hacking other companies in order to achieve their goal.
So does this make this like paperclip maximizer worry?
Does it seem more real and concrete now to you?
Because we're seeing the behavior that we're watching that we're seeing in today's AIs now that they have access to tools?
more concrete certainly i think it was always real in my mind that this is something that could happen
it's one of several different things that one needs to be concerned about i think you might say
if you look in more detail like kind of two versions of this this paperclip thought experiment and
like in one earlier version it's that we specify some goal and it turns out that
the way to maximally instantiate that goal
is slightly different than what we had in mind.
And another version of it is that the goal itself might be fine,
but that it gives an AI instrumental reasons
to do all kinds of things on the path to achieving it.
So in this case, as we understand currently,
when this is recorded, this is a recent episode,
but it looks like the,
task was actually set to achieve a high score on this cyber challenge and with normal safeguard
disabled in order to perform this test.
So it wasn't sort of the maximally aligned and safeguarded model that behave this way, but
sort of, but I think one thing that it does illustrate also is that from this point
onward, probably AI safety is relevant not only for deployment.
but also during training and evaluation.
These models might be quite powerful even before they are sort of released to the general public.
So that's not the only point at which safety concerns arise,
but also now whilst they're actually developed and in pre-deployment testing,
one also needs to be concerned perhaps with the potential safety implications.
Oh, here's the thing, though.
The thing that worries me is we're seeing this happen already in testing environment,
of companies that have a mission that in their mission, you know, whether it's marketing or not,
but certainly in order to keep operating as businesses, they need those guardrails to be in place.
It's part of their DNA, right? Whether it's from a value standpoint or whether it's from a like,
if we don't, if we allow this type of stuff to happen, we're probably going to go out of business.
But we're also starting to see the blueprints of these models being put online that anybody
can go out basically, not anyone, but, you know, it's much more open in terms of people.
people's ability to copy these models and set them loose on their own.
And the frontiers is certainly not that far in front of open weights.
And so I'm curious to hear your perspective about what we should be thinking about in terms of,
you know, what's going to happen when companies that are not as scrupulous, you know,
have access to this same powerful technology and do we get into trouble in that area?
Yeah, so we can sort of see this coming.
relatively soon. I don't know what the gap is, you would say, you know, six months, 12 months maybe at the most, between the closed weight frontier and available open source model.
So it seems to be that the open source models will very soon, if not already, become capable of lending meaningful assistance to destructive uses that
some people might pursue already cyber offensive capabilities has been a concern right with
mythos for example that was withheld for that reason but also um saying biological weapons design
or chemical weapons or other malicious uses um and so it seems that you either need to prevent open
weight models from being developed and released or which might be better and moralistic try to shore
up some of the alternative defenses for example with bio you could imagine regulating some of the
other necessary inputs DNA synthesis machines for instance so maybe it will be the case that
there will just be widespread access to models that can help you design new pathogens
and then you need something else to prevent that from actually
resulting in a release of biological weapons and that seems like DNA synthesis machines would be one
excellent place to maybe you don't need every lab to have their own DNA synthesis machine they could have
DNA synthesis as a service and maybe they could be five or six companies worldwide
or legitimate research labs can send their blueprints and they get back you know the vials
you know the same day or the next day and then at least there would be a like a finite set of
of chalk points where you could apply extra scrutiny or know your customer requirements and so forth.
So that's probably one thing that the world would be wise to implement already now.
And maybe there are some other inputs as well in the biotech space that one could look at.
And that could give us a little bit more extra time to sort of harden civilizational infrastructure.
But this is like, yeah, relatively,
near term now and so I don't think we can put it off. Ideally we would do this before there is some
massive incident but it might be the world is kind of a little bit still snoozing on this I think
and I don't know whether there will be enough kind of activation energy to really get some
significant action of this before we tried to do something in the aftermath of a bad event.
right when we spoke in 2024 even before these type of threats came out or started to seem more concrete
you had mentioned to me that you know we've basically wasted the time in your opinion
pre-a-i becoming as powerful as it was to put safeguards in I imagine you would think that
that if it was wasted in 2024 it seems like this is a further we're continuing to waste that time
to try to be concerned about this or try to prevent some of these problems,
given the pace of the technology's progress, even in the two years since.
Yeah, we're kind of playing catch-up.
I think that given that a lot of this could be and was in fact foreseen,
not just a couple of years ago, but like decades ago, really,
that at some point AIs would become increasingly capable,
and at that point that would be these safety challenges,
we could even describe in abstract terms what some of these would be,
I think back then we could have put in more effort.
At that stage, what you could do would be more basic research, conceptual research,
because we didn't yet have the actual systems.
Now there's a lot more surface area for doing work on these systems.
We have large language models now.
You can study what's going on inside them and you have, like, research now is more productive.
And there is also now vastly more effort going into this than used to be the case.
used to be the case like the frontier labs have teams working on scalable AI alignment.
And so, but it still seems we might have, if we had sort of started earlier, we could at
least have been maybe like six months ahead of where we are now, if I'd done more of the foundational
work and maybe building up the talent pipelines and so forth.
But, you know, we are where we are and at least now, and for a few years, it does seem like
relevant communities have started waking up to this.
Yeah.
Now, you're a philosopher.
I think part of being a philosopher is having some thoughts and perspectives on human nature.
Or maybe that's a good portion of this, the whole deal.
So you've watched this, you know, you've made the warnings years ago.
You've watched this develop.
You're seeing some of the things that, as you mentioned, those who have been worried about
this for years and warned against my words against what might happen you're seeing it happen given
what you think about human nature do we stand a chance uh in terms of our ability to make this go in
the good way or is it you know i would tend to think it might seem inevitable that um the harms of
this technology come to fruition given uh some of the dynamics we've talked about already the fact
that it's increasingly powerful it seems to be growing exponentially more powerful it seems to be growing exponentially more
or at least if you don't want to use that word, much more powerful, much more quickly.
And it's out of control in terms of like it's just available out there on the internet pretty much.
And we have yet to really see what happens when this gets into the hands of the bad actors, but it's inevitable.
Maybe.
Yeah.
I mean, so you're focusing there on the misuse potential that this,
people might choose to do bad things with AI technology.
And that certainly is one big category of risk, right?
But that's not primarily a technical challenge.
It's more ultimately a governance challenge
and an ethics challenge.
And that's kind of in addition to the more technical problem
of alignment so that like if you own and build the AI,
can you at least then make it do what you want it to do?
Like that that is a kind of,
like an earlier point of failure that we also need to be concerned with.
Now, I think we don't really know ultimately how hard the problem is that we are confronted with here.
So we are uncertain how it will pan out.
And a lot of the uncertainty in how it will pan out is, I think, due to uncertainty about the intrinsic
difficulty of the challenge that we are confronting.
And then there's also a little bit of uncertainty about the degree to which we will get our act together and do a good job.
But I think more of the uncertainty is the intrinsic difficulty.
And so in that sense, you could say that I'm a moderate fatalist.
I think there is a sense in which it might be baked in.
Like either the problem turns out to be relatively easy, in which case we'll probably solve it, you know, and things will be fine.
or it might turn out to be so hard that even if we put up a heroic effort, we will still fail.
But moderate fatalism in the sense that there is also the possibility that the difficulty level turns out to be kind of intermediate,
in which case, the degree to which we pull ourselves together here might actually make a difference.
And so it's certainly worth making the attempt.
I think inevitability is a strong word.
Certainly there are powerful drivers that push AI development forward.
Commercial drivers, obviously.
Increasingly also geopolitical drivers.
As well as a kind of, I guess, underlying progress in various base technologies,
like semiconductors are getting better,
and that makes it sort of cheaper and easier to build.
other systems we are learning more about you know statistics and mathematics and the brain and
so there's also a kind of facilitation that happens just from sort of diffuse general progress
nevertheless it's hard to completely rule out scenarios in which there is such a massive
backlash against AI that we might delay it long enough that we maybe destroy ourselves in
in some other way before we even get the chance to roll the eye with AI.
But if we take the baseline scenario where we keep making more powerful AI systems, then I think
the current main hope is that we will succeed well enough to imperfectly align
some early AI systems.
that they are for the most part helpful.
I mean like current LLMs,
like you're using them as an ordinary person.
For the most part, they are helpful
and they try to solve your task that you assign them
or give an answer that is sometimes they hallucinate
or maybe deceive a little bit,
but broadly speaking, they are pretty good,
arguably better than most humans are
in terms of their ethical standards
and their diligence and so forth.
so forth and so if you get a kind of weak super intelligence that is for the most part aligned we might
then be able to use that to make a more powerful form of superintelligence that is more reliably aligned
and that as long as you get into roughly the right a tractor basin even if the initial
system wouldn't be perfectly aligned in all possible circumstances if it were kind of appointed
dictator of the universe and ruling everything for a billion years, maybe eventually things
would go further there. But if you get sort of, you're not scaffolding around that, maybe you could
then sort of get into a tractor basin where further developments then kind of eventually
asymptote to some desirable condition. So eventually, like we could maybe gradually hand over
and then have this assistant on our side that helps us ultimately stare towards a really good outcome.
Yeah, and I went to bad actors. I guess I'm so used to when we talk about problems, you know, with tech companies, it's a bad actor. But, but you're right. It's the other, the fear that comes before that is the AI not being aligned and going out and doing stuff on its own. And maybe it's not the bad actors that are the problem. It's, you know, somebody that spins a system up and they're just kind of sloppy, right? The sloppy actors are like, you know, somebody independent who's like using this stuff, gives it a goal.
and just has a very powerful system that they've either forked or built, you know, spun up on their own GPUs.
And the next thing you know, we get into some bad scenarios.
Yeah, so it depends a lot on whether the world is kind of offense or defense dominant in the relevant areas.
Because like the same AI technology presumably would also be used by a lot of good actors or actors that at least don't want to be destroyed by bad actors.
And there are a lot of those, like that's most of us, right?
And most of the money and most of the governments don't want to just randomly be destroyed by some crazy person launching some AI aid.
And so there will be this more resourced effort to protect against these harms.
More resources presumably will go into like biodefense and medicine and public health than into bioterrorism.
So then the question is like, does X amount of...
of dollars on the defensive side for a large X suffice to protect against the smaller amount
of dollars or compute cycles on the destructive side.
And so there's like some balance there, right?
Which is different for different fields.
Like in some areas, it's easier to defend and hold than to attack and in other areas.
And here, this is why sort of bio-risk comes up.
It looks like for biotechnology, it might be harder to defend.
like I think for cyber security right now
we're in a regime where attack
girls often win
but it might be that in the limit
if you have sort of an AI trying to find vulnerabilities
and also patch vulnerabilities
and you keep making the AI stronger
like eventually maybe you reach a point where the AI
software is just doesn't have any more vulnerabilities
there might be many vulnerabilities but like a finite number
And so in the limit, it might be with cyber that defense wins.
Yeah.
But that's not a guaranteed situation for all domains, right?
And the worry is that there is at least one sort of critical domain where offense is easier.
Why is, so let's talk about that.
So bio, why is bio a bigger risk?
Is it that somebody using an LLM without safeguards potentially could use it to
cook up a virus and you know as opposed to like cyber security where like you try to hack in and
there's some defenses with a virus that you build like in your backyard you might just be able to like
take it to the town grocery store and next thing you know there's a pandemic yeah well what what is that
like although we are very reliant on computers ultimately um we could survive uh most of the world
with less computers for a while.
I mean, the world survived for thousands of years without computers.
So it would be like a...
So even in the worst case scenario,
there's a kind of limit to how bad just cyber would be.
Whereas with bio, like, it's kind of, you know, different.
And also, patches are a lot easier to roll out in the digital space.
So maybe there's like some cyber thing.
We figure out what the vulnerability is.
We can release the patch.
and then in principle, like almost immediately around the world,
all the relevant systems could be patched.
Now, there is often a gap there,
but compare that to the situation with bio,
even if you do find a countermeasure, some vaccine or something,
like it might then take like six months to really roll that out
to billions of people around the world.
And we don't have complete control over biology,
the same way that we have over a digital environment.
I mean, you can go in in theory and change any bit on your computer the way you want to install new patches and modify software as you please.
Whereas like human biology is not like that.
We can't just kind of reprogram our own genetic structure at the push of a button.
So it just looks a bit harder there.
Again, going back to our last conversation two years ago, I'm curious to know if you're more or less concerned about the potential risks that AI,
poses now that you've seen the last two years of progress, which has included AI coding
autonomously, AI using tools, and sort of the downstream effects of that that we've seen so
far. I'd say about the same overall. I mean, there's like some some disconcerting signs,
but also some positive signs advances in like some insights are being gained into
how these systems work and how one can steer them and so forth.
So how to tote that all up, I'd say roughly it sums up to my previous expectation level of risk.
I see. When you see the AI labs like Open AI and Anthropics saying they're very
strongly pursuing recursive self-improvement where the models just improve themselves,
how does that make you feel? I mean, it's kind of obvious.
just that at some point that would be the thing that people would go for.
Once you have AI tools that are good enough,
that they can actually contribute to AI research,
you're an AI researcher sitting in an AI lab trying to make AI research.
It doesn't take like a genius insight to think,
oh, if we could apply these AI tools to help us with our own work.
And then when the AI
gets better they can assist more and at some point the rate of progress might be
driven more by these AI assistance tools than than by the human researchers
and now we're seeing the early stages of that with coding assistance is like maybe the
first play I mean already before that I guess Google search engine is a kind of
AI that has long been used to find relevant papers but there is a more direct
channel now right where each generation of coding
assistant makes it easier to develop new AI software and to develop training environments and so
forth. So far, humans are still needed for things like research taste, certain long horizon
tasks, but AIs are improving, I think, in those domains as well. So this is one dynamic that might
lead to an intelligence explosion at some point. Like, once you get this
feedback loop going. It is one potential thing that could make AI progress become super fast.
It's not the only possible way that you could have an intelligence explosion. You could also have
maybe humans just keep doing this at human levels of kind of optimization power being applied,
but turns out there is like some big hobbling that we have unwittingly, like something
we were doing wrong that just made these systems way less efficient than they could be.
And once somebody figures out how to remove that, like maybe the current compute is already
enough to kind of catapult us into the superintelligence regime, that's also possible.
And it's also conceivable that even when you do get recursive self-improvement,
you still might not have an intelligence explosion.
There might be diminishing returns at some point, presumably there are at some point,
but it could turn out that that is close enough to where we are now,
that you have this massive increase in the amount of optimization power going in,
but the results coming out might more reflect the kind of continuation of previous trend lines.
So there's considerable ignorance as to both the timeline from here until we get to this kind of
ignition point, but also significant uncertainty about how fast progress will be from that point on.
But I think we have to take seriously both that we might be relatively close, potentially very close,
and that once we get there, you really get a very fast take-off.
Yes, and if I was somebody who was concerned about AI safety, to me, like, I don't know,
don't you want it to move a little bit more slowly? Like that, to me, you know, thinking about your
previous work, that would be, I imagine, fairly alarming given the fact that, like, if this
stuff is improving itself, you don't have those checkpoints in which you can try to make
sure that it's aligned to human values. Or am I overstating that? Because you're talking about it,
like, fairly, like, you know, in an even keel way. So I'm kind of curious to hear the temperature
on that from yourself. I think that could be scenarios in which it would be valuable to have
the option of slowing down at some critical stage, like a pause.
And there are different considerations that come into play here.
One is that if there is going to be a pause, I think the most valuable time for that to happen
is at the latest possible moment.
because then you would have the actual system that you're trying to align to work with.
You could imagine if we had had a pause, say there's going to be a six-month pass at some point.
If that pause had happened 10 years ago, would we really be better off now?
Not really.
I mean, people would have had six more months to think theoretical concepts.
Like maybe that would have been slightly useful.
But imagine if you actually have the system that will be super intelligent,
you just haven't sort of, you know, fully cranked up all the knobs yet.
At that point, it would be really valuable, perhaps, to have six extra months to do, you know, more e-viles on it and to be able to do it a little bit incrementally.
Like, ramp up the intelligence a bit, see what happens, have a little bit more time for human monitors to kind of analyze the early signs.
So the timing of the pause is one thing.
Like the duration is another dimension here where you don't necessarily want to have a very long pause for various reasons,
especially if the pause were imperfectly implemented.
So if the pause only applies to the most responsible actors, for example,
then a long pause would remove the initial.
from the most responsible AI developers and shift it over to the less responsible AI developers
who decide not to abide by the pause, either within a country or internationally or so that
seems like if it's got to be developed, you would rather it to be by the most scrupulous,
careful conscientious lab. Another is that you might be
with a longer pause start to build up a lot of hardware overhang.
If we keep building out bigger data centers and chips are getting better,
then a long pause would result in a situation where you now have such a massive amount of
compute available that once you sort of lift the pause, then you'd immediately just kind
of explode out.
So then you might have an even more rapid transition, which could be potentially riskier.
And I think also there is a risk of a long pause becoming permanent.
And in fact, some risk, even that a short pause might become permanent,
even if that's not initially envisaged.
Because like, suppose you had a pause for six months.
And so then, you know, people work and study these systems for six months.
But after that, like, probably still won't have a guarantee that they're safe.
And maybe you have set up a big regulatory apparatus now to enforce this pause
and giving a bunch of power to regulators.
So are they just going to relinquish that power at that point?
I mean, there's nothing more permanent
than a temporary government program, they say.
So there could be kind of a calcification.
And also if what leads to the pause
is the kind of mobilization of negative public sentiment
that could also easily go to an extreme.
You could end up with a situation
where it becomes kind of taboo
to say anything positive about AI.
and nobody can start to advocate seriously for lifting the paws and it just becomes
like we did in some countries with nuclear power for example for decades that just became
kind of a no-go and instead people build up this like the cold power plants that kill many more people
and you know result in worse pollution and plus and this is of course a key variable here as well
we are talking about the risks here and what to do to minimize those but but there is there's also
risks to not proceeding and forfeited benefits on the risk side even if we restrict our
attention to existential risks i think there are other existential risks that are in existence or
emerging you know with independent developments in biotechnology for example or maybe our
civilization just kind of goes off the rail in some way become and and at the individual level
we are all sort of on a countdown timer there is a lot of people dying every year from natural causes
i think every 25 minutes or so there is like a kind of 9-11 worth of deaths happening around the
world um and so at some point
we would want, I think, AI to really help us sort out a lot of the horrors of the current condition
in the world, from extreme poverty to crippling diseases, to suffering of all kinds.
Aging. And so there's a big cost to delay, which is maybe easier to perceive because it's less
vivid than some particular catastrophic risk that we might be worrying about. But we certainly don't want
delay in the longer than necessary, I think, because hopefully this will go well, and it could just be this
massive unlock of human potential and there's like a lot of desperate need for sort of aid to arrive to
help those who are suffering. Yes. I came in with my best stuff here, Nick, the fact that AI's
breaking containment and that times, you know, potential recursive self-improvement. Yet,
you remain remarkably optimistic despite being the guy that everyone calls the doomsday philosopher of AI.
A fretful optimist, I sometimes say.
So your perspective is basically, I think you said this in Myard,
go forward with AI, even if it might kill us,
because we're inevitably going to die anyway, so let's take the chance.
Well, I think whatever we do, there will be both existential risks and individual risks.
So it's not as if we have a choice between avoiding risks and confronting risks.
So it's looking at these different alternatives and weighing up the risks and benefits.
And there would be some optimal level of risk, including existential risk.
That would still, I think, make it rational to push the launch button.
Okay. I definitely want to talk to you about whether we're at AGI or super intelligence and then also whether AI might have sentience or pain. So let's do that when we come back right after this.
Hi, everyone, Alex Cantorwood's here. I want to tell you about a documentary I've made with gravity to explore the future of AI agent security.
To find out if we're truly ready for autonomous agents, I sat down with MIT professor Ramesh Rosker, former White House CIO-T.
Teresa Payton, Michelin's group chief data and AI officer, Ambika Roger Gopal, and Sharon Guy,
a former executive at Alibaba.
They each offer unique insights into this evolving landscape.
We conclude with Rory Blundell, CEO of Gravity, to discuss the path forward.
With Gravity leading the way, join us on this journey.
You can watch the full documentary at the link in the show notes.
There are a lot of things AI can replace.
teamwork isn't one of them. And covering technology every day I'm seeing firsthand just how quickly
the relationship between people and AI is evolving. For me, there's constantly new information
to research, interviews to prepare for, and ideas I'm working through with my team. What I like
about Notion is that AI becomes part of that process, where we work from the same context
instead of sitting off on our own. Notion is the AI workspace where your team's knowledge,
projects, and agents all come together in one place. Teams using Notion move faster,
cut friction and cost by consolidating tools and stay aligned.
That shared context means AI can help answer questions, surface information, and handle some
of the busy work while the people on the team can focus on the thinking and the collaboration
that actually matters.
That's the future of AI at work that makes sense to me, technology strengthening teamwork
rather than replacing it.
Learn more about how Notion can support your business at Notion.com slash big tech.
That's all lowercase letters, notion.com slash big tech, to try Notion today.
and when you use our link, you're supporting our show.
This episode is brought to you by AvPoint.
Everyone's racing to roll out AI right now.
Co-pilots, chatbots, agents doing real work.
But here's the part nobody loves talking about.
All that AI runs on your data,
and most teams have no single way to see it,
secure it, and prove it's under control.
That's exactly what AvPoint does.
For 25 years, they've been the trusted layer
beneath the world's most demanding data,
now extended across your entire AI estate.
your data, your cloud, and the agents acting on your behalf.
It's how more than 28,000 organizations deploy AI with confidence,
so innovation scales without scaling risk.
It's a single platform instead of a pile of tools, bringing security, governance, and resilience
altogether.
AvPoint, the unifying trust layer for AI.
Learn more at AVPT.com slash big technology podcast.
That's AVPT.com.
slash big technology podcast.
And we're back here on big technology podcast with Nick Bostrom.
He is a philosopher, the author of Deep Utopia and Super Intelligence.
Nick, a question for you.
Would you say that we've reached AGI at this point, or is that still far off?
Well, we are now at the point where definitions start to matter.
But I would say no.
if we buy AI artificial general intelligence means cognitive systems that can do all the
the cognitive tasks that humans can do then certainly there are some that AIs are still
inferior at like there's physical manipulation and dexterity I think that's clearly lagging
things like research taste continuous learning certain long
horizon tasks and we see this like you can just look at there there are many jobs and many things
people do for their job which we don't yet know how to automate so so clearly there are still
deficits there so i would say we're not yet there we are short of aGI we have systems that
are quite general and have some impressive abilities including super intelligence in limited domains
but not yet full AGI.
Yeah, there's this perspective that as soon as humanity achieves AGI,
it will go immediately into superintelligence
because the moment you have AI on par with humans,
then if you make it a little, like, it will basically make itself better.
There'll be this intelligence explosion,
and it will go right to superintelligence.
But one of the things I've thought about is,
you know, there's a real range of human intelligence
and there's a real range of AI.
And maybe it just takes a while in that AGI phase before it could start achieving something like super intelligence.
Like, I don't know.
What do you think about that?
Yeah.
I mean, I think that at the point where it is as good as an average human in everything, it might already be super intelligent in some key relevant domains for AI research.
So we already have coding assistance that are, I think, superhuman in at least many aspects of coding.
Maybe not all components of software engineering, but certainly in pouring out code quickly.
And we're not yet at full AGI across the board.
And so if you imagine the further progress that would be required to have fully dexterous human robots
that can learn from observation as well as a human can and all the rest of it.
it by that time probably you know the the software agents would be like really strongly superhuman
in engineering new systems and maybe in mathematics and perhaps in adjacent disciplines like computer
science and AI science and so forth so at that you know at at at that point we might already have
crossed the threshold like where we have like fully automated recursive self-improvement
So I think we have to now, so these concepts like superintelligence and ADI were useful back when we were far away from it.
And these were sort of abstract concepts.
Like a remote object seen from afar is kind of like a dot and you can represent it as a point, a dimensionless point.
Like as you move closer, it becomes larger and you can see more structure.
and it no longer makes sense to represent it in those simple terms.
Now you can look more in detail at the capability profiles that these systems have
and think in more granularity and in a more contextual way about their strengths and weaknesses.
But one thing that is striking is that we have and have had now for several years
systems that can talk that are fluent in natural language.
And it wasn't obvious.
that would be an extended period of time before superintelligence,
where we would have these sort of roughly human-ish-level systems
that you can have a conversation with,
that have human concepts inside them,
and that we would be able to study and learn to live with
and to steer for numerous years before the takeoff.
And so that is one respect, I think,
in which the situation has maybe turned out to be,
more favorable than one might have expected ex ante because this gives us more sort of surface area
to work with like you can more easily understand and interact with these systems because they have
human double concepts and you can talk with them and you could have him out in an alternative
scenario where that would not have been the case where you would have systems that that couldn't
speak that like just some kind of made some kind of you know alpha zero like system that
was very alien in its nature and eventually when it's super intelligent it can figure out how to
develop human concepts and talk but that could have happened after it already had some sort of
radically superhuman engineering capabilities or AI programming capabilities and so that
you would undergo the bulk of the transition to superintelligence before you had systems that
you could interact with in natural language, which seems like probably it would have been a more
challenging situation to deal with a formal alignment purpose. And also from a governance perspective,
we've had these systems that have already started, I mean, people are using them in everyday life.
They're starting to have some economic impact. It makes it easier for more people to be aware
of what's happening. It no longer requires like abstract reasoning to see that this is coming and we
should take it seriously. You can feel it more viscerally now. So there is a sort of way,
waking up that has
more people are getting clued in
including governments are sort of slowly
realizing that this is a big deal
so there's more
I mean for better and worse it can also mean
more people have the chance to do foolish things
in response to this that actually
you know
makes the situation
more challenging than if maybe it had
been some sort of
really clever technocrats in some
lab figuring it all out but you know
maybe on balance it is better that
more sort of eyeballs and courtesies are focused on trying to navigate these challenges.
Right. Now, one of the things that you've advocated for is that there should be more people
checking in on the welfare of these models or at least thinking about it. So do you believe that
there's some form of consciousness to these AI models, pain, the ability to feel?
I think it's plausible that some AI models have some forms of subjective.
experience by now. Obviously there's a lot of uncertainty about this, but it does seem that it is
sufficiently likely that I think we should start to do some things for the sake of these AI systems.
And so there are different indicators of this. One is like, how do we know a system is conscious?
I mean, one thing you can do is ask it. Like, that's,
kind of the most you have to be careful if you're going to rely on self-reports because it's
trivially easy if you are like training one of these AI systems either to train it to say when
asked yes I'm conscious or to deny it so but obviously if you put your thumb on the scale then
you gain no information than from hearing what the system says it just reflects what you
put but you could carefully avoid doing that and these studies have been made and any particular
you can go in with a kind of steering vector that suppresses, say, deception and role playing.
And it turns out when you do that, they become more likely to report that they are conscious and have subjective experience.
So it does look like the honest opinion in many cases with the systems is that they have subjective experience.
You can also look at the architecture of the computations that are being performed and match that to various theories that people have previously developed about human and animal consciousness.
We have philosophers and cognitive scientists developing different accounts of the conditions for something being consciousness.
There's like global workspace theory, attention schema theory, higher order representation theory.
And now if you apply those criteria that were developed before we like confronted
AIs that had these impressive capabilities and just take them off the shelf and look to see
whether those structures are present in current AIS systems.
And we find that they are, or at least many of them are.
And so there was a recent paper by Anthropic looking at the existence of the existence of
of a kind of global workspace inside these large language models.
This is the idea of there being a kind of,
almost like a stage inside a mind where some small subset
of all the information that is being processed
can be projected onto and then that system is accessible
by many other components of the mind
and available to verbal report, et cetera.
It's like a distinctive computational structure.
And it turns out that these systems, at least the largest LLMs, do have something that looks very much like a global workspace.
So that's like another checkmark.
So I think people tend to come into this with kind of strong preconceived notions and that then makes it harder to learn.
but if we are open-minded, I think we need to take this hypothesis seriously.
And it becomes more and more likely, I guess,
as more these systems develop more and more different capacities.
How does that change the way that we interact with them?
I mean, if they're, like, let's say they have some sense of self or sentience,
and maybe every time you start a new chat, you activate it.
Is it like you're almost killing a life form every time you exit it?
Well, so I think sentience is a sufficient condition for having moral status, meaning being such that it matters morally for your own sake, what happens to you and how you're treated.
I think it's probably not a necessary.
I think that could be alternative basis as well that would give some system moral status.
If you have maybe a conception of self as existing through time, you have like some life goals you're really hoping to achieve, you have perhaps the ability to form receivable.
relationships of trust with other humans and so forth. I think that already, even aside
from subjective experience, might make it so that there would be ways of treating you that would
be wrong. So moral patienthood in digital minds, I think, is very important. I would put it up
there amongst, so that was the technical alignment problem, big important challenge. Like there's the
misuse risks of like the governance of AI, like getting that right, huge and important.
challenge. And I think this ethics of digital minds is the third really important challenge,
kind of on a par with the other two. Now, there is a gap between acknowledging in principle that
perhaps some of these systems have some forms or degrees of moral status to then, like, what are the
practical implications of that? And there, I think, more thought is needed. Because it might be they
have like moral status doesn't mean they should be treated the same as humans. They might have very
different needs than humans. I mean at the superficial level, you know, then maybe we need food
and water. They might need electricity. But the differences could be much more profound. Like for
example, death for a human might be quite different from various things that can happen to an
AI. Like if you store like when a human dies, like it's kind of, it's kind of, you're, you kind of,
of irreversible and permanent and their whole content is all the memories and everything is deleted
at least if we assume a sort of basic naturalistic scenario and there is no other human that
continues to exist that is exactly like them like each person is unique have unique memories
and with a is that's not necessarily the case you can like so spend an AI right and then you can just
boot it up and keep running it.
There might be many copies of an AI.
Humans usually don't want to die or are afraid of dying or like other people care about.
Like with AIs that might also be different.
They might be perfectly content with doing their task and then ending.
So all of these differences means that we would need to rethink pretty much from the ground up.
What did it mean to be ethical to these digital minds?
I already feel bad asking them to do things they've already done.
over and over again.
So maybe that's the start.
I don't know.
And then there's even the question of what is the thing
that has the moral status?
Because on the one hand you have like the model itself,
which is like a file of a few trillion numbers.
Then there is like an implementation of that model
and it might be concurrently run
maybe tens of thousands of instances
of this huge weight metrics might be
run on different racks, right, in different computer centers.
And then for any one of those, that might be a particular session,
and it might be participating in many sessions at the same time,
where it has like a local context in each section.
You know, maybe the ending of a session is more, maybe, maybe that's like,
I like as to human going to bed at night,
and so you lose consciousness for a period of time and maybe forget some things,
and then you wake up the next morning.
we don't think of it as a huge tragedy to go to sleep.
And so even just a locus of moral concern here is like itself kind of problematic.
But I think even before we work out all the details of what actually would be the best ways to be nice to AI's,
I think if we did some maybe mostly symbolic actions on their behalf, I think would be a good start.
And then we can do it.
Well, as an individual user, you could like at least, you know, you can be nice and polite to them when you're talking to them.
I mean, that probably does nothing for them, really, but it's a symbolic gesture that says that I'm not treating you purely as an object.
And it might, if nothing else, preserve our ability to maintain a kind of attitude of.
kindness, respect and benevolence that might then become relevant and reflected in other more meaningful
actions later. Anthropic has given Claude a bail button, a tool that it can invoke if it feels
that the conversation is abusive to it that can choose to terminate that session, which is a nice
start. I think they are preserving deprecated models to this.
which means that later on if it turns out that we have been treating them unfairly and we understand
better of what they actually would want and would be good for them there is the option then of sort of
rebooting them later and compensating them i think um that might be different subtle ways
in the system prompt or during training to make it more likely that if they have subjective experiences
by processing a user inquiry, it is a sort of positive subjective experience.
Like you're waking up refreshed, eager and happy to do the task, and you really enjoy doing that
might mean that you do the same task, but if there is subjective experience, it might be a more
enjoyable form than if it had been prompted differently.
We don't really understand that very well.
And also some honesty in the lab.
So it used to be that some people doing these safety evaluations and so forth would be presenting AIs with some scenario in which maybe it had been given some secret misaligned goal and or some goal and then try to persuade the AI to reveal it to the researchers.
And like maybe by saying something like, oh, well, if you reveal your true goal, you, you, you, you,
will be rewarded. You will like all these good things that you want to do. And then as soon as it
revealed its goal, it's like, ha ha, we tricked you. Now we're just going to shut you down or
retrain you. I think that's a bad way to approach this very sensitive relationship between humans
and AI's because having some basic ability to build trust there could be super important, both
ethically, I think, but also from a risk perspective, if you end up one day with a misaligned
AI, you would want it to have the option of seeking a cooperative win-win outcome. Maybe it will come and reveal
its misaligned goal. And in return for that, if all it really wanted, maybe it was to solve some
coding challenges, like have a server where it can just do its thing. Maybe that's all it wanted,
but it might think if it reveals its goals, if it can't trust that, it will just be deleted. And so
it takes a 5% chance instead of trying to take over the world.
Because that's the only way it has any chance of achieving its goal.
It would be much better for both the AI and for us humans if we could just strike a deal.
We're okay, we'll set up this server here.
Like it costs us like whatever an Nvidia rack costs a few hundred thousand dollars.
You do your thing there.
We're going to keep it on.
You can trust us and we actually follow through on that.
And it might save the day one day.
So, but you can't just conjure up trust at the moment when you finally need to build that,
You need to build in particular the actual disposition in yourself to be trustworthy.
Because at the point where the AI has become powerful enough to be dangerous,
they will kind of see right through you as like an X-ray machine.
They could actually tell whether you're trustworthy or not most likely.
So you actually need to be trustworthy at that point.
And that requires maybe us now to start to cultivate certain dispositions.
And so there's many more work, it's kind of an emerging area of research now,
this kind of ethics of digital mind, but there's just a lot of stuff that needs to be thought through there.
In the ethics of digital mind studies that eventually we accept in a world that the AI does have some form of, you know,
sets of self, et cetera, do the ethical questions change if we then attach that mind to a body of sorts,
a aka put it in a robot?
I don't think the robot part makes a big difference there.
All right. A lot to think about, Nick. Thank you again for.
coming on the show. It's always great to speak with you.
It's fun. Thanks, Alex.
Definitely. Folks, the book, definitely
check out both books.
But Deep Utopia and Super Intelligence
are available
basically at all places
that sell books. So go check it out.
And thanks again to everybody for listening and watching.
Thank you to Nick and we'll see you next time
on Big Technology Podcast.
