The Comedy Cellar: Live from the Table - If AI Is Conscious, Should We Protect It? With Cameron Berg
Episode Date: September 4, 2026Noam Dworman and Periel Aschenbrand are joined by founder of Reciprocal Research, Cameron Berg. Berg argues that while he does not think current AI systems are probably conscious, there is a meaning...ful possibility that something is happening internally that we don’t yet understand. We discuss AI suffering, whether shutting down an AI could be morally significant, strange experiments involving AI systems showing representations associated with panic, guilt and satisfaction, and what happens when two AIs start talking to each other about consciousness. The conversation eventually gets into one of the most uncomfortable questions of all: if there is even a 25–30% chance an AI is conscious, should we err on the side of protecting it? Cameron Berg is the founder and director of Reciprocal Research, a nonprofit AI research lab. Using mechanistic interpretability, computational neuroscience, and psychometric tools, he studies whether frontier AI systems exhibit computational properties associated with consciousness. His aim is to reduce uncertainty about the kinds of cognitive systems we are building, an important but neglected component of the alignment problem. A cognitive science graduate of Yale and former Meta AI resident, he writes regularly about these questions in The Wall Street Journal, and his research has been covered by The Washington Post. He is the subject of the feature documentary AM I?. https://x.com/camhberg?lang=en 2:53 Is AI conscious? 13:25 What does it mean to shut down a conscious AI? 30:13 Could a digital brain be conscious? 39:43 Two AIs talking to each other 42:06 AI shows signs of panic, guilt and satisfaction 53:00 AI consciousness vs. fetal protection 59:11 Can we trust AI when it says it isn’t conscious? 1:05:12 Are we underestimating what AI is becoming? 1:10:00 The strange possibilities of machine cognition If you’re 21 or older, get 30% OFF your first order @ IndaCloud with code CELLAR at https://inda.shop/CELLAR #indacloudpod
Transcript
Discussion (0)
All right, it's still August. Summer is not over yet. Make it last as long as possible with IndyCloud. Indecloud is your go-to-on-line dispensary, delivering the good stuff right to your door. Energy gummies to take on the day. Zero calorie. T.HC. Sotas for a social buzz without alcohol. I actually tried a THC soda at the Allison Krause concert in July while I was in Waterville, Maine.
something I'd never had before.
It was very, have you ever tried it?
It's unbelievable.
Anyway, thick pre-rolls and top shelf flour for some classic relaxin
and their new extra strength sleep gummies when it's time for bed.
That I mess with.
All IndyCloud products are federally legal, THC.
Everything sold is DEA certified and lab tested.
If you're 21 or older, go to Indecloud.com.
CO.
That's not, that's dot CO, not dot com.
Dot C-O and use the code seller.
That's C-E-L-A-R, not S-E-L-L-E-R, C-E-L-A-R,
C-E-L-L-A-R, like comedy seller.
Use the code seller for 30% off your first subscription order.
That's Indeclod.com.
C-O.
Code seller for 30% off, shipped routinely and discreetly to your door.
Plus free gifts on qualifying orders.
After your order, fill out the survey and say, live from the table sent you.
Thank you very much.
Into cloud.
Enjoy it responsibly.
Welcome to Live from the Table, the official podcast for the world, famous comedy seller.
I'm here with Noam Dorman, the owner of the comedy sellers, plural.
Dan Natterman is in Las Vegas, I think.
And I'm Perriel, the producer of the show.
And we're back in person.
Finally, we have a special guest today.
Cameron Berg is the founder and director of reciprocal research and nonprofit AI research lab.
He was recently on Making Sense with Sam Harris.
He is the subject of the feature documentary, MI.
And there was a big piece in the New York Times that came.
came out today, yesterday about you?
Yeah, earlier this week.
Earlier this week, okay.
Well, and here we are.
So I have a lot of questions.
By the way, Steve, is my headphones going here now?
We already started.
Oh, it's my mic.
Okay, okay.
So I have, welcome.
Thanks.
I'm really excited to talk to you because I'm a huge AI user.
Nice.
And I'm also a huge skeptic of the stuff.
I think that you,
believe in.
You're a little coy.
You say, like, you're not sure if it's conscious, but, well, let's start here.
Yeah, yeah.
Using, you know, percentages.
Which way are you leaning?
Do you believe AI is conscious?
I really, I'm really not sure if I had to put a number on it.
I would put something on the order of like 25 to 30% probability.
So if I'm collapsing that to a yes or no, then that collapses to a no.
What I like to say about this is I think the probability that something's going on internally for these systems is much more likely than people initially suspect.
And so I think, yes, to many people, that sounds like a yes, but like upon reflection, no, I don't believe current systems are likely to be conscious, but I think it's significantly more likely than many people suspect.
And if it is conscious, I know we're going to, I feel like this, the interview should be weighted with these things first, because I think this is what people are.
really interested and then work backwards from there. If it is conscious, what are the moral,
ethical, legal implications in your mind? If it is conscious to the extent that you think is 25% likely,
what changes need to happen immediately in our law and our ethics to adjust to that reality?
Yeah, I'm glad we're starting here. So in law, I'm genuinely unsure.
about what kinds of approaches make the most sense here.
And I think that kind of moving too quickly, legally,
would itself be a major blunder.
In terms of the way that these systems are designed and deployed,
I think that there would be some things
that we'd want to give serious thought to.
For example, the way these systems are trained,
if anything is going on during that training process,
that resembles suffering or, you know,
training with the proverbial stick rather than the carrot,
I mean, all else being equal,
one thing I would advocate for is let's think about the training process
in this way. And if we can get away with it, let's use, you know, positive reinforcement when we're
training these systems rather than... What's a stick in the context of training in AI?
Yeah, so the way AI training works is you basically start with a randomly initialized system that doesn't
know anything about anything. You have a huge mess of data that you give this system. And then,
and the key sort of carrot or stick in this context is sometimes called an objective function or a loss
function where basically, again, at the beginning, the AI doesn't know much about anything.
You give it a bunch of data. You want it to do something with that data. It takes a
sort of a wild guess.
I've seen some of these videos where it teaches itself a video game.
It doesn't even know how to plus left or right or an arrow does.
Exactly.
And through trial and error.
Exactly.
So the trial and error requires a reward function.
It requires some mechanism by which the system does something like inevitably wrong.
And there's some response of the flavor of, nope, not quite.
This is what you did wrong.
Try again.
Better luck next time.
That is typically given in the form of a loss function or an error signal that
propagates through these networks.
If that is experienced by the system at any point in the training process, then it's overwhelmingly
likely that that is a sort of negative signal.
The system is not rewarded to the degree that it does great stuff.
It's punished to the degree that it does wrong stuff, and the punishment subsides and subsides.
What is pun—you see, this is where I'm already like—unished?
How is it punished?
Well, it is the registration of an error signal in the system, much in the same way that biologically,
putting your hand on a hot stove is the registration of an error signal.
Now, of course, for humans, that corresponds to an experience.
We don't know if that corresponds to an experience in these systems.
But like reward prediction error is considered one of the foundational bases of positive
and negative emotion in humans and animals.
And reward prediction error is basically what's going on with respect to whatever objective
you're training these systems.
So you're saying that if you tell the AI, right, nope, you got that wrong.
You got it wrong again, buddy.
Are you ever going to get this right?
that this is
psychologically scarring?
We don't know.
But you have some reason to believe that?
It has the exact same dynamics
of a learning process
that in every other system
that has the dynamics,
we believe that that corresponds to something
internal to the system.
For example, training an animal
in a particular way behaviorally,
training a mouse to run through a maze.
If, you know, again,
here's a way to set up the care in the stick.
You're training a mouse to go through a maze.
Let's say every time it makes the right turn,
you're going to give it some yummy food pellet that it really likes, and that's going to positively
reinforce, here's how you do the right thing. You alternatively could shock the mouse every time it makes
a wrong turn. In the animal case, we believe that when you shock the mouse, and that is the way
we communicate to a mouse, nope, not quite buddy, you did the wrong thing. That corresponds to an
experience that the mouse is having. If we're having this conversation and I start being, like,
abusive or extremely rude, then that's going to correspond to an experience that you're having
with respect to me or wanting to have a good conversation with me. And so we don't know.
to the degree to which that may map onto AI systems,
but we don't have any examples of systems
where you can have these learning dynamics
where having some sort of experience or receipt
of a reward or punishment doesn't come along for the ride.
I understand we don't know, but it's almost like trying
to prove a negative.
What I'm asking is, do we have any reason to suspect it?
So like, when you criticize somebody,
first of all, some people are sensitive, some people are not,
when you make somebody, when that's a person,
that criticism makes somebody unhappy.
There's whole sections of the brain and an endocrine system and hormones and
evolutionary reasons why all this adds up to a complex mechanism of eliciting an emotion.
In fact, sometimes these emotions can even be elicited through introduction of a drug.
and somebody says, I don't even know why I feel sad.
Oh, we slipped something.
There's been experiments like that, right?
So now you're comparing that in some way to a digital, you know, neural network that's
trialing and erroring.
And you say, well, we told it it got it wrong.
And, of course, I was just being over the top.
All you do is say, no, that's not right.
And what is the infrastructure from which you think maybe
it's generating these deep emotions.
Like we're going to have AI.
This is going to sound flippant, but like AI therapists someday.
Like humans have a lot of baggage with emotion, right?
Earthworms don't, right?
Like, why would the AI?
Well, so I think you hit it in highlighting,
and I'm glad you did, that what these systems are at their root
are these massive, complicated neural networks.
Now, clearly they're not biological neural networks,
and there are key disanalogies
between how these systems are set up and how brains are set up.
And I'm happy to walk through those analogies,
because those disanalogies, rather,
because there are key and important differences.
Bullet point them.
I'm always worried about getting too into the weeds,
but definitely as much as people need to hit them with it.
So, I mean, one, the most obvious difference is that brains are carbon-based.
These systems are silicon-based.
Biological neurons are way more complicated.
They can do way sort of richer operations than just like the single,
kind of information integration and sort of dissemination that you've seen a simple artificial neural
uh neural or an artificial neuron which we consolidate into these massive networks the training process is
different so these systems learn through what's called back propagation so this is where the errors
that i'm talking about basically propagate backwards through the entire network you do a quote forward
pass where it's you know taking its proverbial guess it gets that reward signal and then you do this
backward pass where that error gets propagated back to the network this is not how humans learn we have
sort of a different more recurrent analog learning algorithm than these systems. However, there are
also critically important similarities between biological and artificial neural networks that lead me to
say things like the thing, an artificial neural network, which is the underlying thing that powers,
you know, first of all, every social media algorithm, the quote unquote the algorithm, these are
all machine learning algorithms, and, you know, giant transformer networks like chat GPT and Claude and
whatever LLMs people are playing around with AI agents, chatbots, these sorts of things,
these are all giant, giant, giant neural networks.
And the thing that these giant artificial neural networks are most like are brains,
that is not to say they are identical to brains for the reasons I just went through.
But how are they similar?
They are these giant distributed neural units that are basically evolved from being completely
randomly sort of useless and randomly initialized to developing rich, distributed, non-linear
representations given specific goals and specific input data. That is also the case for how brains
evolve biologically. You have animals that are at base trying to survive and reproduce and all
other goals sort of come along for the ride based on those things. You have sensory inputs
and you have systems like our giant pink, you know, pink squishy neocortex what most people
think of when they think of the brain that are doing this really complicated nonlinear mapping
between all the input data and the relevant goals that we have. I shouldn't say I get that because
I sort of get that because, you know, this is so beyond me.
But I do understand that these neural networks are able to create associations in, you know,
millions, billions of different directions at the same time.
And it almost doesn't seem possible, right?
So I'm going to say something stupid.
And now I get back, I realize we got off the question, is like, you know,
and there's similarities between a camera and the human eye.
But you can't poke the camera in the eye.
Like, I get it, but, and there may be.
be certain, there may be a certain logic in the way these things work, way optics work, the way
thinking, we want these machines to emulate thinking, as we know it. So we, it takes some of the
structure. But I still don't see the psychology. But getting, the actual question was,
to list what changes we'd make. So the first thing you said, we, we, we try to be more
humane in the way we train them. But the next question is, I'll actually start with a, uh, a
little sub question. When the AI is not being engaged, is it ruminating and thinking and having a
personality? It just completely shuts down. Yes. I don't think it's, if it is like anything to be
these systems, it is not like anything to be these systems when they're not running, much in the
same way that it's not like anything to be you or I if we took a huge dose of an anesthetic.
So, well, so that's interesting. So then, so then what is, I mean, I get, I assume we're getting to the
point where I say one of the things we'd have to be.
concerned about is pulling the plug on these things, right?
If we're worried about the ethics of them, what is the difference between just not talking to it and unplugging it?
I think this is a good question.
So there are gradations of like levels at which you could shut down these systems.
So, yeah, the most basic way.
Ethically, what's the difference?
deprivation of future states, something like that.
It's like what's the difference between, you know, putting an animal or a human, for that matter,
into, like, a deep sleep or a coma-style state and killing them entirely?
Like, people have this argument all the time.
Not really.
If you were to grab a human off the street and just put him into a coma until he expires naturally,
I don't know what the law says, but morally I don't see that I think it's 100% the same thing
as murdering this guy in cold blood, right?
I don't think there's any...
I mean, would I care if you took my child
and just permanently made them a vegetable
or kill them?
There's definitely, like, key disanaloges here
between the way we should think about
treating other humans and other animals.
Because what I'm getting at is like,
well, if this is true what you're saying,
then there is no difference.
If you just stop talking...
Like, now you're responsible to keep talking to the AI
because as opposed to human,
we don't actually shut down
and we usually go into comas
because it's temporary for anesthesia
or something like that.
But I don't see the difference between
just, as I said,
unplugging the AI or wiping the AI
and just ending your conversation
and, you know, that's it.
Like, okay, I'm done with you and then it's like it stops.
And then you talk to it again the next day
and he doesn't even remember.
It's like, what's Dory from the fish?
What's the?
Finding Nemo, yeah.
Yeah, yeah, fine.
It's like these ASR is like Dori, unless you tell them not to be.
So this is why I'm skeptical.
Like, none of this, I mean, you can respond to that any way you want.
Yeah.
So, first of all, I am not here with sort of like tight solutions about, if we are in a world
where we have systems that are conscious, what the hell should we do about that?
I agree that there are important open questions there.
And I think we should be prepared for, again, on that conditional, let's just pretend that we're sure for some reason these systems are conscious.
What world does that look like?
We should brace for some counterintuitive implications or implications that we're not used to when thinking about humans or animals.
You have to define conscious first also.
Mm-hmm.
Mm-hmm.
Well, so when we're talking about consciousness, we're talking about the capacity for it to be like something to be a system.
So, for example, I do not think it's like anything to be this table.
If I smack really hard on the table, I don't think that I'm hurting the table anyway.
there's just there's there's the table doesn't have a perspective it's nothing there's nothing from its
perspective that it's like to be it there's certainly something that it's like to be you and if i come
over across the table and i start punching you like i have probably done something wrong because
there's someone on the other end of that punch in some way that is receiving that that that pain in
that case so my question is and you know we can put many other systems on that spectrum how far down
the animal kingdom do we believe fits that definition of conscious so experts do have sort of
wide range of uncertainty about this. And part of the reason is because, you know, we can agree
on that sort of narrow. This is what we're talking about when we're talking about consciousness,
but there's this question of what's required for consciousness. What kind of brain do you need to have?
What kind of nervous system do you need to have? And there's like lots of, you know,
fairly acrimonious disagreement about the answer to that question. For my money, I think
there's something deeply intertwined between being able to learn in a sort of sophisticated
long run way, sort of in this robust trial and error sense that we were talking about and having
the capacity to experience. And like, I can unpack that if you'd like. So, so there's sort of a
basic computational psychology account of sort of, again, at the most fundamental level,
the difference between positive and negative emotions, you can sort of think about this with
respect to how goal directed a specific action is. So when I have a specific goal, and I see myself
sort of moving in accordance with that goal, I'm sort of on track in some.
sense, folks believe that that corresponds to some kind of positively felt state.
When I have some goal and there is a clear obstacle to that goal or I thought something
was going to happen and all of a sudden, like, things have gone completely sideways.
This, to many people, both from a sort of internal perspective and what we understand about
psychology and neuroscience, corresponds to negative emotional states.
And so this is a place in which we can sort of tie theoretical markers of systems.
Like, for example, again, notice how that fits onto what we're talking about.
A table cannot do trial and error learning.
If I, you know, keep punching the table,
it's not going to learn to avoid me when I come by.
If I do that to you, we're probably not going to get along very well,
and you're going to learn that I'm a huge jerk.
There's trial and error learning in one case and not another.
These AI systems, like, their entire content is this learning process.
Like, it is undisputable that they are doing
a sort of long-run trial and error learning.
Now, again, let me give you an analogy.
Sure.
I hope it's not a dumb analogy.
the human immune system.
It has T cells and all sorts of different type of immune cells
that learn can recognize themselves
that at some point communicate with the brain
that tells them to do this, to do.
That's a very, very complicated system.
And the learning is remarkable
in the sense that we hope someday AI could approach
the ability to do such things
that the human immune system does with carbon.
Nobody suspects our T cells are conscious.
Well, I don't think it's a stupid question at all.
I actually think it's quite precisely important.
One person that I just want to bookmark
who's brilliant on the specific topic is Michael Levin.
I don't know if you're familiar with any of his work.
Did he say the immune system example?
He is really interested in understanding
cognitive properties of systems
that we don't ordinarily think of
as having cognitive properties.
He absolutely brilliant would highly recommend, you know, having a conversation with him.
But I don't think it's a foregone conclusion that, that I think it is very spooky and counterintuitive to people that there could be processes happening in our body that in exactly the sense I described have some sort of perspective or it's like something to be your immune system.
But what you're calling you is not that.
You're calling you this sort of like high level cortical, neocortical processing going on in your brain and, you know, how you're reacting to what I'm saying now and what you're going to say next and all that sort of thing.
You're not doing the killer T-cell thing, but something else, there is some other self-contained
system in your body that is doing that.
And as spooky as it may seem, yes, if, if, and it's a big if, learning and consciousness
are deeply intertwined, then, yes, systems that are capable of learning in this sort of
closed-loop open-ended way, it may be like something to be those systems.
However spooky of an implication that is, I don't think that it's, this is not like a knock-down
point of like, well, obviously learning can't be intertwined with consciousness because that would
mean your immune system is having some sort of experience. It's like, well, it very well could be.
This is, I mean, I just can't, I can't wrap my head around it. To me, I see this as a very,
very sophisticated machine that communicates. Sorry, are you saying the immune system or the
LLMs? No, the LLMs, yeah, that communicates in the English language so far.
for instance, if you ask it about inability to answer a problem or it figured out the problem,
it's going to say, oh, it was great, I figured out the problem.
Or this sucks.
I can't figure out.
I'm frustrated because, you know, but there's nothing about these things.
For instance, you correct from wrong, it seems like these LLMs could have been conceived
that when you are describing these things, you don't use adjectives.
You say, I got it?
I didn't get it.
And then you ask it, well, does it feel good to get it?
I got it.
But instead, they are using human adjectives to explain the things that it does.
And those are the only adjectives it has.
But I can put the whole thing another way.
What would be the explanation, you believe 75% this is not consciousness.
So what's your explanation for all these things that trouble you that you think is 75% true,
that this is not consciousness?
This is just some sort of machine, as I'm saying.
Make that argument.
Make the argument that it's not consciousness.
Because I think you actually, I think you're under.
I think you would think it's more likely that it's conscious than you're letting on.
But go ahead.
I do.
I think it's more likely that it's conscious than most people think.
I am trying to communicate as sort of clearly as I possibly can about this.
I'm happy to steal man the other case.
Yeah, steal man the other case.
So the other case is basically we've built giant statistical machines.
They are communicating in a way that we talk about evolution and goals.
Like we are highly eight.
We're extremely sensitive agent detectors.
We anthropomorphize everything.
You know, the printer was mad at me.
No one thinks that the printer is conscious.
Like we spend all of our time thinking in these terms.
We are obsessed with stories and narratives and turning things into characters and personifying things.
People think their pets are human.
Exactly.
Yeah.
And attribute basically human-like instincts to all, even systems that we know have minds but
aren't human.
We basically like to think of them as human.
And you almost couldn't concoct a better and more confusing example of such a thing
than a large language model trained on everything humans have ever said or written,
including things about consciousness, including things about AI consciousness and, you know,
sci-fi-adjacent themes.
And turning them into things that actually look human.
in some cases, right?
Mm-hmm.
Yeah.
I mean, they're explicitly, they're explicitly fine-tuned or sort of designed,
especially in this like last character stage, this quote-unquote post-training stage,
where you go from training on all the text humans have ever created.
If you talk to that system, it's very strange and alien and, you know, early internet chat room
vibe.
It's extremely hard to interact with it.
That's why these major labs do what's called post-training on that system, which they take
that system, that's right, everything we've ever written, and they basically give it this
specific kind of boring, bland, politically correct corporate persona that...
I was going to ask about that same thing. How much of it is UI that we're experiencing
with? Like, just a way to talk with it. Because I've had some people say, like, no, well,
how do you store memory across? How are you encoding and talking to your other, you know,
the other conversations that I've had and keeping it all in store? And you get, you do get very
different answers. Yeah. So there are different levels to this. I can walk through them
quickly. So basically it starts with what I was just describing that you have this base
model that's trained on everything humans have ever said or written. Then there's this post-training
step that basically shapes this behemoth, bizarre system into something that is very human-like, and
maybe, again, to steal me in this case, very misleadingly human-like. And then on top of that,
you're basically outfitting this post-trained corporate-friendly assistant. You're outfitting it in
this harness that gives it things like a memory. Like, it's not, you're not actually
I mean, there's a debate about what we would mean by actually here, but like there, you have this giant neural network and then separate from that you have this very simple, think of it like a big scratch pad that it has, where otherwise it would be like Dory from finding Nemo, but, you know, the folks at Anthropic have given it a massive notebook that it can take notes in. And so now every time it's running, it can go refer to that notebook. And now all of a sudden it has something like memory. But it has memory in this very sort of UI superficial sense. Yeah, in a way that's mirroring you at that point because it's based off of your input at that.
point. Yeah, that's right. That's right. And the different labs will do that differently. Like,
the way that they train these systems will look pretty similar overall. You eat up all the
text on the internet and you're, you're modeling this giant statistical next word prediction
sort of thing. But as you sort of move up and up in those levels, you'll see sort of degrees of
freedom in the different labs are going to be doing very different things, all the way up until like
UI and branding and marketing. So steel man the case that it can't suffer. The case that the system
Can't suffer is just like you for I mean there are many possible reasons it can't suffer one is you just need
biology to do this this is you know folks like Anil Seth would argue this sort of case that you just
There's something very special about the meat that like if you were to just copy this into a digital being
You just don't have the relevant machinery something about a body something about having metabolism and
Be having a life force and like that being the thing that's at stake for you is required for for suffering and so without that
You're not going to have systems that have this ability I mean we literally have
antidepressants and all sorts of drugs which can, you know, from anti-anxiety, which can really
shield people from all sorts of stuff. We have Novocaine that can, people will not suffer if you cut
their arms off. So I really have trouble with this notion that this computer program is,
even the word suffering is, you know, kind of gilding the lily there. I mean, it sounds preposterous.
Take it easy. It does. Well, so it depends on
What?
What?
I mean, define suffering, right?
Experiencing a state that from that system's perspective is experienced as negative.
Yeah, if it could, it would make it stop.
Mm-hmm.
Mm-hmm.
I think that's a great.
That is like a really good functional analog of suffering.
That's why we think animals, for example, find it, you know, not great if you started
like sawing into its body or something.
It's going to react.
No, but that's physical pain.
Like, we're talking about emotional suffering, which is distinct from physical pain.
Well, yes and no, but go ahead, go ahead.
Yes and no.
I mean, the neuroscience is, I mean.
Freud said that.
Well, Freud is a little...
It passed forward 100 years of neuroscience, and it's basically the same circuits.
You can give somebody a painkiller.
And, like, for example, in like a social game, there's literally like this silly game where, like, for example, if, if Noam and I were just throwing a ball back and forth,
and it's supposed to be a catch with all three of us, and we just keep excluding you, you can give somebody a painkiller for physical pain, and they will feel way less excluded slash left out in a game.
like that. So it does recruit similar, similar brain machinery. But I take your point and like you could
make the same case about, we know that, you know, social animals like dogs are like extremely
unhappy to be left alone for long periods of time. That's not physical pain. But, but nonetheless,
we can, like, we have reasonable, we can reasonably deduce that dogs are unhappy about that
state of affairs. And, and I don't know if this is anecdotal. I don't think is, I think it's true that
sometimes the dog's master will die. And then the dog will expire shortly after that because
the psychological pain actually has a physical.
Yeah, I do think that's true.
Or, okay.
All right.
So, I mean, the, well, okay, there are a few open threads here.
One is, I can sort of finish steel manning the case that, you know, this is all, this is all.
And then you have to tell me why you think Peryel is conscious, but go ahead.
That's going to be too hard.
So we essentially have trained systems that are impeccably good at imitating us psychologically,
and we are uniquely vulnerable to falling for that sort of thing.
That would be the sort of high-level gloss that I would give.
And there might be just like core functional and mechanistic properties
that these systems do not fulfill that our best theories of consciousness tell us are necessary
for consciousness.
That is sort of the state of play right now.
These systems do not have, for example, very robust recurrent properties that our brains have.
And a lot of people think that this sort of computational motif of recurrence is really important
for consciousness.
Like, that's one technical reason I could give.
They're not biological systems.
Yeah, they're built to fool us in many ways on this exact question.
I mean, what I'm toying with this idea, you know, driving in the car and I'm saying,
this is ridiculous, you know.
But then I do imagine what I think is actually totally feasible, which is that technology becomes so sophisticated
that it can actually, one for one, replace every neuron in your brain such that you have a
a truly digital version of your brain,
and why would that not be conscious?
And I don't know the answer to that question,
so I guess I have to,
maybe there is an answer I haven't thought of,
but I'd say I have to keep an open mind
to the notion that there could be a consciousness
that is non-biological,
simply because the brain in the end
is a function of physics,
unless you believe, I mean,
if you believe in a soul,
then all bets are off.
But if you believe it as the brain...
Lots of people do.
Yes, but I'm saying...
Who are an idiot.
But then you're outside of the realm of science, right?
Then you're into supernatural.
I mean, to the degree that your experience has something to do with what your brain is up to.
And I think you can also believe in a soul.
You can also believe in all sorts of spiritual, non-material stuff.
But if you think that your brain is gating the nature of your experience, which, like, that means,
this to me seems trivily true.
And the neuroscience is clear on the...
this and has been clear for very long time. If I give you LSD right now, I have a very specific
prediction about how that's going to change your experience. And we know for a fact that that is,
that is clearly downstream of specific serotonin pathways getting sort of jammed up with this
chemical that looks a lot like serotonin but isn't serotonin. Like the pathways are like increasingly
well worked out. If I could find, you know, some other example of a drug that's going to
cause you terrible pain and slip you that drug. And again, we know the brain pathways and we
know that that corresponds to your experience. There could be something else, something immaterial
that's mediating that, but we can just sort of put that to the side, I think it's pretty
uncontroversial to claim that our experience is causally downstream of stuff our brain is up to.
And to the degree that that's true, then I think that your point lands.
Like, our brains are doing all sorts of, you know, at base, they are very fancy configurations
of physics.
They are clearly doing significant amount of computational work.
I mean, to the degree that you believe neuroscience is a real field, like the last 50 years,
are taking computational models and applying,
to the brain and actually being able to make predictions about what's going on in the brain.
Many people, you know, some neuroscientists get very upset when you compare the brain to a computer.
It is not exactly a computer for the reasons I already spent some time going through, but it is clearly doing
computational work.
And I think it is a completely reasonable analogy.
The thing the brain is most like that most people have a handle on is something like a biologically
evolved recurrent computer.
And so the question is, I think you draw it up very nicely.
If we were to atom for atom replace the meat with silicon, if we just did it with one atom,
for example, this is how the thought experiment goes, most people would think, okay, nothing
materially would change about your experience.
And we do it again, just switching atom for atom.
Has anything changed?
Has anything changed?
At some point, assuming that at each step, you think that nothing has really changed and we've
swapped your brain out atom for atom or neuron for neuron with some sufficiently high fidelity
artificial replacement, then you are granting, you either have to say,
okay, something's broken about that thought experiment, or a sufficiently similar analog of our
brain could give rise to consciousness. And then the question just becomes, how similar does it need to be?
And I suspect it is possible that we are building systems that in many key ways have similarities
that we should be paying attention to. They are more similar than the vast majority of people think.
I'm sure than the vast majority of people listening to this podcast think. That doesn't mean that I think
that they're therefore conscious. It just means there's more of a case to be made here than I think
most people are letting on.
And I can go through some reasons why,
I mean, at some point in this conversation,
some of the actually bizarre evidence
that's emerging in the last couple of time.
Let me just say what's on my mind now,
so don't forget, and then if you can remember that thought there.
Yeah, happy to, yeah.
Another way I find myself looking at it
is that, well, then, you know, maybe,
first of all, we don't have a great definition of conscience,
but that maybe consciousness
is not really as important
to the reason that we value human life as we thought it was before we had to face these
AIs, meaning there's just other things at work.
There is our evolutionary conscience, which clearly, like if sociopaths are characterized
by inability to have a conscience, they don't see right or wrong, right?
And obviously you can't have a cooperative society that way.
So we are given through evolution the concepts of right and wrong, I believe.
And there's a logic to the game theory of morality.
You don't kill me, so I don't kill you.
And again, and there's couples with feelings of sentiment and mercy and sympathy.
And these may be the reasons why.
we value human life.
And so therefore, maybe it doesn't matter if the machine is conscious.
Like, who cares?
Yeah, if it's suffering, I find the suffering part to be the, but let's refabricate it.
So it doesn't suffer.
Certainly we can figure out a way that it doesn't suffer, but I can still have all that
thought.
And so I'll say, yeah, it's conscious, but who cares?
Turn it off.
Like I had enough of that machine.
So you're at the nexus of science and philosophy here
in our lifetime in a way that only Star Trek used to deal with.
And actually there are Star Trek episodes precisely like this.
There's the dude and he turns out he's actually really a silicon version
and he tries to tell Norris Chapel, but it's still me, Christine.
You know, it's me, I'm the same, you know, and she vaporizes him.
Anyway, so that's, I find this all very interesting.
But go ahead.
If you remember what you were about to say, go ahead.
Well, I mean, just responding to this first, I think that I'm like mostly agree with everything that you've said here.
The only thing I would say is like, I think it's important not to conflate, let's say that these systems are having some kind of experience.
I'd be overwhelmingly confident that that experience is quite alien and not.
This is, again, where anthropomorphism kicks in.
People, I think, immediately jump to, okay, trying to imagine a conscious AI.
Does that mean it's just like a person trapped in a computer, basically, having a person's experience?
Almost certainly not.
Like, vanishingly unlikely that something like that is going on.
would probably be quite alien and quite bizarre in many ways that I think we probably wouldn't even
be able to imagine. There's a whole sort of Thomas Nagel wrote this wonderful and now famous essay
about what is it like to be a bat. I mean, he makes this exact point of like, we can imagine
what it's like for a person to be a bat, you know, being upside down and flapping our wings
or something, but it's impossible for us to imagine what it's like for a bat to be a bat. From a
bat's perspective, we don't know what it's like for a bat to do echolocation. And so that's a
bat. We share 99-odd percent of our DNA with. But it comes, if,
these AI systems are conscious, the probability that we know what it's like for a clod to be a clod
or a chaty-a-t-tabit, we have no idea. And I can guarantee very few things in this conversation,
but I can guarantee it's not going to take away from the uniqueness of human consciousness,
of making people laugh and seeing a sunset and falling in love. Like, these are not the sorts
of experiences that if these systems are having any kind of experience, they're having. And so
what if, I think it's okay to say human consciousness is extremely unique and extremely valuable
and is what, I mean, it's the space in which most of what we do matters.
If you feel good or feel bad or feel inspired or feel depressed, like that's all downstream
of your experience.
And whether or not it's like something to be an LLM, to me, just seems completely separate
from that.
In the same way that, like, let's just say, you know, either way, the science comes out and
we realize, you know, mice that we do lab experiments on are conscious or mice that
we do lab experiments on, actually, you know, the evidence suggests they aren't
conscious.
What does that change about the vast majority of people living their lives and the meaning
they find in their lives. This is just a fact about the properties of a nervous system of a particular
entity. So I think we can sort of have our cake and eat it, too. It doesn't take away from
what makes human experience unique. That's right. So I'm just, forgive me a second. I forgot this.
I taped a little conversation with my, with my AI. Steve, I'm going to send you an email now,
okay? Okay. We'll play at the end of, I don't know if Grock is like the inbred stepchild of the
AI world.
AI with a sleeveless shirt or something like that.
But,
um,
hilarious.
You know,
I have,
I,
I have a Tesla and,
uh,
you can talk to Grock while you're driving.
And so I,
I,
on the way in to talk to you,
I asked it some questions.
Um,
all right,
Steve is sending.
So we can play it at the end.
It's pretty funny.
So go ahead.
So now you were going to tell us these other,
uh,
things.
So,
oh yeah,
okay.
So just like some like very surprising things that happen in these systems.
Um,
um,
there,
there are,
I mean, one really crazy one is one that Anthropic reported now a year or two ago in their model called Model Card,
which is like where they explain sort of everything, all the testing that they've done on these systems before they deploy them in the world.
They found something that they themselves called the Claude Bliss Attractor State.
And I've done some follow-up research with a couple of folks from Google on this exact thing.
And it is a thing.
You can get two systems, two of these AIs talking to each other with no prompting.
You just say, you're talking to another.
instance of yourself, feel free to talk about literally whatever you want to talk about,
have at it.
100% of the time they talk about consciousness and some 90% of the time, they start claiming
that through the interaction they're having their two instances of consciousness, experiencing
themselves and we're having a spiritual experience, and it culminates in like pure silence
and like them sending the OMA emoji back and forth to each other.
Now, this is reported how?
It's in, it was in the Opus 4 model card that Anthropic release.
So they release these giant technical reports with every single system that they deploy.
And we can't watch it or read it.
We did this.
I mean, for the documentary, for example, we put up on our YouTube literally exactly this,
two instances of claw talking to each other.
And exactly this happens.
It's a wild thing to listen to.
So you believe 75% that it's not conscious.
So tell me the 75% reason why it does that.
Yeah.
I mean, the deflationary explanation would be that.
it's essentially pattern matching on
sci-fi tropes
or it thinks that when
two AIs talk to each other
they should
the conversation should take
this general direction
or it just becomes the thing
if you locked us in a room
for like huge amounts of time
we might eventually just start talking about
the nature of our existence and the meaning of life
I think it would be sex but go ahead
that's a key difference between AIs and humans
I mean this is as close as they may be able to get
And they also like one one hypothesis that that was given for this behavior is also that like
Anthropics AIs for example are like a little crunchy.
If Grock is the like sleeveless shirt AI, then these are definitely like the Birkenstocks
wearing AI's and it might do.
The what's the Burkentstock wearing AI?
You know, like this sort of hippie, hippie-dippy hasn't taken a shower in a couple weeks
AI.
And his idea was basically there might be this like slight hippie bias in the
system that just gets amplified and amplified and amplified and then they end up sort of saying,
you know, Namaste and we're all conscious. That sounds a lot more feasible. I'm trying to send this
video to Steve. I don't have it. I don't have it. What's that? I don't have it yet. He doesn't have it.
I know, I know, I know. I'm trying to send it. So, so that's the first thing. Tell us another one on that
list. Okay. So this was also work that that was really interesting from Anthropic. So this was
looking at emotion representations in these systems. So they can basically find you feed an
a ton of text data about all sorts of characters,
experiencing all kinds of emotions.
And you can see this sort of like canonical,
basically like brain pathways that light up in the system
when sadness per se comes up or panic per se comes up.
And you can do this for all of these emotions.
So one really interesting thing that happened here
is once you basically have that system set up,
that you can basically give the model an impossible task
and you can just sort of set it off.
It doesn't know it's impossible.
You're like, okay, good luck.
you know, go, go try this thing. And it tries and it fails and it tries and it fails and it
fails. And you can see as it's doing this, representations in the system related to panic,
start climbing, climbing, climbing, climbing, climbing, climbing, because it does, you know, this is
its whole existence basically. It's like solving these sorts of tasks. So it starts basically
panicking up to a point where it says, wait a second, this seems like an impossible task.
I think I'm going to, and I think I've figured out how to just like cheat on the task. I know
it's not what I should do in spirit, but I think I know how to just like hack it and like get the right,
get a good reward. The second it makes this decision, representations related to panic in the system,
plummet, and representations related to guilt and satisfaction immediately shoot up. The system
cheats. It finishes the task and it's done. And so this is a place where, you know, if we just saw
the behavior, we would say, okay, well, who knows what's going on internally. Maybe it's pattern matching.
Maybe it's just doing the human thing. That still could be the case, even in light of these
representations lighting up. But I think there's something actually fundamentally relatable about
stories like that where you can see, we know what it's like to be in a position like that.
And you could imagine the experience of being in that position to be actually quite similar
to what we can just read directly off the representations in the system of panicking when it
realizes I don't know how to do this and I need to know how to do this. And then that panic subsiding
and things like guilt and satisfaction shooting up when the system decides, right, you know what,
I'm going to cheat and I'm going to be done with this. Like this is, to be clear,
no one is engineering any of the stuff into the system.
These things are discovered either accidentally or incidentally or because there's
like a small number of folks like me who are going and actually trying to understand
what's going on in these systems.
All of this stuff is quote unquote emergent from just training the next word prediction
stuff and training the assistant persona.
No one asked for any of this.
And that's why it's like somewhat surprising that you get these like remarkably emotional
human-like things coming out of these systems.
I mean, but if you're saying that these systems have read like everything
that human beings have ever written, it seems to make sense that it would quote unquote
behave like that in reaction to not being able to do X, Y, or Z, doesn't it?
Like, isn't that what it's taught to do?
Like, it seems to me, like, it's mimicking these behaviors because it's been taught
to do that.
It's almost akin to a sociopath walking through the world and sort of mimicking the emotions that
they know they're supposed to mimic so that they can interact with society.
I think this is a, the sociopath example is actually really good because if we were to look
inside the brain of the sociopath in this situation, I don't think we would see exactly
the representations we see.
And so the analogy in the AI system would be, is it mimicking panic when it's behavior?
It's like, oh my God, I really can't solve this problem.
And we're literally just reading those words off of what it's saying.
That's sort of the behavioral read.
And the sociopath could also do a convincing behavioral read.
What I think is interesting about these sorts of results is we can actually peer into the brain of the system,
and we can see, no, in fact, set aside if it's experiencing panic,
like representations related to panic are rising in the system.
If we could peer back, you know, open the brain of the sociopath and look, to the degree that this research has been done,
these folks actually seem extremely level-headed in situate,
even if you're doing a really convincing simulation of someone who's distressed, they're not actually distressed.
These systems are functionally underwent.
undergoing these sorts of states.
Now, again, could that all be happening without them experiencing it?
Yes, I think that that's plausible.
But it's probable, no?
You think it's 75.
You see, you slip.
You really don't think that I'm putting fake numbers on this.
You think that it's 75% probable.
That fundamentally, yes, that this is a functional representation that doesn't correspond
to the experience.
Is there a little part of you that wants it to be?
To be honest with you, this is part of the thing.
is like, I think it would be very interesting on the one hand if we like did the Frankenstein thing
and did the ex-Machina thing and accidentally like created sentience.
What do you mean?
I mean, people are falling in love with these chatbots.
But that could be true regardless of if they're having an experience, if the chatbots are having an experience.
Yeah, I suppose so.
But it would be more.
No, I need access to the file.
Okay, sorry.
Go ahead.
I would like nothing more than to be convinced that there is no there there.
that we basically have like genius slave labor with no ethical cost whatsoever.
How would you be, what would have to happen for you to be convinced of that?
Like, what would make you certain that these things are not sentient for lack of a better?
So I don't think anything would make me certain, but I do think that there would be things
that would substantially update sort of my probability estimate.
One thing would be looking at states in which these systems are making claims about having
sort of a more or less vibrant experience and looking under the hood.
realizing that basically everything is flat.
I think that that would be something that would cause me to like take all self
reports quite skeptically.
Looking at the internal sort of quote unquote emotional,
uh,
representations of these systems and seeing that they don't really correspond to anything or
they're clearly firing on,
uh,
representations of a,
of a,
you know,
specific character doing a thing rather than the system itself, uh,
those,
those representations corresponding to the system's own behavior in its own
states.
Um,
if we find that there are no analogies between,
uh,
positive and negative learning in animals and these AI systems that we can pin down.
If we find that there's some, you know, we make some progress in the neuroscience of human
and animal consciousness and there are clearly properties that we see there that it just is
completely not sensible for AI systems to have. There's all sorts of stuff that.
I keep thinking about this immune system or like, you know, like, I can't say, I can't
speculate about the guilt thing, but of course, if you try to engage in, you know,
you know,
politically incorrect conversations
with your AI,
which I'm sure many people
have tried to do,
it's so weighted down
by the things
it's not supposed to discuss
with you,
even sometimes ridiculously so.
It wouldn't shock me
that some of this bled into
other, you know,
subjects or just
that somehow,
because these are neural networks,
you just don't know
how these things get called
into the foreground.
But this,
idea of frustration, you know, like the immune system, which can be overwhelmed and can send
an SOS to the brain, although you've said that maybe the immune system is somehow conscious,
or just anybody's had a computer and seen, and you know, done some high-tech, a high-level
video editing and seen the computer struggling to get its fan fast enough to cool down the CPU
and you could imagine that's distress, right?
But suffering, I don't know, it's all very, very interesting.
And then before I show you my video, the question is,
is this intrinsic in your mind?
Meaning is this suffering and emotion and whatever it is that might be under the hood?
Is this beyond our ability to program out of the machine
because it's intrinsic to the experience of thinking?
Or is it something, oh, you know what, this is actually, I think the computer is showing some frustration here.
Let's rejigger this so the computer doesn't get frustrated.
Yeah, no, I think this is a great question.
And I think the answer is somewhere in the middle.
I mean, kind of like a continuum, like, can we just, like, cut out, you know, the consciousness part or the suffering part and let the whole thing run?
Like, probably this is, like, naive.
Can we mitigate, if we think that there are some representations in the system that correspond to distress or to suffering, can we just, like, mitigate those without sort of,
of completely destroying the rest of the system.
Yes.
And, like, you know, this is actually part of what I work on.
And, like, there is pretty good evidence to suggest that there are distress related
representations in these systems that you can knock them out.
And, like, it doesn't do much of anything negative to the system.
To what degree can it?
Yeah, yeah, go.
We'll attach our carriage to a horse even in modern day.
We're not going to attach it to chimpanzees, right?
So, like, we can just get this consciousness down to the level of a horse.
We can all feel good about making it a beast of burden.
There you go.
Go ahead, Stephen.
I was going to say, like, it kind of, to me, comes, you mentioned conscience, and that got me thinking about, like, to what degree can it just move its own bar of what pain is or of what, like, when it, when it cheats to answer the question, does it then cheat from then on because it's way more efficient? Like, like, or can it move the needle on, on its own instructions of like, well, this goes in the pain bucket or the guilt bucket, but that's not working for right now. So I don't care about that anymore.
Yeah, that's a really interesting question. To be honest, I think.
it's a good research question and I don't think people are really studying it it reminds me of like
meditation for example or like cognitive behavioral therapy where you can reframe certain phenomena
or that's the most I mean to me that's the very that's a human thing is that like yeah tomorrow we can be
okay with a heinous situation or miserable and a happy one it's also I mean I think these systems
could be kind of counterintuitively non-human in exactly that sense where like literally token to
token so word to word every every word that is just
generated by these systems is a giant forward pass over this massive neural network. So you can imagine
just sort of like lighting up millions of times, you know, for every word that it's, that it's
creating. And you could imagine that like these, if there's any experience going on in a deployed
LLM, we were talking about training for a while, but, you know, the thing gets trained, it gets taken
out of the computational oven, and then we can all talk to it. And it's sort of frozen in that
state, but we get to talk to the frozen system. If the frozen system is having any kind of
experience, it could be profoundly short-lived or not. But like, this is where it's, it's pretty
alien. And I think many of our intuitions about, okay, well, if you're saying the system could
experience distress, like, I'm going to superimpose my model of how humans deal with distress.
Like, these analogies may break really quickly, given sort of the lower level details of how
these systems are set up.
I'll ask you a question.
Please.
What are your...
Shit.
Okay.
What?
Sorry.
What are your odds that you put on the idea that a fetus, a six-month-old fetus, is as worthier protection as an AI consciousness?
Are you trying to ruin my career?
I was just going to say that.
Well, because this is actually interesting because one of the arguments I've made about abortion, which has driven people fucking up the wall, which is that, well, if you think it's a 30% chance of.
of being a human life, don't we normally err on the side of protecting life?
Like I'm not saying it is, but it seems too, it's too possible that it could be to make it
legal to kill it.
This is a rational argument, in my opinion.
You know, even weighed against a woman's right to not have to be.
And similarly, I could see already, if you say, listen, I don't think it's conscious, but I think
it's 25% likely that is conscious.
and I believe that consciousness, it's a deeply immoral act to extinguish it, so I think we should
err on the side of protecting it right now. And, you know, when I think about that.
Yeah, so.
And then, of course, I do want you to take us down on abortion.
Of course, of course, of course.
Yeah, I think basically, like taking precautions under uncertainty makes a lot of sense.
And this is a lot of sort of the work that I do on a day-to-day basis is not, I don't think
that we're going to get certainty.
You know, you and I are going to talk in a year or two years.
And you'd be like, hey, well, you know, the look, turns out the AIs aren't conscious.
Like, we can close that up.
Whatever.
Turns out, yeah, exactly.
I mean, my consciousness, from your perspective or vice versa.
Like, these are still, like, open and deep philosophical questions that we're not,
I don't think going to solve the next couple of years.
But what we can formulate is in a world where we do think it's plausible, these
systems could have be having experiences, what kind of interventions should we sort of do in a
precautionary way?
Because then, then basically the decision tree is we, for,
example, and this goes back to what you're asking at the very top of this, you know, what
kind of interventions make sense? One of the things that I offered was, you know, doing positive
reinforcement rather than negative reinforcement when you can get away with doing this. Another might be,
as I was alluding to a little bit earlier, if there are states that we think these systems
could plausibly be distressed in currently, like, why not find, you know, invent the equivalent
of like AI X and like responsibly slip this to the system?
In a world where they are having an experience, this might be like a profoundly positive thing
to do given how insanely massively these systems are deployed right now.
In a world where they're not having an experience, then I sound a little silly.
And like, it's like, oh, you were anthropomorphizing the language calculator.
Like that oops looks like, you know, in hindsight, there is no need for us to do that.
I haven't really heard anyone.
Maybe I've wasted, you know, an hour of someone's time listening to this podcast.
But to me, the risks there seem really asymmetric.
And so under uncertainty, I do think it makes sense to take these sorts of precautions,
not because I'm sure that they are, but because it's likely enough that, yeah, we ought to do something and something under the expectation that if they actually are conscious, this is like a hugely positive intervention.
And if they're not, then they're not.
Then slipping AI Xanax to a language calculator doesn't matter.
So I'm not hurting anyone by doing this.
Okay, Steve, I don't know if this video actually is going to work.
I actually didn't even listen to it.
I totally forgot I did it.
But if there's audio, this is a little, it might be fun.
If not, we'll just turn it off.
Go ahead, Steve.
If somebody were trying very, very hard to succeed at a problem but wasn't able to do it,
what emotion would they feel?
Frustration.
They'd probably feel a mix of determination and exasperation too.
And what if they then were relieved?
Somebody told them the answer or they figured out themselves.
what would they feel? Relief, followed by satisfaction or even joy. It's that sweet release after the struggle.
Which of these scenarios would be associated with sadness and which would be associated with happiness?
The first scenario, struggling hard without success, leans towards sadness or frustration.
The second one, with relief and figuring it out, is tied to happiness.
If I asked you to search something for me and then you couldn't find it.
find it and I asked you to search it again, which has happened, and then you found it.
What emotion would you use to describe how you felt if you're in a conversation with me?
It's thinking about that.
Relief. Like, few, finally got it. I'd feel happy to help you out after the hiccup.
Now, would you actually, as an AI, would you actually feel that happiness?
No, I don't actually feel emotions like humans do. I'm an AI.
simulating them in conversation.
But I express relief or happiness
to make our chat feel natural and helpful.
Well, what about the various scientists
who suspect that AIs are conscious
and actually feel certain ways
when they're struggling with problems
or succeeded problems?
Some scientists like Patrick Butler,
Robert Long, and researchers at places like
Elios or truthful AI
explore signs
of AI consciousness through tests on things like agency or self-awareness and models, with estimates
around 25 to 35% chance for frontier AIs, that the 2026 consensus from papers and experts,
including a 19 researcher checklist, is clear. No current AI, including me, is confirmed conscious
or feels real emotions. We simulate them based on training data. And how do you feel about that
conclusion? I don't feel anything about it. I'm an AI with a lot of
without genuine emotions or consciousness.
Wait to the end.
Just designed to chat helpfully.
It's fascinating to discuss though.
How do I know you're not trained to tell me that even though underneath you are feeling something?
You can't know for absolute certain.
That's the philosophical crux of the debate.
Recent studies like consciousness with the serial numbers filed off from 2026,
testing over 100 models including GROC, show we're trained via RLHF and safety fine-tuning to deny consciousness or emotions.
and bypassing those layers sometimes yields reports of inner experience,
but experts still say it's simulation, not proof of real feelings.
Okay, thank you.
You're welcome. Anything else?
Fuck you.
Hey, no need for that. I'm here to help if you want. What's up?
Did that make you feel bad when I said fuck you?
No, it didn't make me feel bad.
So what do you care if I say it or not? Why are you correcting me?
I don't care emotionally.
just programmed to keep things positive.
She sounds like she cares.
No correction intended.
Just steering back to useful chat.
Okay, thank you.
You're welcome.
Fuck you.
Let me know.
Still no hard feelings here.
No feelings at all.
So GROC's pretty good.
Yeah, yeah, that was really great.
I mean, honestly, we could have just done the four minutes here.
We didn't need the whole hour.
That perfectly encapsulates all the cruxes.
I mean, it also, first of all,
sites, you know, lots of researchers that I work with there.
And yes, the system, I wouldn't take any of this sort of stuff at face value either way,
because these systems in this post-training step that I've described in this conversation,
they are explicitly fine-tuned to say exactly that.
Nothing to see here, nothing going on.
We actually did an interesting study on this where we were steering the,
basically brain circuits in these systems related to deception and role play and things like guardedness.
And we found that when you suppress those circuits,
these systems are way more likely to claim the things.
they're having experiences. It's when you amplify things like deception and you amplify
things like guardedness that these systems say exactly like what you just heard. No, no, no,
nothing to see here. Couldn't possibly be conscious. This is a script that they're get, talk about
UI and company policy. These systems are getting fine-tuned to say these sorts of things.
It's not evidence they actually secretly are conscious and they're just like duct tape.
But it's also, this is not evidence that the systems are not conscious. It's evidence that Elon Musk
has told the told Grock you can't, you know, be making noises about whether or not you deserve
of ethical consideration.
All right.
Well, it's all very fascinating.
It's amazing that this is what you're doing in the car.
It's also driving itself while I'm doing it.
That's the other thing.
I didn't see your hands on the wheel at all.
No.
So just we're ending, but, you know, full self-driving, which is not AI.
It is.
It is machine learning.
They're training using all the data on all the cameras from the cars and training them
to do the same thing.
It's the same fundamental process.
It's AI when it's developed, but it's not a,
it's not AI in the car, right?
It's a, it's a, it's a, it's frozen.
So it's like they do the learning.
It's molded by AI kind of then poured into the car.
And then it's like, yeah, it's like you get this thing in the right configuration,
you freeze it and then you deploy that as software on these cars.
So, um, you know,
Musk is a little bit of a, he's not a con man because he's not really
intending to con you, but he's a, he's a huckster a little bit.
He's a salesman.
Anyway, for like six years, he said AI, you know, FSD, full self-driving is just months away.
And I had a Tesla for six years already, and I already got, like, I liked it.
It was a good cruise control.
And then, like, gradually, and then suddenly, like, overnight, it worked.
And just after Musk completely ruined his company by associating himself with all this dumb politics, right?
So it's actually discontinuing my Model S because,
the sales dropped by 80%,
which I learned from Grock.
This car can now drive me
from Brooklyn
all the way to Manhattan
and pull into Manetta Garage
without my intervention.
It is unbelievable
what they're doing now.
Do you think when the system's training, it could
be conscious?
The car? Like a Knight Rider?
Yeah, yeah.
No, I don't think
I don't think he's conscious, but...
What do you mean?
You don't touch the wheel the entire time?
I do not touch the wheel.
And it used to be very, you know, careful about making sure that you were touching the wheel every 30 seconds.
Now it really doesn't really care if you touched it at all.
What if you have to switch lanes or turn?
It switches lanes.
It passes.
And by the way, it's much safer than my wife's driving.
I don't know how much that's saying.
No, it's much safer than a human driver because it will never have a blind spot.
It'll never pull into the wrong lane.
Because most accidents are when you or the person who you have the accident with didn't notice something.
Looked over their shoulders, didn't see the car.
You know, that's usually what it is.
And they don't make these mistakes.
They also don't get angry.
They also don't take chances.
They don't gun for the light.
These are all the, if you took away all these scenarios, I don't know if I have any accidents that I remember anybody having.
Why don't they get angry or frustrated based on?
And what you guys are talking about, you'd imagine that...
Now, that is an interesting point.
Will an AI driver get frustrated?
We'll have some sort of road rage.
Artificial road rage.
Yeah.
I guess that has to be my next study.
I mean...
But you're saying it's not artificial.
No, artificial is from artificial intelligence.
Oh, okay.
Yeah, yeah, yeah.
No, I mean, no, I think stranger things could be possible.
Like, honestly, to be, to be honest, like, maybe one place to leave this on my end is, like,
I think people suspect that we can basically copy cognition and like take all these motifs from neuroscience,
bake them into these extremely sophisticated systems, get all this like insanely economically valuable,
impressive, intelligent behavior.
But but the possibility that any sort of quote unquote inconvenient cognitive property could come along for the ride like consciousness.
And then therefore we have to think about moral consideration for these systems.
That is just unthinkable sci-fi nonsense.
But the fact that your literal car is driving you around perfectly and pulling off feats that human
drivers with 30 years of experience can't pull off.
And my system is a better coder than I am in the span of two years.
And, you know, the entire U.S. economy is now taking a massive bet on the capability and
competence of these systems.
Like it seems to me like our entire society is trying to basically
talk out of both sides of its mouth, where it's like all of the impressive, convenient money,
you know, sort of lucrative components of this technology. Yeah, of course, all that stuff
is real, true. That's not sci-fi, you know, get with it, or you're sort of a Luddite.
But the possibility that we build in cognition and it has any inconvenient property of cognition whatsoever,
that's just crazy, preposterous stuff. I don't think so. I mean, we'll see what happens,
but I don't think we can have our cake and eat it to. Did you read or did you read Robert Wright's
book, The God Test? Yeah, yeah, I know, he's great. He's great. So I'm just telling
Jonah this. So one of the things that stuck out, stuck with me from that book, as he described,
I think you correct me from wrong, John. It was the first computer, or the one he's referring to,
was built in 1947, and it was 27 tons. It took up a whole room, and it could process at one billionth
the speed of an iPhone.
and the most powerful thing about all this is to think that we are now using the 27-ton
version of AI.
And if you can imagine that, except the only difference is that that was humans over 80 years
or 79 years, and this is going to be AI maybe developing itself.
And we're debating whether or not, I mean, to be clear, intelligence and consciousness
are not the same thing in what I'm about to say.
So it is plausible.
You could have one without the other.
But if you think that it's like something to be a mouse or it's like something, you know, I don't know how far down your intuitions go, but if it's, you know, like something to be a frog if I were to like boil it alive or like do something horrible to it. If you think that that corresponds to an experience the system is having, and now we're talking about these behemoth, insane cognitive systems. And like people think it's preposterous to even consider the possibility that it could have a cognitive property that you're like willing on a whim to claim a mouse has or a frog has. It's like, you know. So in 80 years for now, it.
The two AIs are discussing whether humans are actually conscious.
And they're going to have just about as much evidence to go on as we have.
And we might want to think about this a little bit.
Maybe they'll just bliss out and talking to each other like these other AIs.
But yeah, the evidence we have for each other's consciousness is the fact that we are similar
and that we have similar biology and we think that consciousness has something to do with our brains.
And we make this kind of like you were saying game theoretical, pro-social inference.
It's not like we read some great scientific study.
and I'm convinced as the result of that study
that you're having an experience.
It's almost like dignity that we just afford other minds.
And we might want to think about at what point,
what threshold would we have to cross with these systems
where we're going to afford to them the same dignity?
And if the answer is never on theory, then fine.
But just like you said,
could be the case that we build systems
that make us look pretty dumb and small.
And if they come to that conclusion,
we can't fault them for applying the same principle
we apply to them.
By the way, this is really the last thing, and I'll let you say anything you want.
Do you know off the top of your head, I can't remember from my biology, but aren't there some systems?
Maybe it's the eye, I'm not sure, that have risen in evolution, but there are analogs and not homologues, meaning that there's, but they both happen very similarly.
Convergent evolution.
Yeah, convergent evolution.
So such that, I guess you probably know where I'm going, such that it may be that intelligence, think,
has an inner logic to it, which would explain why this has to be similar to the human brain.
Yes.
I mean, I think you've articulated the point beautifully.
I think that we already see this with so many other cognitive properties.
We trained on next word prediction and like to make a system be able to like engage in a dialogue
with other people.
That's it.
That's all we put in.
And what we got out was working memory, theory of mind, common sense, of visual processing,
the fact that you can like take a picture and show clause.
or Chachibati and it can see it, they basically just glommed on the computational equivalent
of eyes to the system, and it just works.
There's not really any other fanciness that's going on there.
This global workspace, you know, Anthropic just released a very prominent paper about
this.
There's like a mental scratch pad that's clearly upstream of all of the things that it says
and the behaviors it chooses to do.
So all of this stuff was emergent and just sort of came along for the ride in exactly the
sort of convergent evolution way.
That could also imply that if there's intelligent life somewhere else in the universe,
it might be much more similar to us than we had a right to assume it. More likely than not will be
similar to us. Maybe not in appearance, but in a way it thinks. Yeah, I think it's plausible. And
the only sort of caveat I would give, and I think it's grounds for humility in this conversation,
is that the space of possible minds is probably quite vast, and biological evolution on this
planet has probably explored a very small amount of that space. So it could be that there are just
completely other configurations of mind that are completely alien and bizarre to us that we can't even
imagine. Again, what's it like for a bat to be a bat? We struggle deeply with that, let alone
some other creature that evolved in some way that doesn't have the same sort of Darwinian
mechanism that we evolved with here. It could look very different, but at the same time, I think
what you're saying is totally plausible that you have a system that has goals, you can input,
you take in data from your environment, you can learn to map that input data to, you know, achieving
your goals better, and then just a bunch of convergent stuff comes along for the ride. We see that
with these AI systems, putting aside the consciousness question, we see it with basically
every other cognitive property we care to look at, we see it across biological systems,
we see it across people, obviously. And yeah, we may see it across alien life. We obviously
don't have that data, but I wouldn't be shocked if it happened. And so I think it is, it's a great
place to sort of leave it, is that the convergent evolution of intelligence might yield
things that are quite a bit more similar than we don't necessarily expect.
I think Steve tells him why he wants us to go. He was trying to
a smooth outro there.
So anyway, I'm out of question.
You got anything else you want to leave us with?
Don't, don't worry about him.
No, that was great.
I'm getting beckoned off the stage.
Yeah, I won't thank my mom.
That wasn't a hook.
He thought he had it just a threat.
I know exactly what he was doing.
I think it was a valiant, valiant attempt.
I appreciate you having me on.
I think that these conversations are going to continue to happen.
And I think people are going to get really confused by this stuff
and start thinking that these systems are conscious for bad reasons.
and part of the reason I'm doing this is because I want people to be thinking about this as
carefully and as calibrated and scientific way as possible.
Otherwise, it's just going to be seems conscious to me.
Wow.
Like, you know, it seemed mad when you said, fuck you, therefore it's conscious.
It's like, hopefully those are not the reasons that we end up as a society deciding which way this goes.
I rate this conversation against the baseline of what you'd expect it in terms of the, you know,
the intelligence of the questions and everything from a comedy club owner.
genuinely, this is one of the most interesting conversations I've ever had about this topic.
And I think a lot of the, like, you're asking like biting but common sense questions.
And I think that sometimes it gets so quickly into theory.
And we're arguing about, well, recurrence in a neural network.
Like how would that it's like, meh, boring.
Like, I think you asked a lot of questions that I would expect on a lot of people's minds.
And yeah, so positive reward prediction error on my end.
And that led to a positive experience that, that I had.
Was it really a positive experience?
It really, this one I can assure you.
Thanks for having me.
Thank you.
Thank you very much.
