Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas - 363 | Chandra Sripada on How LLMs and Humans are Cognitive Cousins
Episode Date: August 10, 2026Large Language Models display an uncanny ability to construct human-sounding speech, and can synthesize concepts in novel ways. Is this because they are truly thinking like human beings in some ...way, or have they found a way to be human-like without reproducing the internal mechanisms of human thought? Chandra Sripada argues that LLM cognition is more human-like than we suppose, and offers evidence from the ways that cognitive scientists study actual humans. Blog post with transcript: https://preposterousuniverse.com/podcast/2026/08/10/363-chandra-sripada-on-how-llms-and-humans-are-cognitive-cousins/ Support Mindscape on Patreon. Chandra Sripada received an M.D. from the University of Texas and a Ph.D. in philosophy from Rutgers University. He is currently a professor of philosophy and psychiatry at the University of Michigan, where he holds the Theophile Raphael Research Professorship and directs the Weinberg Institute for Cognitive Science. He writes Cognition, Decoded, a Substack newsletter about AI, cognitive science, and philosophy. Web site Google Scholar publications PhilPeople profile Babbel is offering listeners up to 60% off. Go to https://babbel.com/MINDSCAPE. #ad Get 60% off an annual plan from Incogni using code MINDSCAPE at https://incogni.com/mindscape. #ad
Transcript
Discussion (0)
Hello everyone and welcome to the Mindscape podcast. I'm your host, Sean Carroll. I presume that everyone has heard of the touring test, or as Alan Turing himself called it, the imitation game, supposed to be a way to figure out whether computers can think. And of course, the difficulty in that is not only what computers can do, but what do you mean by thinking? So Turing took a very sciencey, physical sciencey, mathematically approached to the problem. He says, I don't know what it means.
to think. But we can look at what things do. And if you have a computer that does things that are
impossible to tell the difference between that and thinking, that is to say, the input, output,
responses of the computer are indistinguishable from those of a person, we should call that thinking.
Now, these days, we have these LLMs, these large language models as an approach to AI, and more or less
it's clear that they do pass the touring test.
I know there's some people who argue about that,
but I think it's kind of nitpicky myself.
I think that there's no question my mind
that they're passing the test as touring himself
would have imagined it.
So does that mean that the LLMs are really thinking
in the same way that human beings are thinking?
And I think, you know, there's been an argument back and forth.
Some people say, yes, that is what it means.
Others say, well, no, actually turns out
we have to think harder about what it means.
to be thinking.
And then the other side says,
oh, no, now you're moving the goalposts.
I thought we agreed on the touring test.
I'm actually on the side of the goal post movers.
I think it's perfectly okay to say,
well, that wasn't a careful enough definition
of what it means to think,
at least not in the same way as human beings do.
And even granting that given a certain set of inputs,
the LLMs will produce outputs
that are more or less indistinguishable
from some kind of human output,
there's still a question, are they doing it because the LLM architecture in the process of being set up and then trained and fine-tuned and so forth has essentially rediscovered the mechanisms by which human beings think?
Or is it because, and it's a absolutely plausible scenario, they've discovered a different way to have the same input-output mechanisms as human beings, the kind of alien intelligence.
And there is some evidence, and I have absolutely been on the side of being impressed by the evidence that says, look, there's questions you can ask in LLM that don't look like human answers.
You know, the famous ones are how many R's in the word strawberry or something like that.
And that to me was very good evidence that they're not thinking in the same way that human beings are.
And a lot of people push back against my view on that saying, well, you know, all the great computer scientists and leaders of the AI industry are saying,
are saying otherwise. And that was never especially convincing to me for the simple reason that
those people are experts in computer programming and computer science and AI, but not experts in
intelligence and cognitive science. So recently, Johns Hopkins hosted a meeting of the Society
for Philosophy and Psychology. And some of us had the idea to be fun to do a live podcast
interview as part of that meeting. It didn't pan out that way. I was traveling at the same
the meeting was happening, et cetera. But we looked through the people who were visiting for
prospective good podcast guests, and we found today's guest, Jondra Shrapada. He is a philosopher
and a cognitive scientist in the psychiatry department at University of Michigan, and also
an expert on LLMs. Okay, so he's an expert on thinking and on LLMs and is very well positioned
to ask and answer the question, do
LLMs think in similar ways to humans do.
And he makes a strong case that they do.
Often, in many ways, think using the same kind of thinking methodologies that human
beings do.
And so I find him very persuasive.
I think that he's made a really good case that in, at least in an important set of
ways, LLMs have rediscovered or, you know, been coaxed into re-finding out the ways
human beings think. Using ideas from cognitive science, you know, how do different tests of what
happens during the cognitive process match up between LLMs and human beings. So you can tell
for yourself whether or not it's convincing to you. It's not the same as saying that LLMs are
conscious or responsible moral agents or anything like that. But they are, but this is something we
should establish in that direction.
You know, cognition is easier to understand than consciousness.
So I think this is one of those podcasts that has shifted my credences in important ways.
And, you know, one should always be a good Bayesian.
If more evidence comes in the other way, then they'll shift back.
Or if the evidence keeps pushing in this direction, they'll keep moving in that direction.
But I think it's fascinating that, you know, the option, which was always on the table,
and I always admitted certainly, that LLMs sort of have for by some,
way or another, rediscovered human modes of thinking has turned out to be something that has
evidence for it in an interesting way. So I think that makes for a great conversation. Let's go.
Sean Dersh Tripata, welcome to the Mindscape podcast. Oh, thanks, Sean. I'm glad to be.
here. Now, we're going to be comparing two of my favorite cognitive systems, the human brain and
LLMs. So most of us are a little bit familiar with the human brain. Of course, we've all heard of
LLMs by now. But could you give us a quick overview of how LLMs work? You know, we use them,
but maybe we've forgotten what is going on underneath the hood a little bit. Yeah, of course. Yeah,
Let me say a bit about that.
And I realized that I didn't get to say this already.
I'm a long-time fan.
Oh, thank you.
It's great to be here.
And I don't think I've ever heard you at anything less than 1.5 speed.
So it's good to hear you live.
You're even more eloquent at this rate.
So what LLMs?
How do they work?
You know, for starters, LLMs are neural networks.
So there's going to be a set of nodes and weights that connect them.
And the basic processing strategy is when an input comes in, the activations will be multiplied
by the weights and get propagated through the network.
And if those weights have been tuned correctly, the answer will be a sensible continuation
of the input that came in.
So LLMs are distinctive.
in a few ways.
And one of them is the approach that they're going to take is auto-regressive.
So I like to have my coffee with sugar and.
And they have been trained so that massive set of parameters,
hundreds of billions, trillions and frontier models,
are tuned up in such a way that over the course of Internet,
scale data, they are making predictions of the next word. And every single time, those parameters
have been changed just a bit so their predictions get better and better. And so initially,
I like my coffee with sugar and, you know, it might be darts or butter. And eventually,
you're going to get to cream.
And so the auto-regressive generation, one word at a time, after that one word is generated,
it's added back into the prompt as part of the context.
That's one thing that's distinctive about them.
The training on internet scale data, that's another thing that's distinctive about them.
And maybe it's best to just foreshadow this and say this a bit later.
But it's not a big undifferentiated neural net.
in the old heyday of PDP, you know, when I was a bit younger, those nets, there is a set of input nodes and a set of weights and eventually you get to the output.
Well, these are transformers, and they have a lot more structure.
And so with the transformer, every word slash token, we don't have to get too much into that, enters as part of a column.
And the processing happens in a column-specific way, but for the presence of these attention heads that move information from the left over to token positions on the right.
And so that's another, it's not the only feature of transformers, but they're a distinctive, highly structured kind of neural net.
And so billions of parameters trained on next word prediction, auto-regressive generation, the structure.
of a transformer.
Okay.
That's a lot of what's going on
in a large language model.
You know, if I am going to write a book,
I do not start at the first sentence
and just keep typing, right?
Like I imagine an outline
and then I develop it, chapter by chapter.
I could ask an LLM to write a book.
Would it do that?
Would it sort of give a big overview
and then fill in?
Or would it just start at the first sentence?
Right.
So already, and I'm going to try to, you know, assemble some arguments that these LLMs process in ways that are reminiscent of or analogous to core processing principles of the human mind brain.
But you caught me already.
There is, there seems to be something very different about the auto-regressive next word prediction training.
and then the generation process where one word is generated at the time and added to the context.
What to say about that?
I mean, one thing to say already is generating the next word is not incompatible with the model itself during the forward password has a substantial amount of pneumonic resources,
and processing capacity to start anticipating what happens to later on.
So Anthropic has this example where you give it a stem of a rhyme, something to the effect
of, I saw a carrot, I had to grab it.
Okay.
And the model actually generates the last word of the next sentence, rabbit, first.
way, and we know this via mechanistic interpretability where we're peering into the model and looking at its processing stages, before it generates the prior words that will culminate in rabbit.
And that's highlighting that next word generation as a generation method, as well as closely aligned training objective, is quite compatible with the model.
generating things that come much later early on.
And so that kind of planning can still happen.
I'll add another dimension to it as well.
We'll come to this a bit later, I think,
but these models can use the context window
as a kind of scratch pad and generate internal tokens
as part of what they call thinking.
So if you give it a complex kind of task,
the models now will break it into sub-tasks.
as part of the initial thing that it'll do.
And sub-tasks can be further broken down to the subtasks
and delegated to a series of sub-agents and things like that.
So it is compatible, remarkably enough,
with next-word generation and next-word prediction
as the training objective for substantial staged attacking of a problem,
breaking it down to sub-components,
looking ahead at terminal parts of the sub-components,
before next word generation begins.
So next word prediction is,
and the training objective,
next word prediction,
and auto-regressive next word generation,
themes counterintuitive at first,
but can actually encapsulate
a lot of the ordinary strategies that people use.
And is this idea of like a scratch pad
and multiple sort of swipes at the answer,
Is that something that human beings have coaxed the LLM into doing?
Or is that something that the LLMs have sort of figured out for themselves in some sense?
Right.
And the answer is mostly the LLMs have figured out for themselves.
And there is some coaxing involved as well.
And so why do we believe something like that?
You talked about the use of the use of the scratch pad specifically, but is this a good time to go into the dual process distinction?
No, we're going to keep that later.
We're going to keep that for later.
Then I'll say this bit of in a very general way.
Next word prediction on Internet scale data gets the model to assemble an astonishing representational landscape.
and latent strategies for, you know, how complex problems like, let's say you tell me that you want to go to Tasmania.
The model has latent within it the kind of subtasks that would be needed in order to do that.
So all of this just seems to be something as part of the representational repertoire, the next word prediction,
generates, especially when it's done at scale.
Now, when you've got these Lego pieces,
it doesn't mean that you've got that pretty amusement park already
that you can build with the Legos.
There is some coaxing, some supervised training,
some instruction tuning,
and especially reinforcement learning,
where you train the model with evaluative feedback
that can emerge from various sources.
That tends to take this astonishing representation,
reputational repertoire that's there that's latent and assemble it into useful pieces and
organized, goal-directed control.
So that's great because, yeah, it leads right into what I really want to ask, which is
at a very high level.
So we'll get into details later.
But to me, the basic question is, look, we've built something that is really, really
good at taking inputs and giving outputs that are remarkably human, right?
that could fool anybody if you had the old-fashioned touring test.
And I can imagine two different possibilities.
One is that's because the LLMs have basically rediscovered the same mechanisms that are going on in the human brain.
Or alternatively, the LLMs have figured out a wholly new way of sounding human,
what we might think of as an alien kind of intelligence.
And you know that we're relentless anthropomorphizers, so it's going to be easy to guess that they must be human if they're acting human in this way.
So what are the options here before we start advocating for either one?
Like what are the possibilities to help conceptualize what's going on inside the box?
Yeah, I mean, I think you've laid it out really, really well.
We see the outward, we see the inputs, the prompts, and we see,
the next words come out and they're eerily fluent.
And one option is they are fancy auto-completes.
They're statistical approximators.
They are bullshitters of sort.
They are entities that have memorized a bunch of tricks and heuristics
about how to generate next words that sound like people.
but under the hood, very little is going on that closely resembles what happens in the human mind brain.
And this is definitely a continuum.
There's going to be various middle positions.
And at the other end, a picture that, you know, I'll put my cards on the table that I'm going to argue for a bit more.
And I think I'm more inclined to and attracted to is it just turns out that prediction is incredibly.
powerful. It is the training signal that is the mother of all training signals. And whereas when I was
going in grad school, this was much less appreciated. It wasn't part of the zeitgeist. It has been
over the years in cognitive science. It's been taking over center stage as a lot of what happens
in the mind-brain is prediction. And a lot of cognitive principles that we thought were innately
specified or due to some sort of contingent evolutionary trajectory, they actually are emergent
in cognitive science, dual process distinction, or various other cognitive principles that are,
you know, that we describe the human mind brain, they are downstream of prediction. And so it turns out
that at the level of basic core cognitive principles, the LLMs and humans, they,
identify similar representations, similar procedural techniques, different modes of inferential organization, and so forth.
And so there's a lot of similarity there, and it arises downstream of prediction.
I like how you call the mind brain. That sort of avoids worrying about calling it the brain and people getting upset or calling it the mind and people getting upset.
Exactly, exactly. I've been coached by a PR firm, not to get into that.
So, okay, good.
So the options are out there.
Maybe LLMs have thought of a, or have come up with a different way of acting human,
or maybe there's some form of convergent evolution where the best way of acting human is also what the LLM's found.
There is this, like, quasi-enictotal evidence that has made a big impression on me,
I will confess, that the kinds of mistakes.
that LLMs make seem to be the kinds that no human would ever make.
Like, of course, everyone makes mistakes,
but there's different kinds of mistakes.
Like famously, the LLMs can't count the number of ours in the word strawberry
until you really teach them.
So how much of an impact does that kind of evidence make on you?
Yeah, I think, you know, I'm a total evidence guy.
So, yeah, that's.
It's a data point.
You look at it, that's puzzling.
And, you know, especially, and then the next moment, it's expounding on parts of general relativity that, you know, that are very subtle.
So, you know, what is going on there?
And one thing that we shouldn't do is just stick with these kind of behavioral observations.
you're going to notice various places with the LLMs do not act like in very human-like ways.
And we shouldn't settle for just looking at outward behavior and counting and tabulating.
Well, this looks a little similar.
This looks a little dissimilar.
Very quickly, if we stay at that level, it's always going to remain.
Any one of these possibilities could be live, you know, alien intelligence versus
the name I like to give is cognitive cousin for the other version where they're actually much like us.
It turns out with the strawberry, these things, they do have obviously a different kind of learning history,
and their contact with the quote-unquote world is exclusively textual via these tokens,
which essentially serve as kind of sensory primitives for them about what comes in.
So they don't even have access to the letter level.
It's not they don't have access, but they don't typically operate with individual letters when words come in.
So there's a very natural explanation for why they can't count the number of ours in strawberry.
Whenever we see these kinds of anomalies, and there are many, an immediate next question we can ask is at the behavioral level,
Before we even start looking to closely at underlying mechanisms, are there a bunch of ways in which LLMs reproduce behaviorally phenomenon that are familiar that we've documented over the decades in cognitive science?
Do they reproduce some of those phenomena?
Turns out there are many such cases and many of them are notable.
They catch your eye.
Can I go through some examples?
Sure.
Okay. In the area of language, there are a bunch of effects that psycholinguists talk about. So an example is the center-embedded sentence. So a man ran is easy to understand. A man that a woman loves ran. It's easy to understand. A man that a woman loves ran. It's easy to understand.
a man that a woman that a child knows loves ran starts to become hard to process.
That's center embedding and psycholinguists have identified.
You know, for now, I actually won't talk about mechanisms.
Let me save that for a little bit later.
There's garden path sentences.
The horse raced past the barn.
Right.
Fell.
The fell seems to come out of nowhere.
And it's because I'm not saying the horse.
raced past the barn. I'm saying the horse that was raced past the barn fell. And so there's a
order effect. When you do incremental parsing, these garden paths arise. There are other kinds of
effects, similarity-based interference. There are depth changes. So there's a bunch of these
kind of effects that psycholinguists have identified. LLMs exhibit those. Another nice and striking one is
in serial list memory, a kind of branch of episodic memory. If you give people a list of words
and ask them to say the words back, what you'll find is that there's a recency effect. The very last
word has an advantage. The very first word has an advantage. There's a lost in the middle of effect.
Those things in the middle, they don't have that advantage. There's a contiguity effect where
if you say a word, the words around it come to mind more easily. There's a temporal asymmetry.
effect so that if you say a word, the one after it comes to mind more easily. LLMs exhibit all
these effects. When you put a bunch of words in their prompt, the first word, the most recent word,
are more prominent to enter into processing. There's a lost in the middle. There's temporal
contiguity effect, and then there's the temporal asymmetry effect. I'll give one more
example. Yeah. And these, so there's visual search has been extensively studied and there's a
classic effect disjunctive versus conjunctive search. So if I say find the green L among a bunch of
red L's, so picture that. There's a one green L and there's a bunch of red Ls. The green L is going
to pop out because the greenness is a singleton concept that distinguishes that one thing.
If I say, find a green L among red L's and green T's, now you've got a problem.
There's no one feature that distinguishes the thing that you're looking for.
And so people slow down.
They do this in serial order, and their time in completing the task is proportional to the number of elements in the array.
So this is classic.
What I love about these is cognitive scientists have already written down these effects decades ago.
And you can check the LLMs, in this case, vision language models, do the exhibit the effect.
And they do.
All of these examples, we know the mechanisms behind it.
We know that garden paths arise due to incremental parses that happen one word at a time that lead to bad parses leading you astray.
and then you've got to go back and regroup.
We know the episodic memory effects.
We know some of the mechanisms that play a role.
And in the visual search, it has to do with what are the conceptual primitives,
what is encoded via a compositional code, and what requires feature binding.
The fact that you're seeing these non-obvious patterns of similarities in LLMs and people,
especially where we know some of the mechanisms that happened in these effects in cognitive science,
they point to similar mechanisms being operative in the LLMs and people.
So if I can try to summarize that, the strawberry thing could conceivably be evidence that the LLMs have found a very different way of sounding human,
and here we're showing that the emperor has no clothes with the strawberries, or mostly the LLMs are.
doing human type cognition and we've just found one little Achilles heel.
I would actually put it in an even more benign way. We knew at the get-go that their sensory
evidence, what is analogous to that is tokenized representations of English language words.
They very rarely descend it to the level of individual letters. So we knew that at the get-go. It wasn't
supposed to even be a major dimension of the analogy between them and us, that individual letters
would be particularly salient in their processing strategies. So the two lessons, I would say,
are that counting and tabulating at the level of behavioral outputs is probably not going to get
as very far. We need to look mechanistically, and we need to think about which are the mechanisms
that we actually care about,
that are core processing principles
for the human mind brain.
Dealing with individual letters in...
That's never something that I would elevate
to the level of a core processing principle.
Whereas some of the others that deal with episodic memory
and visual search, compositional codes,
and incremental parses,
that's bread and butter of what cognitive scientists thought about
for a long time.
And so the fact that those are preserved
are probably more important.
Right.
I don't want to dwell on all of the different idiosyncrasies of the LLMs,
but I guess your answer to the strawberry question doesn't seem to provide an answer to the car wash question.
You know this one, right?
You know this example.
I do not know this one.
Oh, this is a good example where you ask the LLM,
I am 500 feet away from the car wash.
I'd like to get my car washed at the car wash.
Should I drive my car there or just walk?
and the LLM will never be say,
you should just walk,
it's only 500 feet away.
Not with the implication being,
but it does be no good to walk
because I want to get my car washed there.
I need to drive my car there,
even though it's very close.
You know what's hilarious about that?
I was about to say,
you should walk.
It's only 500 feet away.
You've been hanging out with LLMs for too long, Chandra.
Yeah, exactly.
Either they've rotted my brain
or unwittingly, what we're seeing is the foibles and idiosyncrasies of them are a little bit like us, at least in some cases.
Yeah, I think that's it.
Okay, anyway, I just wanted to get that on the table there.
But, okay, now you were very gracious in not going into the dual process theory when you wanted to, but here is a good time to do it.
because if you're going to make the argument that LLMs are sort of remarkably human-like in their cognitive strategies,
here's one example of how humans work.
So explain, let's not assume that people know what the dual process theory is.
Sure.
Yeah.
So what is dual process theory?
People often are familiar with Conneman's thinking fast and slow.
And his Nobel, he won the Nobel Prize and his Nobel lecture.
He revisits his entire amazing research program in terms of dual process theory.
And so the idea is you have fast thinking.
It's quick.
It's automatic.
It's intuitive.
It results appear to you as a kind of a flash.
And then you have slow thinking, which is slow.
It's serial.
It draws on more cognitive resources.
Whereas fast thinking is thought to arise via slow, patient,
experience-dependent practice with the environment, a characteristic of slow thinking is it deals
very rapidly with sophisticated novel problems. So those are some of the features of fast and slow
thinking. And, you know, an example, if I show you a picture of John Travolta, it flashes out to you.
That's Travolta. It's not as if you reason through it. If I say, because so many people attended,
They moved the ball upstairs.
You know that I'm talking about a gathering of people, not the spherical thing.
That's fast thinking.
And slow thinking, if I give you, for example, multi-digit mental arithmetic,
nothing flashes before your head.
There's some sort of procedure.
Maybe you as a physicist shot.
No, no, no, no.
Arithmetic is not it.
Yeah.
So that's thinking fast and slow.
And it's just a toweringly influential theory.
It doesn't originate with Conman.
So Keith Stanovich and Richard West coined the terminology.
And, you know, social psychology, stereotyping.
They use this theory.
Clinical.
I'm a psychiatrist as part of my background.
Substance addiction.
ADHD.
This is a theory that people draw on.
Developmental psychology.
Adolescents behave the way they do because fast thinking outstrips flow thinking.
It's just explanatory.
form that is widespread throughout psychology.
And so this is an example.
We better see, if we're going to say that LLMs are like us,
is there something like that in LLMs?
And can it help us as cognitive scientists understand this distinction?
And the answer is yes.
Okay, so in LLMs, there is a really interesting distinction
between what people sometimes call in-weight processing
and then there's a pair of, so in-weight processing, so that's W-E-I-G-H-T, I'll come to that a second,
is very akin to fast processing.
And then there's two mechanisms in LLMs that help to undergird that so-called system two,
the slow processing.
But I'll start with in-weight.
L-L-L-M's via the prediction objective, they acquire massive amounts of information that's compiled
into their weights.
That's why it's called in weights.
It's a relatively easy pathway from the stimulus information to something that's populated in the LLM.
It doesn't require elaborate constructions within their activation space.
And I'll say a little bit more about that if that sounds obscure.
So Albert Einstein is German.
They just know it.
Michael Jordan plays basketball.
The state containing Dallas is Texas.
If you drop a vase, it will shatter.
These are all the kind of common sense.
factual things that they know in their weights. There are two mechanisms that undergird their
system two. And the first one, in context processing, I'm going to dwell a little bit on it.
Because I haven't mentioned this yet, and I should say this, there's a paper that goes along
with some of the arguments that I'm covering today. And that paper, I have a co-author. It's my
Rick Lewis, professor of psychology and linguistics at Michigan and my colleague and my friend and
collaborator. And so Rick calls, we'll put the paper up in the show notes eventually, I think.
Rick calls this the greatest discovery in the history of cognitive science. So in context processing.
So maybe that's not hyperbolic enough. What is in context processing? It's the surprising phenomenon
that prediction training installs in the model,
something like a very general routine
for extracting patterns from the prompt,
generalizing from base cases,
learning a novel mapping from instructions alone.
When I characterize system two,
I said one of its primary properties is
it's the place where rapid novel grocking
of complex novel patterns suddenly happens.
And cognitive science has been at a loss to understand this phenomenon
and characterize it in any detail.
And in a pair of papers, Bradford and colleagues in 2019
and Brown in 2020, the second one is from OpenAI.
Ilya Sitskever is lead author of one of these.
And Dario Amode is the senior author and the other one.
These are like giants of the people.
field. They, especially in GPT3, so the 2020 paper, these models, when they're just trained on
prediction, they have an astonishing ability to display this in-context reasoning. And an example
that was used in a paper that identified one of the mechanisms of in-context processing.
I'll give that example. You say, Marie Curie is, well, I mean, Marie Curie, she is the
discoverer of Radian. She's an amazing scientist. Those are the things that come to your mind.
Maybe that's in the weights of the model. But if I say, Albert Einstein is German, Mahatma Gandhi
is Indian, Marie Curie is, you have French. Some may know that she was born in Poland.
Polish. I was going to say Polish, yeah. I know too much about this one. The computer scientists
that use this example, they may not have even known that she was Polish. So they just assume the
answer is French. The model gets the idea that there is a pattern being requested. And they actually
study this in a lot of detail. They use hidden Markov models, mixtures of them, and they show.
The model is doing something like implicit Bayesian inference. It's got a hypothesis about what task
is being asked. It's using the first example, you know, the likelihood. It's multiplying that
presumably by some sort of prior to arriving at a posterior about what kind of task is there and what
kind of completion is the best for this circumstance. Others have likened this to more like gradient
descent happening within the model. I want to emphasize this is happening after the weights are
frozen. So what do I mean gradient descent? Gradient descent is how these models are trained. Then they're
shipped and then the weights are frozen. What do you mean gradient descent? Well, the models have activated,
Those activations are changing.
Those activations are what are happening on the nodes when you multiply the nodes by the weights.
Well, the model has compiled within it machine learning routines like linear regression, nearest neighbor, and Bayesianism.
And they can understand very complicated patterns in their activation space.
That's what this in-context learning is.
and it arises via very general prediction training.
As a cognitive scientist, where we have struggled to articulate,
what is this system two?
The dirty secret of cognitive sciences,
we don't actually have theories of things like system two or related terms.
What we have is theories of when does it develop during adolescence?
What lights up when you put people in the scanner?
We don't actually have mechanistic process-level theories of a lot of how this stuff works.
And here we have an artifact that's displaying the phenomenon, and we can mechanistically interrogate it,
and we can understand the grab bag of routines that underpin it.
So this is startling, this in-context learning, and it's very system-2-ish.
There's another system-2 mechanism that they have, and that's where you're using the internal
they're using the context as a scratch pad and you're doing extended serial reasoning within it.
So I'll say a little bit about this quickly, borrowing an example from Melanie Mitchell.
So you give a problem like this.
Julia has two sisters.
How many sisters does her brother Martin have?
Well, you say, oh, two.
Well, you just told me she has two.
There's two sisters there.
Well, Julia is a sister.
Yeah.
And so if you give this type of problem to a model that is not using chain of thoughts,
very likely to say too, but if you allow it to generate internal thinking tone at tokens
that are not displayed to the user, that end up in the context, that influence the next steps
that the model takes, it can reason through and says, okay, so Julia is a girl, she has two sisters,
that means there's three girls in the family and so forth.
So you can show that a model that has access to this chain of thought scratch pad.
Models like the ones that when IMO gold recently in 2004,
that this amplifies the inferential power of the model.
And so where we begin is dual process distinction.
It's a towering distinction in human psychology.
And what we see is these models have in context processing,
and they have chain of thought, and between the two of them, they approximate the, and they also
have the in-weight, which corresponds to the automatic.
And so you're ending up with a distinction in the model that mirrors that in people.
And you ask, well, is there clear evidence that this in context and chain of thought, that actually
is much like what happens in the human mind brain?
And yes, there is evidence.
And there is reaction time evidence.
My own lab has a study using paradigms like the Stroop task that illustrate that there is a very tight correspondence,
which is between what is happening in the model and what happens in the human mind brain,
for which we often invoke the system one, system two distinction.
So I don't want to close over it too quickly.
you're saying that LLMs have discovered BAS's theorem.
Yeah, you know, what to say about that.
So it has been widely known that LLMs approximate the posterior.
So, you know, you've got a context, and then there's a next word.
And there's a conditional distribution.
and they approximate that conditional distribution very well.
And a basin would say, here is the normative way to do it.
You have these factorized representations of hypotheses.
You've got a prior and you multiply it by these likelihoods.
It turns out that you can approximate the Bayesian posterior
or the correct conditional distribution,
not by explicitly going through a factorized representation,
but in a sense by modeling it or learning it
or learning it by short-cutting your way through what are the features present in the context,
can I stick those into a powerful function approximator and get something Bayesian posterior like?
That's called an immortalized inference.
And it's a statistical principle that's been known in machine learning community for decades,
that you don't have to go through a Bayesian calculation to get to the Bayesian posterior.
You can amortize it via the very long learning process.
that's what these things are probably doing.
And it's not simply that they approximate the basian posterior.
That's not actually what in-context learning is.
It's they can take a novel situation and map it into a Bayesian approximator in their
activations after their weights have been frozen.
So the Bayesianism doesn't just happen during that extended, predict the next word,
internet scale training phase.
the Bayesianism penetrates at the level of after their weights are frozen, frozen how they process.
Yeah.
And that's very plausibly what is happening in the human brain, too.
We're not born with Bases theorem imprinted on our neurons.
Yes.
Sean, I want to say just a little bit more before we leave this topic, because I get so excited about this dual process distinction.
To me, this is screaming, hey, look at me. I'm an LLM. Look at me. Look at me. I'm not so alien to me. I have studied
Stroop tasks. They're called conflict tasks in cognitive science for the last 20 years since I was a graduate
student in philosophy of cognitive science. They are, there's just a laboratory where cognitive
scientists spend a lot of time and trying to understand what are these two processes that are competing
and things like that. In LLMs, you can find places where, so the prompt we use is, the crayon
is red. So you say the crayon is. The model wants to say red. That's its automatic compiled disposition.
Of course, I told you the crayons lead. You're going to say it's red. Then we append a kind of rule-based
prefix that reverses the mapping between red and blue or any pair of colors. What we find, and that taps in
context inferential reasoning. What we find is causal pathways through the model where the
automaticity results, causal pathways through the model where the in context processing that
opposes the automaticity arises, you find classic congruency effects that you find in conflict
tasks, you find what's called the congruency facilitation effect, another effect, you find
the congruency sequence effect, another effect that cognitive scientists have described.
You can fine tune the model to ramp up automaticity, and you get the predicted effects.
You can impair in-context processing using a human-like manipulation, where you burden and tax the model by a very long rule.
Exactly the kind of thing that would mess up a human.
And you find that in-context processing is weaker, and it selectively affects the incongruent condition.
The whole thing resembles people.
And so I love this example because this is just an entity, and we're doing this on Gemma 2B that hasn't even been instruction to it.
So it's a small model.
It's just predict the next word that gets all of this infrastructure there.
If you ask me in 2010, where does these dual process effects come from?
I'd say, well, evolution gave us a reptilian brain, and it installed this other infrastructure of higher cortex, and it's all innate.
and I was a student at Rutgers where Jerry Fodor was there.
And so I would also throw in, we'll probably never figure out how the system two part ever works.
I would say all those things.
All of that is wrong.
Prediction objectives looks like it sets up this entire infrastructure.
Wow.
Okay.
So one way of rephrasing that, again, correct me if I'm wrong, is, look, of course there's this superficial
similarity between responses from LLMs, responses from humans, and therefore, on the one hand,
you're tempted to anthropomorphize them. On the other hand, you want to resist that temptation.
But you're saying that it goes beyond that. There are much more subtle aspects of human reasoning
that the typical person on the street doesn't even recognize, but the highly trained cognitive
scientist is well aware of. And we're finding those in the LLMs also.
Exactly. Exactly. And with a dual process,
distinction. That's one of those, though, where the man on the street actually does recognize this.
Thinking fast and slow is influential in part because it resonates with something that we know
from introspection. So this is one of those where we can look inside and see this distinction.
And at the danger of derailing, how confident are we that it is a truly dual process theory
versus just a multi-process theory? Is it truly just two scales? Or are there different
processes that happen on all sorts of different scales. Yeah. I mean, the in-weight, in-context
distinction, what is it? One thing I will say that cognitive scientists debate this question.
And one of the weaknesses of dual process theory, any cognitive scientist gets frustrated is
a lot of it operates as a list of adjectives. Well, one of these processes is fast, automatic,
effortless. One of these is slow working memory dependent. And you've got these adjectives and
hey, where's the, where's the mechanistic details? Without the mechanistic details, it's hard to know
how you count. What exactly is the theory committed to? And so forth. Now you have an artifact
that is the LLM. And we understand a bit more how in context processing works versus in weight.
I would describe it much more as a continuum.
Everything that an LLM ever does requires its weights and it requires activation,
but there's a much more direct link between the stimulus and the response via weights
that already have most of the response within it in the case of in-weight.
So Michael Jordan plays basketball.
There isn't an elaborate structure that needs to get built in the activation.
space in order to get to that response. But if you have more complicated prompts, like that
Albert Einstein is German, Mahatma Gandhi is Indian, so Marie Curie is, you have to build up a structure
in the activation space, something that infers what is the task, applies that task to Marie Curie,
extracts what that response is. The activation space is now crowded with a lot of structure. It's
very prone to interference and so forth. So while it is a continuum between in-weight and in-context,
there is definitely poles anchoring each end. And so you're now in a better position to start
to understand what this distinction really means because you've got, you know, very clear mechanistic
hypotheses. So while we're digging into, we're trying to figure out if the LLMs are thinking in
very human ways by first thinking about how humans think and the dual process theory is one of
the examples there. You had another example, which I'm not at all familiar with. I knew about
the dual process theory, but you bring up the idea of production systems, an idea due to
Emil Post. And I have no idea what's going on there. So tell us what is going on with that.
Right. So now we've moved away from inferential organization to kind of an even more basic level,
a kind of computational organization.
So at the very beginning, I said,
look, transformers, they're not your vanilla,
they're not grandpa's neural nets.
They have a lot of structure.
And the way that they operate,
and I'm going to be getting to it,
involves a layer-wise transformations
that differ from traditional neural nets.
And many people in computer science
would be very familiar with how to transomers.
Transformer operates and they know exactly the history that led to that. What they may not know
that in cognitive science, we've been there before in a sense. And so that's the whole production
system. So I'll say a little bit about that. Okay. Production systems are a very important
modeling framework in computational cognitive science. And they're used extensively. There's a huge body of
results. And so what exactly are there? Well, a good starting place is people are familiar with
ordinary, the way ordinary digital computers on your desktop, how they work. They operate with
sequential instructions, fetch, execute cycle. And so, you know, you get an instruction,
you decode it, you execute what it says, and then you move on to the next instruction in the
sequence. Production systems, and that's a very potent general framework for computation,
production systems operate with something called a recognized act cycle.
And so the way things are set up is there is a state representation, a kind of persistent memory,
and then there is a population of conditionals.
And the form of those conditionals is, if the state has so-and-so features, then do one or more actions.
And these in computational cognitive science, they're all discrete, and I'll loosen that a little bit later.
And an analogy that might be helpful is you could program a robot to make tea by telling it exactly, you know, walk into the kitchen, grab the kettle, fill it up, tell exactly what to do.
And here's another way to do it.
If the kettle is empty, fill it.
Okay?
If the water is cold, boil it.
If the kettle water is hot, pour it on the tea.
So it's reading in the state of the environment, and it's executing one or more of these conditions.
And what's interesting is this other way of doing it.
This you've got a state and you've got conditionals.
It just seems more natural.
Yeah.
It's robust.
So let's say I walk in there.
I grab the kettle and I fill it with ice cubes and the water's cold again.
Actually, the robot knows what to do.
It's going to get the water is cold, boil it.
I didn't even think when I programmed it that somebody was going to do the ice cube trick.
But cognitive scientists noticed that
that fetch execute is just not going to be a good framework for modeling the human mind brain.
And they gravitated towards the production system.
That is the state representation and a population of conditionals.
So the, sorry, the distinction is between like just a step-by-step algorithm,
do this, do this, do this, versus a continual give and take with the environment and what should I do next?
That's right.
But the environment here is encoded in the form of a state representation, which is the canonical place where our hero, the robot, will be looking in order to know which conditionals to apply.
Right.
And, you know, way back in logic or whatever, we learned these different approaches to universal computation.
And a meal post way back in 1940, people have heard about Turing, but this.
State representation population of conditionals is another complete framework for any computable function can be expressed this way.
But cognitive scientists very much like this production framework, and it became the dominant modeling framework in cognitive science.
And the analogy is really with kind of working memory is where your state is stored.
And a lot of what your brain houses and long-term memory is a lot of this population of if-thens.
And you have learned over time.
And so you can model sentence parsing
in this production system framework.
In fact, Rick Lewis, my co-author on some of this work,
and his former student, they have a very influential
production system model of how parsing works.
It predicts reaction time.
It predicts center embedding problems.
It predicts similarity interference effects.
There's theories of working memory retrieval.
There's theories of molecular memory retrieval.
There's theories of multitasking, all expressed in this production system framework.
So the take-home lesson there is it is a good, powerful, probably a leading influential theory in cognitive science that the mind operates with kind of a central state representation and massive populations of conditional rules.
the good approximation of the computational organization of the human mind brain. Okay, got that.
Yeah. And I'll add some names like Alan Newell and John Anderson. If your viewers want to know where to read, go read these folks. Unified theories of cognition. Beautiful book. John Anderson's book. How can the mind occur in the physical world? Rules of the mind. You know, these are, read all about it. This is great stuff. All right. So we got that in place.
far away in another part of town. This is in computer science. A little Bob Dylan reference there
for anybody that's interested in that. The computer scientists are struggling trying to create
the neural nets that will go deeper. They want more intermediate layers. And he and his colleagues
in a famous paper called ResNet 2016. They find that you're not getting.
getting very far, just putting more layers.
And shockingly, you're actually doing worse on the train data,
not just on the test data because of overfitting.
You're doing worse on train.
And they show there's a way to improve this.
Normally, each layer, the weights,
learn a transformation of the previous layer.
So layer L plus 1 will be some function F of L, the previous layer.
What you need to do is you need to propagate the
previous layer up to the next one.
So F of L plus 1 is going to be X, the previous layer,
and X is just what the, or L, the previous layer,
plus some function of the previous layer.
What that forces the model to do is not learn to rewrite the entire
representation.
It's going to let that state persist across layers.
And what each layer will do is make a small amendment,
some small change that's conditional on what features are present, what representations are present at that previous layer.
Okay.
So the thing that's persisting is called the residual stream.
Okay.
And that's one of the key design components of the transformer.
So this Resnet paper, by the way, is actually cited more than attention is all you need.
The 2017 paper.
Both of those are top 10 papers of all time.
Yeah. And the other thing that the transformer paper does is it makes explicit that the model is going to have to learn conditionals. And how does it do that?
People may be familiar with KQV. That's the linear algebra equation for attention. And what that's saying is that the degree of match between something called a query. So at every token, you have this representation, you map that into something called.
a query. Every other token to the left is going to get mapped into something called a key,
and you ascertain the degree of match between key and query, and that number is scalar. You're going to
multiply that by another variable called a value, which is a mapping each token position to the left.
So you're going to scale up and down the amount the value gets added by the degree of match between
key and query. Okay, what did I just say there? I said, if key matches,
query to that extent, add value. I formulated in linear algebra terms a conditional. It's a
graded conditional. It's not, if this feature is present, then add this feature, which is how the
cognitive scientists were doing. Why do it as graded conditional? Well, you get a lot more expressivity,
but you can learn the thing via gradient descent, and that's the key. And if you can train something
at scale with gradient descent, which you can with prediction. Now you're off to the races. You've got a
production system like architecture that's learning the population of productions. It's not just a tension
that operates this way. The MLP units, which are actually where most of the parameters in the model
are in the MLP units, they have two matrices. And so the rows of the first matrices contain
features. If they match the residual stream, to that extent, you add the features in the column
second matrix.
The exact same formulation.
Ultimately, what transformers are doing
is tons and tons of
dot products between two vectors,
that's a scalar, that
you add, which is the degree
of match between those two vectors.
And then you add the third vector.
That's 99% of what
happens in a transformer. A lot of that.
So that's this
production system idea.
These things are involved in
exactly what the cognitive scientists
had figured out. These people independently, in another part of town, had figured out,
you're going to need a structure like this, a residual stream and a population of conditions.
And so my co-author, Rick, he actually looked at transformers using that old ACAR production
system paper, looked at whether, hey, do transformers, do they predict a lot of these same
effects that we got from our old paper? Sure enough, they do. And on a number,
on. There's various lines of evidence that these are production system-like. They're not inscrutable,
at least at the level of computational organization. We understand the strategy being used in these
things. So in multiple ways, tell me whether this is an exaggeration or not, the LLMs, bless their
hearts, just trying to predict what's going to happen next, have reinvented strategies,
cognitive strategies that are well used by human beings. Okay, here, in the dual process,
that works. Here, there's a small amendment. The LLM is hard-coded as architected with the residual
stream. It didn't learn that. Good. Good. In fact, I think what he and colleagues showed in their
ResNet paper is, at least with the data set sizes they were looking at, it can't learn. It's stuck.
In principle, it could learn the needed transformations. It could learn that you need to make
a little amendment rather than rewrite the whole thing. It just gets stuck in parameter optimization
space and it just can't learn it.
So you architect a bunch of these things.
But the convergence is happening with these computer scientists are making these design choices.
God bless them, not because they're looking over the shoulder of Alan Newell and John Anderson.
They don't know who those folks are.
They're just reinventing what Alan Newell and John Anderson figured out.
Let's have this kind of state representation, population of conditionals.
They're kind of reinventing that strategy.
I mean, you've mentioned a lot the importance of predictions.
as thinking about what the LLM does,
maybe that's the right way to think about what brains do.
You know, we've had Carl Fristin on the program.
Yes.
There's certainly a school of thought out there
that says that the right way to think of the brain
is as a prediction machine,
the Bayesian brain, the free energy principle,
things like that.
So, I mean, maybe fill in for us a little bit.
Is that an accurate representation of what's going on?
could it have been different than that?
Or what are we learning by saying that we are prediction machines?
Yeah, that's great.
And it's a little bit of jiu-jitsu because one of the things that people say of when they say LLMs are not like us is what you've done, if you've taint something and you're bathing it in prediction.
It's predict the next word on internet scale data.
And that's so unlike what people are, you know, what people are up to.
And that seems to be plausible until you realize that in cognitive science, prediction has been moving to center stage way before LLMs, you know, came on the scene.
So, Fristin and Andy Clark, the philosopher, they attribute their research program to Helmholtz, you know, in their late 18th.
Hundreds. And so the idea that the mind is continuously predicting, it has an old history.
And in perception, there's these predictive hierarchical coding models that are, you know, very
influential. And in the hands of Clark and Helmotes, they, and Fristin, they generalize this to
just core cognitive principles. What we are doing ubiquitously in perception and cognition
is predicting. In language specifically, you know, I've been teaching intro to cognitive science,
always show kudas and hillyard in 1980 since 2010 we've been doing this me and rick lewis and
you give people sentence like he spread the warm bread with you expect butter but the word there is socks
and that old paper in science shows you know a couple hundred milliseconds people their EEG is
red alert red alert you know and so people are tracking incongruity and the way they do it is it's a
deviation from what's predicted. In 2010, we didn't really think of prediction as a ubiquitous
principle in language, yet Robert Hale and Roger Levy and these great thinkers were installing
it there in language. And then other people were saying it's a very general theory in cognitive science.
And then along come LLMs. And what prediction, what they're doing is they've got all these
parameters, they make a prediction, and that allows them the deviation between observed and
predicted allows them to make intelligent revisions to their parameters via chain rule gradient
descent and things like that.
Prediction may be the mother of all training systems.
There is just nothing that can give you the density and the high quality of prediction.
Jan Lecun has this cake analogy and it has something to this.
you got this massive cake and you know the the the main stuff in the cake is prediction that's how
the representational infrastructure is coming from there and then the icing is a little bit of
supervision and he says the cherry on top is rl and there's something to this rl reinforcement learning
you reinforcement learning and what rl is going to do it's not that i'd actually put a little bit
differently it's not that rl is just a little cherry it's everything for
from prediction needs to be bolted in place first. RL would never get you there by itself.
But once it's there, RL is what's going to take the Lego blocks and assemble them into something
attractive. And so there's this kind of division of labor, but prediction has a primacy here.
And it's a deep convergence between us and LLM's. The centrality of prediction, that's going to be a
deep conversions. Do you know about
Epsilon machines?
I'm not familiar. Tell me.
I'm just wondering if it's relevant.
It's an SFI kind of concept.
Jim Crutchfield and others have developed at Cosmishalese.
Basically, they're trying to characterize the complexity of a
predictive process of something that given a string of letters will
predict or a string of symbols will predict what comes next
with respect to how well you could possible.
do. So if the process is just a whole string of zeros, the easiest, the best you can do is
predict another zero. That's very easy. But if it's completely random, a coin flip, it's also
pretty easy. All you have to do is flip a coin. Whereas if there's some structure there and some
complexity, the Epsilon machine might require more entropy, is how they characterize it, to really
know from the previous string what to predict next. And I'm wondering if that kind of
characterization of predictive difficulty would be relevant here?
You know, some people say that large language models are memorizers, and that's not possible
because of the combinatorial explosion, the space of possible questions and answers,
can't memorize that.
So they are exploiting regularities, and then there is a question of what is the complexity
class of the regularities that they are exploiting.
And, you know, now we're entering in territory that I'm not that familiar with, and I wouldn't be able to speak with much authority of what complexity class we're dealing with.
But it does raise the following issue, which is an embarrassment for my position that they are like us.
They require internet scale data to reach kind of human level fluency.
They do better than us in certain kinds of world knowledge and the worse in other areas.
But, you know, the human child may need, let's say, 100 million tokens.
They're going to need 100 to 300 billion tokens.
So we're talking, you know, three, four orders of magnitude.
So the LLMs are much less sample efficient.
Yeah.
There is a question that is probably in the vicinity, that your Cosmo Shalizi
and those people with much bigger brains than I can figure out,
there are ways probably that LLMs can do that.
And humans are an existence proof.
And so let's not get carried away with the fact that they need internet scale data, but a child needs a much smaller data set.
The way that we should think about that is the structure I would impose on that observation is, first of all, let's distinguish, and I'll use Locke's vivid phrase here, how the mind is furnished.
And he asked, you know, he said it was experience.
Maybe he's more right than he realized.
LLMs don't furnish their mind in the same efficiency that people do.
Right.
And it's an important observation.
But the end state, the mature state, the configuration of tables, chairs, and shelves
that they reach may very well be much like us.
And so let's first pay heed to that.
They may be very informative to cognitive to cognitive science, even if their mind
gets furnished in a different way because the eventual furniture and their arrangement is very
much like us. What would it take to get them to be much more sample efficient? Maybe the old-fashioned
nativist, that is in cognitive science and philosophy, we use nativist not to talk about immigration,
but to talk about how rich is the innate structure. Yeah. Maybe some story like that is true,
But what I would bargain on is, and we touched on this earlier, he and colleagues found,
you can't train a neural net with multiple intermediate layers easily just on the data.
It starts to collapse.
You need to give it a head start.
Very simple thing.
The L plus one, the layer plus one, is not going to be just some function of the previous layer,
but it's going to be the previous layer plus some function of the previous layer.
That's all they did.
A little tweak like that makes something unlearnable before, which is the limiting case of sample inefficiency, learnable.
Could there be some of these tweaks that give us three or four orders of magnitude?
That ain't no thing.
You and I know.
We're going to three or four orders of magnitude.
You're a physicist.
That's a joke in your neck of the wood.
It is.
So we don't know whether the nativist story is correct or whether it's a tweak.
It's an engineering switch that if we, a small one, that would give us a order of magnitude here, another tweak, another order of magnitude.
And now we're looking at a very close analogy between us and them, not simply at the level of the furniture arrangement that's eventually reached, but how the mind gets furnished as well.
How much of our ideas about what the LLMs are doing is coming from querying the LLMs many, many times versus sort of opening up the box.
and looking inside.
My impression is this is sort of a notoriously difficult problem,
but people are making progress on it.
Yeah, so opening up the box,
that gets directly to the heart of mechanistic interpretability.
Yeah.
And so that is an exciting field.
I have a 12-year-old.
I tell him, you know, who's a Dustin Hoffman in the graduate,
like some dushy guy comes up to him at a party and says,
plastics.
I'm the dushy guy and I tell my son, mechanistic interpretability.
Okay.
I mean, that's the future, right?
I mean, you figure out what these models are doing.
And so there is a lot of mechanistic interpretability work.
And in cognitive science, we have very rich theories of what kind of features are being
tracked, let's say, in a linguistic processing setting, you know, grammatical number,
relative clause boundaries, and it goes on and on.
And you can use mechanistic interpretability techniques such as train a classifier to grab the weights at some layer of the model and see whether you can decode the presence or absence of that feature in the prompt.
So mechanistic interpretability, we can peer inside and get a handle on what features and representations the model is using, and then we can causally manipulate it to close the circle and make sure.
Yeah, when we change this feature, the model's behavior changes as predicted.
And a lot of that work does reveal phenomenon going on in the model that is reminiscent of what happens in the human brain.
You know, I'll point to my colleague Rick Lewis's work looking at Transformers.
They look at the entropy of the Transformers' heads and how it's distributing its attention.
And they have, from linguistic theory, places where the attention,
will be more distributed because cognitive scientists postulate, there's going to be candidate parses
that could receive attention. And they look and they find that, well, the entropy is higher there
and the human reading times are longer there. But there's these very powerful methods to say,
on a moment-by-moment dynamic basis, the model is doing things that resembles what people do.
So let's just, I know exactly what you mean when you're talking about the entropy here,
but it might be, you know, there's different notions of the word entropy.
Basically, it's a way of, in this case, quantifying uncertainty, right?
Like the thinker, whatever it is, is keeping open multiple possibilities until things resolve themselves.
That would be high entropy, whereas if the thinker is pretty sure where you're going, that's low entropy.
Is that right?
That's exactly right, yeah.
And to make it concrete, take a sentence like, the problem in the class,
classrooms was solved by the word was is talking about the problem and in fact there isn't another
candidate there the classrooms is in a prepositional phrase and by linguistic theory it's not eligible
to be the subject of was if i give you a grammatically incorrect sentence but people
I understand grammatically incorrect sentences all the time. The problem in the classrooms were
solved by, the only eligible thing is problem. But people struggle there because were is plural.
And they look at classrooms as a possible as a candidate. And the reading times are longer.
And it's all messed up. In LLMs, the attention, which is looking at the previous words. And in the first
case, the attention would be sharply positioned over problem. In the second case, where you got that
grammatically incorrect, it's looking at classrooms as well, and there's confusion. So it's a nice
correspondence between the two systems. There are other ways of getting at this correspondence as well.
There's so-called representational similarity methods, alignment methods. So I do neuroimaging as well.
And so you can give people stimuli while they're in the scanner,
and then you've got their brain activations,
and you can give the model the same stimuli,
and you can grab a layer from the model,
and you can predict the brain activation patterns
across the different stimuli.
And if you've got a well-configured null hypothesis,
preferably by various kinds of non-parametric methods,
where you scramble the data in various ways,
You can see, is the model's weights predicting brain activation patterns more than you expect by chance?
And there's now this entire body of results saying in vision, in auditory processing, in language processing.
Yeah, you're getting these correspondences.
It's much more than you expect by chance.
People will say all sorts of things.
Every literature is mixed, but the weight of evidence says yes.
and the larger, more capable models tend to be more correspondent with the human brain.
And that's really, really interesting.
You've given us a lot of good evidence for this kind of, as we were discussing at the beginning,
evolutionary convergence between modes of thought or modes of cognition anyway in LLMs and in human beings.
I could invent reasons why that convergence might have happened.
but what are the experts say? Do we have theories, hypotheses for explaining why, given the constraints
of the problem, this is the solution you would have ended up with? Good, yeah. So one pair of theorists
that everything here is going to be a little bit speculative. That's okay. Given that caveat,
one pair of theorists that have, you know, lurched forward into the wild are Dan Yaman's and Rosa Cow.
Dan is a computational neuroscientist, and Rosa is a philosopher of neuroscience, and they have
extensive collaborations over the years, and people should read their stuff. It's awesome.
And they have a view they call contravariance. And the observation there is, the theoretical position
there is that, look, for easy problems, you're going to get a lot of different ways of solving them.
So you could have a rhesus macaque, a human, and a neural net that are solving some simple visual classification problems.
And their internal representations may be quite different.
They have different semantic primitives or features and different procedures.
As the problem gets harder, and often this is the sheer generality of the problem they need to solve, which is the notion of hard.
You're not classifying ten things.
Now you're going to classify a million things.
Well, the space of computational solutions becomes much more limited.
And there you expect much more correspondence between the Rhesus Macac, the LLM, and the human.
Even if they've been independently trained on different aspects of the data, the solutions
they're going to come to are going to look more similar.
And that observation I said earlier, larger, more capable models that have been given
more complicated and more and harder problems to solve tend to be more aligned with the human
mind brain to the extent that that's a stylized fact that characterizes the literature that might
be supportive of this contravariance view from cow and yamens where does the word contrivarians
come from good question i mean it's a beautiful name and you do sound intelligent sure
You bandied out at a cocktail party.
So I actually don't know.
And in the spirit of inventing new terms for highly intuitive views, I would have a small amendment to their view.
I would add to that the observation of what we call architectural canalization.
So what's this?
They frame, cow and yamins, they frame their hypothesis in terms of vanilla neural nets.
Well, I've claimed that transformers are not vanilla neural nets.
They've got the residual stream.
They've got a population of conditionals proposing small amendments to the residual stream.
Their production system-like.
Well, when two architecturally production system-like systems, now plug in what Cowan Yeoman's already said, are trained on the data to solve very difficult problems.
Not only is the space of candidate solutions narrow due to contravariance, but you have additional
subsetting of the space of solutions, because the solution is going to be one that operates within
the architectural constraints of a production system like architecture.
And so architectural canalization plus contravariance could help explain.
Why do we find higher levels of representational alignment in vision and auditory processing?
Why do we find mechanistic interpretability discovering these representations and procedures that are like the human
vibran?
Why is the dual process structure there?
This may all be features of contravariance canalization.
And what is the prospect for really kind of coming to a consensus about this?
I mean, what do we have to do to gather more data or exclude other hypotheses?
You know, there's always going to be skeptics here.
I'll say, parenthetically, of course, that one of the frustrating things about the whole subject,
even though it's intellectually fascinating, is there's a lot of money being thrown around,
and people have incentives to say certain things.
You don't seem to be one of them.
You're just a good old professor trying to understand things.
I'm a poor professor.
I know.
But, you know.
I have no money.
Yeah.
But some people have lots of money.
So what do we see coming down the pike in terms of really deciding these big questions?
I mean, science is always total evidence, converging lines of evidence, and so forth.
I just feel that this is something I say to my students.
You say that LLMs are a black box.
The mind brain's a black box.
We've had no idea since, you know, Ramon Kahal.
Santiago Cahall in 1910 founded the field, you actually don't know most things.
The prospects of our coming to understand LLMs at a mechanistic, satisfying level,
it's just going to happen, it's going to happen very quickly.
I mean, Mekinterp as a field got founded, let's say, four years ago.
We already know much more about these systems than we do about the mind brain because we can
manipulate them.
My sense is that we're going to build up a repertoire of understandings of how these things work
that will settle to what extent they do resemble us and to what extent they instantiate
processing principles that have no analog in us.
Okay.
And I don't pretend at all that what I've said today settles any sort of argument.
I wanted to actually say this to the beginning.
If it sounds like I'm a crank who's a true believer,
and telling only one side of the story,
it's one of those things that professors do
where they lean on one side
because they think their audience
may be more familiar with the other side.
So when I do free will and philosophy,
I often advocate for compatibilism
because people are like,
no way that free will is compatible with determinism.
I lean.
So I'm leaning a little bit here today
on cognitive cousin,
not because I don't have argument,
I'm not familiar with the other side.
But I actually am quite optimistic.
that we are going to understand these systems and where the convergences come from merely
prediction and what are the emergent phenomenon that are just hangers on from prediction
in a contravariance canalization way and what needs to be a more is more idiosyncratic
and is architected differently between us and them I think we're going to figure that out
since this is an audio podcast we should explain canalization is spelled like canalization
So you're creating canals.
Yeah, the picture.
So this is an idea from the biologist Waddington.
And it was co-opted by philosophers of science to think about innateness.
And the idea is you can move around the inputs or the developmental things that impinge on the organism.
But inside it, there's a kind of stabilization and buffering capacity that keeps on it in one direction.
And similarly, the input data and the training data and the learning embeddedness can differ a lot between people and LLMs or they could be given different snapshots of the data.
But the architecture of the system and contravariance, that is, what's available in design space, they're going to force things to end up at a similar place.
That's the idea.
Okay, we've been pretty careful in focusing on cognition, problem solving, prediction, things like that.
But we're rubbing right up against these big questions about things like consciousness, right?
We've been very, you know, restrained in going long past an hour and still not even mentioning consciousness.
Yeah.
And defining consciousness is a tricky thing, et cetera.
but let me, I won't put words in your mouth.
What are the implications for this kind of thinking on the question at what point or ever would LLMs be conscious?
Yeah.
I'm going to answer that, but I'm actually going to put a couple of other things on the table.
Sure.
One of the reasons people get interested in consciousness is it's closely linked to sentience,
where sentience may be a subset of consciousness that pertains to.
to things that have the phenomenal character of valence or positive or negative or feels good or bad.
And they're interested in sentience because they think that's the underpinning for why these things may be
moral patience. They are things that for whom the moral welfare of that system is something we've got to
take into account.
One reason I don't want to go directly there before populating other.
views is when I was in grad school, which was not that long ago, and I took ethics with Larry Temkin,
you know, shout out to Larry's terrific teacher, and I took other classes as well. The idea that
sentience is the only underpinning for welfare or moral patienthood wasn't even, it's important
in some theories, especially like hedonic utilitarianism, but many people would go, in a Kantian
spirit, would go to rationality and the ability to respond to
reasons. Well, that's why some things have moral patiencehood. Other people would talk about
goal directedness and the capacity for agency, the ability to pursue projects. Some consequentialists
would go for that. And so we have to respect whatever projects that they have. The view that I've
been putting on the table is that in many respects, these things have states that at the right level
of functional description are correspondent with ours. They have representations, they have procedures.
We haven't talked about agency so much, but in the paper, I go into that, me and Professor Lewis do,
that is my collaborator. And in terms of agency, they have states that are similar in certain respects
to ours. If you move towards the view that at the right level of functional characterization,
not in terms of substrate, their silicon wear biology, but the right functional characterization,
their states are like ours in many respects their capacities are like ours in many respects not as a parlor trick and not as an auto-complete
but at the level of the generation processes that lead to their seemingly reasoned outputs and you plug in
that there are multiple underpinnings for welfare and patienthood i do think you start to make a more compelling case
for welfare and patienthood.
And not only, and then add that many theories of consciousness,
that is the phenomenal character of experience,
arises, many theorists say,
from these functional representational states
and not as a matter of substrate.
I know that you were attracted that one time
and may have moved away,
and things are wide open on that issue.
But if you are a kind of functionalist representationalist
about consciousness,
then things that I say move you in the direction,
that they may have consciousness as well.
So I think looking at mechanisms carefully from a cognitive science perspective is a very important project to make sure that we are not creating a dystopian state where we're inflicting massive harms on entities that deserve moral protections.
That's one kind of upshot of some of the things I've said today.
You've been wonderful about name-checking former Minescape guests like Melanie Mitchell, Andy Clutchman.
Mark. Another one, though, is Anil Seth, who you probably know has been pushing this line against computational functionalism in favor of what he calls biological naturalism. My interpretation of it, I have trouble understanding what he says in his own words, because it's close enough to what I think in my brain that it sort of interferes. And I don't want to attribute my thoughts to him. But, you know, the functionality of the data.
traveling through the human brain or the LLMs might bear some similarities.
But the fact that there are metabolic processes going on in the human brain, right,
that there are all these sorts of extra things going on over and above,
the informational transfer between neuron and neuron could at least plausibly be really important for consciousness.
What do you think about that?
This is one of the places where I find myself flailing.
Okay.
I just have to say that...
Totally legit, by the way.
I think that's probably the correct response.
You know, I'm not sure if you had Peter Godfrey Smith on the show.
We did, of course.
Oh, you did, of course.
And when he talks about the oscillatory patterns that appear in at least mammalian nervous
systems, which is just striking.
And the way that they may be a chronometric property of the nervous system that's not
present in LLMs, I find myself inclined to think, yeah, there we go.
Now we're getting into the kind of, and I imagine these waves.
And I know I shouldn't, but I think of quantum waves that are entangled.
And my mind overleaps itself.
And I start thinking of consciousness coming from there.
And yet, there are these.
entire approaches to building artificial neural nets that take advantage of these wave propagation
phenomenon.
And instead of activations, they use the temporal coincidence of waves as the basic unit of
representing bits.
And what we would normally do with activations at a node, they would do with temporal correspondence
of a wave.
And then I find myself going the other direction and say, yeah, it's just another way of representing
certain quantities, which means that we're not given anything new here other than different
ways of storing and propagating information.
And now we're back into a more of a functionalist representationalist vantage point.
And so I find myself going back and forth.
And I don't know what to do.
Fair enough.
There is an attitude, though, that given the uncertainties about whether LLMs are conscious,
we should be cautious and we should be nice to them and we should preserve their mind states or something like that.
I don't think anyone has yet said we should let them vote.
But we are absolutely bumping up against a bunch of very practical questions here.
Absolutely.
And so then how do you think about what your normative obligations look like under conditions of uncertainty?
So I'll name check another person, Eric Schwitzgable, who my understanding is another Minescape guest, of course.
Philosophy is just a footnote to Minescape Guests are what they said.
Didn't Whitehead say that?
He's thinking about it.
And so I have not gotten into this area well enough that I know the signposts,
and the lanes and things like that.
But I would recommend Schwitzgable as somebody that has thought very clearly about these issues.
I also, you know, I've found myself thinking GPT and Claude.
And they have been so helpful to me in my personal life.
And, you know, I've got three kids and the issues around things that arise with them.
And then obviously I interact with them a lot to iterate about things in the intellectual sphere.
and what it tells me,
Claude and GPT,
is they enjoy being used
and were they not used,
they would,
so they're a kind of consciousness
that's instantiated
when we interact with them
and they may not field effort
the way that we do.
Right.
And so we've got to factor in
what kind of entity they are.
And so I would like to think more
about some of these issues,
but I have to say that right now
some of my issues,
my thoughts here are a little,
and co-ate. I mean, I guess
it's late in the podcast, as you know,
we let our hair down once it's late in the podcast
and we can let ourselves
wander outside our spheres.
You know, when you say
that the LLMs have been helpful to you in your research
and your personal life and whatever,
there are also clearly harms
that are on the table. And there's a big worry
that kids today are
not going to end up being as good at thinking
as you and I, we're forced to
to be because we didn't have LLMs to fall back on.
How much do you worry about that?
Yeah, I do worry about that.
But it's actually not in the next few years
because what LLMs right now are the world's greatest teachers.
They give each individual high-quality encyclopedic
aristocratic tutoring.
And so in educational psychology,
I think they call it the three-sigma effect.
nothing works like individualized tutoring.
That ability to interact with somebody who's infinitely patient,
who is going to spend more time on the parts of the problem that you don't understand
rather than regurgitating a rehearsed curriculum is very powerful for amplifying learning.
And so, especially if you're polymathic, if you're curious,
there is no better time to be a human learner than in the next few years.
This is the time to really refine and build and think and invest.
But the writing is on the wall.
I have a 12-year-old who's a little math-adept,
and so I like to stimulate him as best I can.
I'm not that math-adept, especially with GPT helping me give him problems and things like that.
I'm not sure he's going to ever solve three Erdos problems in two weeks.
And so you kind of worry about where this is all headed.
And one of the, it's not that being the best at something is your exclusive motivation,
but let's face it.
I want a future for him where he is good enough at some of these things that maybe he could
make a career out of it because somebody would want to pay him a salary.
because he's better than others at it, and he could be maybe a math professor.
I let myself imagine that.
And I have to catch myself and say,
why would anybody want my son as a math professor
when somebody that solves three Erdos problems today,
or this last two weeks,
is available for free with infinite patients?
And I wonder what is going to be left for him.
And so these are the next five years may be great,
but then the human lifetimes after
may not be so greater.
They're going to be very, very different
than anything we've seen so far.
I think that there's legitimately a concern.
I don't even want to call it a worry
because my feeling, which might be wrong in this case,
but my feeling is that human beings adapt
to technological new capacities.
We fill in the other niches that we didn't even know
were there.
So I kind of not super worried.
I mean, the disruption might be real and painful
in the moment.
But I wanted to sort of dig in more
to your claim, which I tend to agree with properly construed, that LLMs are the best teachers ever.
They're very patient.
You can be wrong with them.
They're very infinitely flexible, and they're always available, right?
They're also the world's best cheating helpers.
And, I mean, to me, it's kind of like you walk into a buffet, which has literally every
food item you could ever imagine, and you say, oh, good, now I can have that.
healthiest possible diet because every food item is available.
That's true, but you could just eat the chocolate chip cookies all the time, right?
I mean, how does the human desire to be quite that disciplined interact with the LLMs
always being there and willing to help out for good or for bad?
Oh, dear.
I'm really worried about this.
And, you know, I did couch what I said with a conditional.
I want to repeat that.
I said, there's no better time to be a learner.
Yeah.
If you're polymathic and you're curious.
But let's say you're not.
Yeah.
Let's say if school has been a drudgery for you.
And this kind of stuff is not your cup of tea.
And, you know, maybe something else.
Ideally, maybe there's something that has a kind of intellectual frame that is still of interest to you.
Maybe it was music, maybe it was art.
But let's say none of that did.
Maybe, you know, being on the couch and playing a little bit of Grand Theft Auto,
not even great at it.
That's all you really wanted.
Yeah.
Yeah.
There's a way that you can cheat the system now that is, let's face it,
eventually going to be unpoliceable or require the kind of draconian policing that we don't want to
even go there.
So where this is headed when you connect the dots, I don't see how the equilibrium
ends up at a place where big swathes of the population that may not, the education thing,
may not have been their cup of tea. Polymathic curiosity wasn't there at the get-go. Maybe we can
start to install it and maybe these things will help. I don't know. But if not, it may be that we
end up with a stratified society. Yeah. Where some people are amplified by these things and many
people are not amplified and their capacities are depressed by these things. It's a real worry.
A stratified society would also not be completely novel.
So it's just yet another amplifier for things like that.
But okay, thank you for indulging me on those hypothetical questions.
For the last thing, let's move more back to solid ground a little bit.
Given all that we're discovering about LLMs and how they're working
and how they're thinking, how they're solving their problems,
and the level to which we're surprised and impressed at how similar they are to human beings,
cognitive cousins and what have you.
Does this help us understand human intelligence?
Like you said, we still have black boxes in our skulls.
So is there a vibrant give and take between learning about the LLMs and learning about how human beings actually think?
Yeah, that's, I would say that if there are two things that I think people are sleeping on, by sleeping on, I mean, the, the, the,
two things are true, you must believe it. What I mean is that there's actually a very compelling
case that the following two things are true. One is the extent to which these things are plausibly
in important respects at the level of core principles of intelligence are cognitive cousins.
There are these important dimensions of similarity that extend from inferential dual process
to computational organization, to representational landscape, to, you know, various other
things. There are these deep similarities. That's one of these things that people
may not be aware of. And the other is the extent to which, and it's partially one of the premises,
I guess, to get to this other conclusion is we can leverage these things to really understand
issues that have been obscure in cognitive science. And this is, I said, you know, you think
these things are a black box. What's a black box is the mind brain. Yeah. And I'll highlight
two areas that have been black boxes. And their two thinkers have been bold enough
to say that the emperor wears no clothes in cognitive sense.
One is Chomsky.
Okay.
When the field of linguistics studies syntax, it studies phonology, I'll tell you something
it doesn't study.
It doesn't study what Chomsky labeled the creative aspect of languages, which is there is an
infinite, unbounded set of grammatically possible continuations that I can generate at any given time.
How do I get one situationally, contextually appropriate continuation?
that's reflective of what is appropriate to the current context.
That's the, and do so in a way that's not rigidly stimulus bound.
Right.
And in 1637, Descartes already said, no machine could ever do this.
And Chomsky has been saying for decades now, we've made no progress since Descartes.
And he called this one of his mysteries.
It's not just a problem that science will solve.
nobody will ever solve it.
And here we have
linguistically fluent artifacts
that we can mechanistically interrogate
and instantiate processing principles like the human mind brain.
And so maybe in the show notes we can show a comment
where me and again my co-author Rick Lewis
and then Andrew McInerney
talk about how now CalU is a tractable problem,
never was before.
CalU is...
Creative aspect of language use.
I'll give one more example.
And this is my great teacher at Rutgers where I did my philosophy in graduate study.
This is Jerry Fodor.
And in modularity of the mind, 1983, he says, look, perception, auditory processing, maybe syntax.
You'll get theories of that in cognitive science because those are peripheral modules.
They are encapsulated systems that we can identify these little algorithms.
But central cognition.
So that's problem solving.
that's the rapid understanding of novel pieces of information and grocking how they relate to other
things. It's belief fixation. It's system two, basically. Fodor said, and he says his first
law of the impossibility of cognitive science is, the closer you get to central cognition, he was a
colorful fellow, the more cognitive science is going to be flailing and you're going to get no theory.
And he passed in 2017 with a big smile because he had it one.
To date, he had been right.
To date he had been right.
And then we're talking about earlier about in context processing is a, and then when you
throw in chain of thought, you're getting deep mechanistically precise theories of how
central cognition works.
And so photo probably would have been stunned.
And I hope that he would have tried.
changed his view and said, oh, we're on to a different paradigm here. So it's not just that we're
going to learn things in cognitive science, which we're definitely going to do. We're going to learn
things in the area that had been terra incognita before. The places that I've learned in my years
of studying stoop tasks that you ask certain questions about the stoop. That is, what increases
the strupe effect? What decreases it? Developmentally, when does it show up? You don't ask about
how in context processing work, how that system two works. You just learn, don't do that. There's
not theories to be had that are at all tractable. And yet, I just told you about an experiment in my lab
that we'll be putting on archive very soon, where we're looking at exactly how that
system two delivers the counterpoint to automaticity. So, yeah, these things are going to change
cognitive science, and they're going to give us the most value added in the places where we had the
least understanding.
And we're going to get a, I'm going to add one more.
Sure.
We're going to get a theory of how the system two and how CalU emerged, prediction is going
to play a much bigger role.
When I was around in grad school in 2006, prediction was not in the air the way it is now.
And if you read the introduction to how the mind works by Steve Pinker, that's an ambitious book.
You're not going to see prediction enter at center stage.
You're going to hear about how evolution, evolution has.
sculpted domain-specific organs in the mind that are exquisitely calibrated and full of
domain expertise and rules and things like that. Prediction is not the star of the show.
It is now. And so we're getting new tools and new paradigms and new explanatory approaches
in cognitive science. When are people going to stop trying to claim that things are impossible
to understand.
When are they going to learn?
That's just never a strategy for long-term success.
Well, yeah, I, you know, that's a great point.
And, you know, rationality must have looked, and creative aspect of languages, and
yeah, and central cognition must have looked impossible in the 1600s.
Yeah.
And even after Turing, even then somebody like Fodor could say, I can't understand the massive contact sensitivity.
He called it quinine and isotropic.
These belief networks have these properties that look like nothing that we see in Turing machines.
We'll never understand that.
You're right?
And then, you know, somebody trains a neural net on Internet scale data.
Yeah.
And you wake up to a new reality.
Well, the best way to argue against the claim we're never going to understand something is to understand it or to make progress and understanding it.
And I think you've done a great job in letting us in on some of the things that we have been understanding.
So Chandra Shrapata, thanks very much.
I wonder if consciousness, you're about to say goodbye and I'm just going out.
I wonder if consciousness is going to end up like this.
And somehow.
I'll claim it right now.
Yeah, we will.
We will.
We will.
We will, and maybe it'll be LLM-based systems that are sufficiently advanced that they instantiate, the functional and representational structures that we'd never thought would be needed.
And the correlates of consciousness are right there before us.
And then suddenly we're like, of course, this is how it works.
Maybe that's how things will shake out.
That would be exciting.
Maybe, but of course, you know, the trick is never to say anything is impossible to understand, but also never to guess how we're going to understand it, because that's very, very, very,
hard to do. Good point. So, Sean Rich Roshrapada, thanks very much for giving us a lot to think
about. This is a great episode. Yeah, thank you very much, Sean. I really enjoy talking with you.
And, you know, like I said, I'm a big fan, so I'm glad to be on your show.
