Dwarkesh Podcast - Grant Sanderson – AI and the future of math
Episode Date: June 30, 2026Always so much fun to chat with Grant.AI has been making much faster progress in math than in other fields. As a result, mathematics is showing us, very concretely, what AI progress in other fields wi...ll look like. Even within mathematics, there’s a jagged landscape. What does it look like?What is the nature of the most important conceptual breakthroughs in the history of mathematics, and how different are they from what AIs are currently able to do?Does AI (on net) increase or decrease human understanding of the field?How big is the overhang from having AIs systematically try to connect ideas already in the literature?And what advice does Grant have for aspiring mathematicians, coders, and other students who are passionate about fields that are being most transformed upon by AI?Watch on YouTube; read the transcript.Sponsors* Gemini 3.5 Live Translate is what I wished I’d had on my last trip to China. It detects more than 70 languages and translates them in near real-time… and it preserves your original pacing and intonation. If you’re building an app that needs live translation, you should check out Gemini 3.5 Live Translate. Get started at ai.studio/live* Cursor’s harness lets me use models for a huge range of tasks at the podcast. For example, Cursor cuts out the ads from each episode I produce so I can post them on Bilibili. It also helps me prep for interviews — I have a repo full of books and papers that Cursor sorts through to find the exact right file for any given question. Try Cursor yourself at cursor.com/dwarkesh* Jane Street sponsors 3Blue1Brown, so Grant has gotten to spend a lot of time with various Jane Streeters. He actually just recorded an interview with a few of them, so when we sat down for this episode, he told me about some of the things he learned, like how Jane Street keeps their role definitions fuzzy to make sure their people keep learning and growing. Go check out Grant’s full interview at 3b1b.co/janestreetTimestamps(00:00:00) – AI is discovering new proofs. Is that AGI?(00:11:32) – The verification loop on conceptual breakthroughs can be a century long(00:26:12) – Will we understand an AI proof of the Riemann hypothesis?(00:38:08) – Can AI find the hidden bridges between fields?(00:53:48) – Why real-world tasks don’t fit into RL environments(01:07:07) – Good writing requires theory of mind that AI still lacks(01:16:02) – Why learning will still depend on human curation Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
Transcript
Discussion (0)
Today I'm chatting with Grant Sanderson, who runs to Blue and Brown and is now working on a new project documenting the progress AI is making in math.
And I wanted to talk to you about this because AI has been making the fastest progress in mathematics as of any other field.
So whatever is happening here and whatever we were seeing AI progress happened or not happened would tell us about what will happen to the rest of the world as AI gets better and better.
So I wanted to start with this question I asked you when I first interviewed you three years ago.
And I asked you, once we have AIs that can get gold in the international math level,
if you had. Wouldn't that just be HGI?
Wouldn't this just be able to do anything any human can do
given how hard these problems are? And you had
an answer which in retrospect turned out to be
very wise and
correct, which is like it'll be another benchmark, like
all these other benchmarks that they are passing.
Obviously, yeah, has gotten better in general way since then,
but there won't be some
aha moment when this happens.
First, I think I'd be
curious to get your
heuristics on why that turned out to be
true. And second, I'm curious how
long you think this narrowness can continue to be true. So by the point that AI has solved
a million prize problem, do you think it's still possible that at that point there's lots of tasks
that humans are doing that AI still can automate in the economy? It's an interesting question
because it's hard to answer without knowing what the solution looks like ahead of time. I mean,
if we take the IMO, that's something where I think the spirit of your question three years ago
was in looking at how some of the solutions to these problems really seem to require creativity.
And the designers of these problems, they'll try to have them come up with things that you can't train for as easily.
I think the dirty secret with the IMO is that you really can train for a lot of them.
And so with the whole AI and math project undergoing, I think as you point out, one of the reasons it's interesting at all is that there's a spiky frontier to AI.
Math is just right there in one of the spikes.
But there's kind of a fractal nature to that spikiness because when you zoom into the specific progress within math, you have some things are a lot easier than others.
So if we just think about IMO, which is old news at this point,
it's kind of like two years ago that they were really like doing quite well.
They would have gotten a gold in 2024 if or not the following reason.
They're very good.
They're just like cold-solved geometry, basically.
And the IMO has these four categories of problems,
that's geometry, number theory, algebra, and combinatorics.
So like geometry just solves in like 19 seconds in 2024 because it's kind of a brute force
solver.
And the dirty secret is for students, there's also sort of a brute force way that you kind of can go
edit. Combinatorics is the one that's the wildcard of much more like playful, puzzly seeming
problems. And there were two combinatorics problems on that year's test. There's not always.
There's four categories, six different problems. So it's kind of a toss up which one is going to
have two questions. Had it been more geometry questions, they would have gotten a gold that year.
But it struggles on those combinatorics ones. And, you know, someone who's trying to keep that
torch of the last holdout of like math for humanity might say, well,
Those are the ones that require the more creativity.
Even then, though, I think the spirit of your question on, like, if they're solving, you know, a Millennium Prize problem, does that also service a lot of white collar work?
It suggests that whatever the rate limiter is between where we are now and that is the same as the rate limiter for making things better at white collar work.
We can maybe, like, paint a couple different ways that, like, if we focus on, I don't know, Riemann hypothesis, like, what would it look like to solve that?
one possibility would be
these things are extremely good
at a specific domain of knowledge
and just knowing it very deeply
and then knowing another domain
and knowing another domain
and you've pointed this out. It's like bizarre to have something
with this superhuman breadth
that knows all the field so well
that's not just finding those lightning bolts that connect them.
I think we're starting to see sparks of that
of actually finding connection between the things
that it's an expert at. I'm sure we'll talk about it.
If the nature of the solution
to the Riemann hypothesis was something like that, that feels pretty distinct to me than
what's necessary to get good at white color work.
And there's a reason to believe, actually, that that might be the nature of the solution.
I don't know if you know the story of, like, Hugh Montgomery and Freeman Dyson at the IAS.
No.
This is a side tangent, but it's just kind of a fun story on how, I don't know if it was
over lunch or something like that.
Basically, you have this number theorist who is pointing out just trying to understand
the statistical correlation between pairs of zeros of the Riemann's theta function.
So the Riemann hypothesis is all about, like, do all these zeros sit on a straight line?
And he's finding this, like, this quantitative question you could ask about.
And he writes down a formula, it looks like one over sign squared or something like that.
Freeman Dyson, a physicist is like, I know that expression.
That expression comes up and studying the eigenvalues for random Hermitian matrices,
which was something that comes up and studying the energy levels of, like, a nucleus.
and the idea that the statistics of those two seemingly different things were the same
sort of prompted a potential exploration on,
hey,
are there aspects of random matrix theory that might be relevant to like remand Zeta function?
And I think it's a little bit of an open question,
like is there fruit to be had there?
But that kind of bridging together from two different fields,
like if it turned out that the solution to the Riemann hypothesis
was exploring an idea like that even further,
that has this character of kind of how you expect LLMs to be good at math.
It's like they're an expert at the quantum physics, they're an expert at the analytic number theory.
They should be able to see that similarity in a way that doesn't require like Montgomery and Dyson to be having lunch and like happening to talk about that.
That's totally different from white color work, right, in terms of like the extent to which you maybe have a hard time using an AI as an editor.
It's not because they know everything and you just need them to find that lightning gold in between.
Different possibility would be what's the right analogy?
Maybe if we think of Fairma's last theorem
between the moment of Fairma
phrasing the question and then what the solution
itself looks like, where
ultimately the solution involves such
heavy machinery in math, right?
So the beauty of, that problem is you can phrase it so
simply, you ask about, you know, X to the
end plus y to the end equals Z to the end.
Do you have integer solutions for this when
N is bigger than three?
And it's something
you might expect there to be an elementary
number theory approach to it, but just
as far as we can tell, there's just not.
Whereas the actual solution, you know, maybe there is something simpler,
but this might be what it has to be.
There's such a complicated set of ideas that build on like centuries of work
centered around elliptic curves.
And then this other like mountain of ideas centered around these things called modular forms.
And like both of those mountains have to be built before you can ask the right question
that connects it.
So if the solution to the Riemann hypothesis involved building a new mountain,
like that's a kind of skill, like the ability to like come up with the right new ideas
that feels sufficiently different
from like the character
of how they're intelligent right now
that's not like that's what you need
from your hired video editor per se
but that like if it's capable of building mountains
that are the correct new theory
that like crystallizes how we should be thinking
about a subject
that's just such a level of intelligence
that then it starts to feel
like it would be surprising
if that didn't permeate
into other aspects of the economy
besides like just the mountain building
for math itself.
Yeah. Or at the very least
even if it couldn't like
literally do every single thing white-collar humans can do.
Yeah.
It would just have transformative effects in the way that getting gold in the IMO did not
have transformative effects on the world.
First of all, I do want to point out that I'm totally moving the goalposts here.
Because when I interviewed Dario about two, three years ago, I asked this question about
why haven't they able to use their vast knowledge to connect ideas together and come up
with a new discovery that way.
That seems like the kind of thing, even if a moderately intelligent person knew this
much information, they'd be able to like come up with the medical diagnosis from the fact
that like this drug causes migraines and this.
other thing, you know, whatever does this, and maybe it's the same drug that can cure both things.
And yeah, I don't know, from an outsider's perspective, mathematics seems clearly like a
field where finding this counter example to the unit distance problem conjecture was like an example
of this kind of thing. As a total goalpost moving. But then we can ask, okay, what is the next
benchmark now that AI can do this thing that we should have thought they should be able to do,
what is the next thing that would be quite impressive? And there's a couple of candidate ideas here.
So one could be coming up with interesting problems in the first place, and the other is coming up with new kinds of objects or conceptualizations that create or unify fields.
On the first one, right now we just train these models to, like, we have these Millennium Prize problems because, you know, like, mathematicians of notice, like, Riemann came up with this idea of this, like, Remond's data function, and because he thought that it would have some connection with, like, the density of prime numbers, or if the zeros on this function would have some connection to prime number.
And so, like, figuring out that there's, why do we think this is an interesting thing to study in the first place?
Why are we building this object and trying to answer questions about it and answer this particular question about it?
It seems like the kind of thing that would be the next benchmark.
I mean, you highlight two pretty good examples there.
For anyone curious about the unit distance conjecture, there's this really nice video about a math channel called Polylog, where they talk about it.
And one of the people in that, because all of these discussions, it causes people to reflect on, like, the processes.
of doing math, right? They're like, ah, this thing can do these impressive stuff. Like,
what does that mean for us? And he highlights this quote, how good mathematicians prove theorems,
great mathematicians come up with conjectures, and the greatest mathematicians come up with
definitions. And that's more or less exactly your framing here on like those two, like,
we need the conjecture generator, and then like the definition generated, that's the premium
tier mathematician. I don't understand how exactly you'd make that a benchmark in the sense
that usually when I think of the word benchmark, I'm thinking something that you have, like,
it's a goal post. The ball is through the goal or it's not. Like, you can clearly say, like,
yes, this is done. Partly to be able to do things like our LVR, but also partly just to be
able to, like, know that you haven't moved the goalpost and answering. You know,
open AI can have their headline on disproving the unit distance conjecture because it's a clear
distinct. It's like, it did it, right? Whereas imagine trying to have a headline on like,
BPD-5.4, came up with a really good conjecture, right? Like, we promise, everyone thinks it's a good
Conjecture. It just doesn't, it doesn't land the same way. But maybe that doesn't negate the fact that that's the right thing to be thinking about. So I would be surprised if it ever took the form of looking like a benchmark and like we have a score saying that it's past this benchmark because we can quantify how good a conjecture it is. But probably the nature of what it would take is that you would feel a tone shift in conversations with mathematicians about the way that it's useful to work with. Right. And like this series that you referenced that is not at all.
produced yet and probably won't be for a couple months, takes the form of us interviewing a lot of
mathematicians. And what's interesting is we started doing this like over a year ago. And it's fun to
see a little bit of a tone shift in the way that they talk about AI between like mid-2025 and where
we are now in 2026. You know, in the real world, that's a very short amount of time. In the AI world,
that's eons, right? And like we're able to see over those eons like this tone shift, I think the way that
you'd measure conjecture generating ability is going to be more subjective on like that tone shift
where it'll be mathematicians saying they're not just using it to like solve their problems,
but as they step back and decide what their research field should even be,
that a conversation with such and such model like was genuinely helpful for that. I don't,
I don't think it's likely that you'd see it in the form of like a headline saying that like
this was yet another benchmark knocked down. Right. And so it's very interesting. The kinds of things
you can't make benchmarks for are also the kinds of things, at least in the current paradigm,
you can't easily train for, right? Because there's really no fundamental difference between a benchmark
and a training environment. I think it's very easy to come up with some dichotomy of like,
here's a deep reason why AI can't do a certain thing. And then it turns out, well, you're just
thinking about it the wrong way. And actually, I can do it pretty soon thereafter. But I'm going to
come up with... You're going to come up with a couple anyway. And I think that this will probably,
it'll probably turn out that there's ways in which we can train AI to do these kinds of things in the relatively near term.
But it seems like it would have to be different from current or VR training.
So the thing I'm curious about, and the thing it seems to me that drives a lot of the big progress.
And mathematics and science generally is like coming up with a new way to think about a problem or the new way to understand the world that then unifies different fields, spawns entire new fields.
solves problems we weren't even thinking
we were trying to solve in the first place.
Like, the reason
Aynsen was thinking about GR is not because
he wanted to explain white light bends or white black holes exist.
These are phenomena he didn't even know
and needed to be explained in the first place.
But in mathematics, it often seems,
okay, a total outsider,
I don't even know the details of what I'm talking about here.
From the outside, it seems like
there's often ways to, say, prove a specific problem
that can motivate a new conceptualization.
one which results in a whole new field, a whole new way of thinking, which is immensely productive,
and one which doesn't.
I'd be curious to hear you talk about whether Galwa coming up with group theory,
distinguishing his solution to the quintic having no formula for the roots,
and Abel coming up with the different proof a few years earlier that didn't come up with group theory.
But then if you wanted to do a verification loop on, like, is group theory an interesting concept
that was something useful done here?
Why is this proof better?
Potentially that verification loop is 100 years long.
And it involves the cryptography coming around and physics making progress and the ideas in group theory being relevant and understanding like symmetries in physics and all those kinds of things.
There's like a hundred-year verification loop of why is this a productive concept in the first place.
Yeah.
Boy, yeah, you struck a nerve because I had this like project about Galois I was going to do in 2022 that I put on the shelf.
But I spent like a year of my life like thinking a lot about what he did.
So there's a risk of me accidentally talking too long on the specifics.
Hold me back on.
It's a perfect example for your case because describing why it was a valuable insight does not come from immediate utility.
And so certainly if you're thinking about RLVR environments, it's like, okay, this is going to be really hard to do.
But it's interesting to note how even with like human verifiers at the time, like it took a really long time to recognize it as being useful.
Like I think Einstein with GR people sort of felt you can like feel this feels like a good theory right away.
What makes the Galois theory such an interesting example is you have literally this 100-year segment of an idea that flows through many different people's heads before it settles into something that the math community agrees as good.
So to back up a little bit, I mean, do you want the background on the problem at all?
All right.
Well, so we all learn about the quadratic formula in school.
I thought you were going to say we all are about group theory in school.
We all know, we all learn about group theory, about quadratic formula.
So this was known, in some sense, like, Greeks could solve quadratics, but they didn't really
write things in algebra. And so it's really more like the Arabs that, like, wrote down, like,
that formula. There's this delightful story around some, like, dueling Italian mathematicians,
not real duels, just like intellectual challenges who, like, secretively found a formula for the
cubic. And then, very shortly thereafter, found a formula for degree four polynomials.
So the natural open question for, like, mathematicians is, can you find a formula?
find a formula that solves degree five equations. Now, the degree four, it's monsters. It's like,
it would be wild to write it down. You usually don't really write it down in full. You break it up
as like a procedural thing. So you might believe these things have this exponentially increasing
complexity. So many hundreds of years, nobody is like really answering that question. Usually we say
Abel was the first to prove it. He was this young, precocious Norwegian mathematician, and he showed
it's simply impossible. It's not that you can find a quintic formula. He thought he found one,
but he showed it's impossible.
I think the real credit, though,
like, you have to back up a little bit
and talk about Lagrange,
where Lagrange found the right kind of question
to ask about this.
I can go into the details if you want,
but I'll give it a very high level.
He was studying the question,
and he recognized being able to solve these polynomials
is actually very related to understanding
the way that certain algebraic expressions
are, like, symmetric, like more or less so.
Like, if I write down A plus B plus C plus D,
just like adding four variables.
If I permute those, it doesn't change the value of the expression.
Whereas if I write like A plus B multiplied by C plus D, some of the permutations don't change it, but some of them do.
And he had this really, really nice insight about how, if you can find expressions like this that have like four free variables, but all the permutations take on three distinct values, that had this unexpected relationship with being able to reduce degree four into degree three.
So he started approaching the like, can we find a quintic polynomial by saying, I wonder if I can extend that.
And to extend that method, you would have to have an expression that has five free variables
such that as you permute them over all the five factorial permutations, it takes on only four
values or fewer.
So you could put that in a puzzle book.
You could put that in a brain teaser that like a 12-year-old couldn't engage with.
And it's not too hard to find yourself feeling like that's an impossible task.
And so Lagrange is sitting here saying, hmm, here's a strategy that I'm trying to solve this
problem.
Can I find a quintic polynomial?
This strategy doesn't, it seems like it might be impossible, at least from this strategy.
But that was the first time in history that people had the instinct that some kind of question about symmetry
was the right way to be studying these polynomials.
In his mind, it was just a way.
It had yet to be discovered that, like, actually there's a tighter connection.
And also, like, maybe rather than searching for the formula, we should be asking the opposite question.
Can you prove that it's impossible?
So he sort of planted that seed.
Like, around 50 years later, Abel definitely read Lagrange and was influenced by it.
Galois, we know that he loved Lagrange
when he was falling in love with math.
And so it's very hard to imagine that these two young
geniuses, the fact that they both come up
with pretty similar insights around that
problem, it's not like born from Lagrange.
But to your question on like, are you
able to verify that this was a good idea?
There wasn't any like result that Lagrange came
to. There's never like, he solved the problem
and therefore we know that that was like the right
question to ask. He asked it.
There's some like intrinsically interesting thing.
It also wasn't very important for math at the time.
Most people were more interested in like the applications to physics.
This is almost in that like side, almost recreational hobbyist type thing.
Like Abel, you know, he started working on Quintic stuff, but then he was advised to spend more of his efforts studying elliptic functions.
And so more of his work was on that before he died young.
He died at 26 from tuberculosis.
And then Gawa, he pushed both of those ideas like in the right direction where he really understood the nature of
abstraction. And so he had this really nice piece that he wrote while he was in prison,
actually. He was like, we could talk all about his life story. It's pretty wild. But he,
he's like this teenager. He's in prison. He had tried to submit his math papers and they had been
rejected. So again, it's like verifiable reward. The like verifier function that is the academy at that time
is rejecting what he wrote. Because frankly, it was not very coherent. Like it wasn't a complete
proof. He wasn't giving like a clear thought of like what the theory actually was. He was just like
a young fledgling mathematician getting his bearing. So it's like the verified reward there is like
no good. But he has some instinct that there's something there. So he's writing this diatribe on
the nature of like math being something which is, it undergoes these like shifts over time. And he
talks about like the advent of just algebra itself and going from just thinking in terms of numbers
to like having a certain fluency just with like pure algebraic expressions where you're not
tied to interpreting those expressions. And he has this instinct that like there is another layer
of abstraction that seems like what we should be doing where rather than thinking about the
formulas themselves thinking about like what symmetries underlie those formulas. But it was still a pretty
like ill-defined theory. So if you're trying to say, okay, is the verified reward that like he has
solved a problem that other people haven't? It's like, well, Abel proved that quintics are unsolvable.
And you say, what was Galois doing? Well, in principle, the thing that Galois theory will let you do
is take a specific polynomial and it gives you the rules to say, does that specific polynomial have
roots that you could write down? For example, like X to the fifth minus one, you know that a solution
is one or X to the fifth minus two. You can write down.
fifth root of two. So it's not that every quintic
polynomial you can't write down the solution,
but could you find a specific one where you
prove you can't write the solution using radicals?
He also didn't even solve that.
Exactly. He has a much more abstract. He didn't
show for a specific example that he
couldn't. So even describing what problem
did he solve is very tricky. So then
he dies. It's this very romantic story
of he has this duel.
We can get more into it. There's a lot of myth
around supposedly he writes up all his ideas
the night before the duel. Really he tried to get them
published. Working in the credit doesn't seem to be good for
your health. It's very bad. Yeah, yeah, yeah. If you're a young genius, don't work on the quintic.
And so he asks his brother and his close friend, like, get these notes to gouse, get these notes
to, like, the important mathematicians of the day, because I think there's something here.
Even then, it didn't really take. Like, so his brother and his friend, like, tried to get them out.
It wasn't another 20 years until Louisville, like, sees these notes, sees that maybe there's
something in them and tries to, like, clean it up and understand, like, what was Galois
getting at. And then, even then, it was another 20 years or so until Jordan,
actually puts together
something like a modern
treatment of group theory
that they attributed to Gaoua
you could easily imagine history turning differently
where these ideas were kind of coming about
from other points in math and like Gao
could have been forgotten in history if he was the less
like florid character. But between
the time of Lagrange, like having this
inkling of maybe symmetries
of roots as the right way to go to where
at all looks like modern group theory
like you've got this long span
a lot of the time it's like not even passing the
verified reward of human reviewers, because it like gets on someone's desk. They say,
I don't really know if there's anything here. It gets on someone's desk they don't. You have to
have this like one person sort of recognizes it. And then even then, it's not really solving
practical problems at that point. Like you point out cryptography and physics and things like
that. You have to get into the 20th century before you have like Gehman thinking, hmm,
maybe understanding the nature of like how certain groups like break down has this relationship
with what particles are made out of. And like he, he,
he anticipates quarks based on a purely group theoretic question.
And like that's one of the more interesting applications of group theory is that, like,
to even predict the existence of quarks is a group theoretic question.
That's so long after Lagrange before you have anything like that.
And so you have to ask, like, what is the way of measuring progress that's not based on solving
a problem, right?
And that's somehow capturing what is the instinct that's inside Galois's mind when he says,
I think there's something here?
what's the instinct that inside Legrange's mind
when he says, like, I think this is the right way to think about it?
What's the instinct inside Louisville's mind
when he says, hmm, these like scattered notes
from this like long dead youngster might have something to them?
It's so hard to put a finger on that, but
I mean, a different
series of videos I'm making right now is
about, like, you know, the whole compression
is intelligence idea. And even though
this isn't really the angle I'm taking, you know,
there is something to the idea that
the smaller expression
that's more predictive
like feels more intelligent.
And so I wondered the extent to which you can give
some kind of verifiable reward around
not just like did you solve it or what is it solving,
but around the smallness of the concepts required to do it.
I mean, going back to Riemann hypothesis solutions,
what would that look like if an AI solves it?
I think a third way that it could happen
is it just straight up works harder.
Right?
In the same way that you could maybe have an elementary proof
of Fermat's last theorem that's just like spelled out
over like thousands of pages that would be incoherent.
But like the cleaner way to view it,
it is with elliptic curves and all that. Maybe there's some like thousand page proof of
Riemann hypothesis that's like no one's really get anything out of it. And what you actually want is
like what are the succinct like compressed versions of those ideas like they would then
lend themselves to human understanding? Like I don't know, comagora of complexity. Like maybe you
throw that into your like your attempt to quantify what you mean by elegance. But I don't think
it's easy. But I do think it's something you would have to do in order to reward the Galois like
instinct rather than just rewarding have you solved a problem. It's very hard to come up with
the heuristic for science. But it's clearly like humans have been doing this somehow and like
obviously AI will do it at some point. Well it's relevant also not just in terms of verified reward,
but like presumably the end goal is understanding and like human understanding. And so even if you do
have some like a thousand page proof of some math thing or some like grand new physical theory,
the goal is understanding. Yeah. Right. Maybe if the goal is predictiveness, you can just
have automated engineers go off and build rocket chips or something.
We're like, we have no idea how these work, but we can get between stars.
But like, there's going to be a lot of people want to understand.
You're still going to want whatever the concision function is that distills down,
here's this complicated way of thinking into like the right one, like the equivalent of the universal law of gravitation for Newton.
Like you would still want to train AIs to be able to do that and like find the compressed representation.
I grew up in India until I was eight.
And so in addition to English, I also speak Gujarati.
And since Google just released Gemini 3.5 Live Translate, I thought it'd be fun to put it to the test in this midroll.
3.5 live translate automatically detects more than 70 different languages and translates them in almost real time into the target language.
Live translate your original speed and format while speaking, just like it's doing right now.
I visited China back in 2024.
And I remember thinking at the time that this trip would have been so much more productive
if I could have been able to live translate the conversations I'm having with the researchers
and random people I meet on the street.
Now we have that technology.
So if you're building an app that needs live translation, you should 100% check out Gemini 3.5
Live Translate.
It's available now via the Gemini Live API and an AI studio.
Go to AI.studio slash live to get started.
So people have this worry about mathematics in particular that, you know,
the AIs will prove the human hypothesis and our understanding of mathematics won't be any of the better for it.
I have a couple of questions about this.
The first one is whether this is like a thing you should expect.
Like isn't the reason humans come up with general natural
objects and sub-goals and whatever when we're working on a big problem is that it's just like
useful when you're trying to work on a complicated important problem.
And so we can just think about like theoretically, would this even be a simpler way to solve the
around hypothesis as opposed to just coming up with the natural abstractions that are relevant to
thinking about the problem. And then two, empirically, is this what we observe when AI is doing
make progress on problems today? When the AI came up with that counter example to the unit
distance problem conjecture, you can just read its chain of thought. And it seems, it's not
understandable to me, because I don't know anything about mathematics, but it seems to other
mathematicians. It was like understandable. And it made use of like known concepts mathematics
and like proved relationships between them and all in natural language. And as a result,
accelerated our understanding of the connection between this object and this conjecture.
So is this even, like, empirically, is this a thing we should be worried about?
I think it depends on the nature of, yeah.
Like, again, if we sort of break down, like, the three possible ways of, like, solving the
Riemann hypothesis, that one.
And the other, like, big one from this year was, like, a certain erudish problem numbered,
like 1196, but it's a, about these things called primitive sets.
But basically, it had that character of bringing an idea from a seemingly different field,
As soon as you just present the basic idea to a mathematician,
you say, like, what if we, like, use this, like, try the Markov chain process
where we show that this thing is one from the bottom up probabilistically,
rather than the top down, and, like, use the Von Bangold function.
If you, like, say that to someone in the know, they'd be like,
they'd kind of know how to run with it.
So they have this very, like, small idea that has the form of expertise in one field,
expertise in another, draw a little lightning bolt between them.
Like, those are going to be very human parsable, right?
Because all you have to do is just, like, show the,
start an end point of what those connections are. If the character of it is mountain building,
you do have to put in a lot more time to understand that new mountain that was built because it's
like a new thread that's not just like lightning ball between them. And if the nature of the
progress was just like raw hustle, right, it's just like this just super long thing,
no new theories, but it's just like long, long, long chain of reasoning answer, then you would
have that word like, okay, there's this whole digestion process. So I don't think there's
one clear answer. I think it depends on what the like solution there would look like. And
on the mountain building side,
I would actually be really interesting to see.
Like, is it by default a very human,
understandable, like, the way that we, like, see new theories
from, like, great mathematicians?
Or is it, like, like, an alien different kind of mountain being built
where we even have to, like, reprocess the kinds of abstractions that we engage with?
Right.
Well, the closest example here would be, like, the, you know,
the attempted solution of the ABC conjecture that was,
we maybe shouldn't get into that one,
but it is probably not, it just is not,
not a correct solution, but basically it's this whole new way of thinking that this otherwise
reputable mathematician in Japan had like come up with. And it just took mathematicians like a long,
long time to even parse what he was saying, but it had the feeling of just like an alien
bit of mathematics that's theory building. It's not just like long chain of reasoning. It's like,
you called it like inter-university geometry or something. And so the fear that you would have is that like,
yeah, it like does that. The biggest fear would be that it does that. And then much like the ABC
conjecture, like people work for years.
to go up the mountain and they're like, this just isn't right. And like if there's if it
turns out to be wrong, but it like really looked right. But even if it was right, there's just a
lot of effort to like hike up a new mountain. Yeah. If we end up in that situation, David Bessus
had a really great blog post called the fall of the theorem economy. We're talking about this
you know, historically there, as you were saying, mathematics is coming about these
definitions and problems and it's about proving theorems about them. And that really the theorem,
improving stuff is what gets all the credit, but it's like really a parasite on the,
coming with the definition stuff.
And historically, it's not been a problem in terms of credit apportionment because if you
come up with the definition, you're probably going to be the guy who comes up with a theorem.
But now we're in a situation where if the valuable work is the coming up with the insight
and then AI just automates the latter part, it, so, okay, imagine a scenario where we have
AI comes up with like the ABLE-like direct arguments about a bunch of important
conjectures in the world, and then we just have these proofs. And now it's up to humans or to
future AIs to then consolidate. I mean, I'm sure if you had access, again, having no object level
understanding of this argument whatsoever, I'm sure if you had access to it, it would make it
easier for you to then think about like, well, what is going on here? Is there some deeper way in
which you can understand why this proof works that would make it easier to come up with the ideas
behind group theory? Yeah, I think it would be hugely helpful.
helpful, right? Because I mean, so much of, like, trying to discover new math is, like,
mostly being wrong, right? You're, like, trying to solve a problem. It, like,
it doesn't feel like constantly taking the correct step up the mountain. Like, mostly it feels
like a random drunken walk where you're, like, doing a thing and then, oh, you're wrong and, like,
constantly discovering. So if at the very least, you know that trying to digest what you know
is ultimately leading to, like, a correct solution, like, that feels like progress simply because
it's providing, like, a sense of knowing that it leads to a solution. And there's plenty of
Plenty of like instances in the recent history of math where it feels like the reach has sort of exceeded the grasp where there's things that are proven, like long before they're understood.
And I mean, one of my favorite like openings to a paper, it's not even like a research paper, it's more like an expository one, is from this mathematician named Timothy Chow who was trying to understand a concept called forcing.
And so there's this problem called the continuum hypothesis that more or less asks, like you have a size of infinity for the natural numbers.
you have a size of infinity for the real numbers
is there's something in between.
And the answer is both yes and no.
It depends on your axioms.
Like it's sort of outside the scope
of our usual axiom systems,
which is an interesting answer.
But the method to describe it
is just really, really hard to understand.
It's the thing called forcing.
And in the beginning of this paper,
he writes, like, I want to,
like everyone knows the idea
of an unsolved research problem.
Like, I want to propose the idea
of an unsolved expository problem
where like, sure we've proven it,
but we don't really know why it's true.
And so then he proposes
like a partial solution to that expository problem.
You can imagine why I loved that framing,
because this is my whole life.
It's like, I don't do research math.
It's just wholly about like,
what's the most clear way to understand this,
even if it's proven?
There is a difference between proof and explanation.
And so on that side,
I think that you are basically like getting to the importance of that distinction.
Yeah.
And that will be the main incentive for,
or the incentive would have to change
in not just mathematics,
but in other areas of science from,
proving things about the world to consolidating proofs into problems or higher-level insights.
But we have a discussion earlier at lunch about a recent talk you were giving about, you know,
design and how it helps us understand things. And then in the limit, is there really a difference
between the conceptualization for an idea and the idea itself? So, you know, if you think about
special relativity and like space-time diagrams,
and Minskowski space time.
Is it like, yeah, this is like a way in which we illustrate this idea of like
why there's length contraction and time dilation.
But is that like, that is the reality.
So the exposition does seem to be like the explanation in some sense here.
Yeah, I mean, there's a couple interesting things there.
One is it seems like there's a really strong correlation between the people who come up with
genuinely novel insights and also are actually quite clear in their communication of it.
Like, you might imagine, given that the experience of a university student is often that the expert there teaching them is not necessarily the best explainer of that topic because they are so spoiled by their expertise.
But what seems, at least in some cases to be the case, is how the people who are really coming up with something quite novel.
So you've got like Einstein or like Claude Shannon or something there.
You read their papers, they're really lucid papers, right?
It doesn't feel like, oh, this is just for the experts and you have to chop through it with a machete to get.
They're like very good expositors.
Like Feynman has this characteristic too.
Very good expositor.
And so maybe the same part of the brain
that comes up with the correct new way of thinking about it
at a research level also has this knack for like good explanation.
And I think this is pertinent to the AI one
or I kind of used to think that
AIs will become these automated theorem provers
but like the role of the mathematicians
is going to shift towards like my job.
Like explain these things.
I kind of suspect that actually
they'll also be quite good at doing that
and probably just like better than most humans are
at like doing the explanation half and distilling
half and that's actually not what's left
for the mathematicians is like digesting
and explaining what was going on.
Probably the nature of how these things
are going. I could have envisioned
we can talk about like ways this might not be it
but like probably the same thing that is coming
up with like the really good
new idea that solves some new problem
is just also good at explaining it.
That's my new like that's a way
my I think beliefs have changed
What's the last thing you think it will be doing?
Both you and then also with the mathematical community,
the human mathematical community will be doing.
I will probably be doing something like what I am until I die.
Even so like, even...
I have the doer the right.
Maybe that'll be the same.
Exactly.
It'll be for the same reason.
Yeah, yeah.
You know, you're like, build a man a fire and he's warm for one night,
but set a man on fire and he's warm for the rest of his life.
So that's where I am with AI.
No, because some of the function of an explainer or a teacher is to add clarity to a thing that someone's curious about.
That's one thing.
But some of it is like a little bit more relational and a little bit more like providing motivation, providing a sense of curation.
Like one interesting take that I've heard about like what mathematicians will end up being is actually more analogous to art museum curators than anything else where the AIS solved the thing.
So the art exists, right?
They even know how to explain it really well, you know, all there.
But like, you still want someone to help you navigate in this like nearly infinite space of like
what ideas are worth engaging with, like someone kind of doing that.
And that one, even if AIs were in some sense better at that, I think we would always still
prefer like a human that we had a relationship with because the way that we get motivated to be
interested in things is a social phenomenon.
If you have some specific technology you're trying to build, you know, that might be different.
You need to know there.
But I think, like, the people listening to this podcast, they sort of trust your curation on, like, what's an interesting topic in the first place.
It's not that they're landing on here because whatever your next topic is, that's like what they, in a prior sense, wanted to understand.
They're trusting you as a curator.
Yeah.
So my role, and arguably that of, like, other mathematicians might actually just shift subtly into that curation direction of what ideas are worth their spelling.
And that's a lot of my job right now, even now.
It's basically, like, I think people think a lot of the time for a video goes into the visual.
Like, sure, a little bit.
It is, not like immediate, but like, actually a lot of it is just deciding what's worth saying
in the first place or what's worth putting in there.
And because that is, that's just, I want to engage with that.
And I think I have a trust with certain people and they are curious what I would choose
to put forward, even if the AIIs are better than that, in the same way that, like, human
musicians are always going to have a role because of that, like, social function of the
story behind them, even if the, like, objective quality of the MP3 file coming out is, like, better
from some model. That's kind of what I see happening to my job.
Yeah. I want to go back to this question of earlier I was, we were sort of just as AI has crossed
this threshold, this important benchmark of being able to connect existing ideas to come up with a
new discovery or prove or disprove something, just as it's cross the threshold. We're like,
okay, but what's the next thing? I want to just um. There's a lot more to do on that one by the,
like just because a couple lightning bolts have been, I still, I think there's like this flourishing
future over the next couple years of like really connecting.
And so in the limit, you could even say, I don't know this is accurate to say, but
potentially a lot of them, maybe the biggest breakthroughs look like this at some level.
It's just general relativity.
Oh, you're just connecting together like Romanian geometry and special relativity, right?
And so as AIs keep getting better and better at this connection thing, maybe a lot of big
breakthroughs are not really of a different qualitative nature.
I don't know if you have a take on that.
Well, I mean, a lot of the conversation focus has been on problem solving and that nature of math, you know, like taking off erudish problems or something.
I would say it's not even a majority of mathematicians who would maybe characterize their work is like really targeting the next problem to tick down.
Are you familiar with like the Langlands program?
No.
Ah, okay.
So this is like, it's not even a field of math so much it is like a research ethos where Fermat's last theorem is one inkling of this.
on you had like these two different seemingly disparate things
and a connection between them like led to a solution.
So Languins was a mathematician,
he has this like famous letter now,
essentially spelling out how it seems likely
that there's a lot more connections like that
and even got like a little bit more specific
about the nature of the connections such that you might imagine
this like large map and you've got this like valley over here
and this mountain over here and this like set of planes over there.
And there's a lot of mathematicians
who would characterize their work as being part
of like trying to understand the threads, like on this map.
And the progress there, it's not even like,
here's this one specific problem that we know will be solved by that connection.
It's more that there's been enough time and time again cases
where big problems were knocked down by finding connections
that it's almost preemptively finding the connections.
And so you could have...
Yeah, it's actually very interesting.
Like, this...
Anytime you run into a mathematician,
like, ask them whether, you know, the character of their work
is more akin to, like, Lengland's program.
or if it's more akin to like targeting one particular problem, right?
And you get a certain like bifurcated split there.
But the possibility of AI's being supercharged connectors feels like it might be, you know,
an amplifying tool in that pursuit.
It's hard to measure though, right?
Like, because this cuts to what we were saying earlier.
How do you assign a score to say like, yes, you've done it?
If it's knocking down a problem, you have a clear way of saying, yes, you've done it.
You can write the headline.
have your like PR move as the AI company to say we did it. Whereas like if it feels like that
was the right connection drawn, you can like, you can write theorems around it. And this is the
nature of what the papers in that field look like. But I think it will require a lot more like
human in the loop to basically like say what was it a like the kind of connection that we're going
for. But that's my guess on what most of the useful progress from these models will look like
in the next five years is just really filling in that landscape of like connections that you can
draw if you're an expert in multiple fields.
Like you've pointed out, it's kind of surprising we haven't already
had this. Right. And what I'd be curious, like,
I would be curious to know at a technical level
what causes the unlock there.
Because on the one end, you can kind of paint an explanation
in your head for why you could be an expert in all of these things and
not be drawing those connections, which is when the thing
is reasoning, like the method of reasoning is this
auto-regressive chain of thought phenomenon.
Auto-regression is actually like a really, really weird.
way to produce stuff, I think, if you think about it.
Like, you're an intelligent person.
Imagine I've walked you in a box, right?
And then the only way that you have of interacting with the world is that you receive a
slip of paper and then someone says, can you like predict what will come next, right?
And then you predict we'll come next.
And then your memory's wiped, right?
And then you get like another slip of paper and you go, imagine that was done a whole bunch.
And then what comes out on the other end, they're like, look at this essay that you wrote.
You might look at that and be like, this is awful.
That's not the essay that I would have written, right?
Because the process of repeatedly predicting something is just pretty different from how you would think as a writer to, like, compose it and think it through and everything.
And in particular, what would probably happen is you're sort of a slave to your context where you might be answering some question about some particular field.
And so you're like draw on all the context around that and you're going there.
The connection that actually is where all the substance is going to come from is like by its nature a very like unlikely one.
And, you know, you can do all the...
the RL that you want to try to get better in some way, but like, what's the thing that's
specifically upwaiting and incentivizing making these unlikely connections when the vast
majority of them, like, aren't the predictable, you know, next token that would come in there?
And so it's like, it might be the case that you just have this intelligence that sort of
locked in there inside that box, but it's just a weird way of interacting with it.
So the thing I'm curious about is, like, do you ever get any fruit by just like questioning
the premise of how tokens are generated like every now and then in some way?
Right. And I don't think it would be as simple as you like manipulate the temperature or something like that.
But like are there any things that you can do that take like the existing level of intelligence, but like find the right ways of sparking those connections that like unlocks these sorts of things that we're seeing?
Or do you need just a little bit more intelligence such that at the level of prediction, it's kind of predicting that it should be making that lightning bolt to another field?
I think it's more productive to reason instead of architecture or.
even loss function to reason about data.
Like, I don't know, we have diffusion models that do text,
and they're like, not of a whole,
the kinds of things that produce are not of a wholly different character.
They're just not been explored as much.
I think the more relevant thing is what is the data
on which whatever architecture,
whatever loss function you have is incentivizing you to produce.
And it does seem like they're getting better at,
like, okay, forget about math.
I mean, we did have this,
a couple of examples of this kind of thing.
But if you just look at, why are they getting better at being autonomous agents?
It just, I don't know, they have like, they're in an environment where auto-regressively producing
the step that says, let's step back and do a search over the whole code base.
Right.
And then let's step back and, like, assess my mistake.
It's like the thing that works.
I assume what happened in the case of progress in science or maybe in math is you have frontier
math-like problems which require, like mathematicians specifically designed them because they require
connecting together two different fields. And there's all, I'm guessing there's all kinds of clever,
like partially synthetic ways in which to make harder and harder problems like that that require
these kinds of connections, for example, by like eliminating assumptions and still requiring
the AI to continue to get to the answer. And then like, it doesn't really end up mattering what the
loss function is. It's just like, it's really about can you come up with an environment which incentivize
as a stability.
Yeah.
It feels like you should be able to.
Yeah.
I certainly can't speak to the correct ways of doing that,
that like unlock all this.
But it would just be pretty surprising.
Like, don't you think it would be kind of surprising
if over the next three years there's not just like a lot more of those lightning bolts?
So this, I think, is an important thing to think about,
which is we often think about how smart a single system is.
And we don't think about AI's having advantages that are more the result of other facts
about them. So in this context, the key fact about them is that we can just paralyze and arbitrarily
scale them so that whatever level of capability they have, it's not just like one idiosyncratic
genius in the history of mathematics who makes a few connections and then dies in a duel.
It's just universally applying the waterline across all problems that are accessible at the
level of capability. I feel like this is among the many advantages that digital minds inherently
have that we don't think enough about the fact that you can, the other ones being the fact
that you can, like, they can merge all the knowledge together, at least that there will be techniques
that allow this to happen, that you can, that you can spawn off copies with identical levels
of knowledge.
But yeah, I feel like this parallelization is, like, quite an important property.
And I'd be curious about your predictions of even if they're not as smart as your mathematicians,
the fact that they are just, you know, billions of, because for PR reasons that the AI companies
They're just dumping billions and billions of dollars at this,
would have a quantity as a quality all of its own.
That seems in the right direction.
I think, I mean, if we take that, you know,
that conversation between Montgomery and Dyson at the IAS
that, like, suggests some connection between remand hypothesis
or remand Zeta function zeros and random matrices,
that feels like the kind of thing that you could try to, like, automate
and that you have, you know,
agents representing expertise in all these,
and basically having, okay, we all know that an institute is smarter than an individual,
and that, like, the reason for having people all in the same geographic location
is because you want those, like, serendipitous conversations to happen.
What does it look like to sort of engineer those between agents?
I mean, it's interesting, because you sort of point out, like, you can sort of pool all your knowledge.
I actually wonder if one of the advantages is that you can do the opposite of that,
where you have, sometimes when an AI is failing, it's because it sort of gets into a bad chain of
thought and it's really hard to get it out of it, right? So you're like, I'll just like start
again. Same deal with humans, right? Like sometimes you like start thinking about it in a certain
way and actually what's required is to just like back up. Maybe sometimes the form of that,
you know, there's stories about people trying to prove something for a long time. And then at
some point they say, hang on a second. What if I tried to prove that it's impossible? They
like prove the opposite. And that like unwinding your own context and going at it with a
fresh mind, you could imagine systematizing that or like having multiple different agents
deliberately given different pieces of context
and try to like comparing trust there.
Like we don't have the same level of manipulation
on our own context.
In this like AI and math series,
the first episode will be about like when they solved the IMO.
And I want to focus on one specific IMO problem
that they failed on,
which is one that a lot of very smart students failed on.
Terry Tao also failed on it.
And the nature of it is basically that
people were very mad at the problem
because they called it a troll problem.
I almost don't want to spoil it
because I want to construct the episode
around leading someone in
without knowing
that it turns out to have a simple solution
because you can really empathize
with what it's like to be
like a student solving this.
Basically, there's a really elegant way
of going down what you really feel like
is going to be the solution
based on the context of being
the international Math Olympiad problem,
positioned as it is.
The character of this solution is really enticing,
but it's kind of hard to prove
that it's the best.
The reason is that it's not.
There's like this almost brain dead solution that is the best.
And so the like relevance of that to the whole AI story is like for a human, what's required
to answer that question is to like escape your context.
Escape the context that you're in the IMO.
Escape the context of the way you've been trained to solve these like contest math problems.
And if you just approached it like a like a brain teaser that I throw someone off the street,
like they'd probably answer it well.
And you sort of want the same sometimes for like human research.
in other contexts where like sometimes just being able to say refresh your
thinking come at it completely differently so of all the advantages that digital
minds have that might actually be one of them like a little bit more of a
systematic what does it look like to like refresh your thinking try to
answering two separate questions like spin off two agents one who's trying to
prove it one who's trying to disprove it one who tries it like this way
and they like deliberately have different contexts I would be curious to see if
we're having this conversation three years from now how many of the like
significant results that make headlines have that character of basically, like,
erasing the context previously, like trying a bunch of different things, as opposed to
merging the results of like a bunch of different.
It is incredibly interesting because a common concern people have about AI's is this entropy
collapse where they all think the same way because they're trained in similar ways.
This is why they're bad at writing.
They kind of just like go down the same path and have similar patterns of speaking and so forth.
But maybe actually the key advantage AIs have,
is that you can systematically,
it sounded like one of the reasons
the unit distance problem conjecture
took so long to be disproven,
was because people assumed the conjecture was actually true.
So mostly they were trying to figure out ways
in which to prove it.
And so maybe one of the key advantages they guys will have
is actually to increase the entropy
by systematically trying out both the negation
and trying to prove the positive of any given statement.
Or being able to like systematically give,
different agents, different biases.
That's a good point.
Like, it seems like an important thing in the history of human science is that, like,
Einstein is just really motivated by this bias that, like, things should look the same in
different reference frames.
And then he had multiple other biases like these.
But, like, that is just very formative in his thinking.
And you can just, like, systematically survey a bunch of heuristics and see which ones
are being productive at a given problem.
Yeah.
And so you would suggest basically, like, systematically increasing entropy at the prompt level,
even though you have this like inevitable collapse at the like auto-regression level.
Yeah.
Yeah.
And I mean, Einstein would be an interesting example because it's like he's got this bias towards things should be able to.
He also has a bias towards like God should not play dice.
Right.
And it's almost like you want to make sure that you don't accidentally have all of your LLMs or Einstein
because you might halt on quantum mechanics progress, right?
Which actually goes to show you that there's not a correct heuristic.
Exactly.
For science.
You actually just need multiple independent research programs with their own heuristics.
Yeah, yeah.
And that feels like old school software, right?
As long as you're able to like describe that in some way,
you have like old school software that like amplifies that entropy in some way.
And if you're able to like put a clear ontology to the distinct ways of thinking that you want to prompt,
you like explore that full ontology and then each individual one, you know, runs off doing what it is.
But I, you know, I think there's a certain design question there on like,
how exactly do you describe like the different approaches?
The easy one is, are you trying to prove it or disprove it?
the harder one would be to say, what are all the tactics that you could take to prove this?
And make sure that you're like sufficiently applying sufficient breadth to exploring that.
I don't think people appreciate the kinds of things that these models can just go handle for you
when you equip them with a good harness like cursor.
For example, I started publishing my episodes on Billy Billy for a hopefully burgeoning Chinese audience.
But everything I upload there needs the sponsored segments cut out.
Normally, that would have meant that I would have to ask my editors to go back through all the old
episodes, cut out the ads, and re-export everything.
But in about just as much time as it would have taken me to send them that slack message,
I can just tell cursor to do it instead and spare them.
And for research for the podcast, I have a whole repo that I've set up where I've just
put every single book and paper that's been relevant to prepping for any of the recent episodes.
And I've been able to hodgepodge everything because the cursor harness is just extremely
good at helping the model figure out exactly what information to pull, whether that's
from my repo or from the web, in order to answer the questions I have while I'm doing research.
So whatever you happen to be working on right now,
just try pointing cursor at it.
Go to cursor.com slash thore cache to get started.
Obviously, A.F or Math is making a lot faster progress
and everything else, and people point
to verifiability of the domain as the key reason this is happening.
I think that's one of the two important reasons,
but I don't think, I think people really neglect the other one.
And I'm outside the labs, I don't know what's actually going on,
but there's a totally naive theory.
Okay, a tangential question to why AI is making so much progress in math.
Why has it been so slow computer use?
Which is what you with, you know, computers is actually very verifiable?
It's like, you know, is my Etsy package coming or like, it's my event booked, you know, whatever.
These are extremely verifiable things to survey.
What computer use lacks is grindability.
So because websites have like bot detectors and also it takes a tremendous amount of compute to run parallel rollouts,
it's very hard to just run
like a thousand parallel rollouts
at the same checkout flow on Amazon
because you'll get shut down by Andy Jessie, right?
And so you can...
Personally,
presses the like Red X on Dorcasch button.
Exactly.
And so you can try to build clothes
every single website.
This is very labor intensive
and slows you down.
So, and the reason, by the way,
you need to do so many parallel rolls
in order to learn a skill currently
with deep learning is that
we haven't solved sample efficiency.
sucking supervision through a straw, like that's what he says.
Of course, people are working on many different techniques, but fundamentally, there's this big
problem, and there's this big constraint in the way we're training eyes.
With code also, you can containerize a given level of progress in a repository, and then
just spin out thousands of parallel containers or hundreds of parallel containers and say,
try to implement this feature, and it's totally deterministic.
And because it's deterministic, you can solve the credit assignment problem because you know
that whatever caused this rollout to succeed and this one to fail, that did.
is the thing that worked.
And this way you solve the credit assignment problem.
If you have situations that are starting off
at different starting points,
this credit assignment problem
because it's much harder to solve.
But things in the real world
are just very hard to containerize in the same way.
Like coding and math are exceptions to this rule,
but if you're just trying to figure out,
how do I build a new business that succeeds?
How do I like go trade in the markets for a day
and make money?
You can't like, the fact that you had to interact
to the real world and like things change day after day
means that you can't keep replaying and grinding and farming the simulator.
But the math, of course, is the exception.
And I feel like this is actually an important driver of progress in this domain and also in coding.
It's not just verifiable.
It has to be grindable.
The third reason that people point out that AI is making fast progress is they focus a lot on lean and formalization.
Again, I have literally no idea what's going on in the lab.
I feel like lean just doesn't matter that much for like the current level of progress.
in AI or like why is AI able to solve the unit distance problem? Well, they, sorry,
disprove the conjecture by the unistance problem. They release the chain of thought, or at least
a rewrite of the chain of thought. Didn't have any lean in it. I think it's just like the
process-based supervision that lean provides where you know each step is correct. Seems like
less relevant to just having this grindable outcome that is verifiable. That's an interesting
point, like grindability mattering more. I guess I will say on the, yeah, okay, so naively
I think Lean provides something unique for math because you're able to see if it can prove it.
You have old school software that can tell you yes or no.
You use that as your VR.
I mean, so what would corroborate your point is the idea that like the initial attempts,
again, I'll just circle back to IMO.
It's like initially DeepMind basically does that.
It's like everything in Lean.
And then the next year it's all in natural language.
So to your point, not needed.
I think there is a yet to be explored benefit of that formalization domain.
which is at the moment, you still need, you know, ultimately,
like a human is reviewing that,
um,
counter example to the unit distance conjecture to say looks good.
And that,
that provides a certain bound on how like endlessly explorable things are.
Like if you consider like AlphaGo,
alpha zero style stuff where they're just like off in their own universe,
just like playing a bunch of go and exploring themselves,
just completely going potentially off the rails of what any human needs to look at,
but they still have this automated verifiable reward.
It's not just that, hey, you can do RL on that.
It's also, you basically never have to check in,
and you can just, like, pour compute at them, like,
exploring the universe of Go.
What stands to be interesting, like, maybe this won't pan out,
but I think the jury should still be out on, like,
whether this will yield anything.
With Lean, you could imagine having a basically endlessly running program
that's constantly trying to extend Mathlib.
So MathLib, it's this GitHub repository
that's basically like all of math, written in code.
It's very far from all of math,
but they want it to be all of math,
written in code that you can ask, like,
is this proof correct?
It's very labor-intensive to write these proofs.
There's like a whole sub-community around it.
But you could imagine,
what if you just had an AI where you say,
simply try to extend Mathlib?
Maybe it's a fork of it,
so it doesn't have, you know, like, trash in it
because people have certain taste for what they want to be in there.
So you have, like, your fork of, like,
the pure AI math-lib, and it just goes.
and it just like doesn't stop.
It doesn't need anybody to check in on it, right?
It could just keep going, it might come up with its own conjectures,
I might come up with its own theories and like different definitions.
Maybe many of them are useless, but it just has this infinite tree that it can like grow out.
That's a very unique thing that math has that nothing else has, where you could press go
and then just like just pour compute at it and like look away for 10 years and then come back and say like, what do you have?
And there's going to be something, right?
And then there's a question, is it useful or not?
Like, how do you suss that out?
that's just an interesting thing to be able to do.
It would be very surprising if that didn't yield
some sort of interesting mathematical insight from it, right?
So I think that's the real case for...
Okay, there's like two different ways
that lean is important in this story.
That's the first one of them, basically,
is how it's like you could let go,
not even check in, and progress will be made.
You can do that with Go.
I don't think you can do that with natural language math.
This is very interesting.
Did you see Carpathie's auto-research idea?
idea. He wrote this basically one Python file that does basic LLM training and then just had a
repo where LLM agents would like try to make modifications of the file if it sped up the speed run
the modification stays. Eric Jang, who came on to explain how AlphaGo works, did a similar
thing when he was trying to build in a very strong GoBot. And he had interesting observations
about the kinds of, like, it's really good to just go, running an experiment and going down
that path, but it's bad at stopping at dead ends and just doing extremely parallel things.
Anyways, this will probably be changed, this will change in the future.
It's very interesting to think about what it looks like in the limit.
I mean, this is fundamentally like what the human institution of mathematical research is,
right?
It's just like, this is a library, extended it an interesting and useful ways.
And this way you don't have any outcome-based supervision.
No.
There's no outcome that you're trying to incentivize, but you have a process.
You know the steps are correct.
You just don't know if it's going in an interesting direction.
But yeah, if you were doing that, you don't want to completely go off the rails and, like,
do a random walk through the space of logic.
You'd probably want some, like, supervisor model that's trying to provide heuristics on whether it's useful or not.
But, yeah, something of that character.
I mean, you know people are working on it.
And, like, that's one of those, like, five years from now.
I'd be curious to, like, be able to get the future version of us, like, talking about whether,
like, maybe that goes nowhere.
but Terry Tao was talking about one, like, research project that's basically tried to exhaustively search the space of possible, like, algebras.
Like, you could imagine different, like, axioms that you apply to algebraic systems.
And so, like, when we come up with group theory, there's a certain axiom system that, like, has this flavor of, they kind of look like arbitrary rules, unless you know the motivation.
But it's basically, like, what have you tried?
All of them?
Do any of these yield useful things?
And, like, the vast majority of them is just trash in some way.
Like, it all collapses to, like, no interesting results.
But every now and then there would be this little island of like a completely different type of axiom system that at the very least seems rich in terms of like the number of theorems that can come out of it.
And that's like bread and butter for what you would imagine like automated provers being good for.
It's like exploring that space and seeing which one of them turns out to be something.
And like maybe one of those islands actually turns out to be something you can retroactively put motivation on to say this is the kind of structure that's trying to get at in the same way that you could imagine looking at the axioms for a group not knowing that it's about symmetry, but retroactively.
realizing like wow this is very relevant to studying symmetry so you could
imagine results of that flavor but instead of just exploring possible algebra
systems it's like all possible like logical consequences of any kind of
axiom on the point about whether you can provide process-based of provision
without lean so deep seek had their deep seek math model that and they released a
paper on how they trained it and it was quite interesting so they have um the
problem with having natural language proofs is you don't know
know if it's correct or not. And so they have a verifier. And then the verifier is trained by a meta-verifier
that make sure that all the problems that they're training this model to solve in like the
art of problem-solving, that the verifier is giving good feedback on that. And it like, it works. And so
it's just interesting, natural language verification with some sort of meta-verification kind of work,
at least seems to work so far in the published literature. And also it seems to work in the
published products that we're using. Like if you look at coding agents, they're getting better and
better at like writing clean code and refactoring code and stuff like that. And I'm sure that there's
process-based like LLM as judge kinds of things which are saying, trying to provide taste and say,
hey, is this like a clean way to write this function? Are we like, are there duplicates of the same
kind of modular forms and so forth? I feel like that should also work for mathematics, right?
It's like, it doesn't seem. It seems more plausible for math than anything else, even if you're
only working in natural language that you could trust a verifier.
I mean, you and I were talking earlier about why they're bad at writing.
And, you know, I was asking, like, why you can't just have, like, they seem to be good judges.
If I give them two essays that, like, students write, they'd be able to say which one's more, like, accurate and insightful.
So why can't you just have, like, a verifier saying, like, is this a good piece of writing or not?
And, like, maybe the ultimate failure there is, like, even if they're good at discriminating between, like, a B essay and an A essay,
they're not actually good at discriminating between, like, an A essay and, like, a thing you actually want to read that would be, you know, followable on substack and insight.
and all of that, like, they actually end up preferring just uninsightful pieces of writing.
And so on the math front, I guess the question would be like, that step to simply know,
like, is this a correct proof or not, that lends itself to, like, an automated verifier,
even in natural language.
You could probably still make a ton of the progress.
It still doesn't, like, I still like the sort of tree of logic out of lean front, just in that
you can really go off the rails, right?
Like, there's just no constraint on, like, the previous way that things had been phrased,
before in the same way that, you know,
everyone talks about like Move 37 in like AlphaGo and such.
Like, what is the thing that lends itself to just going
outside the prior heuristics?
And it seems productive to have a disconnection
from the rest of the world in that exploration
as like a complementary research pursuit
to the natural language math front.
I mean, the other relevance of lean there would be like,
okay, let's say you have your pure natural language
RL environments, and you have a pure natural language set of proofs, and people have the said,
like, precede AI mathematicians, and they go and they generate, like, 10 papers a day that produce
a bunch of stuff.
If the error rate, if there's, like, any error rate to that at all, so Alex Contrerovich
has talked about this, it becomes insufferable, like, as a mathematician, because you would
basically be like, every single time I see one of these, I kind of don't know if it's worth my time,
even if 99 out of 100 them are right,
I don't know if it's worth my time
to even go through it
because it's really labor-intensive
to find what that error would be.
And it's really frustrating
if it turns out you spent all your time
on a paper that was trash.
And so having anything that's able
to give you that green track mark
that says,
even if this is going to be complicated
to understand, even if it's going to be a pain,
you at the very least know,
it is correct.
Like every other field would kill for that, right?
And like math has that.
If the models are also able
take their natural language proofs
formalize them. And so that seems huge, right? The ability to have that, like, every field would
love to have something like that. And so I think you are right that Lean is maybe overrated on
the side of the importance of it being used as a VR environment for any kind of like just progress
in math generally. But I definitely wouldn't write it out of the story. Yeah. Yeah.
I also love this extension of Mathlev as a metaphor for like what's going to happen to our civilization
pretty soon. Sure. Right? It's just like for millennia, humanity is building this like
corpus of knowledge and understanding and everything that we have now distilled into these models.
And at some point, to the models will just like extend that arbitrarily.
By the way on the writing front, I actually have I have a theory of why writing is making
worse progress than these other domains. So I think one of them is what you said,
that they're bad at judging not only A versus B, but they get like,
just totally derailed by B-star,
which is this like a shitty essay
that just hits all the
bells and whistles that like A is supposed to hit,
and then so the reward hack thing
just like totally goes off the rails.
But I think the other important thing
is that writing is not modular
in the same way that code and math are.
Like, you know, you can write a function
many different ways and they kind of do the same thing
and of course you want it to be very clean and stuff,
but like at the end of the day, it works, it works.
Same with like lemmas and mathematics.
And then, you know, you can like,
have some end product that is different from the way it is produced. So the code is the thing
that produces some end product. And you want a functional end product. Whereas in writing,
the end product is directly the thing the AI is producing. And each paragraph, sentence,
word matters, because that is a thing that is like, like, that is the substance. It's not like
some separate thing that is produced out of the writing. And so any, it's a, it can't just be
It can't be sloped in the way that, like, code can be sloppy and still produce some outcome that you want.
But you were just pointing out how actually we've gotten much better at agents writing not just functional code but clean code.
Why is it not the case that the same progress that allows you to go from merely functional to like clean and like a mergeable PR doesn't also result in like clearer writing?
Yeah.
That's a good point.
I mean, also, has it not?
I agree there's many ways in which they're terrible writers,
but for a lot of writing I consume,
I find it's better to just copy paste it into an LLung
and to say, like, explain this to me.
The explanation will be better
than the thing that is produced by the human.
So it's funny that we say, like,
these are such terrible writers.
And also, my reveal preference is just like,
can I just have an LLN explain it?
Even when I'm talking to a human expert, like, live on a call,
if it's a piece of knowledge they have,
that only they have that's not encoded in,
the distribution, I want them to explain it to me. But then if in order to understand that I need
to understand a more basic concept, I would prefer if it was socially acceptable for me to just
be able to say, let's pause there. I'm just going to ask NLM how that works, and then we can come
back to your special piece of knowledge. Well, it sounds, I mean, that's distillation, right,
and explanation. And so if you're, if I'm thinking of like quality of you as an essay writer,
if it's that I give you a book to read and I want a book report, right? Then I'm
might believe that, okay, the LLM maybe gives me a better book report. But I think what people are
really getting at when they say it's bad at writing, like, what is writing? It's not just distillation
of pre-existing ideas. It's not just like how to explain clearly, because they are good explainers.
It's like, what is the insight? And this is where it gets like, just auto-regression is a very weird
way to generate stuff, because, like, when you're writing, you sort of know, in order for it to
be good, you have to have an element of the unpredictable. And it's not just like,
temperature in your mind or something, right? It's like knowing exactly the correct point when you
want to make an unpredictable move and that that's going to be what's more insightful. And so even if
it's like better at explaining a preexisting thing, it's like, what generated that book that you
wanted distilled in the first place? Right? It wasn't, it wasn't an LLM that like generated it and
you just needed it. It's like some author who, who threw a lot of exploration of ideas in the world
and then deciding what aspects of it were interesting and which ways of presenting it were like
the coherent, well-motivated narrative.
It's like they put that all together in some way.
And, you know, if they're a good author,
it's probably one that actually you would err
on the side of reading their book instead of the distillation.
But so what makes it worthwhile to like explore at all in the first place
and you're uploading it at all,
I think it's all of that side of it that's the, like,
when people will cite them being bad at writing.
And it's that element of unpredictability,
of being deliberately choosing something that's novel.
It's like very directly contradictory,
to like the way that things are being produced.
Yeah, that's a good point.
I think they're also really bad at building really good mental models of people,
which I think is a very important skill in writing.
So Annie Matushak and another collaborator,
whose name I'm forgetting right now,
did an interesting report where they tried to teach LLMs
to write good space repetition prompts.
And I really like this because even though it seems like a really totally random skill,
it's just like people are talking about recursive self-improvement in an era.
Yeah.
And you can't get these things to write good flashcards.
And what's going on there, right?
Right.
They tried many different kinds of techniques, and they're, like, you know, sophisticated people.
Like, they tried to RL open source models.
They tried all kinds of, including chain of thought and the big prompt they sent to the best close source model, et cetera.
And the key constraint, it seemed to me, was that writing a good card is about projecting somebody's mind in three months.
And what is the way in which they will associate the question?
Like what kind of answer we'll be thinking by the moment?
And is that is the elicitation that inspires the detail you actually want to take away from the passage you're trying to make cards about?
I think writing also is similar to this, where if you're writing something, you're like the reason it's such an enervating process that takes so long is each word you should be thinking or each pair of sentence should be thinking, what is happening in my reader's mind right now?
Yeah.
Even if I flip the phrasing around, so the end phrase goes to the beginning.
and like this is the first image that comes to your mind
before you read the rest of the sentence.
That kind of, maybe auto-aggression is bad at that kind of,
this is maybe a more diffusion-like property of considering the whole
rather than going sentence by sentence,
but also I think that requires a lot of mentalizing,
which these models weirdly struggle at.
Well, I mean, interesting question, like,
is it weird that they struggle at that?
So I might butcher this.
You know how when you cite studies that you once read
and it's like maybe the study wasn't real or something?
This is one very memorable one on,
okay so let's say you want to quiz people's EQ
like you show a flash card
of someone's like facial expression
and someone's trying to describe like what's that emotion
it's actually these really good tests online that'll have
like a face and then four possible emotions
and it's like surprisingly hard to
like describe exactly the correct emotion
but you also get the sense there really is a correct answer
and if you try this with like people in your life
you'll notice that the ones who actually are pretty
plugged in socially like do really well on it
and the ones who are a little bit more like left brain
like dumb. Okay. So that is a kind of test you can do. I vaguely remember an experiment to this
effect where they took people who had freshly gotten like Botox in some way and they did like a pre-test
and a post-test and like post-test they were just much worse at like reading people's expressions.
Like that feels kind of weird. They got Botox. So the person taking the test. It's like so you do the
test and then you go and you get Botox and your face is all like frozen and now you are worse at
understanding the emotions of what you see, right? And the thought is that part of
Part of understanding, like, this emotion that you're looking at, is doing it yourself.
That's crazy.
Like at a facial level, like, you know, moving your face muscles, and it's like, you see that,
you mimic that, and you're like, oh, yeah, that's anxiety, right?
At some, like, very subconscious level.
So in that sense, if it is the case that models have bad theory of mind, sure, they know
everything because they've, like, read what everyone wrote.
But at a level of, like, actually able to put themselves in your shoes in the same way
that, like, my face muscles are mimicking your face muscles, that's what helps
me understand how you feel. Not surprising at all. They don't have face muscles. Their brain
works completely different. It's just like, it's like an alien trying to empathize. Like, how could
it have theory of mind? It would be like this very emergent thing to have theory of mind.
Whereas we can just like plug it into our own minds. It's like, we've got the ready made
hardware to just like place it in. And so that's very interesting. It's not that's, from that
lens, it's not that surprising. Okay, Grant, we are both partners with James Street.
I'm sure over the years you've interacted with a lot of James Streeters. What have you found that's
unique about them or their culture?
I mean, I did this interview with them this year that partly was interesting because they
don't usually have anything outward facing.
I mean, in the industry, they're known as having a pretty wild retention rate.
Like people just stay there and just getting an inside view of that.
I remember one of the comments, someone was saying, even though the people have role titles,
like, you know, researcher or trader or engineer, they often don't know what their colleague's
actual role is because everyone's doing a little bit of everything else.
Like, even if you're officially a trader, you're doing a lot of research.
Even if you're officially a researcher, you're doing a lot of coding.
And I suspect maybe that's part of why they have the insane retention that they do,
because anyone who wants to be growing,
they just have the chance to do a lot of different kinds of things.
All right, Grant, I'll do the plug for you this time.
If you want to watch this full sit-down interview that Grant did with some of the folks there,
go to 3B1B.com slash Jane Street.
All right, Grant, let's talk more about AI and math.
What advice do you have about using LLMs to learn?
So as I was describing, for a lot of well-known concepts, I find them very helpful.
But often, just a couple of further messages down, and I'm trying to understand something.
And I just, they're so confused themselves.
They're confusing me, and they don't explain it the right way.
And then I'm just, I know that talking to the right human could clear of my confusion in three minutes.
I don't know.
And then I feel like more and more we're going to want to use these things as somebody who's
taught a lot about education.
Yeah.
and representation and stuff, we're going to want to use these things to learn things.
So, yeah, have you noticed the ways to use them more productively to understand concepts?
I'm curious to hear your take on this.
I mean, I'll give mine.
Even pre-LLM, I feel like a relevant insight in learning was recognizing that like who matters more than what.
So, like, advice to any college student when they're choosing what courses to take,
care a little bit less about your pre-existing interests because they're kind of arbitrary right now
and care a little bit more about whether like the person teaching it is a good educator and someone you resonate with.
I think in choosing what to read, like what books to read, like who the author is maybe matters more than if it's a prior interest.
So if there's a book you've liked before, read what else that author is written rather than reading another thing on that subject.
And I'm getting to like LLMs on this.
So like there's a difference in feel for trying to learn something if you look at a Wikipedia page of it versus if you look at, let's say like it's a philosophy topic and you go to the Stanford Encyclopedia philosophy.
or if it's a math topic, you go to the Princeton Compendium of Math.
The difference there is, like, the articles are deliberately written by one individual
who, like, tries to actually craft a motivation around it and everything.
Whereas Wikipedia, it's this, like, local minimum that's reached where basically every
sentence has to be correct.
And I think a good exposition, you care a little bit less about, like, correctness on the way,
but you can, like, deliberately craft things that are a little bit wrong that you correct along the way.
I think it's like edited out in a crowdsource environment.
So like that LLM explanations feel to me at the moment a lot like Wikipedia,
which is to say, amazing, right?
Like, imagine world before Wikipedia, like how long it would take to like find in like Sussin and everything.
But nevertheless, what's the most useful part of a Wikipedia page?
It's often just the references at the bottom, right?
You look at the like key references and you go to them and you read them.
It's like actually sometimes that gives a much like better overview of it.
So often I like to just ask an LLM.
LLM, like, who should I read?
Right?
Like, and maybe I can even give some specifics on ways I want to learn.
I actually got gas lit by this once where I remember trying to learn about like,
I don't like semiconductors or something.
I was like, this feels very visual.
This is all like text?
I'm like, is there any really good like well-visualized math video?
Or not math, sorry, a well-visualized video kind of like explaining the concepts that you're getting it.
In Klaught, it was like, yeah, here's a couple in the top one.
It was like, here's one from Three Blue and Brown.
I'm like, I can guarantee that there's not.
And it was an actual video and actual link, but it just had to like misattributed someone else's.
And it was good.
And it was like, I had a much better experience clicking over and watching that video to learn about the thing rather than like trying to proceed forward with questions there.
So in that sense, basically using it like a very souped up version of Google on like zero in on the right human written resource.
What about you?
Like what you engage with these a lot?
What's the best way to do you push your finger on it?
The most productive learning sessions I've had is when there's,
some artifact that a human is produced, whether it's an article, a book, a video that organizes
the relevant concepts in the correct way and builds up the motivation of why building up the
next idea would be relevant to solving the next problem you did encounter and the next idea
and the next idea. And using the LLMs to just do a little bit pruning around this, this
branch that the book is identified. So I was actually, I was going through, I think,
You might have recommended Steven Strogetz's textbook on...
The chaos one?
Yeah.
Chaos on non-linear dynamics.
I love that book.
And so I was going through it.
And it was like bliss.
It was like your videos in like a book form.
He's so good.
It was super fun.
And the way I was learning it is like I'd have on one third of the screen,
his like lecture from university.
On one third of the screen, I'd have that part of the textbook.
And on one third of the screen, I have an LLM.
And I was actually thinking if I was back in college and watching this lecture live,
we just totally go over my head.
Like, these kids must be really smart.
Because I'm, like, pausing and, like, reading the text book and talking about LLM's and then restarting again.
But with him curating what is the right order to understand concepts, what is the right problem to motivate understanding a concept.
Oh, also another thing that's a really bad at is a thing a really good human can do is when you ask a question.
They say, like, actually, you're just, like, not really thinking about this topic the correct way.
Yeah.
Like, the question you want to be asking.
Yeah.
The correct way to organize these concepts is X.
Yeah.
And that line just can't really do that.
Yeah, it's a little too placate.
I mean, this is ultimately like the very, like the supplicants.
You know, that's very like, oh, what an insightful question, you know, that kind of thing.
You want to strip that down.
That's a good point.
And I think that cuts to theory of mind a little bit.
Yeah.
Like, recognizing that to ask a certain kind of question reveals that the mental structures are not,
at least they're not the same as what the, like, explainer has.
And sometimes people do this to a fault, right?
Like, I think a really good teacher,
let's say you have like a middle school, like math classroom or something,
if a student, like, asks a question that suggests they're thinking about it in a different way,
it's actually really hard to like take seriously in the moment,
hang on, could you get to a right answer with that before you say,
oh, instead of that, let's do this.
And like the really good teachers are able to like jujitsu,
the, like, creative way that the student was thinking about it
and bring it in.
I mean, LLMs aren't doing that, right,
when they are not reframing your question.
Instead, they kind of run off.
But at the very least, it feels like there's three levels here.
And so, like, LLM is at one,
good explainers at another.
But then, like, the A plus explainer is the one
who can, like, jujitsu your way of thinking
and say, like, oh, that's where that's useful.
And so maybe there is a certain, you know,
cycle all the way around where, again,
five years from now, the LLMs will still be doing that,
but in the better way.
What is your recommendation to students who I'm sure email you this question all the time?
Look, I was curious about doing mathematics.
I'm really passionate about the subject, but seeing all the progress AIs are making,
I don't know if it makes sense for me to pursue this as a career.
And this is not relevant not only to people in mathematics,
but I'm sure to people who are noticing that their field is more and more getting
productivity gains or whatever from AI.
so coding is very adjacent to this.
Yeah, what advice do other people?
I wouldn't trust any advice that I give.
It would maybe be how I'd like couch it.
But even pre-AI, it feels very important for any job that you're going to go into to really understand.
Like if we're talking about a job, right?
We're not talking about like you're a gentleman scientist and you want to like engage with the math world or something.
You should understand where the money's coming from and like what value you're actually adding.
and like the connection between those two.
And I think often like a surprisingly small amount of thought has put towards that,
especially students, they're in this environment where they probably want to go into math
because they've always been good at it.
And they've just been rewarded in life for like proceeding through the next hoop correctly
and next step.
And when they think they want to be a mathematician, it's because it's a version of going to
continue to engage with that.
It's like, well, I'll go like, where do people get to do this?
Rather than thinking like, what value am I adding to other people?
And to what extent is that like the reason that like,
salary is flowing in my direction. It's actually quite different in different cases. In some cases,
it's a very prestigious mathematician, and their presence at a university lends a certain brand value,
and that's why the university wants them. In some cases, it's like the NSF grant is given because
you've got this public good belief that we have that basic science has, and like you've got this
institution around that, and there's going to be this whole bureaucracy around trying to act as a
proxy for what we think that public good is, and a whole song and dance around how to
correctly make them predict that your progress will be in the spirit of that funding.
Sometimes it's just straight up teaching, right?
It's like people like to send their kids to an institute that has experts teaching them,
and like that's what you're doing, and you are providing the brand value by being an expert,
and then the direct value by being a teacher.
So regardless of whether AIs are like proving theorems or not,
or like whether we're talking in 2016 or 2026, like that is a thing that not enough students
thinking I want to be a mathematician think about, but I think it's worth thinking about.
Like for me, I think that, you know, I just like wasn't necessarily thinking about it and kind of stumbled into this career path where basically math exploration can be monetized as entertainment, right?
And I like stumbled into that.
I'm like very grateful that I did, but it was an accident.
It wasn't like this deliberate thing.
And I think I could have avoided relying on serendipity and maybe done that a little bit more by design and had I been like thinking critically about it.
So to your question, if it's the case that you have almost automated theorem proving.
And then let's say it's the case, they're also really good explainers.
So it's like even to get the human understanding.
I think a lot of the like social role that mathematicians serve actually doesn't change that much, right?
You still have a sense of as a public, we sort of feel like there's value to basic science.
And we're trusting in the judgment of mathematicians to determine like where their time is best been.
And the prestige comes from within that community.
It's like other members saying that this was a really good result more than it is like the grant writer who like really
understands algebraic number theory to understand that it's a good result. And so there's going to be
some inner culture of what constitutes like valuable contributions. Maybe it shifts away from theorem
improving and maybe it shifts towards like good definition writing. Maybe it's that like museum
curator idea. But you're going to have that same community. And as long as society as a whole is still
like valuing like the premise of basic science. And if we're in the like abundance world of like
what AI brings, probably there's more funding in that direction in some sense, right? On the side of
prestige to institutions for
who their lecturers are, I mean,
I actually think teaching is one of those stable
post-AGI jobs
that there is, because it's so relational.
It's so, like, this is where
parents want to spend their money if they have an
abundance of wealth, is, like, on
good teaching and good educating.
And it goes so far beyond
explanations, like, even if LLMs are good
explainers, the thing that a teacher is doing
is such a social, like, coaching
mentor type thing that, like,
that's probably the most, one of
most stable careers that's going to exist over the next 50 years. And so insofar as what a lot
of mathematicians role is, like overlaps with that, you know, you as the prospective student going
into it, you could lean into that. Actually, think a lot more students should, like, think about
and give, like, pay credence to the idea of being, like, just a math educator and, like, the value
that that can serve towards the next generation. So I'll couch again on, I don't think I'm the one
to say, here, prospective young mathematician, like, here's how you should,
think about the future, because I'm like a YouTuber, right? I'm someone who is not in the institution
that they are thinking of going into, and so I'm speaking as an outsider looking in. But it feels like
generally good universal advice, know where the money is coming from, know where you plug into that.
And like, if you're just asking those questions, you're actually already like steps ahead of all
of the other, like fledgling prospective mathematicians. Yeah. And in fact, I think in the crazy
world, in the world where within five, ten years, the AIs are coming up with not only solutions
to the Millennium Prize problems, but coming up with, like, just totally novel problems
to be solving in the first place, novel mathematical fields and objects and stuff. It is in that
world where, first of all, there's a ton of abundance, and two, the things that AI minds
will have, like, gone furthest in, where they will have seen, like, furthest beyond our horizons
will be mathematics. And there will be so much.
demand of like what have the AI seen? Can you explain it to us? Yeah. Yeah, I feel like in the in that
world, if there's any jobs whatsoever, surely distilling what the AIs have learned will be one of them.
Also, it's funny because all of this sort of presumes that it's useless, right? Like we're not talking
about the actual practical applications of what math is being is being done. So insofar as there's any
economic utility to it, you would imagine that the people who understand it and are able to like
make the decision of where it should point. Like they actually have a lot of.
more economic value by like being able to make that judgment as curator and point this like
behemoth of like new math like pointed in a useful direction like suddenly that's a much more
levered move to make than it had been previously can I actually ask you about that so obviously
the um one question for AI for math is not only can it do it but is it any good yeah or isn't
any good for anything you were describing all the ways in mission group theory we're trying to solve
this we're trying to figure out random facts about the roots of different kinds of functions
and now it's all these different applications that are practical across many different fields.
Do you have some sense of if we just totally get to a place where mathematics is, the
field of human mathematics is accelerated 10x or 100x that some crazy shit happens, or are we
just actually going to be bottlenecked by other fields?
I think there's some fields that probably
will, I mean, it's super spiky, right? I think, like, progress in algebraic number theory,
it feels unlikely that that then, like, unlocks some things. But I don't know, I remember talking
to this mathematician who does more, like, like, dynamics and, like, PDE-solving type stuff,
and he was referencing, basically, like, his group had some ideas that, let me say if I
summarize this right. It's like the way that Boeing would make planes is they would, like, make it,
And then they would do a bunch of tests, and they had to, like, disassemble it and reassemble it based on those tests.
And they essentially had some insights on how to, like, do more things in simulations such that you don't have to, like, deconstruct and rebuild it.
And it saved Boeing, just, like, billions of dollars or something.
And then they just started funding that, like, group, which is, so that's, it's, like, much more obviously application adjacent because, like, PDEs just sort of are that.
So progress in that domain, you would imagine, like, actually do unlock some things.
And I don't know if it's these like step changes, but maybe it's more on the side of, like, engine design becomes just a little bit more fluid or, you know, like coming up with the right wing shape instead of running a whole bunch of complicated, like, CFD. Or maybe you're able to, like, speed up your, like, CFD simulations because of certain pure math insights, like, makes those more efficient. I bet you'd just see, like, a lot of, like, great incremental improvement there.
it seems less likely that
the massive breakthroughs
in math immediately turned into
this massive economic breakthrough
you solve the Navier-Stokes
problems and then that
unlocks an ability to simulate more things
but you probably will see
at those fringes just
some meaningful
leakage outside of the pure math
insights into other things
also I mean there's a ton of people working on things
like AI engineers like physical
engineers like material science and things like that that would be you have to imagine that like
they would be in a good position to look at the AI math insights and decide if they're relevant in
some way or not and so it's another one of these things where I'm not going to sit here and like
put a flag in the sand like predicting that there will be it would be a little bit disappointing
and a little bit surprising if there weren't over the next five years like economically
valuable improvements that were made that were directly like referable to and
the like AI progress in math.
Like that just would be kind of disappointing
if it was just taking down a bunch of erudish problems
and like none of them actually,
you know, it wasn't doing any of the math
that actually directly touches physical world.
Yeah, I mean to your point about,
well, a lot of history and mathematics
about like building up these like piles of concepts
and connections and whatever.
Yeah.
And sometimes the piles connect with each other
or they're, you discover an application somewhere else.
At the very least, you just build up this huge pile.
And then as a, you know, broader progress
and society happens during the singularity
when we get to the industrial part of the singularity.
You just have all these different ideas
that you can hopefully are useful
in other parts of the world.
I mean, yeah, like I said,
one of the interesting things about what's happening
is it causes people to step back
and ask, like, what is math?
And maybe one of the awkward conclusions of it
will be revealing, like,
ah, man, over the last,
like, it's just become wholly useless.
Like, the kind of questions being asked
to become, like, so divorced from things
that are physically applicable
that, like, that's one of the things
mathematicians have to come to terms with where everyone will look and be like on a second like
weren't you guys supposed to like if there's so much that's like 10 X progress there like why aren't we
seeing it over here and then that church is like every time we wrote those grant proposals and said like
trust us like the elliptic curve progress is going to help with like cryptography like it like
shines a light on the fact that like maybe it doesn't so that's that's one possibility
um grant this is super fun thanks so much for doing it absolutely my pleasure
