a16z Podcast - Daniel Litt: The Mathematician's Guide to AI
Episode Date: September 1, 2026a16z’s Lisha Li sits down with Daniel Litt, Assistant Professor of Mathematics at the University of Toronto, to unpack AI's rapid progress in mathematics, what today's frontier models can actually d...o, and what they're still missing about the way mathematicians think. Daniel explains why some recent AI-generated results are genuinely impressive, including an autonomous solution to the Erdős unit distance problem, but argues that solving problems is only one part of mathematics. Today's models can grind through calculations, combine known techniques, and search enormous spaces, but still struggle with intuition, theory building, identifying the right questions, and developing the kind of big-picture understanding that drives much of mathematical progress. Lisha and Daniel also explore how AI is already changing mathematical research, why an explosion of AI-generated papers could distort academic incentives, and what happens if researchers outsource the work of thinking rather than use AI to deepen it. Ultimately, they ask a question that extends far beyond mathematics: as AI gets better at intellectual work, how do we make sure humans keep getting better at thinking too? Resources: Follow Daniel Litt on X: https://x.com/littmath Follow Lisha Li on X: https://x.com/lishali88 Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
The goal of mathematics is not to produce mathematics papers.
It's to produce some kind of understanding.
Maybe some of that understanding resides in model weights.
To me, that's pretty unsatisfying.
Comparing anthropic with open AI, do you detect any differences in how that is similar to human reasoning?
They definitely are not good at it autonomously.
But with some hints, you can kind of get them to do something interesting.
A lot of progress in mathematics comes from letting, you know, a thousand different flowers bloom,
and people pursue their own curiosity, and then, you know, the boundaries of knowledge of things.
boundaries of knowledge expand in some kind of fairly uniform way.
What has been the most impressive results so far?
My favorite fully autonomous result by an AI so far is the solution to the Earthish unit distance problem.
There was some lemma I wanted to prove none of the front-tier models could do it.
So I worked out a ton of examples on my own and I realized, oh, well, maybe here's some reason
why it could be true.
Once I had that statement, the models were able to very quickly prove that sort of a better statement.
How should the mathematics community best adapt and benefit from this?
AI can increasingly solve math problems that would challenge professional mathematicians,
but solving a problem isn't necessarily the same thing of understanding it.
In this episode, A16Z Infra Partner, Lisha Lee,
sits down with the University of Toronto mathematician Daniel Litt
to separate the headlines about AI and mathematics
from what the models can actually do today.
Daniel explains why recent results,
have changed his views of AI, where frontier models already resemble human mathematicians,
and where they still fall short, particularly when it comes to intuition, developing new theories,
and even figuring out which questions are worth asking.
They also explore what happens to mathematics when generating a proof becomes cheap,
why academic incentives may need to change,
and how mathematicians can use AI without outsourcing the understanding that makes the work valuable in the first place.
And more broadly, they ask,
What's what mathematics can teach us about working with AI as increasingly capable models move into every knowledge profession.
I am so excited to have you on Daniel. And so Daniel, it is a professor of mathematics at the University of Toronto.
Toronto is my hometown, so also very exciting. But the thing that is most special here is Daniel's an actual practicing mathematician.
And in addition, he's been incredibly vocal about his evolving views of AI in math. And so I feel like every, if I just don't check in with you,
you know, like in two weeks, something, you know, different has been revealed, and then you're
very kind of like, what do you call it?
I have a lot of opinions.
You have a lot of opinions, exactly.
So I want to get into that.
So, I mean, one of the things that I'm most interested in
is not just like a discussion of how the capabilities of advance.
I feel like in math, that's definitely the headline, et cetera,
but also you've been very thoughtful about how practicing mathematicians should respond.
And so that kind of gives us a chance and opportunity to talk about actually what is special
about math.
It's not just like, hey, AI has been really making progress here,
but delve into what actually mathematicians do.
And so maybe like with that arc in line, we can start with what has been the most impressive
results so far, given all the recent progress for you. And then maybe, yeah, we'll kind of take it
from there. Yeah. So, okay, so there have now been a lot of results. Some of them produced autonomously,
some produced like semi-autonomously, some who's like where the AI contribution is just not
at all clear. They're in a lot of different areas. So anything I say is kind of, you know, I can only
really comment on things that I can have some expertise on. So it's quite possible that if you talk to
a different mathematician, you'll get different answers here.
So my favorite, like, fully autonomous result by an AI so far
is the solution to the Erdus Unit Distance Problem,
which I think was announced in mid-May.
So what I liked about that is it seemed to me that, like,
it was in some ways a little bit creative.
So I think some of the results we've seen have had kind of the flavor.
You kind of take some known techniques and apply them in maybe a clever way,
or you know they've kind of been at some kind of results I would characterize this, like,
last mile, like where some recent work was done quite deep work done by a group of human
mathematicians and then the AI took the final step.
But yeah, with this sort of distance problem, I think it was something where the result was
like a little unexpected.
So first of all, my sense was like that people working in the area thought it was true and
then there was a counter-re sample found.
But then also it like brought in some techniques from another area.
I think those techniques were like not especially like deep or new.
They were sort of classical ideas from the 60s,
but they were new to this area of studying, you know, point configurations in the plane.
And so that was pretty cool.
And then afterwards we got to see, like, it was kind of fruitful.
So a bunch of mathematicians took those ideas and used them to find counter examples
to a bunch of other interesting open questions.
So, for example, like the song product conjecture for the real numbers.
So that's at least what way I like to think about how cool result is.
Like you look at it post hoc and you see like, oh, well, were whatever new ideas that were
introduced, if any, like, kind of useful to do other things. Did they improve our understanding of
something? And I think that's maybe so far the main example I know of a result of that form.
Yeah, I think it's really meaningful that you're commenting on this because that result came out
to your point in May. And there's been so many headlines so far. And it's kind of probably
hard for somebody who's not a practicing mathematician to appreciate the differences in these
headlines. And so you kind of already started laying out sort of a taxonomy of like what is
different in that proof. And so it would be kind of interesting maybe to use.
that as kind of both an excuse to talk about where you sense the model differences are and what
like mathematicians actually do. So in this case, I mean, the most kind of maybe naive understanding
of what mathematicians do is that we're pushing around symbols in a logical manner. And this is
why RL is so successful at this because you can kind of both verify it somewhat cheaply compared
to other domains and then also because the rules are quite illegible. And so you can kind of like,
if you're superhuman at that, you might be good at math. But I think that, of course, betrays most of
actually what is interesting about mathematics, which is perhaps I think you said this as well,
but I think anybody who's tried to do math is sort of like, it's about the understanding and getting
at truth, remaining confused, and developing intuitions. And I mean, the tool to do that stuff is,
of course, in having really strong abilities to push out logical implications. But maybe if you can
kind of speak to like, when you say it's most impressive and creative, like decoupling just the
inhumane maybe feats of just like logical implication from like where is it being creative
what is it helping engender in terms of like mathematical activity as well yeah okay so first
of all i mean i have characterized as to this like inhuman in some way i actually think the argument
was very human oh i'd love to go into that i or sorry opening i like released some chain of thought
yeah it was very recognizable it was like if i tried to imagine like my chain of thought
and trying to solve a problem like it might look kind of like that yeah we haven't seen the raw
a chain of thought would be maybe that way.
Maybe they cleanse it a little bit, yeah.
Yeah, or maybe the model like to swear a lot in the middle of the chain of thought
and then that up or something would know.
But at least the summary seems pretty recognizable.
And I would say that's actually like kind of typical of most of the results that I've
studied.
Like they don't seem inhuman at all.
They seem absolutely like something a human mathematician could produce.
And they're like typically understandable, if not so well written.
If you just look at problem model output, it's not like there's some move 37 or whatever.
It's like a human mathematician doing math.
It's like a human mathematician doing certain types of math.
So, like, they're definitely, like, the models are still, like, not very good at some mathematical activities.
And here I don't mean, like, by field, but just, like, certain things you do when you try to solve a problem, the models don't seem to be doing.
But certain things they're very bad.
So the ways they might be a little bit unhuman is, like, they don't get tired, they know a lot.
But if you actually just read the final output, it doesn't seem kind of inhuman.
It's interesting because, I mean, obviously, we can't get too much of information.
from the labs who are producing these models,
why of what the training recipes are
or how they're kind of advancing in reasoning.
But at least one of the things we do know,
and I think kind of opening eye spearheaded this,
is just like reasoning in natural language
is actually what they happen to scale up,
and it's not actually pushing a lot of lean verified proofs
as like the training corpus.
And that's kind of amazing.
On another side, what's kind of interesting
is it's not clear that a lot of the mathematical training data
if they use that to a large extent at all
is reflective of how mathematicians think, if that's fair,
because a lot of it is not legible as traces, right?
Like most of the papers are crisp and, like, polished.
The textbooks certainly just show very little motivation
of how something is developed,
which is why it's usually easier to kind of follow a research direction
by actually talking to the researchers
and how they're thinking about it.
So I'm kind of curious, before maybe even going to the taxonomy
as an excuse, like, when you're examining these models
and their results comparing anthropic with open AI,
Do you detect any differences in how that is similar to human reasoning?
And then also if you have any comments on insights on perhaps why natural language scales so well that way, even though it's...
Okay. So first of all, I really like your point, by the way, that they're mostly doing natural language reasoning rather than lean.
Like, I think that that suggests to me that like these...
You hear a lot of people say, like, math is a verifiable domain.
Like, that explains the progress, whatever.
Like, my sense is that because they're primarily scaling informal reasoning, like, probably to take...
techniques are going to generalize to other domains pretty well. That's just my guess. Okay. You asked
a little bit about Claude versus Chachabee. My sense is that they're pretty similar in terms of
capabilities. I've played around a lot more with Chetabee than Claude Fable, but it seems like
there's a lot of cases where Open AI will drop a solution to some problem and Anthropics will
say, oh, you know. We also. Exactly. It actually seems like they're solving a very similar
collection of problems. And it's like kind of a relatively small portion of what human mathematicians do.
So yeah, one thing that is interesting is, like, we see the solutions have a certain flavor, right?
Like, there'll be things that this isn't surprising.
Like, there'll be things that rely on the model strengths, like their ability to ground out along computation
or pull together kind of technical ideas from many areas, or like maybe many papers that, you know,
even my mathematician might not have bred it.
But they're not, like, they seem like weaker in things like intuition or, like, having some big picture
point of view.
Like, a lot of what I do as a mathematician is, like, I have some kind of philosophy that's, like,
very non-rigorous.
Like maybe I think this thing is kind of analogous to this thing.
And then like a lot of what I'm working out is like figuring out how to make that precise
and like, you know, trying to measure the extent to which I've succeeded in understanding that
by like, can I solve a problem or whatever.
Like can't find an interesting phenomenon which I don't understand it.
But now I can understand it.
And like so far I haven't, you know, you can see even the best results the models are producing.
You don't see that much of this kind of reasoning.
It's more like they're very, very good at applying some known techniques.
Which is, to be clear, that's like a very powerful thing to do to be very good at applying like all known techniques.
There are mathematicians who have had great careers doing very high quality work of that flavor.
And I think a lot of what the models are producing is high quality in that way.
But it's like some kind of fairly narrow band of what mathematicians care about.
So far, I do think there's like signs of both, you know, all the frontier models starting to be able to do more fuzzy things.
So, like, I've tried to get both Fable and Chad GPT, 5.6, Saul, I guess, to do some kind of theory building.
And it's like they're not good. They definitely are not good at it autonomously, at least with, like, whatever scaffolding I've set up.
But with some hints, you can kind of get them to do something interesting.
You know, when you do, when you give the models hints, it's always a little hard to tell, like, what part is from the model and what part is from you.
But my experience is that, like, if they can do it with, like, 100 bits.
of hints or whatever in six months, maybe they can do it without hints.
Yeah.
You know, I do think there's signs that they're also kind of somehow picking up some of this
like implicit and unwritten mathematical knowledge.
Okay, I would love to, so go into the intuition part and where it sucks at, to put it
in a very basic way.
But you actually mentioned a small detail, which is that, you know, from the public
you've gleaned that Anthropic and Open Air probably neck and neck, but you personally are
using a lot more chat GBT, like, why is that? Or like, 5.46. Well, I don't know. I mean,
I just, I think it's just in our show. Like, I have a certain. Pretty sure of habit. Yeah.
It got, one thing is that chat GPD got better at math earlier. So, like, for a long time,
the cloud models were just, like, not useful for research math. And I think maybe around opus 4.5
or opus 4.6, they, like, more or less caught up. Yeah. But, you know, some experimentation suggests to me
they're pretty neck and neck.
And so for my own work, you know, except when I'm just experimenting,
I must have just stayed with one.
Yeah.
As people outside the labs like us, it's really interesting just to compare how they
differ on the frontier.
And, I mean, you know, to your point, it might be a little bit of momentum.
I do think that, at least from my anecdotal experience, 5.6 has been like a lot more clear
in exposition.
And it's just like, there's a little bit, and this might not be true.
I mean, obviously the models are just incredibly jagged at the frontier.
But, like, it's in the explanations of results to me, I always find that 5.6 is giving a more accurate theory of mind of what it assumes I know and don't know, whereas Pable might be explaining something very trivial, but then just, like, jump at like, well, you know, obviously you should know these things.
Yeah, I find they're both pretty bad at theory of.
Okay, great.
So then seeing, from your eyes, you're probably asking much deeper questions.
Okay, so the thread that I really wanted to pull on was when you're talking about the models
maybe starting to get better intuitions or even theory building.
So maybe before we even dive into that, it would be useful to kind of talk through like
what is your primary, you know, activity as a mathematician, especially in your area of
like algebraic geometry, probably has a very different flavor than a combinatorialist or, you know,
some other areas.
So if you can give a little, maybe brief way of the land and then kind of explain what you're,
what your mathematical activity was pre-AI and then maybe how it's kind of like changing with AI.
Yeah. So, I think there are a lot of different kinds of mathematicians. There's like a lot of different, you know, spectrum on which one can put in mathematician.
So definitely a lot of mathematicians like solving open problems. And then I think, I'm one of those. Like I like to solve an open problem that's, I think of myself as a problem solver as opposed to like one other.
axonomy you could have as a problem solver versus a theory builder.
At least for me, the point of an open problem is it's supposed to measure your failure to
understand something. So it's kind of like a benchmark. Right. Like, you know, one problem I really
like is the gross and decaps peak or aperture, whatever that is. It measures something about
our failure to understand differential equations. So it's like there's some very basic
object, we would like to understand.
If we can't answer this concept,
we know we don't understand it.
Okay. So in practice, like, how do you
get a problem which is supposed
to be measuring something you don't understand? Well, like, of course,
you try to understand the thing better.
And in practice,
what that means is, like, well,
you try to find the smallest situation
where you can't understand something and
fiddle with it and then you stare at
like once you win, you stare at what you
develop to win and try to turn that into
some theory.
So that's one thing you might do.
You might try to solve a problem, and in so doing, develop some kind of new understanding of the situation.
You might also just, like, have some feeling that, like, this thing is related to this other thing.
And, you know, you might start building a table, like, oh, this property A is related to property A prime, property B is related to property B prime, and so on and so forth.
So, for example, in my work, a lot of it is motivated by some analogy between comology of algebraic varieties and representations of fundamental
groups. So that's some fancy
stuff. But it's just that this analogy
is very, very fruitful.
And like really any phenomenon
that appears on one side, you can find an
analog on the other side. And so
trying to
realize
that dream has
led to a lot of beautiful mathematics the last
like 30 or 40 years
about people like Carlos Simpson and
Chiro Machizuki and others.
And so here it's really
just like someone noticed like here's an analogy.
And then that analogy has led to a huge amount of developments.
There's not really like an open problem at the end,
although of course, like as you develop this,
you come up with lots of open problems.
There's just like a philosophy that you're trying to realize.
And that philosophy is like super not rigorous, actually.
It's not like symbol pushing at all.
Yeah, so that's another kind of activity that I like.
Yeah, beyond that, no, a lot of what you'd,
do when you try to, you know, a lot of the activities, you're, like, you're actually trying to
figure out what the right question is, even. Like, here, you have some object, you feel like you don't
understand it. And, like, figuring out what you don't know is, like, actually a very challenging
thing to do. So, so, you know, there's, like, a list of, you know, you can go online and
list open conjectors or whatever. And this, like, doesn't really capture in a lot of ways what we don't
know. Like, often finding the conjecture is, like, really, really hard. So, I don't know. A good
example of this is like the Birch and Swinertrandyre conjecture, which is one of
melanium problems, is this beautiful relationship between L functions of elliptic
curves and the set of solutions to the corresponding equations of the rank of the
group of solutions, what if that means? It was discovered by like, it was like the first big data
conjecture. So Bershen Swinertan Dyer had like found all these statistics on elliptic curves in
the 60s, like it was one of the first ever computerated bits of mathematics. And they like
graph these statistics and they noticed that, you know, some, so the slope of some line on the graph
was related to some other algebraic invariant they knew and that was the source of this
conductor. So a lot of the time, you're just like working at examples and like trying to, like,
you're doing kind of science, like, you run an experiment and you try to figure out an explanation
for that experiment. And so, I think, going back to AI, I think, like, these are things where
AI seems so far to, you know, help a lot more in some things than other things. So, like, you know,
I think the more vague
a phenomenon is
the less you have a precise question in mind
the less useful it happens to be.
And so you were asking how I use it in my daily life.
Actually, what I've found is that the projects
that I have that kind of predate AI,
like the projects I've been thinking about for three or four
or five years,
it's just not that useful.
Like it's primarily kind of a substitute for Google or something.
I might use it to learn about some related topic
or like something where I would have earlier
or Googled something and then read a paper,
like, maybe I'll discuss it with it instead.
So it'd save some time for sure.
It's like not really doing deep intellectual work for me.
But then, you know, because I, you know,
I'm now like, well, I suck at coding.
And so now I know, you know, my good friend who's really good at coding.
And now I have all these coding projects.
Because like suddenly, well, if I had a question where coding would have been really useful,
I would have procrastinated on it for six months until I.
Tireless PhD student.
Exactly.
Yeah.
So, so, yeah.
Now I picked up all these projects.
It's where, yeah, coding is really useful.
Like, the models are, like, very good for kind of massively parallel things.
Like, if you want to find an example of something, you can just ask it to find, you know, work through a thousand examples in parallel.
It would be 10 examples at a time, you know, 10 different sub-agents.
And, like, that's really useful.
But these are, like, kind of different activities, which are, like, I don't know, on top of what I was doing before.
I'd love to dig in.
Yeah, go on.
Yeah, sorry, because you were mentioning, you know, there's projects that you've been working on for, like, three, four, five years.
And I don't know if it's correct to say, like, those are more of the theory-building aspect of it,
because you did characterize yourself as, like, an open problem solver, or, like, what is the thing that is?
Because, like, the deep thinking part, maybe, maybe, like, you know, kind of for this audience,
it might also be useful to kind of say or explain, you know, most of pure math is, it's not like
it's being motivated by, you know, nothing on applied math, but applied math is like, at least there's
some external motivation for why a certain formal structure might be interesting to study,
Whereas this one, it's like purely, it seems almost sociological.
And to some extent, I think like Thurston made some comment,
or a point about this in the 70s,
that it is a sociological phenomenon of more and more mathematicians start examining something,
then you'll maybe, you know, converge on some interesting structures.
But it's just like it's not, there's some reason why people would prefer to study something
or they think it's beautiful.
Like, what is it that drives maybe you in particular,
and then maybe you can make a more general comment about the profession?
Yeah, I mean, so definitely some people,
are like motivated by like beauty or or kind of aesthetic considerations.
I try not to be motivated by that.
Oh, that's like a controversial.
Yeah.
Yeah.
Like, well, one thing, like that sort of limits you, right?
Like one thing, kind of a failure mode I see among young mathematicians sometimes is like,
you have something and you kind of think, you know how to prove it.
And then like the proof feels really ugly and you decide.
But like, okay, I mean, what if you're wrong?
It's not ugly.
Like, you let me yourself.
For this audience.
What is ugly? Because I have an intuition of what's ugly, but like, what is that for kind of spelling out?
I don't really know. I don't really have aesthetic feelings. But people sometimes feel this way.
Maybe it involves a lot of grinding.
It's calculation that's not illuminating or something like that.
But you should like win by any means necessary in my opinion.
Like I like to think of what I'm doing like doing kind of physics except with concepts.
So, you know, instead of beauty, I try to think about maybe like what is kind of fundamental,
what's going to open up further understanding most.
Yeah.
Will, you know, am I introducing a new idea that will be like broadly useful to understand this object?
Yeah.
I guess, you know, in some sense, there's like some aesthetic consideration there, but it's,
I try to, I think the orientation of like trying to do good science rather than trying to do art.
Yeah.
But there's a huge, I mean, there's a huge variety of opinions here and like lots of mathematicians
think of themselves as being like closer to poets or something.
Yeah, what I do think is broadly true is that like progress comes from people like it seems to come in mathematics from people sort of pursuing their personal curiosity
And that it's sort of been crucial historically that there's like a lot of different people with different views of what's interesting and then the frontier of knowledge expands and expands and then suddenly you know you have these opportunistic
situations where a new idea has been introduced and you can now find you now can suddenly like cascade through a bunch of other things
things that we didn't understand before.
Yeah.
I mean, so then going back to the three, four, five year problems and where are the models
sort of not useful?
Kind of ask it another way.
Like, when you're doing the deep thinking, like, is it just that it's just not clear
that you formulated as a problem?
And it's more that you're thinking about these, like, what are, you know, what are the
fundamental kind of physics of, you know?
So in some cases, I mean, the, the, the, you know, the.
Some cases, there is like a well-stated problem here.
You know, I said, I've been thinking about things for maybe 10 years now at this point for certain problems.
Sometimes there is just like a conjecture that I would like to prove that is there.
I think one reason the models might not be useful for some of these things is like the conjectures are true.
So like I think that certain, you know, for example, you know, this unit distance problem, like the general belief in the community was
that it was true and then it turned out to be false.
And so what that means is that there's like a specific construction you can do to refute it.
On the other hand, I think a lot of the things I think about like, I don't know,
maybe I'm about to, you know, there's someone come up with a found example
to the peak of a curvature conjecture, quarter of the room on hypothesis tomorrow and like,
I'll look like a fool.
But in general, like these conjectures fit into some very broad theoretical framework,
which means that, like, we actually have a lot of evidence that they're true.
And so, yeah, so that's part of it.
Like there's not like a construction you can do to refute it.
You need to somehow kind of, you know, there's, we have this giant framework where certain
pieces of it are only contracture, and you probably need to resolve some of those contractors
to win.
We also have like a pretty good sense, I think, that like very serious new ideas are needed
to resolve those contractors.
So like, of course you can't be sure, like maybe there's sort.
very clever construction that will let you, I don't know, avoid having a big new idea.
I don't know.
Maybe, you know, it's quite possible we'll find this out.
But my sense is that for at least a lot of the things I've been thinking about,
there are, they're just not accessible to, you know, applying known techniques
very, very technically strong way.
So you need to develop a new technique.
I'm just being clear, like I'm not saying the models won't be able to do this.
So far they seem not to...
Actually, that's exactly the point
that I wanted to delve into
because it's, to your point,
it's like, okay, we can try to calibrate
and forecast like why they would get better at this,
but it is true, like, you know,
that just providing construction for a counter example,
they seem to be strong on.
If you have to start developing either new theory
or, to your point, techniques, machinery,
to solve, to prove why a conjecture is true,
it struggles more.
And it's probably because a lot of what it's drawing on
is also just like techniques that have happened in other areas,
and they're porting it over.
And to your point, that's why maybe the unit distance problem
was such a more creative result
because it was maybe doing more of that on its own.
It was like from another. At least it was like growing in something
from the not expected area.
Exactly, exactly. And so like maybe then to kind of ask a question,
it's not that because now we're like, okay, fine,
A, I'm getting so good, so fast, we can't count it out.
But like, why, what do you think it has to,
yeah, I guess it's, you're spelling out what it has to get better at,
but maybe some more kind of feelings on like why it's kind of particularly hard
to then develop that new theory and technique.
Yeah, it's a good question.
I mean, I think you just need a different, like my guess actually is it's probably totally doable
and it just hasn't been done yet.
Like maybe you just need a different RL environment.
I don't know.
Yeah.
So at this point, my expectation is just like that the trajectory will continue upwards.
I'm not a skeptic of continued capabilities growth.
Yeah.
But yeah, you know, I think what is definitely true is that like,
The skill of, like, developing a theory or, like, building your understanding of some poorly understood object is, like, a fuzzier one.
So it might be harder, you know, I guess you can tell it, you know, develop your understanding of Zeta functions.
And then once it proves the real hypothesis, you give it a reward.
But it's, like, kind of harder, I think, to come up with some, like, intermediate things that you can reward.
Yeah.
That said, you know, I do think mathematics as a whole provides a lot of conjectors of varying levels of difficulty.
So, you know, maybe this explains why there seems to be a little bit of progress in these areas.
Like, presumably they are trying to, you know, get it to solve lots of problems and some of those problems develop at least some of the skills of theory.
Humans are able to develop these skills.
You know, I guess sometimes they get rewards from their PhD advisors and their advisor says, oh, that's a good idea or something based on some.
element of human taste or whatever.
Yeah.
And that might be something one can do too.
But yeah, my expectation is that just like as part of continued capabilities,
we'll see growth in these areas too.
Yeah, yeah.
And I like the framing where you're sort of casting these increasingly difficult conjectures
as a, you know, a form of curricula for both humans and, of course, AI.
And it, you know, it may be kind of getting a little bit too philosophical for some people's
taste, it's like, it gets at the question of like, why are we particularly good or maybe
particularly bad at math, too, because it's, you know, what is it that we're either struggling
to do or some people are particularly good at when you develop new theory? Because it's not,
again, it's like not really, or maybe it's related to like, why can we formulate good structures
for physics as well. It's just like it's not obvious. All parts of the world are kind of understandable
and legible that way, but some parts are,
and therefore we try to do it
because there's maybe compressive pressures
on our minds
because we need to,
we can't understand anything
beyond, you know,
compressing stuff more finely.
I don't know if that's also like the correct
interpretation of, like, yeah.
I'm a little bit skeptical as like compression
as a metric of interest,
but it's definitely like as an anthropological reason
for like why we do a certain theory building.
I think it's true.
In fact, I think it's like kind of like our inability to just grind is kind of important to our ability to make discoveries.
So I'm just like as an example.
So I have one paper out so far where the models were kind of useful.
So they like proved a couple lemas.
So this is some situation where like we had to kind of prove in the main result.
And then there were like some lemas I was unhappy with like they seemed not to be optimal.
And so I like worked with actually this was with Gemini Deep Think.
which at the time was also on the frontier, no longer from now,
where I kind of worked with it to improve the lemmas.
And so what happened here was there was some lemma I wanted to prove
and the models couldn't do it.
So none of the frontier models could do it.
And so I worked out a ton of examples on my own
and I realized, oh, well, maybe like here's some reason why it could be true.
So I found a better statement of the lemma.
And then, okay, once I had that statement, like, okay, probably I could.
have done it pretty fast, but also the models were able to very quickly prove that
sort of better statement. So like our inability to prove it led to an improvement in the result.
Okay, now you can take the original lem I had that the models weren't able to prove, put it into
chat GPD 5.6 Pro, and it will output like the worst proof you've ever seen, like 10 pages of like
just brutal calculation with no insight whatsoever. And so, okay, this would have been a perfectly
proof, but it would not have led to this discovery of like, I think, a kind of beautiful conceptual
explanation for why this thing we discovered what's true. So like we found a better proof because we
couldn't do, I mean, okay, I say we couldn't do the calculation, like what actually happened.
And here I understand that I'm being hypocritical about my complaints about ugly proofs before.
It's like I realized like this horrible grind proof would work.
Yeah.
And I like could not bring myself to do it.
And so I look for another argument.
And so I mean, it's not, yeah, it's but now that the models can do for very long technical calculation is pretty reliably.
Yeah, well, I mean, to your credit, I don't know.
I think I understand why you're saying it's like,
don't, you know, just shy away from trying to prove something
just because it seems ugly, because it's like you have to take the first step, right?
And eventually you all try to work towards insight.
And so, you know, you might be shy about admitting it,
but it is still maybe driven by whether it's like aesthetics or just like pure, you know,
it's like, I want to understand.
And if understanding just means it's a little bit simpler or, you know, more compressed
than I've gotten understanding.
I mean, it's just like a deep philosophical question, too, is like what it is.
Right.
Like, I mean, of course, any, you know, there's a reason we do informal mathematics rather than, like, writing out long informal strings of symbols of CFC or whatever.
Yeah, yeah, yeah.
Even though, you know, besides slang, it's just like somehow we're trying to put it in a way where we are actually getting some non-rigorous understanding out.
Yeah.
I mean, and even how that even relates to, like, why it helps.
And again, it's like, is it because of our inability to grind or is it because, like, if we are pushing towards something that's more compressive, it also hopefully also ascends on the, does it?
help understand other fields as well.
And it's just, it is somewhat magical that when we try to optimize for the both,
that it tends to coincide.
I don't know, actually, you know, if that's a fair, it's actually a quantitative statement,
but like why, you know, that is, is kind of magical.
It's at least sometimes true, yeah.
Yeah, exactly.
Or, you know, we certainly biased towards the cases of which it is.
That's why we call that a good theory to build.
But it's, yeah, it's very much, I think, you know,
it's like this is what we can study and understand,
and therefore it's also very convenient that it was rich
in kind of mathematical results there.
Yeah, I mean, okay, so I think there's two things that we can like, you know,
pull on, which is like what, given you've seated that,
or not seated, but you've believed that mathematical ability of AI
is going to continue advancing, it can do some of the things that you're, you know,
doing nowadays, how should, like, the mathematics community kind of best adapt and benefit from this?
I mean, just kind of saying, as somebody who's, you know, doesn't have the time to kind of
practice mathematics anymore, this is kind of very great because I can maybe dabble more.
There's a lot of results that can come out.
But I can also see where, you know, you've made the more precise point of we can not
motivate the right kind of behavior of understanding and development. And so we'd love to kind of
hear more about your views there. Yeah. So, okay, so first of all, like, it is clearly really
exciting, like, that they're increasingly capable models that are, like, producing high-quality
results, at least, you know, some high-quality results. Also, also a lot of thought. Yeah. Yeah, some good
stuff. And so, you know, as the models get really capable, you know, my hope is that they will
answer a lot of the questions that I've been, like, you know, kept up at night.
thinking about, I'm like, I'll get to learn the answers.
I'm not really exciting. And, you know, a lot of people got into math, like, largely because
they enjoyed learning math. Yeah. Right. Like, the first thing you do was a math student is you,
like, learn stuff that other people did. And you do that for, like, 20 years before you,
before you start doing, oh, maybe not quite 20 years, but, you know, 15 years.
If you're lucky, 20 years, yeah. If you're lucky, 20 years, yeah. Yeah. Yeah. Um, yeah. So, um, that said,
like the goal of mathematics is not to produce mathematics papers.
Like it's to produce some kind of understanding.
So maybe, okay, maybe that some of that understanding resides in model weights or something.
To me, that's, like, pretty unsatisfying.
Like, my own, you know, my motivation for doing mathematics is, like,
I would like to satisfy my own personal curiosity.
I think people should be, have the capability to do that.
And, like, that requires a pretty substantial apparatus.
Like, it's, like, simply not the case that you,
you can like study the questions I think are fundamental unless you've invested a huge amount
of time and effort kind of getting to the point where you can meaningfully do so.
And then moreover like that, you know, even the people, like, you know, you have the small
group of people doing like fact that you research math on the frontier or whatever, that relies
on like a huge apparatus of like, you know, thousands and millions or billions of people who
are trying to learn to think mathematically.
Like you need an entire mathematical community to support a small group of people who are
on the frontier.
Like you just, without the pipeline,
then the pipeline doesn't exist.
So if you think that's important,
like development of human capital
that can like meaningfully engage
with frontier mathematics,
I think that, you know,
you still have to incentivize those people
to actually invest their time and effort
getting to the point where they can engage
and then do so in a high-quality, meaningful way.
So right now, I think, like,
the existing incentive structures
for math research do not,
do not encourage people to do that.
So, you know, right now, if you're like a postdoc on the market, you want to get a job,
maybe for the next couple years before the community adapts.
The best way to do that is, like, you want to produce a lot of papers, which maybe prove, you know,
old conjectures or whatever, and you can do that by playing the slot machine until, you know,
the model produces a hopefully correct proof of such a result.
So you don't even have to pick the theorem in advance.
So, like, here's an experiment you can do.
You can take codex.
You can say, go online and find five recent conjectures in algebraic geometry and prove them.
And, okay, I've run this experiment and with some back and forth, I was able to, you know, in an hour, get like three, you know, quite bad papers, but correct papers.
Which, okay, are now sitting on my hard drive waiting for me to email the relevant people.
But, you know, this is not a big use of my time to invest in results.
Yeah.
But yeah, you definitely see people doing this.
So there's, you know, been a huge uptick in post-archive, mostly not very interesting.
Some of it is interesting.
But a lot of it is kind of clearly low quality, even if insofar as the, like, even if the result is something that would have been like highly rewarded a year ago, just like there's sort of no evidence that a human being is actually engaged with it.
Like there's no development of human capital or understanding.
Like sometimes, you know, we've seen examples where like three or four or five papers with the exact same proof of the exact same theorem out within a couple days of each other, which is clearly.
you know, some situation where someone's playing the slot machine.
Yeah.
Chat GPT is kind of consistently finding the same.
That's also interesting.
It's like not, it's kind of mode collapsed on like certain pads of reasoning.
Yeah.
Yeah.
And I think it's like not obvious that this problem goes away as the models get better.
Like maybe it does, but maybe it doesn't.
Like I think a lot of what I was saying earlier is that a lot of progress in mathematics
comes from like letting, you know, a thousand different flowers bloom and people pursue their own curiosity.
And then, you know, the boundaries of knowledge expand in some kind of fairly,
hopefully fairly uniform way and look at this really high dimensional space of mathematics.
And then like opportunistically you you suddenly get some applications or like as is told
questions we found. And it's like not clear to me that if if you know we kind of subordinate
mathematical exploration to what the model want to pursue like if what you're getting is like
one mathematician duplicated a thousand times or like actually you know a million different
mathematicians doing a million different things. Yeah. I think actually that's pretty I mean you know
aside from like, okay, math has to adapt by changing incentive structures,
I think this is actually a pretty maybe dangerous
or something that the labs have to pay attention to.
Because on one side, what's been successful for them
is that this emergent reasoning capabilities
obviously incredibly powerful.
And we've been, I mean, it's just created a lot of great PR headlines.
But to your point and also like where my interest tend to
is that like human mathematicians, they are,
coming from all sorts of weird intuitions that, like, you know, and that is why you end up
developing, you know, to your point, the frontier that is so diverse that you actually can
draw from and then be able to connect and produce a lot more. And so it, if it's true that most
of these like proofs that are being pushed out by the labs are converging on very similar
things, because they are technically drawing on the same body of literature, and that's where
they're strong at now. It's not clear that, I mean, one, you know, where are they developing their
intuitions, it's from practicing mathematicians. It's not clear what that other emergent,
you know, stronger, diverse intuition might be coming from. It might come, but I don't have a good
theory for where it comes from. And it's also just not clear that, you know, just with like
test time compute and various post-training scaling that you can even, I mean, it's a huge debate,
whether you could introduce new capabilities. And it's like precisely, I think you should be studied
in the math context because like where it's good at and where it fails and where human mathematicians
are good is exactly a very kind of like precise question to study it on. So there's like a long
way of saying that I, it isn't clear to me that we'll maybe produce, if we don't incentivize
enough people to interact with it, maybe unlike other domains, it might not actually continue
producing a superior result without a human aid. Yeah. Let me actually push a bit further even. So like
Let's suppose the models become really robustly superhuman, like, even, like, we're not even adding, like, meaningful cognitive diversity.
Okay.
I claim, like, still, actually, we'd be still one.
Okay, great.
Human mathematicians.
So, right, so why?
So, like, it's because the, you know, there's like a, it's like, there's a question on how we've, we're designing society, right?
Like, maybe the optimal situation is, like, you have the models doing all sorts of math research.
And, like, that leads to very, you know, in, you know, and this, like, highly, you know, you know, whatever, non-ununum.
a form diverse way.
Like, maybe you don't need humans to add that kind of diversity.
So, like, maybe that's the optimal thing to do.
And, like, in the end, that leads to lots of applications and lots of improved understanding
and so on.
It's like, despite something being optimal doesn't mean you do it.
Right?
Like, there's no reason to think that, you know, if we hand over control of whatever,
math research to the models and just let them do their own thing, that it will do the optimal
thing.
And in fact, like, if you, you know, if we are, if we kind of instrumentalize what we want
them to do, like we want to say, like, oh, you know, make our life better or whatever, it might not
be the case, like, that, like, what they decide to do that is, like, is, is pursue a wide variety
of interesting research, right? Like, they may just try to take the direct path. We don't know what's
going to happen. So, I don't know, if you believe that there is any value in sort of this broad-based,
like, fundamental research, which I do, look, I think that's one of the most valuable things,
humans or whatever, or the models can be doing. Like, I think the easiest way to guarantee it happens is, like,
to keep a community with like broad interests
where like pushing the models to do
and like helping us to design a society.
Like that's what we're pushing for.
Like I think at least, you know,
my hopeful vision for the future is that like humans are not like totally
disempowered.
Like we have control over where we're going.
And like if that's the case, like what we end up doing
is going to be driven by human interests.
And so you want to have people who have lots of different interests
and also the capabilities to actually pursue them.
Like you want people who are like smart and engaged and like,
you know, well,
can do not just like mathematical thinking, but all sorts of thinking.
Yeah, yeah.
I mean, yeah, a great fear is like as AI advances,
we don't develop the right ergonomic kind of interfaces
to actually encourage us to also be good,
continue to be good thinkers.
And it's so easy to kind of relinquish that
because you just off-land.
And the models aren't even good at that level of, like, you know,
a thinking where it's like the top-level structure,
but despite that, it's so easy to.
And so especially for math, I mean, just like,
you know, another selfish reason is if you kind of
Math Max, you know, I think it's actually a great pedagogical excuse to actually get just really
rigorous at thinking about various things. I mean, this is not why mathematicians do it, but as just
somebody... It's part of what we do. Okay, great, because it was kind of, you know, I thought it just
really helped give me a very good framework to think about many things, not just mathematics. And as
somebody who's also, you know, now a parent of a two-year-old, I kind of think about this a lot, too.
You know, it's not about kind of grinding, even though whatever, it's still good, but like, not to, not the shit on grinding too much.
But, you know, just hearing, I used to collaborate a bunch with some Hungarian mathematicians, Balazeghdi among them.
And I just heard that in Budapest, they would just teach group theory when you're in primary school.
And I'm like, well, we should definitely do that.
We should continue doing that.
And now that AI is so good at, you know, somewhat good at explaining, but it's far more accessible, we should actually, you know,
know, probably proliferate that even more. And so maybe that helps with bring more people to the
frontier rather than, you know, just. Yeah, I mean, this is something I'm concerned about, right?
Of course, I mostly talk about math because that's like where I live. But like I think, you know,
one nice thing about thinking about this is that we're one of the first professions
to kind of, you know, have a significant impact of high quality models. Although I think
maybe we're one of the first
professions for it to happen to publicly
but you know my sense is that there are plenty
of third professions that are
coding 100% coding but also
just like I mean I think that you know
anything you do at a computer like probably a huge amount
is being done by the models at this point and then like
there's not having a public reckoning about it
but you know because
math capabilities are useful for the company
for the labs to talk about I think
it would be more publicly than everyone else
but yeah I mean I think
one reason to try to maintain
like human capital in this area is just like it's
a model for all professions like presumably
we still want people who are like meaningfully
engaging with the world and like experts and
you know have like talents and
trained skills and so on like
yeah so I was
it's actually very convenient that
the math profession is so inclined with education
here because I think we're also
seeing like you know some amount
of challenges
you know among
college and
and like secular education coming from AI2.
So as you said, I mean, it's also an amazing tool to learn.
Yeah.
You know, people have, have, I've heard people start talking about a bimodal distribution
in their classes where there's some people who are really like figuring out how to take advantage
of new tools and other people who are just like letting them do their homework and then bombing everything else.
Unfortunately, I don't think that adapts fast enough, but it's like we definitely want to be living in a world where we're producing better thinkers.
I think it's just, you know, even if we talk about just the pure kind of optimization game,
I think that's better for us.
So, but as, you know, just like a human being,
I'm like, that would be pretty inconvenient if we became worse thinkers just as AI
sends.
And it's too easy to let that happen.
So we should kind of be thinking hard on how to actually take advantage of this
and harness it for our own improvement as well.
Yeah, I think it's like it's sort of interesting to observe, like, the models
at their current level capabilities, let you do a lot more.
They let you do a lot of things you wouldn't have done more cheaply than,
you know, cheaply enough to do them now.
But it's not clear to me that, like, in many cases,
they're actually improving the quality of outputs.
Yeah.
And I think this is common.
Like, you have a new technology that's doing something a little bit worse
than was previously done, but much cheaper.
And so you get a lot of suddenly a lot of, like,
low quality outputs that are displacing previous high quality outputs.
But I think it's possible to use the tools in a way
that actually, like, improves, you know, the quality among,
along every dimension.
It just requires some thoughtfulness and some redesign of institutions.
should actually incentivize that.
Yeah.
Well, hopefully capitalism works there.
I do feel like that the most high value things do require people to use it effectively.
And right now the models are not good enough without like the human experts to actually
participate.
But to your point, there's a vast majority of maybe like more junior and entry level.
And if, you know, it's harder for those roles to adapt as well.
And so the thing that would be a mistake is to use the AI models in a way that doesn't, basically, you need to be ascending and using the models to deepen your understanding.
And it's so easy for human nature just to be lazy.
And you have to resist that because that is the moment that you will kind of lose, basically.
And so you kind of, you know, kind of forgive the very competitive language.
But it really is that like it's just so easy to kind of relinquish thinking to the models.
The models can't really think.
And so as things are ascending so fast,
it's like critical that you continue developing those facilities
and actually leverage it to improve those facilities
rather than relinquish, yeah.
Yeah.
Yeah, I mean, one thing I think has been nice about,
you know, sort of this vast increase in like semi-expert attention
or like model attention on math problems is like now there's been, you know,
okay, I've been complaining about slot papers or whatever.
Like, you know, people who are not, you know, producing really high quality stuff.
Often that's actually coming from professionals.
Like, it's not saying, like, you know, there are people who, you know, like within academic mathematics,
there are incentives to produce, like, a lot of stuff.
And that's, you know, that's one place to slop coming from.
So there's definitely also, like, stop coming from non-experts.
But that I kind of actually don't see as a net negative.
Like, okay, there's a lot of, now there's a lot of, like, documents on the internet.
One might have to come through to figure out if a problem has been solved or not.
But to me, it seems like just the fact that there's lots of people excited about math is like kind of a positive.
So that's like a nice thing.
I totally agree.
I know I get to talk to you.
And not even kind of a positive, obviously.
Yeah, yeah, no, exactly.
It's like suddenly there's a spotlight on it and I can nerd out about math more.
Yeah.
I was actually kind of curious you've had any comments on the, I guess this is another constructive result, but the elliptic curve of rank 30.
That just came out yesterday.
So if that's right.
Yeah, yeah.
We don't have any details about it yet.
I know, there's rumors.
Yeah, we don't know how it, I mean, so it was, it's due to, I guess,
Claude Fable, prompted by Levent-Alpoche and a collaborator whose name I unfortunately forget.
Maybe you can settle on it on it.
Yeah, I just know the Twitter handle.
Yeah, we can have.
So, yeah, I mean, so a lot of these nice recent results have come from Levent with unclear amounts of autonomy.
So my sense is that some of them are semi-autonomous,
doesn't fully autonomous.
Yeah, with this, we don't know anything about the methods.
So, yeah, this is a fun construction.
Without knowing about the methods, it's very hard to say how significant it is.
What I would say is that if you want to understand
kind of historically how such results have been proved,
or sorry, have been understood by the community,
they're like cool, but I wouldn't say they're like a big deal.
So like the typical place of result like this might go is like someone's website of records.
It's on an annals level result.
Yeah, it's not an animal.
But that said is cool.
And it's like, you know, it definitely, there were a few very, very, very, very talented
mathematicians who like these kinds of questions.
So know Malkies being maybe the most famous example.
So Alchies and Clagsbrun are the ones who kind of have been pushing this record for a while.
And they recently found a ranked 29 example.
One up.
Which was, you know, the previous record.
Yeah, yeah.
You know, so those are mathematicians.
And I think, you know, people like it.
And it's cool now that the models can.
do this sort of thing.
But yeah, it's very, you know, one thing I always say about a model result is, like,
you cannot evaluate it except in retrospect.
And like, this is also true of human mathematics.
Like, sometimes a problem we thought was really important or would require really deep,
new ideas, does not?
And sometimes it does.
And you can't really know in advance.
So it's always exciting when a problem like that gets solved.
But then, like, to figure out how significant is, you start to look and like, well,
Levent and Claude and I think Ava Howell, maybe is the third collaborator,
have not yet told us how they did it.
Have they been mostly more secretive?
I think they released some traces for stuff.
Yeah, so for this one, I think they haven't yet, unless I missed it.
Yeah, they, you know, Levent likes to tweet out his results.
But yeah, he has been, you know, slowly releasing some kind of PDF-writups too.
I think it's, you know, he's just having fun on the internet.
Yeah, yeah.
It reminds me of when you're saying, you know, you have to evaluate how.
the results came. So I had Mark Selke and Metab Swani on from Open AI recently, and they were saying how,
you know, what's kind of been the most charming or delightful is just that the proofs have been
relatively short. They're not like 200 page. And, you know, maybe corresponding to your grinding
point as well. But maybe this is kind of optimized for, in retrospect, it was picked that it was
short or maybe do you find that on average
stuff that you throw at
GPT, you know, Sol or Fable
tends to be shorter and more
legible to the human or versus
it might just go haywire and
just grind it out. Yeah, so
I mean, I think it is nice
that will sometimes produce short, clever proofs.
Yeah. So that's, of course, everyone likes
a short clever proof. I think my sense is
that the reason they're not producing
long, complicated proofs is that they cannot.
Just like the ability to check correctness is not yet there.
So, you know, I mean, even actually, you know, if you ask the models to produce short proof,
you can then often ask them to, you also just to ask, is that correct?
And they will often say no.
So, like, they're much more reliable than they were six months ago for sure.
But, like, it's still, you know, they will still sometimes produce things that are just wrong.
and they know their wrong.
Yes.
Yeah, yeah.
I think the problem with producing a very long thing is they might not know they're wrong.
And so what I wonder if, you know, presumably internally, Open AI and Anthropic have probably
solved a lot more problems than they've released.
And I imagine quite a few of them are they're just not sure if they're true.
So, for example, you know, with this recent list of 10 problems released by Open AI, those
were all formalized in Lean, which is, of course, a very good evidence that they're true.
I have no doubt that there were a lot more
that they could not formalize and leave
because the prerequisite options
have not been put into Massua yet, for example.
There were probably more that were longer,
which also makes it challenging to check.
So yeah, this is my guess.
And we do actually see very long
AI generated proofs on the archive.
So, for example,
someone recently posted
a claimed proof of resolution of singularities
and positive characteristic,
which was 800 AI-generated pages.
It's definitely,
I mean, I'm sorry, I haven't read it.
I haven't been an error, but there's no way it's correct.
Like, this will be a major result.
It's just not with the capacities of the current models
that you're reasonably well calibrated.
Yeah, yeah, yeah.
And it's definitely no human has read it.
Definitely the models are not able to check this kind of thing, yeah.
So, I mean, yeah, so this is my expectation
is that the reason it's producing short, clever things
is just like that's what we can check.
Yeah, no.
And you can get it to produce long things that are grinding or hard to check,
but then.
Yeah, yeah, yeah.
to get there. This is actually more what it says about capabilities, frontier capabilities,
rather than in a perhaps more negative way, rather than, hey, it's just so good at these clever.
No, I think that's totally fair. I mean, especially if you look at an adjacent domain like code,
right? Remember reading something that cursor put out about testing their long horizon,
you know, harness. In this case, they were trying to reproduce SQL light in Rust. And it was
just so telling, like, how far we are from, you know, and that seems. And that seems,
It seems like a very comparable task of like it just, it's very long.
You have to make sure there.
There's something to verify that it's a correct implementation by testing a suite of, you know, cases.
But, you know, you actually need a harness there in this case.
It's not just like the raw models.
And it takes a while.
And you can actually compare with different frontier versus, you know, non-frontier models,
who's the planner, et cetera, like differences in their capabilities as well.
I was thinking in practice, like, to elicit a long proof, you kind of need a harness.
And when you make a harness whose goal is to elicit a proof, I think it often decreases reliability because you're just trying to produce output.
So like chat chaddbd5.6 pro is like very, it really tries not to say wrong stuff, for example.
And although it happens and then you'll ask it like, oh, was that correct?
And it'll say no.
But, you know, when you're trying to get it to ID8, like you try to get it.
like you try to get it to be creative,
you try to get it out of this very, like, you know, rigorous rut
in order to, like, you know, actually get it somewhere.
And I think if your interest is in, I mean,
there are a lot of people who are, you know,
try to arbitrage the prestige mechanics of academic mathematics
and trying to elicit a lot of proofs,
which are not necessarily being checked.
And in order to do that,
I think you just decrease the reliability in order to get a lot of stuff.
So, you know, I think any harness that can elicit a 250-page paper
is probably not being very,
careful about what is producing.
And you're saying that it's decreasing, or it's not reliable just because the capability
isn't there yet. And so it's just kind of forcing a longer horizon task on it.
Yeah.
I mean, in practice, like how does, like, a human checking, like, can't check a 250 page paper
either. Like, like, you, you know, you can't reliably check it line by line.
What you try to do to understand it is you try to understand the overall global structure
of the argument and you, like, stress tested in various ways.
Like, would this argument apply something else that I know to be?
false, like, oh, like, does it work in the special case?
Blah, blah, blah.
And also the models seem not to be able to do that kind of, like, more, I don't know,
fuzzy, like unit testing of proof very well yet.
So, actually, like, one of my favorite tests for the models, which they haven't succeeded
at yet, is there's a paper, I won't name it, that came out a couple years ago that was
wrong.
And it was, like, very hard to find, like, in close enough area to myself that I, like,
immediately, like, it was claiming some big result.
like I immediately downloaded and like started reading it.
And it was like very hard to find the actual specific error.
But it was also very clear from the structure of the argument that it couldn't work.
So it's like, you know, me and a bunch of other arguments, well, he was proving something that was too strong to be true.
He took a little bit further.
Yeah.
So me and a bunch of other experts like immediately like realized it was wrong.
And like we, you know, we emailed the author and like we kind of went back and forth until someone figured out what the precise specific error was.
And so far the models seem not to have been able to do this.
Like the specific area is quite subtle,
but they're also not able to do this kind of overall kind of big picture check checking,
which is kind of like how people in practice check papers.
Yeah, yeah, which is kind of it mirrors, you know, why in code.
It's like it's so clear it's good at the syntax, but, you know, higher level architectural stuff.
Still very weak.
Maybe you'll ascend there, but probably need some harness help.
Who knows?
I mean, people have, you know, evolve.
opinions on how much the harness and the model have to co-evolve and which one is necessary,
but the next model requires less.
So I feel like in math, it would be very interesting to see how you, if you do any experiments
with the harness there and how that improves, because it is a marker of general reasoning.
Yeah.
Yeah, I mean, I do, you know, I have my own sort of bad little harness and codex.
Oh, yeah.
But, yeah, it's, but, you know, I personally, I do not enjoy auto-autonomous mathematics
very much. So I mostly do not use the harness. I mostly try to use it to help me understand stuff.
Okay. No, that's totally fair. Yeah, exactly. You don't want to automate your job away because that does involve you being in the loop to understand it, which is necessary to participate. Maybe to finish off. I'd love to, this could be just something you haven't thought about or actually thought a lot about. I think you also have a toddler, right? Yeah. Yeah, yeah. So how have you?
Yeah, a three-year-old. Great, great. So you're one year, more advanced.
and probably, you know, have more thoughts on this.
Like, how are you thinking about, is it her education or his education?
Her, yeah.
Or how are you thinking about her education in math and, you know, not to grind,
but, you know, really just like, you know, pass on the love of it
and how to how to react to AI.
Yeah, so she's three.
She's never used AI.
Good.
She is starting to add.
That's about as far as you are in math.
That's more far as much than my deal.
She can add single digit numbers, like, you know, by counting on her fingers.
and count up
so maybe 30 reliably
and 50 semi-reliably
so I'm very proud of it.
Yeah, I definitely
encourage that.
We talk about shapes and stuff.
Actually, a couple days ago,
I woke her up and she was like
hiding under the blankets.
And I was like,
oh, I'm doing some math.
Oh, I'm doing some math.
Oh, I think he tweeted about that.
That was like more.
Yeah, it was great.
So, you know, I think she has like some sense
that I like math and she's into it because of that.
Yeah, I mean,
I would say I don't know, you know, I think the world is probably going to look pretty different in, you know, 20 years or whenever she's kind of fully adult and doing her own thing. But, you know, I think a lot of what we educate people for is like pretty robust changes in the nature of the world. Like, I think the reason to learn math has always been like to think clearly and like better understand the world. And like presumably that's something you want to do even if there are sort of extremely capable AI's.
And this is also true.
You know, I personally like math a lot, but also the reason to, like, read a lot of books and do the humanities and so on.
Right.
So, I mean, I think, like, the actual values of the math profession and, like, education more broadly are, like, things we definitely try to want to try to preserve.
That I hope to instill in my daughter.
You know, how much institutions have to change to make sure that happens is maybe an open question, I think, a lot.
But, yeah, I mean, at least at a personal.
level, I'm definitely trying to, you know, convince my three-year-old that math is super cool.
Oh, yeah, 100%.
One of her first words was Icosahedron, so.
Was it what?
She has a, she, my parents gave her little icosahedron toy when she was one.
That was the first word.
She, yeah, not literally first word, but she learned the platonic solid squatterly.
Oh, very good, very good.
It was very fun.
Next stop group theory.
I mean, it's like very natural.
That's right, yeah.
I mean, it's actually, when I, when I teach her about a, uh, a, you know,
addition and subtraction, for example, we'll do it in the context of a general group.
Oh, very good.
Well, at least you give the motivation.
I think a lot of probably skip that part.
Yeah, yeah.
And, you know, maybe this is how math grad students can also, you know, focus on training the next,
much younger generation to use AI in service of actually getting better at math rather than
just fucking understanding.
Wonderful.
Okay, well, thank you so much, Daniel.
Thank you.
It was a lot of fun.
Yeah, yeah.
It was a lot of fun.
And yeah, I mean, I think there's going to be a lot more progress very soon.
I'd love to maybe catch up and chat again.
Sounds great.
Thanks for listening to this episode of the A16Z podcast.
If you like this episode, be sure to like, comment, subscribe, leave us a rating or review
and share it with your friends and family.
More more episodes, go to YouTube, Apple Podcast, and Spotify.
Follow us on X and A16Z and subscribe to our Substack at A16Z.com.
Thanks again for listening.
and I'll see you in the next episode.
As a reminder, the content here is for informational purposes only.
It should not be taken as legal business, tax, or investment advice,
or be used to evaluate any investment or security
and is not directed at any investors or potential investors in any A16Z fund.
Please note that A16Z and its affiliates may also maintain investments
in the companies discussed in this podcast.
For more details, including a link to our investments,
please see A16Z.com forward slash disclosures.
You know,
