a16z Podcast - OpenAI Researchers on the Future of Mathematical Reasoning
Episode Date: September 8, 2026a16z Infra Partner Lisha Li sits down with OpenAI mathematicians Mehtaab Sawhney and Mark Sellke to discuss how quickly AI’s mathematical capabilities are advancing, what recent results reveal about... model reasoning, and what happens when AI begins making progress on problems mathematicians have struggled with for decades.Mehtaab and Mark unpack several recent results from OpenAI’s models, including advances in sphere packing and the construction of a non-sofic group. They explain why the surprising part isn’t simply that models can search more possibilities or work longer than humans: in many cases, the reasoning traces look remarkably similar to the work of an expert mathematician, including choosing promising approaches, backtracking when they fail, and combining ideas from across the literature.They also explore what this means for mathematics itself: how the role of human taste and judgment may change, whether AI could produce far more mathematics than humans can absorb, and why models that accelerate discovery may also make sophisticated results easier to understand.Resources:Follow Lisha Li on X: https://x.com/lishali88Follow Mehtaab Sawhney on X: https://x.com/mehtaab_sawhneyFollow Mark Sellke on X: https://x.com/MarkSellke Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
Often as a practicing mathematician, you have an idea,
and then you kind of think it might work.
Then you try for a few hours, a few weeks,
and at some point you give up.
Whereas for GPT, like, okay, human told me to do this,
like, let's just do this.
And so that's why we're sort of in this renaissance
of, like, reachable results.
This is the best part about this problem,
which is really nobody had any idea.
Is the model just guessing in some insane way?
It doesn't seem like there's a limit so far,
but it doesn't have that context yet.
It'd be nice for the world if we'll find math not next one along.
faster. The ceiling for difficulty of a math problem is pretty high. Even if AI continues getting
exponentially better at math, plausible will never solve something like P versus N. What's the ideal
way that this is being taken up by the math community? Probably most at this point are like,
okay, AI is obviously doing some non-trivial stuff. So AI isn't just getting better at math
benchmarks. It's beginning to make progress on mathematical problems that have resisted humans for
decades. In this episode, A16Z-infra-partner, Leisha Lee, sits down with open-AI
mathematicians Metabswani and Mark Selke to understand what's actually changing. They walk
through recent results in sphere packing, coding theory and group theory, and explain why these
advances can't be reduced to brute force. The models try different approaches, abandoned dead
ends, connect ideas across fields, and in some cases produce reasoning that reads surprisingly
like the notes of a human mathematician.
They also tackled the bigger question.
What happens to mathematics
when proving a result becomes less of a bottle leg?
From mathematical taste and human judgment
to understanding an explosion of new results,
Leisha, Metab and Mark
explore how AI could change
not just what problems get solved,
but what it means to practice mathematics.
Well, thank you guys for coming.
This is really exciting
because I think math has been moving so fast.
with AI. We just love to get to both practicing mathematicians and who work at Open
AI to chat on some of these results. So we have with us Mark Selke and Natalzwani.
We're connected, actually, because Ufa was actually your advisor. So both of you guys have worked
much more deeply in math since I have quit many, many, like over a decade ago. So this is
very exciting to kind of hear your download of your thoughts on how open AI has been sort of
approaching this and also just like where you think math is going with the incredibly
rapid advance of how AI has been helping. We can start off with some very basic questions.
What do you do to the extent that you can of course share? And how did you come from being a
practicing mathematician to working at Open AI? Yeah, I mean, I guess we both broadly got
excited last year when the models started to really take off in math. So I joined a little bit
Before Metaub, I saw the IMO gold medal last summer, basically.
I thought, this is amazing.
I want to see what the heck they did.
Let me go see.
And then, yeah, I guess in the fall, Mark gave me a GPG5 account.
And then I started playing with the models and very quickly became convinced that, yeah,
it was extremely exciting to play with them.
And you two were collaborating before them.
Yeah, we've known each other for a while.
Yeah.
We have one paper we actually wrote jointly.
Yeah.
Yeah.
So GPG5 was your conversion?
Yeah, yeah.
What was the magic that sort of, what question do you throw at it?
What process?
Yeah, so I think, actually, yeah, so I think how this started was, at least for me, the starting moment was something like there's a collection of problems called.
So Paul Ardish is a very famous mathematician.
He posed a bunch of problems.
And so they've now all been collected on this site.
And so I specifically worked in combinatorics and a lot of these questions are among the most important.
So it's always fun to flick through this light.
But one thing that often happened to me that was extremely frustrating was,
I would look at a question, see that it's marked as open, and then not actually know if it's correct, not actually know if it was still unsolved, because the literature is often quite hard to search.
And one instance, I just plugged it into GPD5, and five minutes later, it found a reference.
And this was a case where a few of my friends actually started thinking about the poem on the site, I was talking with them.
And, I mean, we had spent a few hours.
It wasn't clear if the poem was within reach.
And it was very nice.
Okay, to be told, yes, this is in reach.
Here's how you do it.
And, yeah, GPD5 told me this.
And then I told Mark about this.
Yeah, this is sort of, yeah, this for me was quite a surprising moment.
Yeah.
And then we looked more into it and we found 10 more cases sort of like this.
At the time, I feel like being better at maybe making connections between, as you're saying,
like the search for whether there's been a result or a related thing earlier is just kind of humanely hard,
but maybe better for machine.
But I imagine as the progress has happened in the last year, what has been impressive,
kind of reach beyond that.
And maybe through talking about it more abstractly
or if it's more natural to talk about it
through one of the problems that has been recently announced
through, you know, Astra,
you can kind of enlighten me as to like
how the recent progress has been a lot more
than just searching through more areas,
making these connections between the field
and perhaps just actually deeper,
more mathematical reasoning.
That's a similar to a working mathematician.
Yeah, I mean, I think this like
search point of being,
familiar with everything is still definitely like a relative strength that maybe informs, like,
the types of problems that AI is solving now. I think there are some other relative strengths and
weaknesses. Another relative strength that's pretty noticeable is just, like, it's very good at
executing on some, like, idea once it has it. Whenever you have an idea, there's usually some amount
of getting everything lined up. Is epsilon, like, smaller than delta, this kind of thing? You have to
get everything correct. And like for a human, it's easy to get lost in these kinds of details.
And the AIs just kind of always nail these kinds of arguments I find. Yeah. I feel like you guys
will know more detail on this. But like for the unit distance problem, it was just like the approach,
there was definitely contributions from open AI. But like the approach perhaps was suggested even
originally by Airdush. And then it's just that the actual reasoning was a very, very like,
momentous feat. And so for human, you're like, well, I only have a limited amount of time. And
if after so many steps, it is still not clear. I mean, maybe you're like Andrew Wiles and you
actually spend 10 years alone and do something, but like it's not clear that the risk reward
is not good enough. Whereas for GPT, okay, like a human told me to do this, let's just do this.
And so that's why we're sort of in this renaissance of like reachable results. Does that track?
And did you feel like with the astro results, is that sort of like where the
strengths have been primarily, or there's an extra ingredient or magic here.
I feel, I think the unit distance example is about it's quite telling.
In the sense of maybe the exact construction, you can make it look very similar to what
people had tried before.
But I think, I mean, often as a practicing mathematician, you have an idea, and then you
kind of think it might work, then you try for a few hours, a few days, a few weeks, and
at some point you give up.
And then a not so uncommon experience is that you find out a, a you find out a, a you
year or two later that somebody else got the idea to work that you thought that didn't work.
So somehow getting an idea to work can even be a large portion of the battle.
And I think, especially in the case of the unit distance conjunctured, there's just a lot of
extraordinarily finicky details.
And very often when you're doing mathematics, you're kind of gambling against the problem.
You're like, maybe I should try this approach, but it seems really unlikely and just not
worth my time.
And the model, I think in several of these cases, both by combining what it knew and sort of
having good taste, kind of made the correct balance.
And you can kind of see this in the summarized chain of thought we released.
You can sort of look at it.
It's reasoning like a mathematician, and because it knows a few very correct fits,
it makes the right decisions and eventually able to prune the search tree.
It's not really trying everything.
It tries a lot of different things.
It's extremely dogged.
I mean, it can't try every idea.
It has to try a limited set of ideas.
And it's able to kind of use its knowledge plus good mathematical judgment
and find the right path to go along.
So, I mean, that for me was like, because this was a problem
which a lot of people have thought about.
And I mean, the fact that the idea is not so foreign
probably indicates that a lot of people have tried it,
or at least a few very serious mathematicians have tried.
And I think that's what made it really interesting to see.
I think something else that I feel when I see these proofs is like,
if I have an idea and I'm trying to execute,
it might be that I have some wrong plan for like how to get things to work.
And as a human, if you have some wrong path you go down for a while,
it can be hard to like rewire your brain to start over
and try a different path.
The initial idea is kind of linked in your brain
with these other things that ended up not working.
It's sort of like your context window is like a little polluted
and you can't just make another clone of yourself
from last week and say,
don't do this, try something else,
build your intuition in another direction.
But it's very easy to do this with an AI.
So I think this is another reason that it's like
getting the details right
once you have some good general direction
is much less of a barrier all of a sudden.
And when you say it's much easier to do with AI,
It's like it's not actually being directed with human interference, too.
It's, as you were saying, in the reasoning traces, it's like making these choices.
Maybe it backtracks, but then it's able to not be distracted by maybe the context in which it's thinking about the problem, like, by these machinery.
And it's like, do you see it kind of go back as well?
Or is it just like good choices?
Like, is it a lucky sample or is it actually reasoning like a mathematician or, okay, it doesn't do well in this path, but it goes back, but then it doesn't let that pollute?
No, I mean, it definitely makes mistakes and then it goes back and thinks about it.
I think it's somehow very calculating, very correct.
I mean, as human mathematicians, you're not always perfect in making these decisions.
The first time something doesn't work, you automatically kind of downgrade how likely this approach is to work,
and you keep doing this a few times.
The model somehow is much better able to, like, it seems, for several of the solutions we've seen,
somehow it seems much better able to update, like, how likely the path is to work,
like, versus rejecting a path versus a human doing it.
But I think even if it weren't, the fact that you could just start another model session
over.
It's always going to be the case that...
Got it, got it.
So in some sense, it is, like, still leveraging the fact that you could, like, run
kind of parallel, you know, agents on the problem.
But if it were kind of backtracking, then does make it seem much more like a human,
you know, mathematician.
And perhaps it is kind of doing some of that stuff, too.
Because, like, obviously, like, we have to, like, make mistakes in order to, like,
even gain intuition for, like, why that solution space is, like, not, you know, not in the set
of paths that it could be in.
I mean, I think this kind of thing happens with humans, too, where, like, if you get stuck on some approach, you might tell another human your kind of general idea, and then they'll come back and, like, figure out how to get it to work.
And, you know, it's just, it takes more time to do this, like, with humans.
I wonder, I mean, you know, maybe this gets to extend that you can actually talk about sort of like, obviously don't talk about the training recipes or whatever.
But, like, it's interesting that if you're just studying, for instance, from math papers, it's like a very poor training set, like a priority for math, because.
I mean, maybe math textbooks are even a pure example of this.
It's, like, really bad at actually reconstructing the motivation for why things were.
It's like, I mean, maybe some people like it, but don't learn real analysis from Roodin.
It's just like it's very clean already and crisp.
And I think that's bad because it doesn't show the struggle that made us formulate definitions in a certain way.
Like, why do we even need to have real numbers be defined in this, like, super abstract way?
et cetera. And so, you know, I think papers also, I mean, unless you're, most people don't write papers with the, the
context of I need to educate somebody to be a mathematician. And so like the actual maybe curriculum
of like learning math is not inherent in like a lot of our artifacts as mathematicians. So maybe
another way to ask this question is if the reasoning tracers are actually producing things, I was like,
okay, this is actually more close to mathematical thought. Like how does that arise?
I mean, yeah, I guess Open AI has been like the pioneer of reasoning models and, you know, teaching AI to reason in this way.
So, you know, we're doing a lot of work at kind of kind of all possible directions on, you know, teaching models to reason better and for longer and, you know, in all kinds of different domains.
Yeah, I mean, I think we're training general purpose reasoning models and kind of one.
a lot of these behaviors that we're describing mathematically,
like backtracking or kind of starting again,
I mean, these are not really specific to mathematics.
I mean, we're seeing them specifically in mathematics in these examples,
but kind of their general purpose tools for reasoning.
And I think if you work hard at reasoning,
you should see these patterns eventually.
So it's a submergent because it's,
I mean, I do think that's why the opening eye approach was so,
I mean, it's like it doesn't rely on, you know,
doing auto-formalization in all.
order to like guide the reasoning. I think that's like obviously more like us. But it's just like
so not obvious that if you're just like training on say a corpus of like math proofs, maybe
auto-formalized and lean, that you get, get the sort of like projection of like how to think well.
Like put another way like with maybe maybe if we think about it with code. Like code is such a good
corpus to train on because it's one of the few datasets that has such large context. You just like,
I mean, maybe you see this kind of with books,
but they're less structurally interconnected.
There's just less structure there, I think.
It's safe to say, like on average in a book
compared to like a piece of code.
And so, like, with math papers, I feel like,
but maybe what we're still bad at with coding models
is stuff that that data set doesn't contain,
which is like kind of the semantics.
Like the syntax is there,
but there's a little bit of like the higher level
semantics of what produced, like, why do I have to write it this way? It's not, I'm kind of getting
too much in the philosophical good. It is just, like, really interesting how it's still emergent
that it's doing good mathematics. And we'll probably get into this more detail if you guys,
you know, wanted to talk in more detail about some of the problems, which is just like it's,
it's not just doing like the expected, like, we'll push the brute force thing. Like, you clearly
are impressed with some of the reasoning traces. And it's just not obvious that's gleaned from,
you know, what we would imagine would be the easy training set here.
Yeah, absolutely. I mean, I think this kind of thing is one reason we decided it was important to release like these summarized chains of thought for these kinds of results. Because if you, if you've never seen these and you just see all these proofs coming out, you're kind of, you're not sure what it means. Like, is the model just guessing in some insane way? Like, is it thinking in some totally foreign, like what's going on? But actually, it's reasoning kind of shockingly like an expert human would.
Yeah, yeah.
Yeah, it's very much like reading colleagues' like notes.
I mean, it's a little more disorganized in some way,
they're kind of, like, especially you work close enough with the collaborative,
sometimes you'll just see them like spill out their thoughts and an email to you.
And it kind of, it feels like reading a lot of those change together.
So it's, yeah, it's very, it's quite surprising the first few times.
Were you two sort of very involved in choosing the problems to release in this, like,
last 10 problem set?
that Astros applied to?
Which was your favorite?
Yeah, we have definitely involved.
Do you want to start?
Yeah, I mean,
spearpacking maybe?
Yeah, I guess, yeah.
So I guess my personal favorite among these problems is the following.
It's extremely simple question,
which is just like,
it's just about how efficiently can you put a bunch.
My circles are not very good,
and they're not all the same size.
But we're assuming they are.
Yeah, so the question is just like,
how dense can you place a bunch of,
So you have a bunch of spheres.
You have a bunch of spheres of radius 1 and D dimensions.
So the question is, how densely can they pack?
And so, yeah, so in two dimensions, it's kind of like the, so, so D equals 1, this is not an interesting question, kind of, it's just the real line.
Yeah, you can cut it up in a sphere in dimension 1 is just a unit segment, so, okay, you can cover everything.
So in D equals 2, it's kind of the picture that you know that everybody loves.
It's just like a bunch of spheres which sort of form like a hexagonal lattice.
Hopefully I've drawn it well enough that I can draw the hexagon.
Kind of betraying my naivete on this problem.
Is that like obvious?
It's like a very elegant proof that it's a regular lattice?
Yeah, it's not so obvious that this should work.
It was only proven in the 60s, I think.
There's a short argument, but it's not so easy.
Yeah.
Where's the intuition?
Like, what is kind of like the machinery of the argument?
I mean, it definitely looks like it should work.
That's why I'm going to go.
It's so good.
Yeah, I think this is the best part about this problem,
which is really nobody has any idea of this problem.
Yeah.
So, yeah, so, I mean, honestly, the best intuition that I have for this is that, like,
bees do this.
And if there was a more efficient way, then probably bees would pack potty comb some other way.
Because evolution is efficient.
Yeah.
I think beyond that, like, I don't have a great argument.
I mean, and I think how little we know is demonstrated by the fact, so, okay, D equals three, the answer is just like, it's how you pack like oranges in a grocery store.
And this was only, this was proved by Hales sometime in the 2000s.
And like, and we don't have a short proof of this.
Like, I think the shortest proof is like a few hundred pages.
What area does it, like, draw from in order to it?
So it's a lot of linear programming arguments and it's very delicate, like, genre.
It's quite ugly, actually.
Yeah, it's like...
This is like a famously ugly argument.
Oh, no.
And then the two most famous results are D equals eight and 24.
Eight and 24.
It must be some weird subspace thing.
Yeah, exactly.
So this was done in some...
Like gluing something.
Yeah, so this was done in 2017.
Sounds slightly prettier, though.
Yeah, so the reason it works out in these two very special dimensions is that,
so this is called a lattice.
packing, so it's like kind of very regular. And it turns out in these two dimensions,
there are two very special lattices. They're called the E8 and leach lattice. And they're very nice
and they're like unusually dense. Like kind of, they're just very, very pretty structures coming
from other areas of math. And it turns out that they're the optimal structures. But they're still
like regular. Yeah, they're very regular. But I mean, beyond this, so we don't know any more
exact dimensions. We know these five dimensions and we kind of don't know anything else.
And, I mean, like, to give an indication of how little we know,
so there are two very surprising things about this.
So you can define, like, delta D to be, like,
the densest sphere packing in D dimensions.
So there's kind of an easy lower bound of, like,
two to the minus D.
Basic, yeah, this is not so hard to show.
Basically, any packing where you can't put in another sphere
has to have this density.
So, okay, it's not ridiculous.
small. And we know that it has to decay exponentially. So it has to decay, like, it grows like
one minus C for some, at least for some constant. So in large dimensions, you can only cover like
a vanishingly small portion. But we know like basically nothing else. And so for-
That's just because like the high dimensional sphere like thing where they occupies, um, it's just like,
yeah, the volume behavior is weird. Yeah. So basically, I mean, basically they don't want to touch next
I don't know, I don't think there's a particularly short way to see that it's like exponentially small, but it's known to be exponentially small.
And for a long, long time, the best bound was something like this funny number like two to the minus point 599 D.
And this was proved by two mathematicians in the 70s.
Kapitiansky.
Okay.
That's a weird number.
Where's that spit out of it from?
It's like commentatorial.
It's it.
It's the answer to some extremely ugly optimization problem.
There's like a nice underlying strategy.
Okay, that's nice.
Yeah, yeah.
I'll say one last thing about this, yeah.
So, yeah, I mean, these were these two Russian mathematicians in the 70s.
It's thinking very hard to find their paper.
Like one page, it's like two pages long.
Yeah, they don't write very many details because paper was far.
But, yeah.
And the two negative Ds, there's just like the square lattice, like the dumb one?
So yeah, it's actually not so easy to...
So the argument for this is as follows.
Basically, imagine that you have a set of spheres,
and you can construct a set of spheres
so that, like, you can't put down another sphere.
Because if you could put down an extrasphere,
you just keep putting it down.
So you have a set of sphere
so that there's no other sphere,
which you can put down.
That's here?
It's like almost like a...
Just, yeah, just take any such packing.
Okay, yeah, yeah.
And I claim that this has to cover
at least two to the minus D.
The reason is that, like, if you blew up each of these spheres by a factor of two,
then, like, they have to cover every point in space.
And the reason is otherwise you could put down, if there was any empty space, you could put down
a sphere there.
So I guess if you take the usual lettuce, there are actually, like, more places you can put
things kind of diagonally.
Okay, yeah, yeah, so that's actually not.
It's like a worst bound.
Yeah, you can just keep plopping things in.
Yeah, this is like more, okay.
Yeah, this is related to this, like, really funny fact where you, like, put a sphere,
on every point.
If you take a cube in high dimensions,
you put a sphere on every point.
It's like vanishingly small.
It's so small that you can put
another sphere in the middle.
Yeah, yeah.
And it fits.
Another high dimensional sphere behavior.
Yeah, it's very weird.
And so, okay.
So the great part is that
so the model shows the following.
So I'll write two things.
So this is Astra, I guess.
Probably the right way to refer to this.
So, okay, I'm going to be.
to write something called the LP bound.
I'll explain this in a second,
and it shows that it's smaller than
this very nice
number. Do you want to say equals?
Yeah, it's equals, actually.
D to the
2 pi
was little 1 to the
D. And if you can work out what this
number is, it's like,
it's like roughly something like
2 to the minus 0.6.
Close about.
Yeah, it's surprising.
one, you know, you're like, oh, maybe there's some like nicer kind of like structure
there that fell out.
And it's a most, this is like roughly something like two to the minus point six zero one dot dot dot
dot D.
That's the numerics.
I thought it was six or four.
Great.
Yeah, this shows mine.
Yeah.
So, okay.
So there are a couple of things.
So first, what is this LP of D?
So Viazazzo's work actually builds on some earlier work.
It turns out that there's a way to attack sphere packing via what's called.
a linear programming bound.
So LP just stands for linear programming.
So Conan Elke's
gave an approach for sphere packing
based on linear programming.
So it's like a linear optimization problem
over a convex set,
but it's all kind of infinite dimensional here.
And basically what this reduces down to
is you try to understand the follow.
So what you try to show is you,
basically you construct a function F.
So this is in D dimensions and it's mapping to R,
and it has the following properties.
So first, F of X, this is a function in D dimension,
so it's always less than zero if the size of X is bigger than one.
And you second have that the Fourier transform of X,
this is always non-negative.
So this is just, this is a linear program
because the Fourier transform is a linear operator.
and night race.
So you're taking just some arbitrary F that satisfies this property?
Yeah, so you can take any F that satisfies these properties.
And what they prove is that Delta D is bounded by the ratio of the Fourier Transformer
at zero to its to the the for a transformer at zero,
an F of zero to over the Fourier Transformer of zero times the volume of the ball of radio.
just one half in D dimensions.
And so, okay, this proof is not so short
for experience about petition.
It's like it's half a paragraph to prove it,
but it's a little bit tricky.
And the point is, so it turns out,
so this is a relaxation problem.
There's no guarantee that taking the optimal left
will give you a good bound on Delta D.
But, so what Vazasasca did,
and this was sort of the key,
I mean, a large part of the reasons you want
of fuels metal, and
in 2020 was that, or in 2022,
was that she constructed a function in 8 in 24 dimensions
such that this upper bound matches exactly these two very special lives.
And these are kind of miracles of nature
that both you can construct this function
and that it gives you the optimal bound.
But you can just, this is a very, very natural problem.
It's a function with two very simple properties
and you just want to understand how this behaves
in for large dimensions D.
And that was a big mystery.
There was a numerics paper by Cohn and several others,
which conjectured that just based on doing numerics,
that this was the answer,
but they had no idea why this would be the answer.
And what the model shows is that actually,
the linear programming bound in large dimensions
had this extremely nice asymptotic behavior.
And the proof kind of explains where this is called from.
And because you understand this LP bound perfectly,
this actually just gives a better bound on delta D.
It turns out that this old bound
can be kind of reinterpreted in this framework.
And what the model does is it shows you the best possible bound
you can get by this framework.
So the model sort of made the connection.
And what is the sort of like, what do we?
So I think, so the model gives a function F, which,
so first it constructs a function F, which gives you this down.
And then it shows that there's no function F
which does any better.
So it's inequality,
which is quite strong.
So, like, sort of,
we now understand this problem
in high dimensions very well.
And that's pretty remarkable.
And the model was just kind of told,
like, analyze this linear program in high dimensions,
you know, go have fun.
Got it.
And to give an indication of, like,
how it was known,
I think this conjecture was based basically only by doing numerics.
Extremely clever numerics, but numeric.
And so, yeah,
You have to kind of figure out why this is the right thing to aim for, and it does.
And that was pretty remarkable.
Yeah, I mean, I had actually thought about this problem for about six months at some point when I was a graduate student.
And yeah, I remember making like absolutely zero progress on it.
So it was very nice to be like explained why it was, yeah, why it was true.
So that was a positive experience.
I think also in general it was one of these solutions which I knew several people who had tried the problem.
It's pretty remarkable because like the model solution, especially for the,
this being like, the LP can't do better than this was like quite short. It's a few pages of
like complex analysis, but it's kind of exactly the right approach. Like once you see it,
it's kind of, it's like unbelievable. Like why, why hasn't somebody done this before? It was,
it was like, there are many types of good mathematics, but I think one of them is just like,
you see it and you're like, oh man, why didn't I think of this? And it was really fun.
And it was, I mean, I sort of knew why I didn't think of it, but it was quite nice to see it.
And it was fun to see. That's why I like this problem a lot.
Yeah. So this is.
is the first of the 10 problems that Astrosol.
But the second is actually closely related.
So this was sphere packing.
The second one is spherical and binary codes.
So what's like you draw a code?
You drew a packing.
It's going to be the same picture.
Okay, sure.
Otherwise, we're going to have his picture.
The hexagonal packing in our minds.
Yeah, a spherical code is literally just a sphere packing,
but on another sphere.
Yeah, so I mean, a spherical code is basically just a sphere packing on the surface of another sphere.
So, yeah, sure.
Okay, so same picture as before, except you're kind of on a curved surface.
Okay, okay, okay.
Okay, so why is it called a code?
Well, you can, I guess the reason is because of binary codes, which is, again, the same sort of thing,
but now it's on a cube.
Okay, yeah, fine.
Let me draw a picture of a cube
and some, like, simplest possible code on it.
So, like, when you're, like, sending...
So this is really, like, about error-correcting codes.
So, so what are error-correcting codes?
So, you know, it's like, I send you some string of bits, right?
and maybe I'm worried that some of the bits I send you get corrupted, right?
So maybe like just because of some errors in my system, like this one gets changed.
And we want some communication protocol so that like you can decode this like small amount of error
and like recover what I was trying to tell you.
And you know, like normal English language kind of has this sort of property, right?
if I make a few typos, you're going to be able to understand what I'm saying.
But if we have some really brittle communication scheme, it's not going to work.
So codes are kind of the way you solve this.
And mathematically, it just means, like, you know,
what's a binary string like this is a fixed length?
It's like a point on some hypercube.
And we want a dictionary of allowable code words that are, like, separated from each other.
So, like, in this case, if I don't want any two to be adjacent,
I would kind of take these four vertices,
kind of the like even ones,
if you sum up the digits, right?
And like, okay, I guess,
okay, in this case, I guess if I have an error,
you can't tell which one it's from,
but at least you can tell it's like not,
at least you can tell there was an error.
Oh, I see, I see.
Yeah, yeah, because it's like kind of sparse in the,
yeah, it's like it's not too adjacent
so that like, is like when the hand you're just like,
yeah, yeah, yeah, right, right.
Yeah, so you want like a large hamming distance between any distinct in your dictionary.
And yeah, I guess if you take two opposite corners, then if I have like a single bit error,
I can always like recover which point it was coming from.
One that it's definitely closest to.
Yeah.
So there's kind of the, you know, same question in both of these cases, like in a very high dimensional
setting, what kind of rate can you get?
And like for binary codes, it's really like, you know, an extremely practical question.
it's sort of like if I send you
like an N-bit string
and there's like you know
1% error rate
like how much longer does my message have to become
to tolerate that amount of errors
right this is like some fundamental
information theoretic limit
of like like you know
communication
and
you know but you can see like
certainly this
this spherical case
is like it looks very much like
sphere packing for example if you like
If you make all these little spheres really small,
then the curvature of the big sphere is kind of not going to matter so much,
and it looks like just packing spheres in full space.
And in fact, yeah, like these problems turned out to be very related.
So for these problems, there were similar bounds coming from the KL authors,
and like there's like something for the sphere and something for the cube,
but it's all kind of the same stuff.
And our models found better bounds for these cases as well.
And, like, it, I mean, the techniques look pretty different, actually,
if you, if you, like, write them out.
So this full-space analysis of this linear programming was using, like,
just complex analysis.
But if you, like, the method for these cases,
we're using representation theory.
Like both the sphere and the cube have a lot of symmetry.
And basically the idea of the proof was to really leverage this symmetry.
Like there's some amount of this in the previous existing method,
and really the improvement is to lean into the representation theory really hard
and kind of make the algebraic symmetry, like, in turn, a more sophisticated way.
And then it, like, turns out that from the representation theory formulas,
as if you kind of take this like, like small sphere limit in the spherical code case,
you recover like part of this result and you recover this value.
So like this result isn't a special case.
You kind of only went one direction of the bound from looking at it from the code's point
of view, but like there's like a very close connection.
Okay.
Yeah.
So you guys let this like run in parallel.
So it's like kind of discovering or because you're not sort of like feeding it.
So actually this was the one case.
where there was some interactivity involved.
Oh, interesting.
So for all of, so except for this pair,
it was just, you know, we had some problems.
We fed them in and we, you know,
the model came back with some solutions.
What happened here is actually pretty interesting.
So we first asked it to improve the bounds for the codes.
And it came back with an improvement that like used some amount of representation theory.
And then we kind of asked it,
can you like push this further?
Like, you know, what happens?
And then it came back with some like much more sophisticated representation theory.
And like it turned out that you got this conjectured value for full space sphere packing
like out of that method by pushing it as far as it can go.
So then we kind of asked to directly analyze the sky and try to complete the picture.
Okay.
Yeah.
So the relationship like isn't a coincidence.
Yeah, yeah.
It's like interesting when you're saying the first prompt, which is, you know, maybe so basic,
which is like, can you push this further? It does require some judgment for mathematicians,
but like eventually you would imagine by scaling the models, you don't need to do that,
or there's another view that the harness actually does matter, and this is kind of part of the harness apparatus.
Do you guys have any views on that with your working with Astra, especially generations of models
and how much do you have to kind of input or how much the harness matters versus not?
I mean, yeah, I guess there have been some, like, funny quirks like this that just come from, like, exactly what you asked the model to do, basically.
Yeah.
Like, in this case, what the model was asked to do originally for codes was to improve the bounds by, like, some exponential factor.
So it really, like, shows up in this, like, leading constant up here.
And, you know, it improved the bounds, and it didn't try to push things too much further.
Like, sometimes you see it do, but sometimes it just doesn't bother.
But yeah, you know, you just ask it again and it goes further.
So it wasn't like a capabilities issue.
It just didn't feel like it at the time.
Do you call that judgment or like what is the?
Because like there is a, yeah, what do you call that?
Models tend to be pretty task oriented.
If you tell it to do a task, it accomplishes that.
It's pretty happy.
So yeah, the task oriented in this, it's like,
but do we expect that level to kind of ascend up to?
It's not that they will be less good at,
being task-oriented, it's like, they'll ascend to the level of like, okay, no, let's go in this
direction. You'll have the judgment, too, because you guys have the judgment, too, like, okay,
this is pretty promising. It looks like you're using a lot of representation theory. It doesn't
seem like there's a limit so far, but it doesn't have that context yet. But, like, I guess what
I'm trying to say is, like, this one, it's hard to, maybe harder to extrapolate, but from, like,
previous generations, when you had to give it more, maybe prompting, more of that harness work,
but eventually probably had to give it less. So it probably gives you some confidence that
there's this like really fast ascension.
And do you see, yeah, like, where?
Somehow solving a harder math problem is, like, you have to solve many smaller,
like, somewhat less hard math problems.
And the fact that the math problems are getting harder is kind of an indication that
the model is able to take on more and more work in like a single continuous unit.
And I think that's the thing that looks very promising.
Somehow, like any of these solutions, it's not like one idea.
And then you're kind of home free.
You need several, you need several pieces to kind of,
interact and talk to each other.
The model doesn't come up with all the ideas at once, right?
It doesn't pull everything out in an instance.
So kind of the fact that it needs to sort of see how this piece interacts with another
piece that's kind of like solving a problem in itself or piecing together many problems
in itself.
It could just be that, okay, when you're telling it, okay, push this even further, that was
of the same order of like magnitude as like all the smaller things it's solving as well in
between and so you don't think that this kind of like a privileged direction. It's just sort of like,
hey, let's give it like one more help. Or you actually think that there's, I guess what I'm trying
to get out of a bigger question is like, is there a good sense of like, you know, taste? Because like
when people talk about, for instance, how well the models are getting at like doing research,
for instance, and that's what we want a little bit of RSI. And sort of like, there's surprising
things about how that improves. And then there's like the, oh, you know, maybe right now it's at a level
still like a junior researcher. It's like not really asking like the right problems. And so I'm just
trying to get like a maybe a sense of like where you're seeing that progress through the model
advancements each generation. I mean, what is taste even? Yeah, I think I tend to be pretty
utilitarian in my view of taste. If you're able to solve problems faster by making better judgments,
I think that's like the best like general proxy I have for a taste and somehow the fact that it's solving harder problems means it has kind of by definition means it has better taste.
I think there are these no yeah I think occasion because they are task oriented you do occasionally get these these symptoms of like oh it clearly has made a breakthrough it kind of understands it's made a breakthrough and then it doesn't kind of push all the way to the limit because that's not what you asked but that seems yeah that seems seems rather minor.
compared to the state of progress we've seen so far.
Okay, yeah.
I think that's pretty clear.
Like, I think it's like maybe you're liable to get confused
if you're trying to like do a concrete long horizon task
and show taste kind of at the same time.
But like, you know, if you have like one model that's responsible for taste
and one model that's responsible for going out and like, you know,
working for a long time at solving a hard problem,
kind of as the like, you know,
underling of the supervising AI,
I feel like that's kind of going to be fine currently.
Oh, interesting.
Because that is like saying that these two things are somewhat,
if not separate, at least they shouldn't kind of pollute each other's context,
which is a little bit, I mean, it could be potentially like a strong,
stronger statement than, I guess, you know,
it's just kind of interesting,
because it might just be, like, to your point,
it's, you know, let's take the huge hillitarian answer,
is solving harder and harder problems.
It's doing a lot more than just, like, you know,
brute forcing something.
It's making choices.
It's like pruning, you know, a vastly large space
of possible paths into something that's like really,
you know, it's both tractable,
but then ends up being like it's a diminishingly small path within that space.
But, like, having, like, why would,
would be like a separate model,
is a separate generation or something that's a different version of the model that would contribute to taste,
or maybe that's totally like it's too abstract, doesn't make any sense.
You know, we should just let the actual.
This might, like, the related question would be like, you know, what is the thing that gets us to a better version of intelligence, the harness in the model?
Is it just the model?
And it's like we see this in, you know, at least in applied AI or, you know, startups where it's like it's a continual battle.
of like you need the harness,
but then the harness adapts very poorly to a new model
because sometimes like a very, very minimal harness
is still the best way to expose to the raw power of the model.
But then now we also have these like training regimes
where we require the harness to be trained with,
I mean part of this is to keep things more proprietary
and harder to, harder for other people to use it.
But I think partially it's maybe actually that it helps
have more control on like the reasoning traces you care about.
It's a long, rambling way of saying it's like, yeah, I don't actually.
This is so interesting to see how the models have gotten better at math.
And maybe something that's like very abstract and hard to describe, like, taste is a way to tease out, like, what is actually necessary here.
I think my only, like, non-tremial thought here is that, like, when you're working, I mean, just when you're doing any tasks, occasionally you get pigeonholed and you, like, work really hard and just having a friend look over your shoulder and be like, what are you doing?
And then just like, just having that one bit of, like, step back for 10 seconds.
Like, this is often very useful.
Yeah, yeah.
I mean, I see no reason why humans would be so different than models somehow.
Yep, yep.
Having, or models would be so different than humans.
Having a few humans working together is often more powerful than just having one.
Yeah.
It's like in this kind of collaborative thing, you actually, you kind of, yeah, artificially created it, but it's very similar in dynamic.
But I think a lot of taste is also, like, having a sense of what problems you or like some method you have in mind are going to be good at solving.
Mm-hmm.
Like, it's, I mean, certainly there's.
some amount of like absolute aesthetic point, right?
But there's also just like, you know,
having a nose for what you might want to pursue
because you'll be able to make progress.
And, you know, I think for that, like,
there's, you know, you would expect that
as a side product of being good at completing tasks
you would get there sort of, right?
Let me know if we still want to do like a section on soft at groups
because I think, you know, up to you guys,
it's definitely super interesting.
So maybe the first question is what is a group?
Let's remind ourselves.
So a group is a set of elements with some multiplication operation.
And basically, this is how mathematicians think about symmetry.
So you're like, basically like if G and H are elements of elements of your,
group, then GH has some other well-defined element of your group.
And you have like associativity and you have an inverse.
So for every G there's some inverse.
And there's some like specific element in the group that is kind of the identity.
Okay, so it's some like abstraction of like composing operations.
So these could be like numbers, they could be like multiplying matrices.
They could be like rotating something, which is a special case of multiplying matrices.
And a group is so thick.
If, well, there's some, you know, precise definition.
but
you know
roughly
it means
it
so I should say
like groups that can be
finite or infinite
so like
you know
if you have like a square
like all the rotations of it
form a group with like four elements
if you have like a circle
then the rotations form a group with like
uncountably any many elements
and
so phic groups are either finite
or countable
you should think of
as being countably infinite,
so there's like the same number of elements
as like the integers.
And if it's suffolk,
if in some sense,
it can be approximated by finite groups.
So we didn't know if there was a non-sophic group.
So the result that Astor proved
is simply that there exists a non-sophic.
Yeah, and without like, I mean,
we can, you know,
before going to that prove,
It is like, you know, I feel like a lot of the programs of math is like, okay, we are such finite creatures.
Let's see how well our finite approximations are, you know, do.
And in this case, especially for the countable case, it's like, maybe you'll be relating it to like the Aldous Leon's thing.
It's just like it helps kind of anchor the picture of like it seems like such a, I mean, it's a nice result if it were true, but it's not.
And it seems almost like reasonable.
And so, yeah, I actually didn't, um, didn't, uh, go.
I would love to hear the explanation of how it found a counter example.
Yeah, I mean, I would say that, like, you know, the hope that there was no non-Sophic group,
so every group has this kind of approximation, like maybe this is sort of like people hoping
that there's a miracle because it turns out that groups like this have a lot of nice properties
because you can run certain proofs for finite groups and then, you know, kind of approximate them
in whatever way, the definition of being Sophic
lets you approximate them and get the result.
So there's this notion of being a surjunctive group.
So there's some fact that any group,
which is Sophic, is also surjunctive.
Surjunctive is some property of like dynamical systems on the group.
And I guess the original question was,
whether every group is surjunctive.
This is some question of gotchalk from the 70s.
And this fact that follows this pattern of prove it for finite groups
and then do this approximation is what motivated the question about if there's a non-sylphic group.
Yeah, maybe I'll say a little bit about this.
I'll just lion's conjecture.
Yeah, sure.
Yeah, yeah.
So I guess I had heard of this a little bit beforehand because there's a related stronger conjecture.
in probability that was made popular by Aldous and Lyons.
This conjecture, roughly what it says, is like,
any infinite graph with some nice property called unimodularity.
A unimodular random graph can be
approximated by large finite graphs.
So maybe the way to like explain what these kinds of things are trying to say without getting into technical weeds is to say what they mean about the integers.
So, so like how would I draw the integers as a graph?
So this is called like heli graph.
You're just going to connect nearest neighbors.
Okay.
So there's some kind of canonical way in which this is like the graph that represents the integers.
Okay.
And there's some sense in which you can approximate this by finite graphs.
Why?
Well, if you look at integers mod n, then you kind of get the same picture, but you have like a big circle instead of an infinite line.
And the point is if you look at any point here and any point here,
like in nearby things look the same.
You have to go very far away to kind of see this global geometric structure
that you have a circle and not a line.
And in fact, the integers and integers mod n are both groups
just by like adding numbers or adding numbers mod n.
So these integers mod n are like
sophic approximations for the it full integers.
So, like, this approximation is kind of why the integers are a Sophic group.
So, um, so the sophisticity, the statement that every group is sophic is sort of a generalization
of the fact that you can do this approximation with groups.
And this, uh, Aldous Lyons conjecture is kind of a broader conjecture that, like any network,
you can do this.
Um, and it, like, you don't require as much algebraic structure roughly.
So it's, it's kind of a broad.
our conjecture. So this conjecture was disproved earlier, like two years ago, and it was kind of a
really tour-to-force work. It was like 250 pages building on another 200 pages. It used as like
quantum complexity theory. So it really builds this like, you know, very complicated bridge and like,
you know, I think not really people could understand this. Right. So since this is a stronger
conjecture, the disproof, like, is weaker than disproving this statement that all groups are
Sophic. But it turns out that the direct proof that there's a non-Sophic group was, like, much
shorter and easier than this really amazing disproof of all those lines conjecture. It's like,
like, 15 pages, maybe. And it doesn't have any of this, like, very complicated connection
with quantum complexity. It just kind of stays in group theory land. I mean, it uses some important, like,
you know, existing results by other mathematicians,
like Coon and Coon and Tom,
but it's like,
it's like a very reasonable,
normal kind of proof.
Yeah.
And it's kind of spelling maybe the obvious,
but like the connection between the Sophac group,
statement is just you take the Kali graph,
and that's the one that is like what they use or for the Aldous Leon?
Yeah, yeah.
And so that's why it's like a, you know,
a subset of.
Right.
So, yeah, basically what happens is, yeah.
So for a group,
you can take exactly a Kali graph.
So you take some like elements that like generate the group
and you kind of connect elements that are adjacent.
So in this case, like this is a Kali graph of the integers.
So right, so when you do that from a group,
you get like a deterministic graph.
Right, you just get like a single graph.
You fix some set of generators.
So this conjecture is stronger basically because it allows
a broader set of graphs that aren't deterministic.
It allows them to be random, but have some
extra, you know, you know, modularity property that constrains exactly how it can be random.
But, yeah, basically that's the difference.
Like, here you kind of have to give a deterministic network instead of a random one.
Yeah, anything kind of interesting, surprising about the results.
I mean, you mentioned some things, which is like it stayed within group theory, the techniques.
I mean, I think maybe it's like a nice example of this general pattern
that theorems produced by AI have generally been like the proofs are pretty short,
generally.
They're like...
Like with a counter examples so far.
Yeah, but this one, it's like, okay, it's sort of a counter-example,
but like there's some, you know, there's some like stuff you have to do to analyze things.
The difficult part here is that, like, the property of being a SOFIC group is not so easy to get your hands on.
So you have to find, like, a concrete way of saying, like, producing a way of saying this group cannot be SOFIC.
And like, and the proof is actually, it's very short.
It's like, it's almost a, it's a cognitive argument.
But it's like a very delicate commonatorics argument.
Like, somehow you need to both have the right statement and know what piece of the literature and then execute it correctly.
And that's very nice.
Like, the difficulty in this problem is that.
that like it's just a really, really hard.
It's like very hard to get your hands on like being approximated by any possible finite
through.
Yeah.
It's like what is happening at that countable infinity that's like resisting this approximation?
Like do you guys kind of, did it give a sense of, like, when you're like, okay,
Astra, explain to me.
Like what is the?
Oh yeah, we did that.
Yeah, yeah, yeah.
Yeah.
What was the good explanation you got out of it?
I think there's like some concrete like combinatorial obstruction.
Basically it's like, it's hard to explain, but there's like,
some concrete combinatorics of structure, which if you read the previous papers, you realize that
that's what they couldn't rule out. And Astor found a way to kind of say, okay, no, no, if you add
this one extra algebraic fact, this, this like weird conspiracy can't happen. It's like very clearly
trying to rule out a conspiracy in the sort of previous authors had implicitly written about.
And those were the actual suspects.
Yeah, yeah.
It turned out. So they were sort of on the right track. And then this did the last mile of, well,
whatever, however you quantify that.
But I think it's like, like a year ago, I would have been very surprised to learn like
all of these AI proofs are like very short and elegant.
Like they're, you know, you're kind of like afraid that they're going to like generate all
these thousand page things.
Yeah.
Yeah.
I don't know.
I would never be able to understand it.
But it's been kind of the opposite.
Yeah.
Like only humans can generate like 200 page proofs right now.
Yeah.
Well, and also I was like asking him like, if you do that.
post-mortem, it ends up usually engendering more mathematics. Because when you do that with humans,
like that's what, you know, breeds new mathematics. So maybe, maybe if you kind of alter the prompt
a little bit and be like, how would you, you know, generalize this or something? Like, yeah,
I don't know if that's been a technique for you guys to like have it explore and exploit what it has
already developed. Well, there has been some, there has been follow up on this already, actually,
by Kun and Tom, who this was always built on. So they like. So that community is coming.
Yeah, yeah, yeah, which is kind of what we're hoping.
Yeah.
You know, we don't want to be, you know, writing lots of follow-up papers ourselves,
but if there's some interesting follow-up that, you know, it's like we're very, very happy
that there's some follow-up building out these ideas more and giving, like, more examples
of non-Sophic groups in this case.
Yeah, well, actually, maybe that's a great segue into, like, how, you know, what's the ideal
way that this is being taken up by the math community?
Because I feel like there's a spectrum of answers from working mathematicians,
sense of like some, you know, probably most at this point are like, okay, AI is obviously doing some
non-trivial stuff. It would be a disadvantage not to admit that in my workflow. I've definitely
heard some stories where people are kind of, you know, would find it hard to either take AI as a
co-author or like how do you even do kind of attribution this way. But I don't know, like what,
maybe to paint the more optimistic picture, so you're saying you want the mathematicians to be building
on these results, it definitely generates a lot more results to be verified. So it puts pressure
on the community and the profession. Like, how do you, how do you kind of expect the evolution
of kind of uptake and collaboration with mathematicians? I mean, given that the fact that the
models can produce sophisticated mathematics means they can help you understand, like,
sophisticated mathematics. I mean, like, I don't know, occasionally, like, I enjoy looking at the
archive and I want to understand some proof. And, like, I could read the introduction.
But in practice, it's just much faster, take the PDF,
put it into my favorite model,
and then get an output of like, what is the rough proof strategy?
And so I have this, like, along, I mean, of course,
models are going to help us produce exponentially more mathematics,
but they also make it much easier to absorb it.
And right now, okay, it's still a bit of a challenge back and forth,
but I think it's, for me, at least,
much, much faster at understanding.
It's much, much faster to understand it,
mathematics with a model than without it.
It's helping solve the problem it creates anyways.
Yeah, I feel like that at least.
And it's, you know, I don't, I don't view it as like creating much more wrong.
But again, like, I don't have such, you know, high stakes in like, okay, I'm going to get,
I'm not going to get tenure, et cetera.
So like, I agree, like making it more accessible.
Like, if I'm not spending so much time absorbing an area, I can like put it into chat
GPT and then expect to, I mean, you guys have an even more powerful model, hopefully releasing,
for other people to enjoy as well.
But like it's, I think like the positive version of that
is actually more people can participate in mathematics.
It's like people might be coming with other intuitions
and they could actually maybe generate good mathematics.
Is that sort of like closer to the vision of what you're hoping
this is, you know, pushing towards?
Or like what things do you think mathematics should be wary of
to kind of adapt fast enough to take advantage of AI?
Yeah.
I think certainly there will be a lot of changes, right?
Like, I guess in math, like, there are a lot of things that are kind of important for,
for, like, a given result, right?
You need someone to come up with it, but you also need people to understand and absorb it
and, like, you know, internalize it enough to do more with it and, like, figure out where it
fits into, like, humanity's understanding, right?
And like a couple of years ago
Like the proving the result was like
So hard that kind of the other stuff was just kind of coming along for the ride right
You know like if you if you manage to like prove this thing yourself
You're automatically going to understand it quite well
You're kind of responsible for like maintaining it in some sense and like you know explaining it to other people
And yeah now this kind of what was the main bottleneck
before is kind of much less of a bottleneck.
And, you know, these other kind of constraints come into play.
So, yeah, the, like, optimal structuring for, you know,
organizing the knowledge could look rather different.
Yeah.
How does that look?
I mean, does this make the field a lot more kind of empirical?
Will people do sort of the hard, like the first thing that was scarce,
which is like all the reasoning and then more.
I mean, not that it's like a bad thing to make it empirical,
but it's almost like it functions as a very different discipline.
Like a lot of the fun stuff is understanding, you know.
And so understanding, communicating, maybe assembling,
having still the human taste,
does that sort of remain rarefied?
And that's how, you know, current mathematicians need to adapt
and reward, you know, contributions.
Or is this too much of a caricature that's like something else?
I think certainly understanding how to put, as we get more and more mathematics, put it in like a proper framework and sort of how sort of like being able to explain it to other humans so that they can also appreciate it. I mean, so implicitly we valued this, but it was usually because you were the person proving the results that gave everybody else the understanding. But I think increasingly would be a function of like you're sort of helping, you're the human who can sort of give this understanding to other people and sort of help them with it. I think that more, sort of, sort of,
that communal understanding will, I think, become.
It was much more implicit in how we viewed math and gene,
but I think it will be an increasingly more explicit
and valuable part of the subject.
I mean, a nice thing about math is that the ceiling
for difficulty of a math problem is pretty high.
So even if, you know, kind of,
even if AI continues getting exponentially better at math,
like it might, you know, plausible will never solve something
like P versus NP.
And it could be.
be that like the the field kind of becomes more you know attached to like like these big mysteries
and less to like smaller mysteries that are more like routine now yeah yeah I think that's a
positive vision of that I mean also like I don't know there are things I spent like months or years
of my life wondering about not getting to know and hopefully we get that yeah some portion of them
I'll get to know the answer to it.
I'm pretty happy about that.
No, exactly.
No, I'm excited about this, like,
renaissance of results and understanding,
and I feel like, I mean, this is such an infinite, you know, field.
Like, okay, no pun intended.
But, like, it's just, like, it's just,
there's so much that you can actually create here.
So, I mean, especially for somebody like me,
who's not going to have the time to actually, like,
practice mathematics.
Now there's, like, a lot more that you can actually do
in the activity of now.
So, yeah.
Yeah, I think the, like,
The ability of someone who's not working on math is like their literal job all the time to like understand what's going on and like, you know, learn about some of the mysteries they might have wondered about will go up quite a lot.
Also, you know, if you're like, if you're working on something that requires some math, you know, suddenly you don't need to like find a world expert on this topic to be able to, you know, use it in your own work.
Sorry mathematicians.
No, it's true.
I mean, I think there was just like a dearth of actual, like, people who could do that.
And so I think this is helpful.
Maybe it's helpful for theoretical physics, like we'll see.
But a lot of other applied areas as well.
It'd be nice for the world if applied mathematics on a lot faster.
Yes.
I mean, I'm over that.
Well, thank you guys for joining.
This is a lot of fun.
And I'm, you know, just so excited for how much the models are advancing.
So maybe we'll have you guys back soon.
Yeah.
Thanks so much for having us.
Yeah.
Thanks for having us.
Thanks for listening to this episode of the A16Z podcast.
If you like this episode, be sure to like, comment, subscribe,
leave us a rating or review, and share it with your friends and family.
For more episodes, go to YouTube, Apple Podcasts, and Spotify.
Follow us on X and A16Z and subscribe to our substack at A16Z.com.
Thanks again for listening, and I'll see you in the next episode.
As a reminder, the content here is for informational purposes only.
be taken as legal business, tax, or investment advice, or be used to evaluate any investment or
security and is not directed at any investors or potential investors in any A16Z fund.
Please note that A16Z and its affiliates may also maintain investments in the companies discussed
in this podcast. For more details, including a link to our investments, please see A16Z.com
forward slash disclosures.
