a16z Podcast - OpenAI Researchers on the Future of Mathematical Reasoning

Episode Date: September 8, 2026

a16z Infra Partner Lisha Li sits down with OpenAI mathematicians Mehtaab Sawhney and Mark Sellke to discuss how quickly AI’s mathematical capabilities are advancing, what recent results reveal about... model reasoning, and what happens when AI begins making progress on problems mathematicians have struggled with for decades.Mehtaab and Mark unpack several recent results from OpenAI’s models, including advances in sphere packing and the construction of a non-sofic group. They explain why the surprising part isn’t simply that models can search more possibilities or work longer than humans: in many cases, the reasoning traces look remarkably similar to the work of an expert mathematician, including choosing promising approaches, backtracking when they fail, and combining ideas from across the literature.They also explore what this means for mathematics itself: how the role of human taste and judgment may change, whether AI could produce far more mathematics than humans can absorb, and why models that accelerate discovery may also make sophisticated results easier to understand.Resources:Follow Lisha Li on X: https://x.com/lishali88Follow Mehtaab Sawhney on X: https://x.com/mehtaab_sawhneyFollow Mark Sellke on X: https://x.com/MarkSellke Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Transcript
Discussion (0)
Starting point is 00:00:00 Often as a practicing mathematician, you have an idea, and then you kind of think it might work. Then you try for a few hours, a few weeks, and at some point you give up. Whereas for GPT, like, okay, human told me to do this, like, let's just do this. And so that's why we're sort of in this renaissance of, like, reachable results.
Starting point is 00:00:17 This is the best part about this problem, which is really nobody had any idea. Is the model just guessing in some insane way? It doesn't seem like there's a limit so far, but it doesn't have that context yet. It'd be nice for the world if we'll find math not next one along. faster. The ceiling for difficulty of a math problem is pretty high. Even if AI continues getting exponentially better at math, plausible will never solve something like P versus N. What's the ideal
Starting point is 00:00:40 way that this is being taken up by the math community? Probably most at this point are like, okay, AI is obviously doing some non-trivial stuff. So AI isn't just getting better at math benchmarks. It's beginning to make progress on mathematical problems that have resisted humans for decades. In this episode, A16Z-infra-partner, Leisha Lee, sits down with open-AI mathematicians Metabswani and Mark Selke to understand what's actually changing. They walk through recent results in sphere packing, coding theory and group theory, and explain why these advances can't be reduced to brute force. The models try different approaches, abandoned dead ends, connect ideas across fields, and in some cases produce reasoning that reads surprisingly
Starting point is 00:01:23 like the notes of a human mathematician. They also tackled the bigger question. What happens to mathematics when proving a result becomes less of a bottle leg? From mathematical taste and human judgment to understanding an explosion of new results, Leisha, Metab and Mark explore how AI could change
Starting point is 00:01:40 not just what problems get solved, but what it means to practice mathematics. Well, thank you guys for coming. This is really exciting because I think math has been moving so fast. with AI. We just love to get to both practicing mathematicians and who work at Open AI to chat on some of these results. So we have with us Mark Selke and Natalzwani. We're connected, actually, because Ufa was actually your advisor. So both of you guys have worked
Starting point is 00:02:10 much more deeply in math since I have quit many, many, like over a decade ago. So this is very exciting to kind of hear your download of your thoughts on how open AI has been sort of approaching this and also just like where you think math is going with the incredibly rapid advance of how AI has been helping. We can start off with some very basic questions. What do you do to the extent that you can of course share? And how did you come from being a practicing mathematician to working at Open AI? Yeah, I mean, I guess we both broadly got excited last year when the models started to really take off in math. So I joined a little bit Before Metaub, I saw the IMO gold medal last summer, basically.
Starting point is 00:02:53 I thought, this is amazing. I want to see what the heck they did. Let me go see. And then, yeah, I guess in the fall, Mark gave me a GPG5 account. And then I started playing with the models and very quickly became convinced that, yeah, it was extremely exciting to play with them. And you two were collaborating before them. Yeah, we've known each other for a while.
Starting point is 00:03:12 Yeah. We have one paper we actually wrote jointly. Yeah. Yeah. So GPG5 was your conversion? Yeah, yeah. What was the magic that sort of, what question do you throw at it? What process?
Starting point is 00:03:23 Yeah, so I think, actually, yeah, so I think how this started was, at least for me, the starting moment was something like there's a collection of problems called. So Paul Ardish is a very famous mathematician. He posed a bunch of problems. And so they've now all been collected on this site. And so I specifically worked in combinatorics and a lot of these questions are among the most important. So it's always fun to flick through this light. But one thing that often happened to me that was extremely frustrating was, I would look at a question, see that it's marked as open, and then not actually know if it's correct, not actually know if it was still unsolved, because the literature is often quite hard to search.
Starting point is 00:03:59 And one instance, I just plugged it into GPD5, and five minutes later, it found a reference. And this was a case where a few of my friends actually started thinking about the poem on the site, I was talking with them. And, I mean, we had spent a few hours. It wasn't clear if the poem was within reach. And it was very nice. Okay, to be told, yes, this is in reach. Here's how you do it. And, yeah, GPD5 told me this.
Starting point is 00:04:18 And then I told Mark about this. Yeah, this is sort of, yeah, this for me was quite a surprising moment. Yeah. And then we looked more into it and we found 10 more cases sort of like this. At the time, I feel like being better at maybe making connections between, as you're saying, like the search for whether there's been a result or a related thing earlier is just kind of humanely hard, but maybe better for machine. But I imagine as the progress has happened in the last year, what has been impressive,
Starting point is 00:04:48 kind of reach beyond that. And maybe through talking about it more abstractly or if it's more natural to talk about it through one of the problems that has been recently announced through, you know, Astra, you can kind of enlighten me as to like how the recent progress has been a lot more than just searching through more areas,
Starting point is 00:05:05 making these connections between the field and perhaps just actually deeper, more mathematical reasoning. That's a similar to a working mathematician. Yeah, I mean, I think this like search point of being, familiar with everything is still definitely like a relative strength that maybe informs, like, the types of problems that AI is solving now. I think there are some other relative strengths and
Starting point is 00:05:29 weaknesses. Another relative strength that's pretty noticeable is just, like, it's very good at executing on some, like, idea once it has it. Whenever you have an idea, there's usually some amount of getting everything lined up. Is epsilon, like, smaller than delta, this kind of thing? You have to get everything correct. And like for a human, it's easy to get lost in these kinds of details. And the AIs just kind of always nail these kinds of arguments I find. Yeah. I feel like you guys will know more detail on this. But like for the unit distance problem, it was just like the approach, there was definitely contributions from open AI. But like the approach perhaps was suggested even originally by Airdush. And then it's just that the actual reasoning was a very, very like,
Starting point is 00:06:16 momentous feat. And so for human, you're like, well, I only have a limited amount of time. And if after so many steps, it is still not clear. I mean, maybe you're like Andrew Wiles and you actually spend 10 years alone and do something, but like it's not clear that the risk reward is not good enough. Whereas for GPT, okay, like a human told me to do this, let's just do this. And so that's why we're sort of in this renaissance of like reachable results. Does that track? And did you feel like with the astro results, is that sort of like where the strengths have been primarily, or there's an extra ingredient or magic here. I feel, I think the unit distance example is about it's quite telling.
Starting point is 00:06:55 In the sense of maybe the exact construction, you can make it look very similar to what people had tried before. But I think, I mean, often as a practicing mathematician, you have an idea, and then you kind of think it might work, then you try for a few hours, a few days, a few weeks, and at some point you give up. And then a not so uncommon experience is that you find out a, a you find out a, a you year or two later that somebody else got the idea to work that you thought that didn't work. So somehow getting an idea to work can even be a large portion of the battle.
Starting point is 00:07:23 And I think, especially in the case of the unit distance conjunctured, there's just a lot of extraordinarily finicky details. And very often when you're doing mathematics, you're kind of gambling against the problem. You're like, maybe I should try this approach, but it seems really unlikely and just not worth my time. And the model, I think in several of these cases, both by combining what it knew and sort of having good taste, kind of made the correct balance. And you can kind of see this in the summarized chain of thought we released.
Starting point is 00:07:49 You can sort of look at it. It's reasoning like a mathematician, and because it knows a few very correct fits, it makes the right decisions and eventually able to prune the search tree. It's not really trying everything. It tries a lot of different things. It's extremely dogged. I mean, it can't try every idea. It has to try a limited set of ideas.
Starting point is 00:08:05 And it's able to kind of use its knowledge plus good mathematical judgment and find the right path to go along. So, I mean, that for me was like, because this was a problem which a lot of people have thought about. And I mean, the fact that the idea is not so foreign probably indicates that a lot of people have tried it, or at least a few very serious mathematicians have tried. And I think that's what made it really interesting to see.
Starting point is 00:08:24 I think something else that I feel when I see these proofs is like, if I have an idea and I'm trying to execute, it might be that I have some wrong plan for like how to get things to work. And as a human, if you have some wrong path you go down for a while, it can be hard to like rewire your brain to start over and try a different path. The initial idea is kind of linked in your brain with these other things that ended up not working.
Starting point is 00:08:50 It's sort of like your context window is like a little polluted and you can't just make another clone of yourself from last week and say, don't do this, try something else, build your intuition in another direction. But it's very easy to do this with an AI. So I think this is another reason that it's like getting the details right
Starting point is 00:09:06 once you have some good general direction is much less of a barrier all of a sudden. And when you say it's much easier to do with AI, It's like it's not actually being directed with human interference, too. It's, as you were saying, in the reasoning traces, it's like making these choices. Maybe it backtracks, but then it's able to not be distracted by maybe the context in which it's thinking about the problem, like, by these machinery. And it's like, do you see it kind of go back as well? Or is it just like good choices?
Starting point is 00:09:30 Like, is it a lucky sample or is it actually reasoning like a mathematician or, okay, it doesn't do well in this path, but it goes back, but then it doesn't let that pollute? No, I mean, it definitely makes mistakes and then it goes back and thinks about it. I think it's somehow very calculating, very correct. I mean, as human mathematicians, you're not always perfect in making these decisions. The first time something doesn't work, you automatically kind of downgrade how likely this approach is to work, and you keep doing this a few times. The model somehow is much better able to, like, it seems, for several of the solutions we've seen, somehow it seems much better able to update, like, how likely the path is to work,
Starting point is 00:10:04 like, versus rejecting a path versus a human doing it. But I think even if it weren't, the fact that you could just start another model session over. It's always going to be the case that... Got it, got it. So in some sense, it is, like, still leveraging the fact that you could, like, run kind of parallel, you know, agents on the problem. But if it were kind of backtracking, then does make it seem much more like a human,
Starting point is 00:10:27 you know, mathematician. And perhaps it is kind of doing some of that stuff, too. Because, like, obviously, like, we have to, like, make mistakes in order to, like, even gain intuition for, like, why that solution space is, like, not, you know, not in the set of paths that it could be in. I mean, I think this kind of thing happens with humans, too, where, like, if you get stuck on some approach, you might tell another human your kind of general idea, and then they'll come back and, like, figure out how to get it to work. And, you know, it's just, it takes more time to do this, like, with humans. I wonder, I mean, you know, maybe this gets to extend that you can actually talk about sort of like, obviously don't talk about the training recipes or whatever.
Starting point is 00:11:02 But, like, it's interesting that if you're just studying, for instance, from math papers, it's like a very poor training set, like a priority for math, because. I mean, maybe math textbooks are even a pure example of this. It's, like, really bad at actually reconstructing the motivation for why things were. It's like, I mean, maybe some people like it, but don't learn real analysis from Roodin. It's just like it's very clean already and crisp. And I think that's bad because it doesn't show the struggle that made us formulate definitions in a certain way. Like, why do we even need to have real numbers be defined in this, like, super abstract way? et cetera. And so, you know, I think papers also, I mean, unless you're, most people don't write papers with the, the
Starting point is 00:11:48 context of I need to educate somebody to be a mathematician. And so like the actual maybe curriculum of like learning math is not inherent in like a lot of our artifacts as mathematicians. So maybe another way to ask this question is if the reasoning tracers are actually producing things, I was like, okay, this is actually more close to mathematical thought. Like how does that arise? I mean, yeah, I guess Open AI has been like the pioneer of reasoning models and, you know, teaching AI to reason in this way. So, you know, we're doing a lot of work at kind of kind of all possible directions on, you know, teaching models to reason better and for longer and, you know, in all kinds of different domains. Yeah, I mean, I think we're training general purpose reasoning models and kind of one. a lot of these behaviors that we're describing mathematically,
Starting point is 00:12:44 like backtracking or kind of starting again, I mean, these are not really specific to mathematics. I mean, we're seeing them specifically in mathematics in these examples, but kind of their general purpose tools for reasoning. And I think if you work hard at reasoning, you should see these patterns eventually. So it's a submergent because it's, I mean, I do think that's why the opening eye approach was so,
Starting point is 00:13:05 I mean, it's like it doesn't rely on, you know, doing auto-formalization in all. order to like guide the reasoning. I think that's like obviously more like us. But it's just like so not obvious that if you're just like training on say a corpus of like math proofs, maybe auto-formalized and lean, that you get, get the sort of like projection of like how to think well. Like put another way like with maybe maybe if we think about it with code. Like code is such a good corpus to train on because it's one of the few datasets that has such large context. You just like, I mean, maybe you see this kind of with books,
Starting point is 00:13:42 but they're less structurally interconnected. There's just less structure there, I think. It's safe to say, like on average in a book compared to like a piece of code. And so, like, with math papers, I feel like, but maybe what we're still bad at with coding models is stuff that that data set doesn't contain, which is like kind of the semantics.
Starting point is 00:14:05 Like the syntax is there, but there's a little bit of like the higher level semantics of what produced, like, why do I have to write it this way? It's not, I'm kind of getting too much in the philosophical good. It is just, like, really interesting how it's still emergent that it's doing good mathematics. And we'll probably get into this more detail if you guys, you know, wanted to talk in more detail about some of the problems, which is just like it's, it's not just doing like the expected, like, we'll push the brute force thing. Like, you clearly are impressed with some of the reasoning traces. And it's just not obvious that's gleaned from,
Starting point is 00:14:36 you know, what we would imagine would be the easy training set here. Yeah, absolutely. I mean, I think this kind of thing is one reason we decided it was important to release like these summarized chains of thought for these kinds of results. Because if you, if you've never seen these and you just see all these proofs coming out, you're kind of, you're not sure what it means. Like, is the model just guessing in some insane way? Like, is it thinking in some totally foreign, like what's going on? But actually, it's reasoning kind of shockingly like an expert human would. Yeah, yeah. Yeah, it's very much like reading colleagues' like notes. I mean, it's a little more disorganized in some way, they're kind of, like, especially you work close enough with the collaborative, sometimes you'll just see them like spill out their thoughts and an email to you. And it kind of, it feels like reading a lot of those change together.
Starting point is 00:15:26 So it's, yeah, it's very, it's quite surprising the first few times. Were you two sort of very involved in choosing the problems to release in this, like, last 10 problem set? that Astros applied to? Which was your favorite? Yeah, we have definitely involved. Do you want to start? Yeah, I mean,
Starting point is 00:15:49 spearpacking maybe? Yeah, I guess, yeah. So I guess my personal favorite among these problems is the following. It's extremely simple question, which is just like, it's just about how efficiently can you put a bunch. My circles are not very good, and they're not all the same size.
Starting point is 00:16:02 But we're assuming they are. Yeah, so the question is just like, how dense can you place a bunch of, So you have a bunch of spheres. You have a bunch of spheres of radius 1 and D dimensions. So the question is, how densely can they pack? And so, yeah, so in two dimensions, it's kind of like the, so, so D equals 1, this is not an interesting question, kind of, it's just the real line. Yeah, you can cut it up in a sphere in dimension 1 is just a unit segment, so, okay, you can cover everything.
Starting point is 00:16:38 So in D equals 2, it's kind of the picture that you know that everybody loves. It's just like a bunch of spheres which sort of form like a hexagonal lattice. Hopefully I've drawn it well enough that I can draw the hexagon. Kind of betraying my naivete on this problem. Is that like obvious? It's like a very elegant proof that it's a regular lattice? Yeah, it's not so obvious that this should work. It was only proven in the 60s, I think.
Starting point is 00:17:04 There's a short argument, but it's not so easy. Yeah. Where's the intuition? Like, what is kind of like the machinery of the argument? I mean, it definitely looks like it should work. That's why I'm going to go. It's so good. Yeah, I think this is the best part about this problem,
Starting point is 00:17:19 which is really nobody has any idea of this problem. Yeah. So, yeah, so, I mean, honestly, the best intuition that I have for this is that, like, bees do this. And if there was a more efficient way, then probably bees would pack potty comb some other way. Because evolution is efficient. Yeah. I think beyond that, like, I don't have a great argument.
Starting point is 00:17:36 I mean, and I think how little we know is demonstrated by the fact, so, okay, D equals three, the answer is just like, it's how you pack like oranges in a grocery store. And this was only, this was proved by Hales sometime in the 2000s. And like, and we don't have a short proof of this. Like, I think the shortest proof is like a few hundred pages. What area does it, like, draw from in order to it? So it's a lot of linear programming arguments and it's very delicate, like, genre. It's quite ugly, actually. Yeah, it's like...
Starting point is 00:18:10 This is like a famously ugly argument. Oh, no. And then the two most famous results are D equals eight and 24. Eight and 24. It must be some weird subspace thing. Yeah, exactly. So this was done in some... Like gluing something.
Starting point is 00:18:26 Yeah, so this was done in 2017. Sounds slightly prettier, though. Yeah, so the reason it works out in these two very special dimensions is that, so this is called a lattice. packing, so it's like kind of very regular. And it turns out in these two dimensions, there are two very special lattices. They're called the E8 and leach lattice. And they're very nice and they're like unusually dense. Like kind of, they're just very, very pretty structures coming from other areas of math. And it turns out that they're the optimal structures. But they're still
Starting point is 00:18:57 like regular. Yeah, they're very regular. But I mean, beyond this, so we don't know any more exact dimensions. We know these five dimensions and we kind of don't know anything else. And, I mean, like, to give an indication of how little we know, so there are two very surprising things about this. So you can define, like, delta D to be, like, the densest sphere packing in D dimensions. So there's kind of an easy lower bound of, like, two to the minus D.
Starting point is 00:19:26 Basic, yeah, this is not so hard to show. Basically, any packing where you can't put in another sphere has to have this density. So, okay, it's not ridiculous. small. And we know that it has to decay exponentially. So it has to decay, like, it grows like one minus C for some, at least for some constant. So in large dimensions, you can only cover like a vanishingly small portion. But we know like basically nothing else. And so for- That's just because like the high dimensional sphere like thing where they occupies, um, it's just like,
Starting point is 00:20:00 yeah, the volume behavior is weird. Yeah. So basically, I mean, basically they don't want to touch next I don't know, I don't think there's a particularly short way to see that it's like exponentially small, but it's known to be exponentially small. And for a long, long time, the best bound was something like this funny number like two to the minus point 599 D. And this was proved by two mathematicians in the 70s. Kapitiansky. Okay. That's a weird number. Where's that spit out of it from?
Starting point is 00:20:33 It's like commentatorial. It's it. It's the answer to some extremely ugly optimization problem. There's like a nice underlying strategy. Okay, that's nice. Yeah, yeah. I'll say one last thing about this, yeah. So, yeah, I mean, these were these two Russian mathematicians in the 70s.
Starting point is 00:20:51 It's thinking very hard to find their paper. Like one page, it's like two pages long. Yeah, they don't write very many details because paper was far. But, yeah. And the two negative Ds, there's just like the square lattice, like the dumb one? So yeah, it's actually not so easy to... So the argument for this is as follows. Basically, imagine that you have a set of spheres,
Starting point is 00:21:11 and you can construct a set of spheres so that, like, you can't put down another sphere. Because if you could put down an extrasphere, you just keep putting it down. So you have a set of sphere so that there's no other sphere, which you can put down. That's here?
Starting point is 00:21:26 It's like almost like a... Just, yeah, just take any such packing. Okay, yeah, yeah. And I claim that this has to cover at least two to the minus D. The reason is that, like, if you blew up each of these spheres by a factor of two, then, like, they have to cover every point in space. And the reason is otherwise you could put down, if there was any empty space, you could put down
Starting point is 00:21:46 a sphere there. So I guess if you take the usual lettuce, there are actually, like, more places you can put things kind of diagonally. Okay, yeah, yeah, so that's actually not. It's like a worst bound. Yeah, you can just keep plopping things in. Yeah, this is like more, okay. Yeah, this is related to this, like, really funny fact where you, like, put a sphere,
Starting point is 00:22:03 on every point. If you take a cube in high dimensions, you put a sphere on every point. It's like vanishingly small. It's so small that you can put another sphere in the middle. Yeah, yeah. And it fits.
Starting point is 00:22:13 Another high dimensional sphere behavior. Yeah, it's very weird. And so, okay. So the great part is that so the model shows the following. So I'll write two things. So this is Astra, I guess. Probably the right way to refer to this.
Starting point is 00:22:32 So, okay, I'm going to be. to write something called the LP bound. I'll explain this in a second, and it shows that it's smaller than this very nice number. Do you want to say equals? Yeah, it's equals, actually. D to the
Starting point is 00:22:49 2 pi was little 1 to the D. And if you can work out what this number is, it's like, it's like roughly something like 2 to the minus 0.6. Close about. Yeah, it's surprising.
Starting point is 00:23:05 one, you know, you're like, oh, maybe there's some like nicer kind of like structure there that fell out. And it's a most, this is like roughly something like two to the minus point six zero one dot dot dot dot D. That's the numerics. I thought it was six or four. Great. Yeah, this shows mine.
Starting point is 00:23:24 Yeah. So, okay. So there are a couple of things. So first, what is this LP of D? So Viazazzo's work actually builds on some earlier work. It turns out that there's a way to attack sphere packing via what's called. a linear programming bound. So LP just stands for linear programming.
Starting point is 00:23:38 So Conan Elke's gave an approach for sphere packing based on linear programming. So it's like a linear optimization problem over a convex set, but it's all kind of infinite dimensional here. And basically what this reduces down to is you try to understand the follow.
Starting point is 00:24:01 So what you try to show is you, basically you construct a function F. So this is in D dimensions and it's mapping to R, and it has the following properties. So first, F of X, this is a function in D dimension, so it's always less than zero if the size of X is bigger than one. And you second have that the Fourier transform of X, this is always non-negative.
Starting point is 00:24:29 So this is just, this is a linear program because the Fourier transform is a linear operator. and night race. So you're taking just some arbitrary F that satisfies this property? Yeah, so you can take any F that satisfies these properties. And what they prove is that Delta D is bounded by the ratio of the Fourier Transformer at zero to its to the the for a transformer at zero, an F of zero to over the Fourier Transformer of zero times the volume of the ball of radio.
Starting point is 00:25:03 just one half in D dimensions. And so, okay, this proof is not so short for experience about petition. It's like it's half a paragraph to prove it, but it's a little bit tricky. And the point is, so it turns out, so this is a relaxation problem. There's no guarantee that taking the optimal left
Starting point is 00:25:22 will give you a good bound on Delta D. But, so what Vazasasca did, and this was sort of the key, I mean, a large part of the reasons you want of fuels metal, and in 2020 was that, or in 2022, was that she constructed a function in 8 in 24 dimensions such that this upper bound matches exactly these two very special lives.
Starting point is 00:25:46 And these are kind of miracles of nature that both you can construct this function and that it gives you the optimal bound. But you can just, this is a very, very natural problem. It's a function with two very simple properties and you just want to understand how this behaves in for large dimensions D. And that was a big mystery.
Starting point is 00:26:07 There was a numerics paper by Cohn and several others, which conjectured that just based on doing numerics, that this was the answer, but they had no idea why this would be the answer. And what the model shows is that actually, the linear programming bound in large dimensions had this extremely nice asymptotic behavior. And the proof kind of explains where this is called from.
Starting point is 00:26:31 And because you understand this LP bound perfectly, this actually just gives a better bound on delta D. It turns out that this old bound can be kind of reinterpreted in this framework. And what the model does is it shows you the best possible bound you can get by this framework. So the model sort of made the connection. And what is the sort of like, what do we?
Starting point is 00:26:51 So I think, so the model gives a function F, which, so first it constructs a function F, which gives you this down. And then it shows that there's no function F which does any better. So it's inequality, which is quite strong. So, like, sort of, we now understand this problem
Starting point is 00:27:08 in high dimensions very well. And that's pretty remarkable. And the model was just kind of told, like, analyze this linear program in high dimensions, you know, go have fun. Got it. And to give an indication of, like, how it was known,
Starting point is 00:27:23 I think this conjecture was based basically only by doing numerics. Extremely clever numerics, but numeric. And so, yeah, You have to kind of figure out why this is the right thing to aim for, and it does. And that was pretty remarkable. Yeah, I mean, I had actually thought about this problem for about six months at some point when I was a graduate student. And yeah, I remember making like absolutely zero progress on it. So it was very nice to be like explained why it was, yeah, why it was true.
Starting point is 00:27:50 So that was a positive experience. I think also in general it was one of these solutions which I knew several people who had tried the problem. It's pretty remarkable because like the model solution, especially for the, this being like, the LP can't do better than this was like quite short. It's a few pages of like complex analysis, but it's kind of exactly the right approach. Like once you see it, it's kind of, it's like unbelievable. Like why, why hasn't somebody done this before? It was, it was like, there are many types of good mathematics, but I think one of them is just like, you see it and you're like, oh man, why didn't I think of this? And it was really fun.
Starting point is 00:28:22 And it was, I mean, I sort of knew why I didn't think of it, but it was quite nice to see it. And it was fun to see. That's why I like this problem a lot. Yeah. So this is. is the first of the 10 problems that Astrosol. But the second is actually closely related. So this was sphere packing. The second one is spherical and binary codes. So what's like you draw a code?
Starting point is 00:28:46 You drew a packing. It's going to be the same picture. Okay, sure. Otherwise, we're going to have his picture. The hexagonal packing in our minds. Yeah, a spherical code is literally just a sphere packing, but on another sphere. Yeah, so I mean, a spherical code is basically just a sphere packing on the surface of another sphere.
Starting point is 00:29:13 So, yeah, sure. Okay, so same picture as before, except you're kind of on a curved surface. Okay, okay, okay. Okay, so why is it called a code? Well, you can, I guess the reason is because of binary codes, which is, again, the same sort of thing, but now it's on a cube. Okay, yeah, fine. Let me draw a picture of a cube
Starting point is 00:29:46 and some, like, simplest possible code on it. So, like, when you're, like, sending... So this is really, like, about error-correcting codes. So, so what are error-correcting codes? So, you know, it's like, I send you some string of bits, right? and maybe I'm worried that some of the bits I send you get corrupted, right? So maybe like just because of some errors in my system, like this one gets changed. And we want some communication protocol so that like you can decode this like small amount of error
Starting point is 00:30:31 and like recover what I was trying to tell you. And you know, like normal English language kind of has this sort of property, right? if I make a few typos, you're going to be able to understand what I'm saying. But if we have some really brittle communication scheme, it's not going to work. So codes are kind of the way you solve this. And mathematically, it just means, like, you know, what's a binary string like this is a fixed length? It's like a point on some hypercube.
Starting point is 00:30:59 And we want a dictionary of allowable code words that are, like, separated from each other. So, like, in this case, if I don't want any two to be adjacent, I would kind of take these four vertices, kind of the like even ones, if you sum up the digits, right? And like, okay, I guess, okay, in this case, I guess if I have an error, you can't tell which one it's from,
Starting point is 00:31:22 but at least you can tell it's like not, at least you can tell there was an error. Oh, I see, I see. Yeah, yeah, because it's like kind of sparse in the, yeah, it's like it's not too adjacent so that like, is like when the hand you're just like, yeah, yeah, yeah, right, right. Yeah, so you want like a large hamming distance between any distinct in your dictionary.
Starting point is 00:31:43 And yeah, I guess if you take two opposite corners, then if I have like a single bit error, I can always like recover which point it was coming from. One that it's definitely closest to. Yeah. So there's kind of the, you know, same question in both of these cases, like in a very high dimensional setting, what kind of rate can you get? And like for binary codes, it's really like, you know, an extremely practical question. it's sort of like if I send you
Starting point is 00:32:07 like an N-bit string and there's like you know 1% error rate like how much longer does my message have to become to tolerate that amount of errors right this is like some fundamental information theoretic limit of like like you know
Starting point is 00:32:22 communication and you know but you can see like certainly this this spherical case is like it looks very much like sphere packing for example if you like If you make all these little spheres really small,
Starting point is 00:32:37 then the curvature of the big sphere is kind of not going to matter so much, and it looks like just packing spheres in full space. And in fact, yeah, like these problems turned out to be very related. So for these problems, there were similar bounds coming from the KL authors, and like there's like something for the sphere and something for the cube, but it's all kind of the same stuff. And our models found better bounds for these cases as well. And, like, it, I mean, the techniques look pretty different, actually,
Starting point is 00:33:16 if you, if you, like, write them out. So this full-space analysis of this linear programming was using, like, just complex analysis. But if you, like, the method for these cases, we're using representation theory. Like both the sphere and the cube have a lot of symmetry. And basically the idea of the proof was to really leverage this symmetry. Like there's some amount of this in the previous existing method,
Starting point is 00:33:46 and really the improvement is to lean into the representation theory really hard and kind of make the algebraic symmetry, like, in turn, a more sophisticated way. And then it, like, turns out that from the representation theory formulas, as if you kind of take this like, like small sphere limit in the spherical code case, you recover like part of this result and you recover this value. So like this result isn't a special case. You kind of only went one direction of the bound from looking at it from the code's point of view, but like there's like a very close connection.
Starting point is 00:34:21 Okay. Yeah. So you guys let this like run in parallel. So it's like kind of discovering or because you're not sort of like feeding it. So actually this was the one case. where there was some interactivity involved. Oh, interesting. So for all of, so except for this pair,
Starting point is 00:34:37 it was just, you know, we had some problems. We fed them in and we, you know, the model came back with some solutions. What happened here is actually pretty interesting. So we first asked it to improve the bounds for the codes. And it came back with an improvement that like used some amount of representation theory. And then we kind of asked it, can you like push this further?
Starting point is 00:35:02 Like, you know, what happens? And then it came back with some like much more sophisticated representation theory. And like it turned out that you got this conjectured value for full space sphere packing like out of that method by pushing it as far as it can go. So then we kind of asked to directly analyze the sky and try to complete the picture. Okay. Yeah. So the relationship like isn't a coincidence.
Starting point is 00:35:28 Yeah, yeah. It's like interesting when you're saying the first prompt, which is, you know, maybe so basic, which is like, can you push this further? It does require some judgment for mathematicians, but like eventually you would imagine by scaling the models, you don't need to do that, or there's another view that the harness actually does matter, and this is kind of part of the harness apparatus. Do you guys have any views on that with your working with Astra, especially generations of models and how much do you have to kind of input or how much the harness matters versus not? I mean, yeah, I guess there have been some, like, funny quirks like this that just come from, like, exactly what you asked the model to do, basically.
Starting point is 00:36:08 Yeah. Like, in this case, what the model was asked to do originally for codes was to improve the bounds by, like, some exponential factor. So it really, like, shows up in this, like, leading constant up here. And, you know, it improved the bounds, and it didn't try to push things too much further. Like, sometimes you see it do, but sometimes it just doesn't bother. But yeah, you know, you just ask it again and it goes further. So it wasn't like a capabilities issue. It just didn't feel like it at the time.
Starting point is 00:36:35 Do you call that judgment or like what is the? Because like there is a, yeah, what do you call that? Models tend to be pretty task oriented. If you tell it to do a task, it accomplishes that. It's pretty happy. So yeah, the task oriented in this, it's like, but do we expect that level to kind of ascend up to? It's not that they will be less good at,
Starting point is 00:36:58 being task-oriented, it's like, they'll ascend to the level of like, okay, no, let's go in this direction. You'll have the judgment, too, because you guys have the judgment, too, like, okay, this is pretty promising. It looks like you're using a lot of representation theory. It doesn't seem like there's a limit so far, but it doesn't have that context yet. But, like, I guess what I'm trying to say is, like, this one, it's hard to, maybe harder to extrapolate, but from, like, previous generations, when you had to give it more, maybe prompting, more of that harness work, but eventually probably had to give it less. So it probably gives you some confidence that there's this like really fast ascension.
Starting point is 00:37:31 And do you see, yeah, like, where? Somehow solving a harder math problem is, like, you have to solve many smaller, like, somewhat less hard math problems. And the fact that the math problems are getting harder is kind of an indication that the model is able to take on more and more work in like a single continuous unit. And I think that's the thing that looks very promising. Somehow, like any of these solutions, it's not like one idea. And then you're kind of home free.
Starting point is 00:37:55 You need several, you need several pieces to kind of, interact and talk to each other. The model doesn't come up with all the ideas at once, right? It doesn't pull everything out in an instance. So kind of the fact that it needs to sort of see how this piece interacts with another piece that's kind of like solving a problem in itself or piecing together many problems in itself. It could just be that, okay, when you're telling it, okay, push this even further, that was
Starting point is 00:38:22 of the same order of like magnitude as like all the smaller things it's solving as well in between and so you don't think that this kind of like a privileged direction. It's just sort of like, hey, let's give it like one more help. Or you actually think that there's, I guess what I'm trying to get out of a bigger question is like, is there a good sense of like, you know, taste? Because like when people talk about, for instance, how well the models are getting at like doing research, for instance, and that's what we want a little bit of RSI. And sort of like, there's surprising things about how that improves. And then there's like the, oh, you know, maybe right now it's at a level still like a junior researcher. It's like not really asking like the right problems. And so I'm just
Starting point is 00:39:04 trying to get like a maybe a sense of like where you're seeing that progress through the model advancements each generation. I mean, what is taste even? Yeah, I think I tend to be pretty utilitarian in my view of taste. If you're able to solve problems faster by making better judgments, I think that's like the best like general proxy I have for a taste and somehow the fact that it's solving harder problems means it has kind of by definition means it has better taste. I think there are these no yeah I think occasion because they are task oriented you do occasionally get these these symptoms of like oh it clearly has made a breakthrough it kind of understands it's made a breakthrough and then it doesn't kind of push all the way to the limit because that's not what you asked but that seems yeah that seems seems rather minor. compared to the state of progress we've seen so far. Okay, yeah. I think that's pretty clear.
Starting point is 00:40:00 Like, I think it's like maybe you're liable to get confused if you're trying to like do a concrete long horizon task and show taste kind of at the same time. But like, you know, if you have like one model that's responsible for taste and one model that's responsible for going out and like, you know, working for a long time at solving a hard problem, kind of as the like, you know, underling of the supervising AI,
Starting point is 00:40:30 I feel like that's kind of going to be fine currently. Oh, interesting. Because that is like saying that these two things are somewhat, if not separate, at least they shouldn't kind of pollute each other's context, which is a little bit, I mean, it could be potentially like a strong, stronger statement than, I guess, you know, it's just kind of interesting, because it might just be, like, to your point,
Starting point is 00:40:55 it's, you know, let's take the huge hillitarian answer, is solving harder and harder problems. It's doing a lot more than just, like, you know, brute forcing something. It's making choices. It's like pruning, you know, a vastly large space of possible paths into something that's like really, you know, it's both tractable,
Starting point is 00:41:12 but then ends up being like it's a diminishingly small path within that space. But, like, having, like, why would, would be like a separate model, is a separate generation or something that's a different version of the model that would contribute to taste, or maybe that's totally like it's too abstract, doesn't make any sense. You know, we should just let the actual. This might, like, the related question would be like, you know, what is the thing that gets us to a better version of intelligence, the harness in the model? Is it just the model?
Starting point is 00:41:43 And it's like we see this in, you know, at least in applied AI or, you know, startups where it's like it's a continual battle. of like you need the harness, but then the harness adapts very poorly to a new model because sometimes like a very, very minimal harness is still the best way to expose to the raw power of the model. But then now we also have these like training regimes where we require the harness to be trained with, I mean part of this is to keep things more proprietary
Starting point is 00:42:08 and harder to, harder for other people to use it. But I think partially it's maybe actually that it helps have more control on like the reasoning traces you care about. It's a long, rambling way of saying it's like, yeah, I don't actually. This is so interesting to see how the models have gotten better at math. And maybe something that's like very abstract and hard to describe, like, taste is a way to tease out, like, what is actually necessary here. I think my only, like, non-tremial thought here is that, like, when you're working, I mean, just when you're doing any tasks, occasionally you get pigeonholed and you, like, work really hard and just having a friend look over your shoulder and be like, what are you doing? And then just like, just having that one bit of, like, step back for 10 seconds.
Starting point is 00:42:49 Like, this is often very useful. Yeah, yeah. I mean, I see no reason why humans would be so different than models somehow. Yep, yep. Having, or models would be so different than humans. Having a few humans working together is often more powerful than just having one. Yeah. It's like in this kind of collaborative thing, you actually, you kind of, yeah, artificially created it, but it's very similar in dynamic.
Starting point is 00:43:11 But I think a lot of taste is also, like, having a sense of what problems you or like some method you have in mind are going to be good at solving. Mm-hmm. Like, it's, I mean, certainly there's. some amount of like absolute aesthetic point, right? But there's also just like, you know, having a nose for what you might want to pursue because you'll be able to make progress. And, you know, I think for that, like,
Starting point is 00:43:37 there's, you know, you would expect that as a side product of being good at completing tasks you would get there sort of, right? Let me know if we still want to do like a section on soft at groups because I think, you know, up to you guys, it's definitely super interesting. So maybe the first question is what is a group? Let's remind ourselves.
Starting point is 00:43:56 So a group is a set of elements with some multiplication operation. And basically, this is how mathematicians think about symmetry. So you're like, basically like if G and H are elements of elements of your, group, then GH has some other well-defined element of your group. And you have like associativity and you have an inverse. So for every G there's some inverse. And there's some like specific element in the group that is kind of the identity. Okay, so it's some like abstraction of like composing operations.
Starting point is 00:45:06 So these could be like numbers, they could be like multiplying matrices. They could be like rotating something, which is a special case of multiplying matrices. And a group is so thick. If, well, there's some, you know, precise definition. but you know roughly it means
Starting point is 00:45:34 it so I should say like groups that can be finite or infinite so like you know if you have like a square like all the rotations of it
Starting point is 00:45:46 form a group with like four elements if you have like a circle then the rotations form a group with like uncountably any many elements and so phic groups are either finite or countable you should think of
Starting point is 00:45:59 as being countably infinite, so there's like the same number of elements as like the integers. And if it's suffolk, if in some sense, it can be approximated by finite groups. So we didn't know if there was a non-sophic group. So the result that Astor proved
Starting point is 00:46:17 is simply that there exists a non-sophic. Yeah, and without like, I mean, we can, you know, before going to that prove, It is like, you know, I feel like a lot of the programs of math is like, okay, we are such finite creatures. Let's see how well our finite approximations are, you know, do. And in this case, especially for the countable case, it's like, maybe you'll be relating it to like the Aldous Leon's thing. It's just like it helps kind of anchor the picture of like it seems like such a, I mean, it's a nice result if it were true, but it's not.
Starting point is 00:46:54 And it seems almost like reasonable. And so, yeah, I actually didn't, um, didn't, uh, go. I would love to hear the explanation of how it found a counter example. Yeah, I mean, I would say that, like, you know, the hope that there was no non-Sophic group, so every group has this kind of approximation, like maybe this is sort of like people hoping that there's a miracle because it turns out that groups like this have a lot of nice properties because you can run certain proofs for finite groups and then, you know, kind of approximate them in whatever way, the definition of being Sophic
Starting point is 00:47:32 lets you approximate them and get the result. So there's this notion of being a surjunctive group. So there's some fact that any group, which is Sophic, is also surjunctive. Surjunctive is some property of like dynamical systems on the group. And I guess the original question was, whether every group is surjunctive. This is some question of gotchalk from the 70s.
Starting point is 00:48:05 And this fact that follows this pattern of prove it for finite groups and then do this approximation is what motivated the question about if there's a non-sylphic group. Yeah, maybe I'll say a little bit about this. I'll just lion's conjecture. Yeah, sure. Yeah, yeah. So I guess I had heard of this a little bit beforehand because there's a related stronger conjecture. in probability that was made popular by Aldous and Lyons.
Starting point is 00:48:36 This conjecture, roughly what it says, is like, any infinite graph with some nice property called unimodularity. A unimodular random graph can be approximated by large finite graphs. So maybe the way to like explain what these kinds of things are trying to say without getting into technical weeds is to say what they mean about the integers. So, so like how would I draw the integers as a graph? So this is called like heli graph. You're just going to connect nearest neighbors.
Starting point is 00:49:35 Okay. So there's some kind of canonical way in which this is like the graph that represents the integers. Okay. And there's some sense in which you can approximate this by finite graphs. Why? Well, if you look at integers mod n, then you kind of get the same picture, but you have like a big circle instead of an infinite line. And the point is if you look at any point here and any point here, like in nearby things look the same.
Starting point is 00:50:14 You have to go very far away to kind of see this global geometric structure that you have a circle and not a line. And in fact, the integers and integers mod n are both groups just by like adding numbers or adding numbers mod n. So these integers mod n are like sophic approximations for the it full integers. So, like, this approximation is kind of why the integers are a Sophic group. So, um, so the sophisticity, the statement that every group is sophic is sort of a generalization
Starting point is 00:50:49 of the fact that you can do this approximation with groups. And this, uh, Aldous Lyons conjecture is kind of a broader conjecture that, like any network, you can do this. Um, and it, like, you don't require as much algebraic structure roughly. So it's, it's kind of a broad. our conjecture. So this conjecture was disproved earlier, like two years ago, and it was kind of a really tour-to-force work. It was like 250 pages building on another 200 pages. It used as like quantum complexity theory. So it really builds this like, you know, very complicated bridge and like,
Starting point is 00:51:26 you know, I think not really people could understand this. Right. So since this is a stronger conjecture, the disproof, like, is weaker than disproving this statement that all groups are Sophic. But it turns out that the direct proof that there's a non-Sophic group was, like, much shorter and easier than this really amazing disproof of all those lines conjecture. It's like, like, 15 pages, maybe. And it doesn't have any of this, like, very complicated connection with quantum complexity. It just kind of stays in group theory land. I mean, it uses some important, like, you know, existing results by other mathematicians, like Coon and Coon and Tom,
Starting point is 00:52:07 but it's like, it's like a very reasonable, normal kind of proof. Yeah. And it's kind of spelling maybe the obvious, but like the connection between the Sophac group, statement is just you take the Kali graph, and that's the one that is like what they use or for the Aldous Leon?
Starting point is 00:52:25 Yeah, yeah. And so that's why it's like a, you know, a subset of. Right. So, yeah, basically what happens is, yeah. So for a group, you can take exactly a Kali graph. So you take some like elements that like generate the group
Starting point is 00:52:37 and you kind of connect elements that are adjacent. So in this case, like this is a Kali graph of the integers. So right, so when you do that from a group, you get like a deterministic graph. Right, you just get like a single graph. You fix some set of generators. So this conjecture is stronger basically because it allows a broader set of graphs that aren't deterministic.
Starting point is 00:53:00 It allows them to be random, but have some extra, you know, you know, modularity property that constrains exactly how it can be random. But, yeah, basically that's the difference. Like, here you kind of have to give a deterministic network instead of a random one. Yeah, anything kind of interesting, surprising about the results. I mean, you mentioned some things, which is like it stayed within group theory, the techniques. I mean, I think maybe it's like a nice example of this general pattern that theorems produced by AI have generally been like the proofs are pretty short,
Starting point is 00:53:41 generally. They're like... Like with a counter examples so far. Yeah, but this one, it's like, okay, it's sort of a counter-example, but like there's some, you know, there's some like stuff you have to do to analyze things. The difficult part here is that, like, the property of being a SOFIC group is not so easy to get your hands on. So you have to find, like, a concrete way of saying, like, producing a way of saying this group cannot be SOFIC. And like, and the proof is actually, it's very short.
Starting point is 00:54:12 It's like, it's almost a, it's a cognitive argument. But it's like a very delicate commonatorics argument. Like, somehow you need to both have the right statement and know what piece of the literature and then execute it correctly. And that's very nice. Like, the difficulty in this problem is that. that like it's just a really, really hard. It's like very hard to get your hands on like being approximated by any possible finite through.
Starting point is 00:54:33 Yeah. It's like what is happening at that countable infinity that's like resisting this approximation? Like do you guys kind of, did it give a sense of, like, when you're like, okay, Astra, explain to me. Like what is the? Oh yeah, we did that. Yeah, yeah, yeah. Yeah.
Starting point is 00:54:46 What was the good explanation you got out of it? I think there's like some concrete like combinatorial obstruction. Basically it's like, it's hard to explain, but there's like, some concrete combinatorics of structure, which if you read the previous papers, you realize that that's what they couldn't rule out. And Astor found a way to kind of say, okay, no, no, if you add this one extra algebraic fact, this, this like weird conspiracy can't happen. It's like very clearly trying to rule out a conspiracy in the sort of previous authors had implicitly written about. And those were the actual suspects.
Starting point is 00:55:19 Yeah, yeah. It turned out. So they were sort of on the right track. And then this did the last mile of, well, whatever, however you quantify that. But I think it's like, like a year ago, I would have been very surprised to learn like all of these AI proofs are like very short and elegant. Like they're, you know, you're kind of like afraid that they're going to like generate all these thousand page things. Yeah.
Starting point is 00:55:45 Yeah. I don't know. I would never be able to understand it. But it's been kind of the opposite. Yeah. Like only humans can generate like 200 page proofs right now. Yeah. Well, and also I was like asking him like, if you do that.
Starting point is 00:55:55 post-mortem, it ends up usually engendering more mathematics. Because when you do that with humans, like that's what, you know, breeds new mathematics. So maybe, maybe if you kind of alter the prompt a little bit and be like, how would you, you know, generalize this or something? Like, yeah, I don't know if that's been a technique for you guys to like have it explore and exploit what it has already developed. Well, there has been some, there has been follow up on this already, actually, by Kun and Tom, who this was always built on. So they like. So that community is coming. Yeah, yeah, yeah, which is kind of what we're hoping. Yeah.
Starting point is 00:56:28 You know, we don't want to be, you know, writing lots of follow-up papers ourselves, but if there's some interesting follow-up that, you know, it's like we're very, very happy that there's some follow-up building out these ideas more and giving, like, more examples of non-Sophic groups in this case. Yeah, well, actually, maybe that's a great segue into, like, how, you know, what's the ideal way that this is being taken up by the math community? Because I feel like there's a spectrum of answers from working mathematicians, sense of like some, you know, probably most at this point are like, okay, AI is obviously doing some
Starting point is 00:56:59 non-trivial stuff. It would be a disadvantage not to admit that in my workflow. I've definitely heard some stories where people are kind of, you know, would find it hard to either take AI as a co-author or like how do you even do kind of attribution this way. But I don't know, like what, maybe to paint the more optimistic picture, so you're saying you want the mathematicians to be building on these results, it definitely generates a lot more results to be verified. So it puts pressure on the community and the profession. Like, how do you, how do you kind of expect the evolution of kind of uptake and collaboration with mathematicians? I mean, given that the fact that the models can produce sophisticated mathematics means they can help you understand, like,
Starting point is 00:57:45 sophisticated mathematics. I mean, like, I don't know, occasionally, like, I enjoy looking at the archive and I want to understand some proof. And, like, I could read the introduction. But in practice, it's just much faster, take the PDF, put it into my favorite model, and then get an output of like, what is the rough proof strategy? And so I have this, like, along, I mean, of course, models are going to help us produce exponentially more mathematics, but they also make it much easier to absorb it.
Starting point is 00:58:13 And right now, okay, it's still a bit of a challenge back and forth, but I think it's, for me, at least, much, much faster at understanding. It's much, much faster to understand it, mathematics with a model than without it. It's helping solve the problem it creates anyways. Yeah, I feel like that at least. And it's, you know, I don't, I don't view it as like creating much more wrong.
Starting point is 00:58:35 But again, like, I don't have such, you know, high stakes in like, okay, I'm going to get, I'm not going to get tenure, et cetera. So like, I agree, like making it more accessible. Like, if I'm not spending so much time absorbing an area, I can like put it into chat GPT and then expect to, I mean, you guys have an even more powerful model, hopefully releasing, for other people to enjoy as well. But like it's, I think like the positive version of that is actually more people can participate in mathematics.
Starting point is 00:59:00 It's like people might be coming with other intuitions and they could actually maybe generate good mathematics. Is that sort of like closer to the vision of what you're hoping this is, you know, pushing towards? Or like what things do you think mathematics should be wary of to kind of adapt fast enough to take advantage of AI? Yeah. I think certainly there will be a lot of changes, right?
Starting point is 00:59:24 Like, I guess in math, like, there are a lot of things that are kind of important for, for, like, a given result, right? You need someone to come up with it, but you also need people to understand and absorb it and, like, you know, internalize it enough to do more with it and, like, figure out where it fits into, like, humanity's understanding, right? And like a couple of years ago Like the proving the result was like So hard that kind of the other stuff was just kind of coming along for the ride right
Starting point is 01:00:00 You know like if you if you manage to like prove this thing yourself You're automatically going to understand it quite well You're kind of responsible for like maintaining it in some sense and like you know explaining it to other people And yeah now this kind of what was the main bottleneck before is kind of much less of a bottleneck. And, you know, these other kind of constraints come into play. So, yeah, the, like, optimal structuring for, you know, organizing the knowledge could look rather different.
Starting point is 01:00:37 Yeah. How does that look? I mean, does this make the field a lot more kind of empirical? Will people do sort of the hard, like the first thing that was scarce, which is like all the reasoning and then more. I mean, not that it's like a bad thing to make it empirical, but it's almost like it functions as a very different discipline. Like a lot of the fun stuff is understanding, you know.
Starting point is 01:00:57 And so understanding, communicating, maybe assembling, having still the human taste, does that sort of remain rarefied? And that's how, you know, current mathematicians need to adapt and reward, you know, contributions. Or is this too much of a caricature that's like something else? I think certainly understanding how to put, as we get more and more mathematics, put it in like a proper framework and sort of how sort of like being able to explain it to other humans so that they can also appreciate it. I mean, so implicitly we valued this, but it was usually because you were the person proving the results that gave everybody else the understanding. But I think increasingly would be a function of like you're sort of helping, you're the human who can sort of give this understanding to other people and sort of help them with it. I think that more, sort of, sort of, that communal understanding will, I think, become.
Starting point is 01:01:48 It was much more implicit in how we viewed math and gene, but I think it will be an increasingly more explicit and valuable part of the subject. I mean, a nice thing about math is that the ceiling for difficulty of a math problem is pretty high. So even if, you know, kind of, even if AI continues getting exponentially better at math, like it might, you know, plausible will never solve something
Starting point is 01:02:13 like P versus NP. And it could be. be that like the the field kind of becomes more you know attached to like like these big mysteries and less to like smaller mysteries that are more like routine now yeah yeah I think that's a positive vision of that I mean also like I don't know there are things I spent like months or years of my life wondering about not getting to know and hopefully we get that yeah some portion of them I'll get to know the answer to it. I'm pretty happy about that.
Starting point is 01:02:46 No, exactly. No, I'm excited about this, like, renaissance of results and understanding, and I feel like, I mean, this is such an infinite, you know, field. Like, okay, no pun intended. But, like, it's just, like, it's just, there's so much that you can actually create here. So, I mean, especially for somebody like me,
Starting point is 01:03:06 who's not going to have the time to actually, like, practice mathematics. Now there's, like, a lot more that you can actually do in the activity of now. So, yeah. Yeah, I think the, like, The ability of someone who's not working on math is like their literal job all the time to like understand what's going on and like, you know, learn about some of the mysteries they might have wondered about will go up quite a lot. Also, you know, if you're like, if you're working on something that requires some math, you know, suddenly you don't need to like find a world expert on this topic to be able to, you know, use it in your own work.
Starting point is 01:03:40 Sorry mathematicians. No, it's true. I mean, I think there was just like a dearth of actual, like, people who could do that. And so I think this is helpful. Maybe it's helpful for theoretical physics, like we'll see. But a lot of other applied areas as well. It'd be nice for the world if applied mathematics on a lot faster. Yes.
Starting point is 01:03:57 I mean, I'm over that. Well, thank you guys for joining. This is a lot of fun. And I'm, you know, just so excited for how much the models are advancing. So maybe we'll have you guys back soon. Yeah. Thanks so much for having us. Yeah.
Starting point is 01:04:11 Thanks for having us. Thanks for listening to this episode of the A16Z podcast. If you like this episode, be sure to like, comment, subscribe, leave us a rating or review, and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on X and A16Z and subscribe to our substack at A16Z.com. Thanks again for listening, and I'll see you in the next episode. As a reminder, the content here is for informational purposes only.
Starting point is 01:04:43 be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see A16Z.com forward slash disclosures.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.