The a16z Show - Daniel Litt: The Mathematician's Guide to AI

Episode Date: September 1, 2026

a16z’s Lisha Li sits down with Daniel Litt, Assistant Professor of Mathematics at the University of Toronto, to unpack AI's rapid progress in mathematics, what today's frontier models can actually d...o, and what they're still missing about the way mathematicians think. Daniel explains why some recent AI-generated results are genuinely impressive, including an autonomous solution to the Erdős unit distance problem, but argues that solving problems is only one part of mathematics. Today's models can grind through calculations, combine known techniques, and search enormous spaces, but still struggle with intuition, theory building, identifying the right questions, and developing the kind of big-picture understanding that drives much of mathematical progress. Lisha and Daniel also explore how AI is already changing mathematical research, why an explosion of AI-generated papers could distort academic incentives, and what happens if researchers outsource the work of thinking rather than use AI to deepen it. Ultimately, they ask a question that extends far beyond mathematics: as AI gets better at intellectual work, how do we make sure humans keep getting better at thinking too?   Resources: Follow Daniel Litt on X: https://x.com/littmath Follow Lisha Li on X: https://x.com/lishali88 Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Transcript
Discussion (0)
Starting point is 00:00:00 The goal of mathematics is not to produce mathematics papers. It's to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me, that's pretty unsatisfying. Comparing anthropic with open AI, do you detect any differences in how that is similar to human reasoning? They definitely are not good at it autonomously. But with some hints, you can kind of get them to do something interesting. A lot of progress in mathematics comes from letting, you know, a thousand different flowers bloom,
Starting point is 00:00:26 and people pursue their own curiosity, and then, you know, the boundaries of knowledge of things. boundaries of knowledge expand in some kind of fairly uniform way. What has been the most impressive results so far? My favorite fully autonomous result by an AI so far is the solution to the Earthish unit distance problem. There was some lemma I wanted to prove none of the front-tier models could do it. So I worked out a ton of examples on my own and I realized, oh, well, maybe here's some reason why it could be true. Once I had that statement, the models were able to very quickly prove that sort of a better statement.
Starting point is 00:00:57 How should the mathematics community best adapt and benefit from this? AI can increasingly solve math problems that would challenge professional mathematicians, but solving a problem isn't necessarily the same thing of understanding it. In this episode, A16Z Infra Partner, Lisha Lee, sits down with the University of Toronto mathematician Daniel Litt to separate the headlines about AI and mathematics from what the models can actually do today. Daniel explains why recent results,
Starting point is 00:01:27 have changed his views of AI, where frontier models already resemble human mathematicians, and where they still fall short, particularly when it comes to intuition, developing new theories, and even figuring out which questions are worth asking. They also explore what happens to mathematics when generating a proof becomes cheap, why academic incentives may need to change, and how mathematicians can use AI without outsourcing the understanding that makes the work valuable in the first place. And more broadly, they ask, What's what mathematics can teach us about working with AI as increasingly capable models move into every knowledge profession.
Starting point is 00:02:06 I am so excited to have you on Daniel. And so Daniel, it is a professor of mathematics at the University of Toronto. Toronto is my hometown, so also very exciting. But the thing that is most special here is Daniel's an actual practicing mathematician. And in addition, he's been incredibly vocal about his evolving views of AI in math. And so I feel like every, if I just don't check in with you, you know, like in two weeks, something, you know, different has been revealed, and then you're very kind of like, what do you call it? I have a lot of opinions. You have a lot of opinions, exactly. So I want to get into that.
Starting point is 00:02:38 So, I mean, one of the things that I'm most interested in is not just like a discussion of how the capabilities of advance. I feel like in math, that's definitely the headline, et cetera, but also you've been very thoughtful about how practicing mathematicians should respond. And so that kind of gives us a chance and opportunity to talk about actually what is special about math. It's not just like, hey, AI has been really making progress here, but delve into what actually mathematicians do.
Starting point is 00:03:02 And so maybe like with that arc in line, we can start with what has been the most impressive results so far, given all the recent progress for you. And then maybe, yeah, we'll kind of take it from there. Yeah. So, okay, so there have now been a lot of results. Some of them produced autonomously, some produced like semi-autonomously, some who's like where the AI contribution is just not at all clear. They're in a lot of different areas. So anything I say is kind of, you know, I can only really comment on things that I can have some expertise on. So it's quite possible that if you talk to a different mathematician, you'll get different answers here. So my favorite, like, fully autonomous result by an AI so far
Starting point is 00:03:38 is the solution to the Erdus Unit Distance Problem, which I think was announced in mid-May. So what I liked about that is it seemed to me that, like, it was in some ways a little bit creative. So I think some of the results we've seen have had kind of the flavor. You kind of take some known techniques and apply them in maybe a clever way, or you know they've kind of been at some kind of results I would characterize this, like, last mile, like where some recent work was done quite deep work done by a group of human
Starting point is 00:04:06 mathematicians and then the AI took the final step. But yeah, with this sort of distance problem, I think it was something where the result was like a little unexpected. So first of all, my sense was like that people working in the area thought it was true and then there was a counter-re sample found. But then also it like brought in some techniques from another area. I think those techniques were like not especially like deep or new. They were sort of classical ideas from the 60s,
Starting point is 00:04:31 but they were new to this area of studying, you know, point configurations in the plane. And so that was pretty cool. And then afterwards we got to see, like, it was kind of fruitful. So a bunch of mathematicians took those ideas and used them to find counter examples to a bunch of other interesting open questions. So, for example, like the song product conjecture for the real numbers. So that's at least what way I like to think about how cool result is. Like you look at it post hoc and you see like, oh, well, were whatever new ideas that were
Starting point is 00:04:59 introduced, if any, like, kind of useful to do other things. Did they improve our understanding of something? And I think that's maybe so far the main example I know of a result of that form. Yeah, I think it's really meaningful that you're commenting on this because that result came out to your point in May. And there's been so many headlines so far. And it's kind of probably hard for somebody who's not a practicing mathematician to appreciate the differences in these headlines. And so you kind of already started laying out sort of a taxonomy of like what is different in that proof. And so it would be kind of interesting maybe to use. that as kind of both an excuse to talk about where you sense the model differences are and what
Starting point is 00:05:33 like mathematicians actually do. So in this case, I mean, the most kind of maybe naive understanding of what mathematicians do is that we're pushing around symbols in a logical manner. And this is why RL is so successful at this because you can kind of both verify it somewhat cheaply compared to other domains and then also because the rules are quite illegible. And so you can kind of like, if you're superhuman at that, you might be good at math. But I think that, of course, betrays most of actually what is interesting about mathematics, which is perhaps I think you said this as well, but I think anybody who's tried to do math is sort of like, it's about the understanding and getting at truth, remaining confused, and developing intuitions. And I mean, the tool to do that stuff is,
Starting point is 00:06:14 of course, in having really strong abilities to push out logical implications. But maybe if you can kind of speak to like, when you say it's most impressive and creative, like decoupling just the inhumane maybe feats of just like logical implication from like where is it being creative what is it helping engender in terms of like mathematical activity as well yeah okay so first of all i mean i have characterized as to this like inhuman in some way i actually think the argument was very human oh i'd love to go into that i or sorry opening i like released some chain of thought yeah it was very recognizable it was like if i tried to imagine like my chain of thought and trying to solve a problem like it might look kind of like that yeah we haven't seen the raw
Starting point is 00:06:56 a chain of thought would be maybe that way. Maybe they cleanse it a little bit, yeah. Yeah, or maybe the model like to swear a lot in the middle of the chain of thought and then that up or something would know. But at least the summary seems pretty recognizable. And I would say that's actually like kind of typical of most of the results that I've studied. Like they don't seem inhuman at all.
Starting point is 00:07:14 They seem absolutely like something a human mathematician could produce. And they're like typically understandable, if not so well written. If you just look at problem model output, it's not like there's some move 37 or whatever. It's like a human mathematician doing math. It's like a human mathematician doing certain types of math. So, like, they're definitely, like, the models are still, like, not very good at some mathematical activities. And here I don't mean, like, by field, but just, like, certain things you do when you try to solve a problem, the models don't seem to be doing. But certain things they're very bad.
Starting point is 00:07:43 So the ways they might be a little bit unhuman is, like, they don't get tired, they know a lot. But if you actually just read the final output, it doesn't seem kind of inhuman. It's interesting because, I mean, obviously, we can't get too much of information. from the labs who are producing these models, why of what the training recipes are or how they're kind of advancing in reasoning. But at least one of the things we do know, and I think kind of opening eye spearheaded this,
Starting point is 00:08:06 is just like reasoning in natural language is actually what they happen to scale up, and it's not actually pushing a lot of lean verified proofs as like the training corpus. And that's kind of amazing. On another side, what's kind of interesting is it's not clear that a lot of the mathematical training data if they use that to a large extent at all
Starting point is 00:08:26 is reflective of how mathematicians think, if that's fair, because a lot of it is not legible as traces, right? Like most of the papers are crisp and, like, polished. The textbooks certainly just show very little motivation of how something is developed, which is why it's usually easier to kind of follow a research direction by actually talking to the researchers and how they're thinking about it.
Starting point is 00:08:45 So I'm kind of curious, before maybe even going to the taxonomy as an excuse, like, when you're examining these models and their results comparing anthropic with open AI, Do you detect any differences in how that is similar to human reasoning? And then also if you have any comments on insights on perhaps why natural language scales so well that way, even though it's... Okay. So first of all, I really like your point, by the way, that they're mostly doing natural language reasoning rather than lean. Like, I think that that suggests to me that like these... You hear a lot of people say, like, math is a verifiable domain.
Starting point is 00:09:17 Like, that explains the progress, whatever. Like, my sense is that because they're primarily scaling informal reasoning, like, probably to take... techniques are going to generalize to other domains pretty well. That's just my guess. Okay. You asked a little bit about Claude versus Chachabee. My sense is that they're pretty similar in terms of capabilities. I've played around a lot more with Chetabee than Claude Fable, but it seems like there's a lot of cases where Open AI will drop a solution to some problem and Anthropics will say, oh, you know. We also. Exactly. It actually seems like they're solving a very similar collection of problems. And it's like kind of a relatively small portion of what human mathematicians do.
Starting point is 00:09:53 So yeah, one thing that is interesting is, like, we see the solutions have a certain flavor, right? Like, there'll be things that this isn't surprising. Like, there'll be things that rely on the model strengths, like their ability to ground out along computation or pull together kind of technical ideas from many areas, or like maybe many papers that, you know, even my mathematician might not have bred it. But they're not, like, they seem like weaker in things like intuition or, like, having some big picture point of view. Like, a lot of what I do as a mathematician is, like, I have some kind of philosophy that's, like,
Starting point is 00:10:21 very non-rigorous. Like maybe I think this thing is kind of analogous to this thing. And then like a lot of what I'm working out is like figuring out how to make that precise and like, you know, trying to measure the extent to which I've succeeded in understanding that by like, can I solve a problem or whatever. Like can't find an interesting phenomenon which I don't understand it. But now I can understand it. And like so far I haven't, you know, you can see even the best results the models are producing.
Starting point is 00:10:45 You don't see that much of this kind of reasoning. It's more like they're very, very good at applying some known techniques. Which is, to be clear, that's like a very powerful thing to do to be very good at applying like all known techniques. There are mathematicians who have had great careers doing very high quality work of that flavor. And I think a lot of what the models are producing is high quality in that way. But it's like some kind of fairly narrow band of what mathematicians care about. So far, I do think there's like signs of both, you know, all the frontier models starting to be able to do more fuzzy things. So, like, I've tried to get both Fable and Chad GPT, 5.6, Saul, I guess, to do some kind of theory building.
Starting point is 00:11:31 And it's like they're not good. They definitely are not good at it autonomously, at least with, like, whatever scaffolding I've set up. But with some hints, you can kind of get them to do something interesting. You know, when you do, when you give the models hints, it's always a little hard to tell, like, what part is from the model and what part is from you. But my experience is that, like, if they can do it with, like, 100 bits. of hints or whatever in six months, maybe they can do it without hints. Yeah. You know, I do think there's signs that they're also kind of somehow picking up some of this like implicit and unwritten mathematical knowledge.
Starting point is 00:12:02 Okay, I would love to, so go into the intuition part and where it sucks at, to put it in a very basic way. But you actually mentioned a small detail, which is that, you know, from the public you've gleaned that Anthropic and Open Air probably neck and neck, but you personally are using a lot more chat GBT, like, why is that? Or like, 5.46. Well, I don't know. I mean, I just, I think it's just in our show. Like, I have a certain. Pretty sure of habit. Yeah. It got, one thing is that chat GPD got better at math earlier. So, like, for a long time, the cloud models were just, like, not useful for research math. And I think maybe around opus 4.5
Starting point is 00:12:44 or opus 4.6, they, like, more or less caught up. Yeah. But, you know, some experimentation suggests to me they're pretty neck and neck. And so for my own work, you know, except when I'm just experimenting, I must have just stayed with one. Yeah. As people outside the labs like us, it's really interesting just to compare how they differ on the frontier. And, I mean, you know, to your point, it might be a little bit of momentum.
Starting point is 00:13:07 I do think that, at least from my anecdotal experience, 5.6 has been like a lot more clear in exposition. And it's just like, there's a little bit, and this might not be true. I mean, obviously the models are just incredibly jagged at the frontier. But, like, it's in the explanations of results to me, I always find that 5.6 is giving a more accurate theory of mind of what it assumes I know and don't know, whereas Pable might be explaining something very trivial, but then just, like, jump at like, well, you know, obviously you should know these things. Yeah, I find they're both pretty bad at theory of. Okay, great. So then seeing, from your eyes, you're probably asking much deeper questions.
Starting point is 00:13:47 Okay, so the thread that I really wanted to pull on was when you're talking about the models maybe starting to get better intuitions or even theory building. So maybe before we even dive into that, it would be useful to kind of talk through like what is your primary, you know, activity as a mathematician, especially in your area of like algebraic geometry, probably has a very different flavor than a combinatorialist or, you know, some other areas. So if you can give a little, maybe brief way of the land and then kind of explain what you're, what your mathematical activity was pre-AI and then maybe how it's kind of like changing with AI.
Starting point is 00:14:23 Yeah. So, I think there are a lot of different kinds of mathematicians. There's like a lot of different, you know, spectrum on which one can put in mathematician. So definitely a lot of mathematicians like solving open problems. And then I think, I'm one of those. Like I like to solve an open problem that's, I think of myself as a problem solver as opposed to like one other. axonomy you could have as a problem solver versus a theory builder. At least for me, the point of an open problem is it's supposed to measure your failure to understand something. So it's kind of like a benchmark. Right. Like, you know, one problem I really like is the gross and decaps peak or aperture, whatever that is. It measures something about our failure to understand differential equations. So it's like there's some very basic object, we would like to understand.
Starting point is 00:15:15 If we can't answer this concept, we know we don't understand it. Okay. So in practice, like, how do you get a problem which is supposed to be measuring something you don't understand? Well, like, of course, you try to understand the thing better. And in practice, what that means is, like, well,
Starting point is 00:15:31 you try to find the smallest situation where you can't understand something and fiddle with it and then you stare at like once you win, you stare at what you develop to win and try to turn that into some theory. So that's one thing you might do. You might try to solve a problem, and in so doing, develop some kind of new understanding of the situation.
Starting point is 00:15:50 You might also just, like, have some feeling that, like, this thing is related to this other thing. And, you know, you might start building a table, like, oh, this property A is related to property A prime, property B is related to property B prime, and so on and so forth. So, for example, in my work, a lot of it is motivated by some analogy between comology of algebraic varieties and representations of fundamental groups. So that's some fancy stuff. But it's just that this analogy is very, very fruitful. And like really any phenomenon that appears on one side, you can find an
Starting point is 00:16:23 analog on the other side. And so trying to realize that dream has led to a lot of beautiful mathematics the last like 30 or 40 years about people like Carlos Simpson and Chiro Machizuki and others.
Starting point is 00:16:40 And so here it's really just like someone noticed like here's an analogy. And then that analogy has led to a huge amount of developments. There's not really like an open problem at the end, although of course, like as you develop this, you come up with lots of open problems. There's just like a philosophy that you're trying to realize. And that philosophy is like super not rigorous, actually.
Starting point is 00:16:58 It's not like symbol pushing at all. Yeah, so that's another kind of activity that I like. Yeah, beyond that, no, a lot of what you'd, do when you try to, you know, a lot of the activities, you're, like, you're actually trying to figure out what the right question is, even. Like, here, you have some object, you feel like you don't understand it. And, like, figuring out what you don't know is, like, actually a very challenging thing to do. So, so, you know, there's, like, a list of, you know, you can go online and list open conjectors or whatever. And this, like, doesn't really capture in a lot of ways what we don't
Starting point is 00:17:38 know. Like, often finding the conjecture is, like, really, really hard. So, I don't know. A good example of this is like the Birch and Swinertrandyre conjecture, which is one of melanium problems, is this beautiful relationship between L functions of elliptic curves and the set of solutions to the corresponding equations of the rank of the group of solutions, what if that means? It was discovered by like, it was like the first big data conjecture. So Bershen Swinertan Dyer had like found all these statistics on elliptic curves in the 60s, like it was one of the first ever computerated bits of mathematics. And they like graph these statistics and they noticed that, you know, some, so the slope of some line on the graph
Starting point is 00:18:17 was related to some other algebraic invariant they knew and that was the source of this conductor. So a lot of the time, you're just like working at examples and like trying to, like, you're doing kind of science, like, you run an experiment and you try to figure out an explanation for that experiment. And so, I think, going back to AI, I think, like, these are things where AI seems so far to, you know, help a lot more in some things than other things. So, like, you know, I think the more vague a phenomenon is the less you have a precise question in mind
Starting point is 00:18:49 the less useful it happens to be. And so you were asking how I use it in my daily life. Actually, what I've found is that the projects that I have that kind of predate AI, like the projects I've been thinking about for three or four or five years, it's just not that useful. Like it's primarily kind of a substitute for Google or something.
Starting point is 00:19:06 I might use it to learn about some related topic or like something where I would have earlier or Googled something and then read a paper, like, maybe I'll discuss it with it instead. So it'd save some time for sure. It's like not really doing deep intellectual work for me. But then, you know, because I, you know, I'm now like, well, I suck at coding.
Starting point is 00:19:23 And so now I know, you know, my good friend who's really good at coding. And now I have all these coding projects. Because like suddenly, well, if I had a question where coding would have been really useful, I would have procrastinated on it for six months until I. Tireless PhD student. Exactly. Yeah. So, so, yeah.
Starting point is 00:19:39 Now I picked up all these projects. It's where, yeah, coding is really useful. Like, the models are, like, very good for kind of massively parallel things. Like, if you want to find an example of something, you can just ask it to find, you know, work through a thousand examples in parallel. It would be 10 examples at a time, you know, 10 different sub-agents. And, like, that's really useful. But these are, like, kind of different activities, which are, like, I don't know, on top of what I was doing before. I'd love to dig in.
Starting point is 00:20:04 Yeah, go on. Yeah, sorry, because you were mentioning, you know, there's projects that you've been working on for, like, three, four, five years. And I don't know if it's correct to say, like, those are more of the theory-building aspect of it, because you did characterize yourself as, like, an open problem solver, or, like, what is the thing that is? Because, like, the deep thinking part, maybe, maybe, like, you know, kind of for this audience, it might also be useful to kind of say or explain, you know, most of pure math is, it's not like it's being motivated by, you know, nothing on applied math, but applied math is like, at least there's some external motivation for why a certain formal structure might be interesting to study,
Starting point is 00:20:38 Whereas this one, it's like purely, it seems almost sociological. And to some extent, I think like Thurston made some comment, or a point about this in the 70s, that it is a sociological phenomenon of more and more mathematicians start examining something, then you'll maybe, you know, converge on some interesting structures. But it's just like it's not, there's some reason why people would prefer to study something or they think it's beautiful. Like, what is it that drives maybe you in particular,
Starting point is 00:21:03 and then maybe you can make a more general comment about the profession? Yeah, I mean, so definitely some people, are like motivated by like beauty or or kind of aesthetic considerations. I try not to be motivated by that. Oh, that's like a controversial. Yeah. Yeah. Like, well, one thing, like that sort of limits you, right?
Starting point is 00:21:20 Like one thing, kind of a failure mode I see among young mathematicians sometimes is like, you have something and you kind of think, you know how to prove it. And then like the proof feels really ugly and you decide. But like, okay, I mean, what if you're wrong? It's not ugly. Like, you let me yourself. For this audience. What is ugly? Because I have an intuition of what's ugly, but like, what is that for kind of spelling out?
Starting point is 00:21:43 I don't really know. I don't really have aesthetic feelings. But people sometimes feel this way. Maybe it involves a lot of grinding. It's calculation that's not illuminating or something like that. But you should like win by any means necessary in my opinion. Like I like to think of what I'm doing like doing kind of physics except with concepts. So, you know, instead of beauty, I try to think about maybe like what is kind of fundamental, what's going to open up further understanding most. Yeah.
Starting point is 00:22:10 Will, you know, am I introducing a new idea that will be like broadly useful to understand this object? Yeah. I guess, you know, in some sense, there's like some aesthetic consideration there, but it's, I try to, I think the orientation of like trying to do good science rather than trying to do art. Yeah. But there's a huge, I mean, there's a huge variety of opinions here and like lots of mathematicians think of themselves as being like closer to poets or something. Yeah, what I do think is broadly true is that like progress comes from people like it seems to come in mathematics from people sort of pursuing their personal curiosity
Starting point is 00:22:45 And that it's sort of been crucial historically that there's like a lot of different people with different views of what's interesting and then the frontier of knowledge expands and expands and then suddenly you know you have these opportunistic situations where a new idea has been introduced and you can now find you now can suddenly like cascade through a bunch of other things things that we didn't understand before. Yeah. I mean, so then going back to the three, four, five year problems and where are the models sort of not useful? Kind of ask it another way. Like, when you're doing the deep thinking, like, is it just that it's just not clear
Starting point is 00:23:21 that you formulated as a problem? And it's more that you're thinking about these, like, what are, you know, what are the fundamental kind of physics of, you know? So in some cases, I mean, the, the, the, you know, the. Some cases, there is like a well-stated problem here. You know, I said, I've been thinking about things for maybe 10 years now at this point for certain problems. Sometimes there is just like a conjecture that I would like to prove that is there. I think one reason the models might not be useful for some of these things is like the conjectures are true.
Starting point is 00:23:54 So like I think that certain, you know, for example, you know, this unit distance problem, like the general belief in the community was that it was true and then it turned out to be false. And so what that means is that there's like a specific construction you can do to refute it. On the other hand, I think a lot of the things I think about like, I don't know, maybe I'm about to, you know, there's someone come up with a found example to the peak of a curvature conjecture, quarter of the room on hypothesis tomorrow and like, I'll look like a fool. But in general, like these conjectures fit into some very broad theoretical framework,
Starting point is 00:24:31 which means that, like, we actually have a lot of evidence that they're true. And so, yeah, so that's part of it. Like there's not like a construction you can do to refute it. You need to somehow kind of, you know, there's, we have this giant framework where certain pieces of it are only contracture, and you probably need to resolve some of those contractors to win. We also have like a pretty good sense, I think, that like very serious new ideas are needed to resolve those contractors.
Starting point is 00:25:01 So like, of course you can't be sure, like maybe there's sort. very clever construction that will let you, I don't know, avoid having a big new idea. I don't know. Maybe, you know, it's quite possible we'll find this out. But my sense is that for at least a lot of the things I've been thinking about, there are, they're just not accessible to, you know, applying known techniques very, very technically strong way. So you need to develop a new technique.
Starting point is 00:25:32 I'm just being clear, like I'm not saying the models won't be able to do this. So far they seem not to... Actually, that's exactly the point that I wanted to delve into because it's, to your point, it's like, okay, we can try to calibrate and forecast like why they would get better at this, but it is true, like, you know,
Starting point is 00:25:47 that just providing construction for a counter example, they seem to be strong on. If you have to start developing either new theory or, to your point, techniques, machinery, to solve, to prove why a conjecture is true, it struggles more. And it's probably because a lot of what it's drawing on is also just like techniques that have happened in other areas,
Starting point is 00:26:07 and they're porting it over. And to your point, that's why maybe the unit distance problem was such a more creative result because it was maybe doing more of that on its own. It was like from another. At least it was like growing in something from the not expected area. Exactly, exactly. And so like maybe then to kind of ask a question, it's not that because now we're like, okay, fine,
Starting point is 00:26:22 A, I'm getting so good, so fast, we can't count it out. But like, why, what do you think it has to, yeah, I guess it's, you're spelling out what it has to get better at, but maybe some more kind of feelings on like why it's kind of particularly hard to then develop that new theory and technique. Yeah, it's a good question. I mean, I think you just need a different, like my guess actually is it's probably totally doable and it just hasn't been done yet.
Starting point is 00:26:46 Like maybe you just need a different RL environment. I don't know. Yeah. So at this point, my expectation is just like that the trajectory will continue upwards. I'm not a skeptic of continued capabilities growth. Yeah. But yeah, you know, I think what is definitely true is that like, The skill of, like, developing a theory or, like, building your understanding of some poorly understood object is, like, a fuzzier one.
Starting point is 00:27:11 So it might be harder, you know, I guess you can tell it, you know, develop your understanding of Zeta functions. And then once it proves the real hypothesis, you give it a reward. But it's, like, kind of harder, I think, to come up with some, like, intermediate things that you can reward. Yeah. That said, you know, I do think mathematics as a whole provides a lot of conjectors of varying levels of difficulty. So, you know, maybe this explains why there seems to be a little bit of progress in these areas. Like, presumably they are trying to, you know, get it to solve lots of problems and some of those problems develop at least some of the skills of theory. Humans are able to develop these skills.
Starting point is 00:27:51 You know, I guess sometimes they get rewards from their PhD advisors and their advisor says, oh, that's a good idea or something based on some. element of human taste or whatever. Yeah. And that might be something one can do too. But yeah, my expectation is that just like as part of continued capabilities, we'll see growth in these areas too. Yeah, yeah. And I like the framing where you're sort of casting these increasingly difficult conjectures
Starting point is 00:28:14 as a, you know, a form of curricula for both humans and, of course, AI. And it, you know, it may be kind of getting a little bit too philosophical for some people's taste, it's like, it gets at the question of like, why are we particularly good or maybe particularly bad at math, too, because it's, you know, what is it that we're either struggling to do or some people are particularly good at when you develop new theory? Because it's not, again, it's like not really, or maybe it's related to like, why can we formulate good structures for physics as well. It's just like it's not obvious. All parts of the world are kind of understandable and legible that way, but some parts are,
Starting point is 00:28:56 and therefore we try to do it because there's maybe compressive pressures on our minds because we need to, we can't understand anything beyond, you know, compressing stuff more finely. I don't know if that's also like the correct
Starting point is 00:29:12 interpretation of, like, yeah. I'm a little bit skeptical as like compression as a metric of interest, but it's definitely like as an anthropological reason for like why we do a certain theory building. I think it's true. In fact, I think it's like kind of like our inability to just grind is kind of important to our ability to make discoveries. So I'm just like as an example.
Starting point is 00:29:33 So I have one paper out so far where the models were kind of useful. So they like proved a couple lemas. So this is some situation where like we had to kind of prove in the main result. And then there were like some lemas I was unhappy with like they seemed not to be optimal. And so I like worked with actually this was with Gemini Deep Think. which at the time was also on the frontier, no longer from now, where I kind of worked with it to improve the lemmas. And so what happened here was there was some lemma I wanted to prove
Starting point is 00:30:06 and the models couldn't do it. So none of the frontier models could do it. And so I worked out a ton of examples on my own and I realized, oh, well, maybe like here's some reason why it could be true. So I found a better statement of the lemma. And then, okay, once I had that statement, like, okay, probably I could. have done it pretty fast, but also the models were able to very quickly prove that sort of better statement. So like our inability to prove it led to an improvement in the result.
Starting point is 00:30:33 Okay, now you can take the original lem I had that the models weren't able to prove, put it into chat GPD 5.6 Pro, and it will output like the worst proof you've ever seen, like 10 pages of like just brutal calculation with no insight whatsoever. And so, okay, this would have been a perfectly proof, but it would not have led to this discovery of like, I think, a kind of beautiful conceptual explanation for why this thing we discovered what's true. So like we found a better proof because we couldn't do, I mean, okay, I say we couldn't do the calculation, like what actually happened. And here I understand that I'm being hypocritical about my complaints about ugly proofs before. It's like I realized like this horrible grind proof would work.
Starting point is 00:31:10 Yeah. And I like could not bring myself to do it. And so I look for another argument. And so I mean, it's not, yeah, it's but now that the models can do for very long technical calculation is pretty reliably. Yeah, well, I mean, to your credit, I don't know. I think I understand why you're saying it's like, don't, you know, just shy away from trying to prove something just because it seems ugly, because it's like you have to take the first step, right?
Starting point is 00:31:32 And eventually you all try to work towards insight. And so, you know, you might be shy about admitting it, but it is still maybe driven by whether it's like aesthetics or just like pure, you know, it's like, I want to understand. And if understanding just means it's a little bit simpler or, you know, more compressed than I've gotten understanding. I mean, it's just like a deep philosophical question, too, is like what it is. Right.
Starting point is 00:31:52 Like, I mean, of course, any, you know, there's a reason we do informal mathematics rather than, like, writing out long informal strings of symbols of CFC or whatever. Yeah, yeah, yeah. Even though, you know, besides slang, it's just like somehow we're trying to put it in a way where we are actually getting some non-rigorous understanding out. Yeah. I mean, and even how that even relates to, like, why it helps. And again, it's like, is it because of our inability to grind or is it because, like, if we are pushing towards something that's more compressive, it also hopefully also ascends on the, does it? help understand other fields as well. And it's just, it is somewhat magical that when we try to optimize for the both,
Starting point is 00:32:30 that it tends to coincide. I don't know, actually, you know, if that's a fair, it's actually a quantitative statement, but like why, you know, that is, is kind of magical. It's at least sometimes true, yeah. Yeah, exactly. Or, you know, we certainly biased towards the cases of which it is. That's why we call that a good theory to build. But it's, yeah, it's very much, I think, you know,
Starting point is 00:32:52 it's like this is what we can study and understand, and therefore it's also very convenient that it was rich in kind of mathematical results there. Yeah, I mean, okay, so I think there's two things that we can like, you know, pull on, which is like what, given you've seated that, or not seated, but you've believed that mathematical ability of AI is going to continue advancing, it can do some of the things that you're, you know, doing nowadays, how should, like, the mathematics community kind of best adapt and benefit from this?
Starting point is 00:33:30 I mean, just kind of saying, as somebody who's, you know, doesn't have the time to kind of practice mathematics anymore, this is kind of very great because I can maybe dabble more. There's a lot of results that can come out. But I can also see where, you know, you've made the more precise point of we can not motivate the right kind of behavior of understanding and development. And so we'd love to kind of hear more about your views there. Yeah. So, okay, so first of all, like, it is clearly really exciting, like, that they're increasingly capable models that are, like, producing high-quality results, at least, you know, some high-quality results. Also, also a lot of thought. Yeah. Yeah, some good
Starting point is 00:34:10 stuff. And so, you know, as the models get really capable, you know, my hope is that they will answer a lot of the questions that I've been, like, you know, kept up at night. thinking about, I'm like, I'll get to learn the answers. I'm not really exciting. And, you know, a lot of people got into math, like, largely because they enjoyed learning math. Yeah. Right. Like, the first thing you do was a math student is you, like, learn stuff that other people did. And you do that for, like, 20 years before you, before you start doing, oh, maybe not quite 20 years, but, you know, 15 years. If you're lucky, 20 years, yeah. If you're lucky, 20 years, yeah. Yeah. Yeah. Um, yeah. So, um, that said,
Starting point is 00:34:49 like the goal of mathematics is not to produce mathematics papers. Like it's to produce some kind of understanding. So maybe, okay, maybe that some of that understanding resides in model weights or something. To me, that's, like, pretty unsatisfying. Like, my own, you know, my motivation for doing mathematics is, like, I would like to satisfy my own personal curiosity. I think people should be, have the capability to do that. And, like, that requires a pretty substantial apparatus.
Starting point is 00:35:16 Like, it's, like, simply not the case that you, you can like study the questions I think are fundamental unless you've invested a huge amount of time and effort kind of getting to the point where you can meaningfully do so. And then moreover like that, you know, even the people, like, you know, you have the small group of people doing like fact that you research math on the frontier or whatever, that relies on like a huge apparatus of like, you know, thousands and millions or billions of people who are trying to learn to think mathematically. Like you need an entire mathematical community to support a small group of people who are
Starting point is 00:35:48 on the frontier. Like you just, without the pipeline, then the pipeline doesn't exist. So if you think that's important, like development of human capital that can like meaningfully engage with frontier mathematics, I think that, you know,
Starting point is 00:36:01 you still have to incentivize those people to actually invest their time and effort getting to the point where they can engage and then do so in a high-quality, meaningful way. So right now, I think, like, the existing incentive structures for math research do not, do not encourage people to do that.
Starting point is 00:36:16 So, you know, right now, if you're like a postdoc on the market, you want to get a job, maybe for the next couple years before the community adapts. The best way to do that is, like, you want to produce a lot of papers, which maybe prove, you know, old conjectures or whatever, and you can do that by playing the slot machine until, you know, the model produces a hopefully correct proof of such a result. So you don't even have to pick the theorem in advance. So, like, here's an experiment you can do. You can take codex.
Starting point is 00:36:45 You can say, go online and find five recent conjectures in algebraic geometry and prove them. And, okay, I've run this experiment and with some back and forth, I was able to, you know, in an hour, get like three, you know, quite bad papers, but correct papers. Which, okay, are now sitting on my hard drive waiting for me to email the relevant people. But, you know, this is not a big use of my time to invest in results. Yeah. But yeah, you definitely see people doing this. So there's, you know, been a huge uptick in post-archive, mostly not very interesting. Some of it is interesting.
Starting point is 00:37:18 But a lot of it is kind of clearly low quality, even if insofar as the, like, even if the result is something that would have been like highly rewarded a year ago, just like there's sort of no evidence that a human being is actually engaged with it. Like there's no development of human capital or understanding. Like sometimes, you know, we've seen examples where like three or four or five papers with the exact same proof of the exact same theorem out within a couple days of each other, which is clearly. you know, some situation where someone's playing the slot machine. Yeah. Chat GPT is kind of consistently finding the same. That's also interesting. It's like not, it's kind of mode collapsed on like certain pads of reasoning.
Starting point is 00:37:56 Yeah. Yeah. And I think it's like not obvious that this problem goes away as the models get better. Like maybe it does, but maybe it doesn't. Like I think a lot of what I was saying earlier is that a lot of progress in mathematics comes from like letting, you know, a thousand different flowers bloom and people pursue their own curiosity. And then, you know, the boundaries of knowledge expand in some kind of fairly, hopefully fairly uniform way and look at this really high dimensional space of mathematics.
Starting point is 00:38:20 And then like opportunistically you you suddenly get some applications or like as is told questions we found. And it's like not clear to me that if if you know we kind of subordinate mathematical exploration to what the model want to pursue like if what you're getting is like one mathematician duplicated a thousand times or like actually you know a million different mathematicians doing a million different things. Yeah. I think actually that's pretty I mean you know aside from like, okay, math has to adapt by changing incentive structures, I think this is actually a pretty maybe dangerous or something that the labs have to pay attention to.
Starting point is 00:38:55 Because on one side, what's been successful for them is that this emergent reasoning capabilities obviously incredibly powerful. And we've been, I mean, it's just created a lot of great PR headlines. But to your point and also like where my interest tend to is that like human mathematicians, they are, coming from all sorts of weird intuitions that, like, you know, and that is why you end up developing, you know, to your point, the frontier that is so diverse that you actually can
Starting point is 00:39:23 draw from and then be able to connect and produce a lot more. And so it, if it's true that most of these like proofs that are being pushed out by the labs are converging on very similar things, because they are technically drawing on the same body of literature, and that's where they're strong at now. It's not clear that, I mean, one, you know, where are they developing their intuitions, it's from practicing mathematicians. It's not clear what that other emergent, you know, stronger, diverse intuition might be coming from. It might come, but I don't have a good theory for where it comes from. And it's also just not clear that, you know, just with like test time compute and various post-training scaling that you can even, I mean, it's a huge debate,
Starting point is 00:40:03 whether you could introduce new capabilities. And it's like precisely, I think you should be studied in the math context because like where it's good at and where it fails and where human mathematicians are good is exactly a very kind of like precise question to study it on. So there's like a long way of saying that I, it isn't clear to me that we'll maybe produce, if we don't incentivize enough people to interact with it, maybe unlike other domains, it might not actually continue producing a superior result without a human aid. Yeah. Let me actually push a bit further even. So like Let's suppose the models become really robustly superhuman, like, even, like, we're not even adding, like, meaningful cognitive diversity. Okay.
Starting point is 00:40:47 I claim, like, still, actually, we'd be still one. Okay, great. Human mathematicians. So, right, so why? So, like, it's because the, you know, there's like a, it's like, there's a question on how we've, we're designing society, right? Like, maybe the optimal situation is, like, you have the models doing all sorts of math research. And, like, that leads to very, you know, in, you know, and this, like, highly, you know, you know, whatever, non-ununum. a form diverse way.
Starting point is 00:41:12 Like, maybe you don't need humans to add that kind of diversity. So, like, maybe that's the optimal thing to do. And, like, in the end, that leads to lots of applications and lots of improved understanding and so on. It's like, despite something being optimal doesn't mean you do it. Right? Like, there's no reason to think that, you know, if we hand over control of whatever, math research to the models and just let them do their own thing, that it will do the optimal
Starting point is 00:41:34 thing. And in fact, like, if you, you know, if we are, if we kind of instrumentalize what we want them to do, like we want to say, like, oh, you know, make our life better or whatever, it might not be the case, like, that, like, what they decide to do that is, like, is, is pursue a wide variety of interesting research, right? Like, they may just try to take the direct path. We don't know what's going to happen. So, I don't know, if you believe that there is any value in sort of this broad-based, like, fundamental research, which I do, look, I think that's one of the most valuable things, humans or whatever, or the models can be doing. Like, I think the easiest way to guarantee it happens is, like,
Starting point is 00:42:09 to keep a community with like broad interests where like pushing the models to do and like helping us to design a society. Like that's what we're pushing for. Like I think at least, you know, my hopeful vision for the future is that like humans are not like totally disempowered. Like we have control over where we're going.
Starting point is 00:42:26 And like if that's the case, like what we end up doing is going to be driven by human interests. And so you want to have people who have lots of different interests and also the capabilities to actually pursue them. Like you want people who are like smart and engaged and like, you know, well, can do not just like mathematical thinking, but all sorts of thinking. Yeah, yeah.
Starting point is 00:42:43 I mean, yeah, a great fear is like as AI advances, we don't develop the right ergonomic kind of interfaces to actually encourage us to also be good, continue to be good thinkers. And it's so easy to kind of relinquish that because you just off-land. And the models aren't even good at that level of, like, you know, a thinking where it's like the top-level structure,
Starting point is 00:43:02 but despite that, it's so easy to. And so especially for math, I mean, just like, you know, another selfish reason is if you kind of Math Max, you know, I think it's actually a great pedagogical excuse to actually get just really rigorous at thinking about various things. I mean, this is not why mathematicians do it, but as just somebody... It's part of what we do. Okay, great, because it was kind of, you know, I thought it just really helped give me a very good framework to think about many things, not just mathematics. And as somebody who's also, you know, now a parent of a two-year-old, I kind of think about this a lot, too.
Starting point is 00:43:37 You know, it's not about kind of grinding, even though whatever, it's still good, but like, not to, not the shit on grinding too much. But, you know, just hearing, I used to collaborate a bunch with some Hungarian mathematicians, Balazeghdi among them. And I just heard that in Budapest, they would just teach group theory when you're in primary school. And I'm like, well, we should definitely do that. We should continue doing that. And now that AI is so good at, you know, somewhat good at explaining, but it's far more accessible, we should actually, you know, know, probably proliferate that even more. And so maybe that helps with bring more people to the frontier rather than, you know, just. Yeah, I mean, this is something I'm concerned about, right?
Starting point is 00:44:17 Of course, I mostly talk about math because that's like where I live. But like I think, you know, one nice thing about thinking about this is that we're one of the first professions to kind of, you know, have a significant impact of high quality models. Although I think maybe we're one of the first professions for it to happen to publicly but you know my sense is that there are plenty of third professions that are coding 100% coding but also
Starting point is 00:44:42 just like I mean I think that you know anything you do at a computer like probably a huge amount is being done by the models at this point and then like there's not having a public reckoning about it but you know because math capabilities are useful for the company for the labs to talk about I think it would be more publicly than everyone else
Starting point is 00:44:59 but yeah I mean I think one reason to try to maintain like human capital in this area is just like it's a model for all professions like presumably we still want people who are like meaningfully engaging with the world and like experts and you know have like talents and trained skills and so on like
Starting point is 00:45:15 yeah so I was it's actually very convenient that the math profession is so inclined with education here because I think we're also seeing like you know some amount of challenges you know among college and
Starting point is 00:45:32 and like secular education coming from AI2. So as you said, I mean, it's also an amazing tool to learn. Yeah. You know, people have, have, I've heard people start talking about a bimodal distribution in their classes where there's some people who are really like figuring out how to take advantage of new tools and other people who are just like letting them do their homework and then bombing everything else. Unfortunately, I don't think that adapts fast enough, but it's like we definitely want to be living in a world where we're producing better thinkers. I think it's just, you know, even if we talk about just the pure kind of optimization game,
Starting point is 00:46:02 I think that's better for us. So, but as, you know, just like a human being, I'm like, that would be pretty inconvenient if we became worse thinkers just as AI sends. And it's too easy to let that happen. So we should kind of be thinking hard on how to actually take advantage of this and harness it for our own improvement as well. Yeah, I think it's like it's sort of interesting to observe, like, the models
Starting point is 00:46:24 at their current level capabilities, let you do a lot more. They let you do a lot of things you wouldn't have done more cheaply than, you know, cheaply enough to do them now. But it's not clear to me that, like, in many cases, they're actually improving the quality of outputs. Yeah. And I think this is common. Like, you have a new technology that's doing something a little bit worse
Starting point is 00:46:44 than was previously done, but much cheaper. And so you get a lot of suddenly a lot of, like, low quality outputs that are displacing previous high quality outputs. But I think it's possible to use the tools in a way that actually, like, improves, you know, the quality among, along every dimension. It just requires some thoughtfulness and some redesign of institutions. should actually incentivize that.
Starting point is 00:47:04 Yeah. Well, hopefully capitalism works there. I do feel like that the most high value things do require people to use it effectively. And right now the models are not good enough without like the human experts to actually participate. But to your point, there's a vast majority of maybe like more junior and entry level. And if, you know, it's harder for those roles to adapt as well. And so the thing that would be a mistake is to use the AI models in a way that doesn't, basically, you need to be ascending and using the models to deepen your understanding.
Starting point is 00:47:39 And it's so easy for human nature just to be lazy. And you have to resist that because that is the moment that you will kind of lose, basically. And so you kind of, you know, kind of forgive the very competitive language. But it really is that like it's just so easy to kind of relinquish thinking to the models. The models can't really think. And so as things are ascending so fast, it's like critical that you continue developing those facilities and actually leverage it to improve those facilities
Starting point is 00:48:06 rather than relinquish, yeah. Yeah. Yeah, I mean, one thing I think has been nice about, you know, sort of this vast increase in like semi-expert attention or like model attention on math problems is like now there's been, you know, okay, I've been complaining about slot papers or whatever. Like, you know, people who are not, you know, producing really high quality stuff. Often that's actually coming from professionals.
Starting point is 00:48:28 Like, it's not saying, like, you know, there are people who, you know, like within academic mathematics, there are incentives to produce, like, a lot of stuff. And that's, you know, that's one place to slop coming from. So there's definitely also, like, stop coming from non-experts. But that I kind of actually don't see as a net negative. Like, okay, there's a lot of, now there's a lot of, like, documents on the internet. One might have to come through to figure out if a problem has been solved or not. But to me, it seems like just the fact that there's lots of people excited about math is like kind of a positive.
Starting point is 00:49:00 So that's like a nice thing. I totally agree. I know I get to talk to you. And not even kind of a positive, obviously. Yeah, yeah, no, exactly. It's like suddenly there's a spotlight on it and I can nerd out about math more. Yeah. I was actually kind of curious you've had any comments on the, I guess this is another constructive result, but the elliptic curve of rank 30.
Starting point is 00:49:20 That just came out yesterday. So if that's right. Yeah, yeah. We don't have any details about it yet. I know, there's rumors. Yeah, we don't know how it, I mean, so it was, it's due to, I guess, Claude Fable, prompted by Levent-Alpoche and a collaborator whose name I unfortunately forget. Maybe you can settle on it on it.
Starting point is 00:49:40 Yeah, I just know the Twitter handle. Yeah, we can have. So, yeah, I mean, so a lot of these nice recent results have come from Levent with unclear amounts of autonomy. So my sense is that some of them are semi-autonomous, doesn't fully autonomous. Yeah, with this, we don't know anything about the methods. So, yeah, this is a fun construction. Without knowing about the methods, it's very hard to say how significant it is.
Starting point is 00:50:05 What I would say is that if you want to understand kind of historically how such results have been proved, or sorry, have been understood by the community, they're like cool, but I wouldn't say they're like a big deal. So like the typical place of result like this might go is like someone's website of records. It's on an annals level result. Yeah, it's not an animal. But that said is cool.
Starting point is 00:50:29 And it's like, you know, it definitely, there were a few very, very, very, very talented mathematicians who like these kinds of questions. So know Malkies being maybe the most famous example. So Alchies and Clagsbrun are the ones who kind of have been pushing this record for a while. And they recently found a ranked 29 example. One up. Which was, you know, the previous record. Yeah, yeah.
Starting point is 00:50:48 You know, so those are mathematicians. And I think, you know, people like it. And it's cool now that the models can. do this sort of thing. But yeah, it's very, you know, one thing I always say about a model result is, like, you cannot evaluate it except in retrospect. And like, this is also true of human mathematics. Like, sometimes a problem we thought was really important or would require really deep,
Starting point is 00:51:07 new ideas, does not? And sometimes it does. And you can't really know in advance. So it's always exciting when a problem like that gets solved. But then, like, to figure out how significant is, you start to look and like, well, Levent and Claude and I think Ava Howell, maybe is the third collaborator, have not yet told us how they did it. Have they been mostly more secretive?
Starting point is 00:51:27 I think they released some traces for stuff. Yeah, so for this one, I think they haven't yet, unless I missed it. Yeah, they, you know, Levent likes to tweet out his results. But yeah, he has been, you know, slowly releasing some kind of PDF-writups too. I think it's, you know, he's just having fun on the internet. Yeah, yeah. It reminds me of when you're saying, you know, you have to evaluate how. the results came. So I had Mark Selke and Metab Swani on from Open AI recently, and they were saying how,
Starting point is 00:52:03 you know, what's kind of been the most charming or delightful is just that the proofs have been relatively short. They're not like 200 page. And, you know, maybe corresponding to your grinding point as well. But maybe this is kind of optimized for, in retrospect, it was picked that it was short or maybe do you find that on average stuff that you throw at GPT, you know, Sol or Fable tends to be shorter and more legible to the human or versus
Starting point is 00:52:32 it might just go haywire and just grind it out. Yeah, so I mean, I think it is nice that will sometimes produce short, clever proofs. Yeah. So that's, of course, everyone likes a short clever proof. I think my sense is that the reason they're not producing long, complicated proofs is that they cannot.
Starting point is 00:52:53 Just like the ability to check correctness is not yet there. So, you know, I mean, even actually, you know, if you ask the models to produce short proof, you can then often ask them to, you also just to ask, is that correct? And they will often say no. So, like, they're much more reliable than they were six months ago for sure. But, like, it's still, you know, they will still sometimes produce things that are just wrong. and they know their wrong. Yes.
Starting point is 00:53:20 Yeah, yeah. I think the problem with producing a very long thing is they might not know they're wrong. And so what I wonder if, you know, presumably internally, Open AI and Anthropic have probably solved a lot more problems than they've released. And I imagine quite a few of them are they're just not sure if they're true. So, for example, you know, with this recent list of 10 problems released by Open AI, those were all formalized in Lean, which is, of course, a very good evidence that they're true. I have no doubt that there were a lot more
Starting point is 00:53:49 that they could not formalize and leave because the prerequisite options have not been put into Massua yet, for example. There were probably more that were longer, which also makes it challenging to check. So yeah, this is my guess. And we do actually see very long AI generated proofs on the archive.
Starting point is 00:54:07 So, for example, someone recently posted a claimed proof of resolution of singularities and positive characteristic, which was 800 AI-generated pages. It's definitely, I mean, I'm sorry, I haven't read it. I haven't been an error, but there's no way it's correct.
Starting point is 00:54:21 Like, this will be a major result. It's just not with the capacities of the current models that you're reasonably well calibrated. Yeah, yeah, yeah. And it's definitely no human has read it. Definitely the models are not able to check this kind of thing, yeah. So, I mean, yeah, so this is my expectation is that the reason it's producing short, clever things
Starting point is 00:54:40 is just like that's what we can check. Yeah, no. And you can get it to produce long things that are grinding or hard to check, but then. Yeah, yeah, yeah. to get there. This is actually more what it says about capabilities, frontier capabilities, rather than in a perhaps more negative way, rather than, hey, it's just so good at these clever. No, I think that's totally fair. I mean, especially if you look at an adjacent domain like code,
Starting point is 00:55:01 right? Remember reading something that cursor put out about testing their long horizon, you know, harness. In this case, they were trying to reproduce SQL light in Rust. And it was just so telling, like, how far we are from, you know, and that seems. And that seems, It seems like a very comparable task of like it just, it's very long. You have to make sure there. There's something to verify that it's a correct implementation by testing a suite of, you know, cases. But, you know, you actually need a harness there in this case. It's not just like the raw models.
Starting point is 00:55:33 And it takes a while. And you can actually compare with different frontier versus, you know, non-frontier models, who's the planner, et cetera, like differences in their capabilities as well. I was thinking in practice, like, to elicit a long proof, you kind of need a harness. And when you make a harness whose goal is to elicit a proof, I think it often decreases reliability because you're just trying to produce output. So like chat chaddbd5.6 pro is like very, it really tries not to say wrong stuff, for example. And although it happens and then you'll ask it like, oh, was that correct? And it'll say no.
Starting point is 00:56:12 But, you know, when you're trying to get it to ID8, like you try to get it. like you try to get it to be creative, you try to get it out of this very, like, you know, rigorous rut in order to, like, you know, actually get it somewhere. And I think if your interest is in, I mean, there are a lot of people who are, you know, try to arbitrage the prestige mechanics of academic mathematics and trying to elicit a lot of proofs,
Starting point is 00:56:33 which are not necessarily being checked. And in order to do that, I think you just decrease the reliability in order to get a lot of stuff. So, you know, I think any harness that can elicit a 250-page paper is probably not being very, careful about what is producing. And you're saying that it's decreasing, or it's not reliable just because the capability isn't there yet. And so it's just kind of forcing a longer horizon task on it.
Starting point is 00:56:56 Yeah. I mean, in practice, like how does, like, a human checking, like, can't check a 250 page paper either. Like, like, you, you know, you can't reliably check it line by line. What you try to do to understand it is you try to understand the overall global structure of the argument and you, like, stress tested in various ways. Like, would this argument apply something else that I know to be? false, like, oh, like, does it work in the special case? Blah, blah, blah.
Starting point is 00:57:18 And also the models seem not to be able to do that kind of, like, more, I don't know, fuzzy, like unit testing of proof very well yet. So, actually, like, one of my favorite tests for the models, which they haven't succeeded at yet, is there's a paper, I won't name it, that came out a couple years ago that was wrong. And it was, like, very hard to find, like, in close enough area to myself that I, like, immediately, like, it was claiming some big result. like I immediately downloaded and like started reading it.
Starting point is 00:57:45 And it was like very hard to find the actual specific error. But it was also very clear from the structure of the argument that it couldn't work. So it's like, you know, me and a bunch of other arguments, well, he was proving something that was too strong to be true. He took a little bit further. Yeah. So me and a bunch of other experts like immediately like realized it was wrong. And like we, you know, we emailed the author and like we kind of went back and forth until someone figured out what the precise specific error was. And so far the models seem not to have been able to do this.
Starting point is 00:58:16 Like the specific area is quite subtle, but they're also not able to do this kind of overall kind of big picture check checking, which is kind of like how people in practice check papers. Yeah, yeah, which is kind of it mirrors, you know, why in code. It's like it's so clear it's good at the syntax, but, you know, higher level architectural stuff. Still very weak. Maybe you'll ascend there, but probably need some harness help. Who knows?
Starting point is 00:58:39 I mean, people have, you know, evolve. opinions on how much the harness and the model have to co-evolve and which one is necessary, but the next model requires less. So I feel like in math, it would be very interesting to see how you, if you do any experiments with the harness there and how that improves, because it is a marker of general reasoning. Yeah. Yeah, I mean, I do, you know, I have my own sort of bad little harness and codex. Oh, yeah.
Starting point is 00:59:03 But, yeah, it's, but, you know, I personally, I do not enjoy auto-autonomous mathematics very much. So I mostly do not use the harness. I mostly try to use it to help me understand stuff. Okay. No, that's totally fair. Yeah, exactly. You don't want to automate your job away because that does involve you being in the loop to understand it, which is necessary to participate. Maybe to finish off. I'd love to, this could be just something you haven't thought about or actually thought a lot about. I think you also have a toddler, right? Yeah. Yeah, yeah. So how have you? Yeah, a three-year-old. Great, great. So you're one year, more advanced. and probably, you know, have more thoughts on this. Like, how are you thinking about, is it her education or his education? Her, yeah. Or how are you thinking about her education in math and, you know, not to grind,
Starting point is 00:59:51 but, you know, really just like, you know, pass on the love of it and how to how to react to AI. Yeah, so she's three. She's never used AI. Good. She is starting to add. That's about as far as you are in math. That's more far as much than my deal.
Starting point is 01:00:06 She can add single digit numbers, like, you know, by counting on her fingers. and count up so maybe 30 reliably and 50 semi-reliably so I'm very proud of it. Yeah, I definitely encourage that. We talk about shapes and stuff.
Starting point is 01:00:21 Actually, a couple days ago, I woke her up and she was like hiding under the blankets. And I was like, oh, I'm doing some math. Oh, I'm doing some math. Oh, I think he tweeted about that. That was like more.
Starting point is 01:00:33 Yeah, it was great. So, you know, I think she has like some sense that I like math and she's into it because of that. Yeah, I mean, I would say I don't know, you know, I think the world is probably going to look pretty different in, you know, 20 years or whenever she's kind of fully adult and doing her own thing. But, you know, I think a lot of what we educate people for is like pretty robust changes in the nature of the world. Like, I think the reason to learn math has always been like to think clearly and like better understand the world. And like presumably that's something you want to do even if there are sort of extremely capable AI's. And this is also true. You know, I personally like math a lot, but also the reason to, like, read a lot of books and do the humanities and so on. Right.
Starting point is 01:01:16 So, I mean, I think, like, the actual values of the math profession and, like, education more broadly are, like, things we definitely try to want to try to preserve. That I hope to instill in my daughter. You know, how much institutions have to change to make sure that happens is maybe an open question, I think, a lot. But, yeah, I mean, at least at a personal. level, I'm definitely trying to, you know, convince my three-year-old that math is super cool. Oh, yeah, 100%. One of her first words was Icosahedron, so. Was it what?
Starting point is 01:01:50 She has a, she, my parents gave her little icosahedron toy when she was one. That was the first word. She, yeah, not literally first word, but she learned the platonic solid squatterly. Oh, very good, very good. It was very fun. Next stop group theory. I mean, it's like very natural. That's right, yeah.
Starting point is 01:02:06 I mean, it's actually, when I, when I teach her about a, uh, a, you know, addition and subtraction, for example, we'll do it in the context of a general group. Oh, very good. Well, at least you give the motivation. I think a lot of probably skip that part. Yeah, yeah. And, you know, maybe this is how math grad students can also, you know, focus on training the next, much younger generation to use AI in service of actually getting better at math rather than
Starting point is 01:02:32 just fucking understanding. Wonderful. Okay, well, thank you so much, Daniel. Thank you. It was a lot of fun. Yeah, yeah. It was a lot of fun. And yeah, I mean, I think there's going to be a lot more progress very soon.
Starting point is 01:02:44 I'd love to maybe catch up and chat again. Sounds great. Thanks for listening to this episode of the A16Z podcast. If you like this episode, be sure to like, comment, subscribe, leave us a rating or review and share it with your friends and family. More more episodes, go to YouTube, Apple Podcast, and Spotify. Follow us on X and A16Z and subscribe to our Substack at A16Z.com. Thanks again for listening.
Starting point is 01:03:11 and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. It should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments,
Starting point is 01:03:36 please see A16Z.com forward slash disclosures. You know,

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.