The Joy of Why - Are We Thinking Correctly about AI Intelligence?

Episode Date: August 20, 2026

When an LLM answers a question, is it reasoning like humans, or just producing text that looks like reasoning? The distinction isn’t just philosophical, this determines what we can trust AI... to do, how closely we need to supervise it, and ultimately what its real-world impact will turn out to be.Melanie Mitchell at the Santa Fe Institute argues that we lack adequate methods for measuring machine cognition, and that AI is a form of “alien intelligence” that operates through non-human cognitive mechanisms. In this episode of The Joy of Why, Mitchell tells Steven Strogatz how methods that psychologists use to study cognition in other kinds of “alien intelligence” — babies and animals — can be adapted to probe AI, and she lays out six principles for better assessing machine cognition. Their conversation ranges from the challenge of interpreting what’s happening inside these systems, to recent AI-assisted breakthroughs in mathematics, to why a math-performing horse from the early 1900s offers a cautionary tale for how we assess intelligence.

Transcript
Discussion (0)
Starting point is 00:00:00 Okay, here we go. I'm Steve Strogetz. And I'm Jan O'Levin. And this is the Joy of Why. A podcast from Quantum Magazine where we explore some of the biggest unanswered questions in math and science today. Well, hello, hello. This is unsurprisingly yet another show about AI. I'm telling you, it's a topic people can't seem to get enough about.
Starting point is 00:00:30 And I'm becoming reluctant to pontificate anymore. It's changing too quickly. It's true. It is moving very fast. Anything we say could be obsolete by next week. Oh, yeah. As we speak, it's July 23, 2026. And it feels different to me than it did in July 23rd, 2025. That's for sure. That's actually relevant as talking about timelines, because our guest today, Melanie Mitchell, who is a cognitive scientist and computer scientist at Santa Fe Institute, is someone that we had on the
Starting point is 00:01:04 previously. She and I spoke about five years ago, and that is before chat GPT. Right. And was she interested in AI then? Oh, yes. Okay. So it wasn't just cognitive science. Absolutely. I mean, yes, I should say. Melanie has been thinking about AI for a long time, and she'll tell us about that. But the thing that's going to be so interesting, I feel, for us to discuss today is Melanie's point of view. Which is to think about the problem of AI from the standpoint of fields like developmental psychology. Like, how does a baby or a young child get to be as intelligent as they soon become? Oh, I think that's so interesting because we're so excited about the artificial mind when we have very little comprehension of the human mind.
Starting point is 00:01:53 Exactly. Right? So we're trying to skip a step. Well, that's right. And not just human mind, but also animal minds. Right. So there's the field of comparative psychology where we look at intelligence. in birds or dogs or dolphins, whatever.
Starting point is 00:02:07 Yeah. We have a lot to learn about thinking about intelligences other than our own adult human intelligence. Yeah, and this idea that we're going to somehow simply understand a mechanism to generate an artificial intelligence when we, again, don't understand the mechanism that brings a baby to have its level of intelligence when it's born or when it's developing. I mean, I think that's really interesting to combine those two. So I'm looking forward to this one. Well, great. So then let's dive in with Melanie Mitchell. Here she is.
Starting point is 00:02:40 Hi there, Melanie. Hey, Steve. Very excited to see you again. This is going to be fun. We talked a few years ago back when this show was called The Joy of X. And I think you may be our first return champion. Oh, boy. I'm honored. Well, you should be. And I have you back because so much feels like it's changed in artificial intelligence. We talked, I think it was maybe 2021, and ChatGPT tidal wave hit the world at something like November of 2022. Is that right?
Starting point is 00:03:15 That's right. So everybody knows that AI is everywhere. We seem to be talking about it. People are worrying about it. Some people are excited about it. It's certainly very widely used. I suppose I'd like to start by asking what has surprised you the most about the past few years. Oh, wow.
Starting point is 00:03:32 So much has surprised. surprised me. Just the thought that we could get to where we are now, just by training these models on huge amounts of human-generated language and images and so on, I never would have dreamed it. So I've just been really surprised by what's happened in AI. Also, just the kind of polarized reaction that appeared in the AI community, and just society at large, I think, has been a little surprising to me too. Polarized in terms of, like, sometimes people will distinguish AI doomers and AI optimists. Is that the kind of thing you're talking about? There's that dimension, then there's the dimension of people who believe that AI is smarter than humans and people who
Starting point is 00:04:22 think that it's far, far from being anywhere near human-like intelligence. I guess related to that is sort of the love it and hate it. And these are separate dimensions, but maybe they're correlated. Well, right. And the love it and hate it can be also tied to things like the impact on the environment versus, you know, the economic prosperity for certain companies. But then again, what about job loss? There's so many dimensions to this. Oh, there's so many, yeah. But the thing that I really want to focus on with you today is complex systems, cognitive science, artificial intelligence. You have a lot of different hats,
Starting point is 00:05:01 but I'm really very curious about the work that you've been doing to look at AI through the lens of either developmental psychology, like the way that we try to think about the alien intelligence of human babies, or comparative psychology with the alien intelligence of our pet dogs or smart birds or dolphins or that kind of thing. I mean, it's a really interesting take on, this alien intelligence of AI. Yeah, many people have described AI as an alien kind of intelligence because it's
Starting point is 00:05:34 very different from humans, even though it's been trained on human language and books and everything on the internet and so on. But the way that these systems work, the way that they learn, the way that they reason, the way they do what they do is just really different from the way humans do it. And this theme was actually picked up by people in developmental psychology, especially Mike Frank at Stanford, who wrote this paper about how AI people should take some inspiration from the study of babies and young children, developmental psych. And then other people have extended that to what about animal intelligence? And I guess one of the things that people in Cox-I have been urging is that people in AI actually adopt some experimental methodologies
Starting point is 00:06:29 that would make AI more like a science. Yeah, I really like this point of view, and I think it may not be so familiar to our listeners. I have to admit it wasn't that familiar to me. You know, I never studied cognitive science or never took a course in developmental psychology, and people in those fields have been thinking about these issues for, well, I don't know, you tell me. Yeah, at least 100 years. Yeah, 100 years now. Wow.
Starting point is 00:06:55 And I was thinking on the way over, we constantly talk about AI as a black box, that we can't read the weights on the neurons very easily, or even if we can, we don't know what they tell us. But for that matter, couldn't you say that our own intelligence is in a lot of ways of black box? Absolutely. I mean, we have different ways to penetrate the black box. One is neuroscience, where we actually stick probes into neurons or we use fMRIV. or other imaging techniques. There's also psychology where you actually look at just the behavior
Starting point is 00:07:27 of a person or an animal and try and infer from that underlying mechanisms. And those two traditions have for a long time been quite separate, but the field of cognitive science tried to integrate them. And originally, the fields of cognitive science also included AI. Somehow that integration didn't work. You mean it didn't catch on sociologically, or what do you mean?
Starting point is 00:07:57 You know, originally it was thought we're going to program them the way that humans work. And there was a very close connection between human psychology and people trying to build human psychology into AI. And then that actually didn't yield success in AI. the way that we've seen neural networks and learning from data rather than trying to program it in. I see. And neural networks itself was originally inspired by neuroscience. But the way that neural networks work today has diverged considerably from that original inspiration. So I think the field of machine learning has gone much more in the direction of statistics,
Starting point is 00:08:43 which is quite separate from how cognitive science works. So at this point, I guess I'd like to talk a bit about benchmarks because they do seem to be a big part of the discussion broadly in society these days. There was something that got a lot of people chattering. In the world of math, one of the latest frontier models did something that looked like a kind of creativity. Solved an old longstanding math problem, one of the problems that Paul Erdash, the great Hungarian math. mathematician, he left lots of problems for people to think about, and one of them, that they call the unit distance problem, was recently solved in a very clever way by AI, and it involved putting two parts of math together in a way that hadn't really been tried before. And so I bring that up because the last time we spoke, we were talking about an old AI that was learning to play some Atari game or something. And you talked about how it was so good at playing, but then if you move the paddle, a couple pixels up or something, it had to relearn all over again.
Starting point is 00:09:52 It didn't know how to play the slightest variation on the original game. So the thing you said at the time that stuck with me, the strange thing is that these machines don't seem to be able to transfer their brilliance to any other domain than the one they've been trained on. So that was five years ago. Now, I guess I wonder, what do you think? Is that still true? Yeah, I mean, that particular model was not a large language model. It was a specific model to play this Atari game, whereas now we have large language models that are trained on everything.
Starting point is 00:10:23 So in some sense, they don't have to transfer anything. They're already trained. But people in AI or machine learning talk about things that are in distribution and out of distribution, and that means that is this thing that we're asking the model to do similar to things that it's seen in its training data? Or is it wholly different? And I think it's hard to know. We don't know what it's been trained on.
Starting point is 00:10:49 The model that solving these problems has certainly been trained on a lot of math because there's a lot of math out there on the internet. It's been trained on textbooks. It's been trained on all of Steve Strogatz's videos that are on YouTube. And these models are pretty good at taking things from one area and putting them together with another area. Yeah. But, you know, I don't know how to talk about this notion of transfer
Starting point is 00:11:18 when something's been trained on everything. Oh, okay. Especially in a field like math. Huh. Where, you know, trained on everything, I think has some meaning in a way. If you say it's been trained on everything that has to do with being human, clearly that's not the case. Right.
Starting point is 00:11:39 But if you say it's been trained on everything, having to do with math or with code, I don't know, is all of mathematical knowledge out there in some kind of textual or video format? Well, you're asking me. So the thing that is roiling our community in math lately as we try to make sense of what just happened is we used to think, okay, these machines are very good at searching, or these programs are good at searching big spaces. They have a tremendous amount of knowledge, because as you say, they've ingested the whole
Starting point is 00:12:10 internet and the Library of Congress and anything you can read, they've read. So anything where knowledge and the ability to search and to compute very fast and to not forget, all that, that plays into their strength. But to spot a connection between different branches that hadn't been noticed before and to exploit that to solve a longstanding problem. If a human being did that, we would consider that an aesthetic high point. You know, mathematicians love it when an idea from topology gets used to solve a problem in geometry or when an idea from algebra helps.
Starting point is 00:12:43 But then again, maybe it's sort of easy. If you know everything that's been done and you can look for a lot of possible connections, maybe you'll occasionally get lucky. So that's what it sort of seems like happened here. Yeah. No, I think that's right. I don't, you know, who knows how it happened because we can't really look at the innards of these models very well for many reasons.
Starting point is 00:13:04 But it is creative to bring two unexpected things together. and have something that's actually working. I consider that creative, but it sort of reminds me in a way, there was a math discovery program way back in the 70s, maybe, done by this guy, Douglas Lanat. It was called Eurisco, I think. And basically, it was trying to find new ideas in math. And it explicitly tried to bring together things
Starting point is 00:13:36 and stick them together, and it would generate hundreds and hundreds and hundreds and hundreds of these things. Most of them were just junk. But occasionally it would come up with something interesting. A human had to go in and look and say, is this interesting? The machine couldn't figure it out itself. So how much of that is going on here? I don't know.
Starting point is 00:13:57 I think here the difference is that the machine obviously is at a much bigger scale. And I don't know how many tokens of reasoning trace that it generates, in the course of solving this problem and how many kind of wrong paths it went down and how it figured out that it was on the right path. I mean, these are things that I think are part of the science of AI that not enough people are kind of pursuing right now. Yeah, let's get into that now because that's really where I wanted to go with you. It's a nice phrase, the science of AI. I'd like to encourage people to look at this article of yours, Melanie, about the six principles to assess cognitive capacity of AI. But just as a teaser, could you enunciate what are those six and say a little
Starting point is 00:14:45 about them? Sure. So the first one is to be aware of your own anthropomorphic cognitive biases. So we tend to project human likeness onto things that talk to us in fluent English. So people very much think that these models have human-like qualities when maybe they actually don't. The The second one's a very common sense, one for scientists, be skeptical of hypotheses and develop control experiments. That's just like Science 101, although I'm not sure how often it's really followed through in science. People tend to like their own hypotheses.
Starting point is 00:15:25 The third is to develop novel variations of your stimuli or your benchmark items in order to test robustness and generalization. The fourth one is these systems don't have to be black boxes. You can probe them in many different ways, and we need more people who are very curious about why they're getting the results that they do get. Fifth principle is to consider performance versus competence, sort of what you can show that you can do
Starting point is 00:15:53 versus what you actually can do. And in the paper, I give some examples of that. The sixth is to analyze failure types and to embrace any negative results. We tend to put papers with negative results in a drawer and forget about them, but actually, they can be incredibly enlightening. We all have very direct experience with number six, don't we? When we see the hallucinations, it starts to make you wonder what's really going on with these systems. It's true.
Starting point is 00:16:28 You learn a lot from the errors. Yeah, people celebrate their positive results, and they try to explain away their negative results, but it's important to really understand what's going on by looking at where it fails. So one example that you give in your article, this is not about AI, but this is about the kind of lesson from biology or from psychology that subtle things can be happening, that you need to have an alert and skeptical mind to notice what might really be going on. So could you just regale us with the old story of Clever Hans?
Starting point is 00:17:01 So Clever Hans was a horse who lived in early 1900s in Germany, and Clever Hans was able to answer arithmetic questions. So you'd say, like, what's 14 plus 12? And he would tap his hoof that many times. It looked like a genius horse. And people, including many scientists living back then, were very convinced that this was an animal who could do mathematics, who could count, who could reason about simple problems in the way that humans do. And people were very excited. But then a psychologist named Oscar Fulksd came along
Starting point is 00:17:45 and said, well, let's do some controlled experiments here, this notion of controlled experiments in psychology being kind of a new idea, I think. Let's see what happens if he can't see the person who's asking the question. Okay. And then he fails. And And it turns out that what he's doing is he's reading subtle cues on the face of the person who's asking the question. It turns out that if the person who's asking the question doesn't know the answer already, he also fails. Because what the person is doing is they're reacting to his hoof taps.
Starting point is 00:18:23 And when he gets to the answer, there's some unconscious signal they're sending that he's reading. So he is a genius horse. just not at the things that people thought he was a genius at. Instead, he's a genius at reading social signals in human faces. And so in this parable, then, as far as like when we are impressed by something seemingly genius that AI is doing, what is our lesson? That we should be doing controlled experiments or what? Right.
Starting point is 00:18:54 So an AI system was shown to be really good at reasoning about diagrams. in scientific papers, let's say. I think this is actually a real example, and could answer questions about them. But then the control experiment was, give the questions without showing the diagrams. Seems crazy, right? How could you answer questions about a diagram
Starting point is 00:19:19 without seeing the diagram? And it turned out that the AI could do this task because somehow there was some kind of spurious association between the words in the questions and the correct answer. So that seems like a case of poor experimental design on whoever was doing the benchmark attempt. In retrospect. In retrospect. And in retrospect, this happens all the time in psychology and other fields, I'm sure, too, poor experimental design.
Starting point is 00:19:49 Experimental design is a very hard thing. And there's all kinds of confounding possibilities. So this is why the notion of replication. in science became so important. If one group doesn't experiment, then they get a result. We shouldn't necessarily believe that result. That result might be due to some other aspect of their experimental design that wasn't intended. That's why it's very important for independent groups to replicate studies. This isn't something that people in AI do very much. No. And why not? Is it that the replication is not very glamorous because you're coming in second. Like there's no incentive. That's true in all
Starting point is 00:20:29 parts of science, right? I think that's true in all parts of science. But it's also because I think most of AI research is done by people whose background is in computer science or a related field that's not focused on experimental methodology. I'm a computer scientist. I never had to take a course in experimental methodology. No such course was ever offered to me in my department. It wasn't seen as part of what computer science was all about. And I think that's one of the things that's lacking in today's AI discussion. How can we trust the results of these experiments and studies that are done that show that AI can do all these different things? Fascinating. So it seems to me that there's this cognitive science version of the interference of the observer that everyone talks about in quantum mechanics, right?
Starting point is 00:21:23 The observer themselves is interfering with the experiment or the outcome of the experiment. And that is such an interesting role. Of course, is clever Hans very famous. And I agree that that is a very clever horse for being able to read the social cues. But how interesting if this is also happening with AI, that it's not just the role of the experimenter that's interfering. It's actually the role of the psychology of the experimenter that's interfering. Yeah, it's a whole dimension that many of us in the theoretical sciences and math don't get trained in, as Melanie freely admits. Right.
Starting point is 00:22:03 You know, I never took a course in experimental design. You as a physicist, I assume you had to take some experimental physics, but... Yeah, it doesn't really weigh in my actual work. It's really not experimental. Yeah. So I would not be a very good architect of a good experiment. Well, and it seems like it is something that's a very live issue, because I don't know. These days, the AI companies frequently use benchmarks to show how, well, to assess how far along are there systems on this quest for either artificial general intelligence or superhuman intelligence, that sort of thing, or even just to outcompete the other AI companies. We would like to know what the capacities are of these new machine learning systems and other AIs. Well, I think it might be that it's just, I don't think we really know how to evaluate human intelligence.
Starting point is 00:22:53 or to really know what somebody's doing when they're thinking. I don't think we know about ourselves. I don't think we can self-report very well. I can't say to you, oh, this is how it's working in here right now as I'm constructing the sentence. I listen to it and this was the process. I don't know. Right?
Starting point is 00:23:10 Yes. It's just natural. It just comes out. And I'm not that privy to the inner workings. And I feel the AI similarly. A lot of people have said, I've had conversations on our show before with other cognitive scientists and computer scientists. They say it's really hard for the AI to answer questions because a lot of people say,
Starting point is 00:23:28 why don't you just ask it? And it can't self-reflect either in an accurate way. This whole thought, the mystery of the black box, we use the term black box so often for the AI, but of course our own intelligence is a black box. Absolutely. Not just for mine to you, but even me to myself, as you're emphasizing. But it makes me wonder if there's a role for magicians. because, you know, magicians or sleight of hand people are so good at showing us our own
Starting point is 00:23:55 psychophysical limitations, how easily we're fooled or the sorts of cognitive errors we tend to make. And there are people who are analogous to the magicians who show the deficits in common sense of the AIs, right? They're sort of playing games that are almost like magic tricks on the AIs. I wonder how revealing those will be, you know, in a serious scientific way. Well, Melanie has a lot more to say about the depth of AI cognition and understanding and also how it might change whole fields of science, including math. We will be hearing more about that after the break. Welcome back to the Joy of Why. We're joined today by Santa Fe Institute, computer scientist Melanie Mitchell.
Starting point is 00:24:58 You have been a college professor for much of your life. When you're working with students, they can get the answers right, but as you start to probe what they actually understand, you start to realize that they might be getting the right answers for the wrong reasons. They don't really know what they're doing. And that's important if you want to be a helpful teacher. This brings up another point, competence versus performance. Can you expand on this idea and what would it mean in the AI context? So competence versus performance is kind of an old distinction from psychology and linguistics. The idea is that you might have the competence for a particular cognitive capacity,
Starting point is 00:25:39 but there might be some reasons why you can't perform the task that I'm giving you. Like they have the competence, they could solve the problems, but they're just emotionally frozen. There's some performance block. But then there's the other way around, which is performance without competence. So if the student in your office hours, say, had memorized a problem from the textbook and the solution, but they didn't understand the general principle. So if you gave them a slightly different version of the problem, they couldn't do it. That's performance without competence.
Starting point is 00:26:15 Okay. So if we would say that we're trying to work out ways of testing whether the AI understands, what would count as evidence? Suppose that you're an AI advocate who said that these new systems, because we've scaled them up, or because we have some nice new architecture with world models or social models or whatever, we've now crossed a threshold where they actually understand. It's not just that they can compute, they understand. What would count as evidence of understanding?
Starting point is 00:26:43 Oh, gosh. I hate to get pedantic about understanding. But there's so many different meanings of it. We had a talk here at Santa Fe Institute from a philosopher who broke down understanding into 25 different types. I didn't know what I was getting myself into with the question. There's like P understanding and G understanding, and there's this very long typography of understanding. And I'm not sure there is any sort of single notion of real understanding. One of the recent things that I and my collaborators have been working on is looking at different dimensions of understanding.
Starting point is 00:27:22 One example is you can get one of these language models or chatbots to generate a story. just generate a short story about something. And they will. They'll generate a very beautiful, little coherent short story. But then if you start asking them questions about the story, they often will fail in weird ways, even though they generated it. And I think the same thing is true in a lot of different tasks that they understand along one dimension, but not along another dimension.
Starting point is 00:27:53 And in some sense, deep understanding might be just you understand across many different of these dimensions. Uh-huh. That sounds like a promising direction. Let's talk about tasks a little more, because that's a phrase or a term that I've seen in some of your writing, the phrase, the tyranny of tasks. What's that about?
Starting point is 00:28:13 I first heard that from Shannon Valor, a philosopher. The idea is that in AI, the world is divided in terms of tasks. So when you think about what AI systems can do, People say, oh, they can make summaries. Let's test their ability to summarize articles. Or let's test their ability to answer questions about diagrams or, I don't know, some other benchmark. Well, I mean, these days, they've been benchmarked a lot on international mathematical Olympiad, very hard high school problems. Then there were research level problems. Now there's open problems that are unsolved in math. These are all like three levels of math. benchmarks that are out there. Right. Their capabilities are defined in terms of these benchmarks. You know, one benchmark might be the bar exam for law students, and they do really well in the bar exam. And so we say, oh, lawyers, you should be afraid. Your job is threatened because these
Starting point is 00:29:16 AI systems are getting as good as you are. But the way that we're defining that is by looking at how well they do on a specific set of questions or a task. And jobs as a whole are not the same as just one independent task after another. This is, I think it's almost like a fallacy that if an AI system can do a bunch of tasks, it can do the job of a person who's associated with those tasks. So just one example of this. So there's a famous quote from Jeffrey Hinton, where he said something like AI systems are incredibly good at diagnosing or interpreting radiology images. Nobody should go to school anymore to be a radiologist.
Starting point is 00:30:04 AI is going to take all the jobs within five years. Well, that was 2016. That was 10 years ago. Now we actually have a shortage of radiologists. I don't know if that's because he said that, but it turns out that even though AI systems can beat. human doctors on these benchmarks, that's not the same as doing this job out in the real world, which is much more open-ended, which is not just a series of well-defined tasks. Huh. Still, it does leave you wondering, like in the case of radiology, you could imagine,
Starting point is 00:30:39 if they are really good at that task, then what's left for the human radiologist? Should we still be in that part of the game? Like, in my own world of math, You know, if they're very good at proving theorems, but they're not so great yet at coming up with new concepts, or as we sometimes speak of it, theory building, right? There's this big distinction between problem solving and theory building. So is it that we're sort of going to find our niche, that we can do the parts that they don't do? So like in the case of radiology, they have the open-ended part, but not the scan reading part? I guess that's what I'm wondering. Is that how it's going to go?
Starting point is 00:31:18 Maybe. I wouldn't be at all surprised if jobs like yours change quite a bit because of these new tools. Yeah. These are going to become incredibly useful tools for mathematicians. So it might change your job, just like when personal computers came out. But there's a fantastic book by George Lakoff and Rafael Nunez about math and where ideas in math come from, the metaphors, and they feel that human embodiment is a very important part of understanding and mathematics. Exactly. I think that's our only hope, because right now they, the machines,
Starting point is 00:31:59 don't have great embodiment. And you're right that a lot of great ideas in math are inspired by experience with the world. And that's what I was going to say about applied math, that I feel like that's even more so than pure math, where we get so much inspiration from nature and from engineering and society and all that, that I think we have a lot more chance of being useful as humans in applied math. But I do think pure math will expire before applied math does. And maybe neither will. Maybe we'll just keep going forever. What does it look like to you? I mean, math is often thought of as some kind of gold standard. Like the AI companies have a lot of use for math, right? They can demonstrate how good their systems are because they can verify
Starting point is 00:32:44 that they've solved a problem or not? Well, that's a big question I have, which is suppose that your prediction comes right and pure math expires in some sense for humans. What does that mean for other fields? Does that mean that these machines are on their way to taking over everything? Or is it more like 1997 or whatever it was that Deep Blue beat Kasparov? And that actually beating the best of, human at chess did not necessarily mean that was going to go anywhere in other fields.
Starting point is 00:33:20 I don't know. What do you think? It feels to me like science is much more open-ended than math in that respect. Yeah, I believe that. I don't think that solving all the Erdisch problems means that the average person has to fear for their job. Okay, now we have many different things on the table at that point. But even just in the world of pure brainiacs, whether it's scientists or mathematicians, just the fact that biology, there are so many things to be measured, we have so much data that we could collect, that we haven't collected, so many new ways of observing. I mean, that seems very inexhaustible to me compared to math. I agree. And even in physics, I think, which is maybe closer to math, there's so much,
Starting point is 00:34:07 you know, open-ended questions that aren't well formulated, that don't have something like a proof that can be constructed. So I do feel like the hope for math is to continue to take inspiration from the real world. And von Neumann had said something like that, too, that when math becomes too much art for art's sake, when it drifts too far from the source, which for him the source was nature or reality, if it becomes too far removed, it becomes sterile, said von Neumann. So I think this could be a really good era for pure. math if it starts taking more inspiration from nature.
Starting point is 00:34:48 That's been less so in the 20th and 21st century, but I think if we go back to that, we can probably eke out a few more centuries of human pleasure in math. Well, just say there's this dictum in AI, which is that easy things are hard and hard things are easy. Right. And pure math is seen by humans as like the most exalted exhibition of intelligence and brilliance. It's the hard thing.
Starting point is 00:35:15 And yet we know that hard things are easier for machines and easier things are harder. Yeah, and there's the word soft also, right? In science, we talk about the hard sciences and the soft sciences. And the soft sciences of economics and psychology and anthropology. Those are the really hard ones. Right. Well, so if we meet again in five years. The joy of gamma or something.
Starting point is 00:35:38 Yes, the joy of Omega by then, right. What do you hope we would understand about AI systems by then or what kinds of tests would we want to be able to do that we can't do today? Yeah, I mean, what I really hope will go well in the science of AI is this field called mechanistic interpretability, which is the neuroscience analog, where you're actually looking at the activations and the weights. Yes. and the, you know, all the messy innards of the system and understanding at a higher level, sort of what they are doing. These days, it's kind of a smallish subfield where people are trying to develop tools
Starting point is 00:36:23 that do that analogous to things like fMRI or whatever. And I don't think anybody's really figured out exactly how to do this the right way yet, but I'm hoping that's something that we can accomplish. And then we would have a genuine way of understanding sort of their limitations, what they can do, what they can't do, what kinds of mistakes they're likely to make and maybe how to fix them. Interesting that you put your finger on that because the first time I became aware of you, it was in connection with that in a broad sense. So what I'm thinking of is back when you used to work on something that in the jargon was called GAs for CAs. genetic algorithms for cellular automata, you and Jim Crutchfield were looking at this problem of evolving algorithms that could solve a certain class of problems, hard computer science
Starting point is 00:37:22 problems, and you were using this evolutionary algorithm to select better and better algorithms that kept improving through a kind of selection process. But then the part that you did that I found so creative is once you've got a really good system, you looked at it in what felt to me like an analog of mechanistic interpretability. You tried to see what was making that system so smart, analyzing it in terms of particles that were colliding with each other according to certain rules in the diagrams. I don't know if I've summarized it reasonably well, but it seems like this is a longstanding interest of yours.
Starting point is 00:38:01 Yeah, that's true. I hadn't made that connection exactly, but it is interesting, though, right? It's interpretability. It's interpretability. And it's also, I think, in the field of complex systems, people talk about this notion of emergence. Yeah. And we thought of that as a kind of emergent computation. And I think these AI systems also have emergent computations that are not easy to find, but they're there.
Starting point is 00:38:30 And if we understood them better, we would understand how the system is actually working, doing what it does. Yeah, it's an interesting. It feels, honestly, to me, very sweet and very old school. This hope that, okay, you're chuckling because you see where I'm going. It's a mean thing I'm saying. But this conceit that we with our limited minds can keep doing science, you know, and we're going to figure out how these AIs are doing what they're doing. And that's what our game will continue to be just like it always has been in science.
Starting point is 00:39:03 And the dark side of me thinks our days are numbered to be able to do that. As these gadgets get bigger and bigger, who says we can keep doing science on them and figuring them out? What's your reaction to that? We have nothing else to do. We have to try. That's an interesting question. Why do we do science in the first place? I mean, you know, we do science because we want to solve problems.
Starting point is 00:39:29 That's one thing. But we also do science because we're driven to understand things. Yes. You see this in little children. They're driven to understand. Often, one of their first words is why. They ask it constantly. So I think that's a human drive, and it's hard to fight against that.
Starting point is 00:39:52 And that's why you and I both went into science. It's important to us. Now, I was a little despairing when I was. went to a panel discussion at a conference on the role of AI in science. And there were a bunch of famous people on the panel talking about how AI was going to revolutionize weather prediction and genetics and cosmology and you name it. And I asked them, at the end, well, like, is this going to contribute to human understanding of the world? And they're like, why should we care about that. Yeah. To me, this is the bifurcation that we're all thinking about now, because science has
Starting point is 00:40:35 this double-edged aspect, that it gives us pleasure, we like figuring things out, there is the joy of why, and as you say, it's deep in our species. So yes, we're curious, but then there's the other side that for so long science has been this instrumental thing that helps us in technology and medicine. And I guess the question I have, and I think a lot of us have, is will we continue to take pleasure in the joy of curiosity when we are no longer the best at solving the important problems? But let me ask you one last thing. For people who haven't heard our earlier conversation, what was your draw to this field? And if you were starting out today, do you think you'd have the same kind of curiosity.
Starting point is 00:41:21 Yeah, that's a great question. When I was a child, I loved logic puzzles, like the knights and the naves, the knights who always told the truth and the naves who always lied. There's a fun, several books by Raymond Smolian, a mathematician who wrote a bunch of puzzles in this genre that I absolutely loved. When I got to college, I read Douglas Hofstadter's book, Gerdl Escher-Bach, which was was the real world version of these in a way. I mean, he was talking about Gerdell's theorem
Starting point is 00:41:55 and paradoxes in mathematical logic and how all this related to cognition and thinking and creativity and so on. And I was just completely blown away and that this is what I want to do in my life. I didn't exactly know what it was, but it seemed like it might be artificial intelligence. So I pursued Doug as an advisor and got
Starting point is 00:42:19 to join his group and was studying analogy via a new set of puzzles, which were analogy puzzles. And I was very entranced by all of that. If I were that age today, I would be worried. In fact, I have a son who is getting a PhD in machine learning, and he wants to do research in machine learning, but he's actually quite nervous that there will be no more roles for humans doing research in machine learning, because AI will be doing all the research in machine learning and improving itself and so on and so forth. And I wonder if I think the same thing. I don't know. Maybe we do have to revisit this in five years, because we may know by then, given how fast everything is going, who knows? I really appreciate your spending time.
Starting point is 00:43:14 with us. This has been a wide-ranging, a little bit amorphous conversation, but it's just wide open, and I can't think of a better guide to it. Thank you very much for joining us. Thanks, Steve. It's been great. Hmm. Hmm. I just remember being a student and learning Newton's laws for the first time, and then Kepler's laws, which really make Newton's laws beautiful, this application to the celestial cycles. I didn't think, oh, I'm not the best at this, therefore I shouldn't learn it. Nor did I think, unless I one day become the best at this, I cannot feel pleasure or joy in my experience of acquiring this information. Of course, lots of people study things that other people already know and are better at. So I sort of wonder if maybe the AI will know things before us,
Starting point is 00:44:12 but we will still need to acquire the understanding ourselves. And in that acquisition is a similar experience. Instead of maybe the AI will be a filter between us and interrogating nature directly, but we'll still be acquiring, I don't know, the knowledge and having that experience. I'm not sure. Maybe it's all going to pass us by. Well, let's explore this a little more. I like especially your emphasis on not being the best. And how, in a way, unfraught, that is.
Starting point is 00:44:46 I learned as soon as I went to college what it means to not be the best. You know, this fixation with being the number one, especially in an age of optimization, there's so many optimization algorithms, we talk about faster, cheaper. But in our own lives, very often we're not the best. I'm certainly not the best tennis player. I love to play tennis. I'm not the best chess player, and I'm still happy to play chess. and try to be the best dad, but I may not be.
Starting point is 00:45:14 But still, all these things are worth doing for their own sake, right? They give us pleasure. I do feel very philosophical and almost religious about this. Like we get a little time on earth alive. And these questions about AI do tap into questions about the meaning of life. What are we trying to do? If the meaning of life is that you're going to be the best in some domain, or you're going to make a discovery that's going to change the world,
Starting point is 00:45:42 then most people will have a meaningless life. And I just don't want to believe that's the correct version of the meaning of life. It was not for my dad. He didn't even get to go to college. You know, he grew up in the Depression. That was not an option. His life was being a good parent and taking care of the people that bought shoes at the shoe store that he had. Yeah.
Starting point is 00:46:02 And he knew everyone's shoe size in our little town. And he left a good name when he died. People remembered him well. Right. So, okay, what is that doing on our show here about science? Well, I think that, let's say the meaning for some people of life has to do with acquisition acquiring wealth. They're going to love this stuff, right? Because there's going to be this new tool that simply leverages all kinds of buttons that they now have faster access to and can exploit and acquire more wealth. There are people who found meaning in singing songs or writing poetry or being novelists. or doing math. And I think all of those fields are a little more nervous, right, about reevaluating what the place is going to be for them and how to secure that place and how to think about it. If I'm playing games of what may or may not happen, I mean, there is still a world in which
Starting point is 00:46:58 AI is like a supercomputer. And we've talked about this before, Steve, just because a supercomputer can crunch all of these numbers if it presents it to us as a string of symbols. Even though it has in some sense an answer, it's not a meaningful answer for us, and none of us value it. We still, as human beings, have a very important role between us and a supercomputer, rendering an image of a galaxy, or looking at an image of a biomedical neural map. It hasn't actually robbed scientists of their work. And so it might be that it really will continue to be a tool and not simply something that overtakes. and discards us?
Starting point is 00:47:39 Well, that's the question, right? I think there are two plausible scenarios. One is that it continues to be a tool, and we always have some essential role in science and math at the cutting edge. The other option is, and actually in my heart, I believe this is the case, that we will not be at the cutting edge, and that will happen very soon. And so then what is the point? Then I feel like it's still meaningful, just like when I was in high school and I discovered
Starting point is 00:48:08 things about math, they were discoveries to me. They were not discoveries to the world. That's what I mean, yeah. You know, I think we may have to all settle for that. We're not going to be making genuine discoveries for the world. The AIs will be doing that. I really do believe that's going to happen very soon. I may be wrong. I mean, there may be fundamental reasons why the AIs won't be able to do that. For instance, they don't have bodies. They don't have social life. You know, there's a lot, But I just think all that stuff will be solved before long. Anyway, what's your take? Well, I think there's a difference between making discoveries and understanding.
Starting point is 00:48:45 And I guess that's kind of what I mean in example. In some sense, maybe the supercomputer made the discovery before the person did, but we still say the person did. Because the discovery didn't count as a discovery until they rendered it in a way that human beings could comprehend. But I honestly don't know. I am not incredibly saddened or pessimistic. So I guess I would have to say that in my heart, intuitively, I am not terrified of this prospect. Maybe I should be, but maybe it's just sort of a bliss of being naive, and I'm just going to wait for it to sneak up on me. There's one thing I think we can be very optimistic about and hopeful about, which is I think we're going to have a glorious golden age of science where we will understand.
Starting point is 00:49:28 and discoveries by the AIs or by people in conjunction with AIs, that's all going to be happening in the next whatever, five, 10, 15 years. And it's going to be a spectacular fireworks time for science. And I think we'll hopefully, with any luck, we'll be alive to see all that. Yeah, there's definitely going to be a transition period where people are moving it fast and furious and they're part of the story and there's great accomplishment and it will be exciting to see. I know people very accomplished who are very excited about using it.
Starting point is 00:50:01 Use it every day. They have multiple things going on, and they just feel like their productivity has doubled or more. And they're excited. They're enjoying themselves. I think there's really nothing we can do but chime in and participate in this at least transition phase before we're obsolete. Well, I'm getting choked up just thinking about it. Thanks, Janice. It's always great to see you.
Starting point is 00:50:26 and we'll see you next time on The Joy of Why. Thanks, Steve. If you're enjoying The Joy of Why, and you're not already subscribed, hit the subscribe or follow button wherever you're listening. You can also leave a review for the show. It helps people find this podcast. Find articles, newsletters, videos, and more at Quantamagine.org. The Joy of Why is a podcast from Quantum Magazine,
Starting point is 00:50:52 an editorially independent publication supported by the Simons Foundation. Funding decisions by the Simons Foundation have no influence on the selection of topics, guests, or other editorial decisions in this podcast or in Quanta magazine. The Joy of Y is produced by PRX Productions. The production team is Caitlin Foles, Jade Abdul Malik, Genevieve Sponsler, and Merritt Jacob. The executive producer of PRX Productions is Jocelyn Kond. Gonzalez. Edwin Ochoa is our project manager. From Quanta magazine, Simon France and Samir Patel provided editorial guidance with support from Samuel Velasco, Kit Sudol, Simone Barr, and Michael
Starting point is 00:51:37 Canyon Golo. Samir Patel is Quanta's editor-in-chief. The episode art is by Chanel Nibbling, and our logo is by Jackie King and Christina Armitage. Special thanks to Garth Avery at the Cornell broadcast studio. I'm your host, Steve Strogatz. If you have any questions or comments, please email us at Quanta at simonsfoundation.org. Thanks for listening.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.