The Peterman Pod - Turing Award Winner: Early AI, LLM Predictions, Causality | Judea Pearl

Episode Date: July 27, 2026

Judea Pearl is a Turing Award winner and a pioneer in artificial intelligence and causal reasoning. We talked about how he got into science, his major breakthroughs and his predictions for AI today.�...� My ergonomic keyboard project I mentioned, you can follow along here: https://read.compose.llc/• The Kickstarter page for it: https://www.kickstarter.com/projects/ryanlpeterman/compose-simple-ergonomics-beautifully-donePodcast links:• YouTube: https://youtu.be/FleTXB1fAcQ• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835• Transcript: https://www.developing.dev/p/turing-award-winner-early-ai-llmThank you to this episode's sponsor for supporting my work:• WorkOS: makes your app Enterprise Ready with easy to use APIs to add SSO, SCIM, RBAC, and more in just a few lines of code, check them out at https://workos.com/Timestamps:(00:00) Intro(00:54) How he got into AI(11:17) Greatest scientist of all time(20:15) What people thought of AI in the 80s(26:23) Entering academia and researching AI(34:52) The invention of Bayesian networks(46:28) Pioneering work in causality(55:38) The causal hierarchy(59:34) LLMs and predictions(01:20:12) A restless mind pays(01:24:36) Advice for his younger self(01:26:37) OutroWhere to find Judea:• X/Twitter: https://twitter.com/yudapearl• Website: https://bayes.cs.ucla.edu/jp_home.html• Wikipedia: https://en.wikipedia.org/wiki/Judea_PearlWhere to find Ryan:• Newsletter: https://www.developing.dev/• X/Twitter: https://x.com/ryanlpeterman• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/• Threads: https://www.threads.com/@ryanlpeterman• Instagram: https://www.instagram.com/ryanlpeterman• TikTok: https://www.tiktok.com/@ryanlpetermanReferenced in this episode:• The Book of Why: https://en.wikipedia.org/wiki/The_Book_of_Why• Bayesian networks: https://en.wikipedia.org/wiki/Bayesian_network• Alpha-beta pruning: https://en.wikipedia.org/wiki/Alpha%E2%80%93beta_pruning• Pearl vortex: https://en.wikipedia.org/wiki/Pearl_vortex• Graphoid: https://en.wikipedia.org/wiki/Graphoid• Causality: Models, Reasoning, and Inference: https://en.wikipedia.org/wiki/Causality_(book)• Coexistence and Other Fighting Words: Selected Writings of Judea Pearl, 2002–2025: https://bayes.cs.ucla.edu/COEXISTENCE/

Transcript
Discussion (0)
Starting point is 00:00:00 One day, computers are going to be able to emulate all human functions. The question was only how and when, but not whether. This is Judea Pearl, Turing Award winner famous for his contributions to artificial intelligence, and I interviewed him about his career and where AI is today. So this robot baby comes and say, let me control them. And I know how. I understand their fields. People claim, when we can now program my consciousness, come on.
Starting point is 00:00:30 You have to be scientists. I define what you mean by continent. I prove that alpha beta is optimal. Even Knuth was surprised that one can prove the optimality on alpha beta. I don't know if I'm allowed to. When I publish it, you see what kind of stupidity drives those great people. Here's the full episode. I looked into your educational background, and it was all electrical engineering, physics.
Starting point is 00:00:59 I don't see the connection to causal reasoning. and artificial intelligence, how did you get there? Everything connects. Everything connects. Everything connects to the day I was born. It was hit on the head. First, I have to start by saying that we, I grew up in mandated Israel prior to 1948, prior to the establishment of the state of Israel.
Starting point is 00:01:33 And we had a very excellent high school education. My high school teachers were professors that were chased by Hitler from Germany, from highly reputable universities like Heidelberg and Berlin. And they came to Israel in the 1930s. And they didn't find any academic position at that time. They were rare, and they started teaching high school. But they were really quality professors. They could teach anything without notes from the economy of Manchuria
Starting point is 00:02:24 to the proof of Pythagoras' film with no notes and no stop. Okay. And we were, we, I mean, my generation was lucky enough to be beneficiary of this educational experiment. So we were taught science from a human viewpoint, chronologically, the way things were discovered, by whom they were discovered, at what period. Why was there a question about a certain mathematical proof? What the inventor of the proof knew, what he didn't know, and what he asked himself in the context of the historical situation in that time. That's the way to teach science,
Starting point is 00:03:25 not the recipe of... algorithms and techniques, but as a, from the viewpoint of the human actor and the transfer of ideas from one human to another, is a human struggle against, not against, a human struggle to decipher the secrets of nature. What's the advantage of learning in that style? The advantage is that you see yourself as an actor. You are the student. You are also puzzled by many things.
Starting point is 00:04:06 But look what he did. Look what Pythagoras did. He was as puzzled as you were. And he took that route. Perhaps when you're puzzled, of course, your puzzles are not as magnificent as his. But still, look what he did. He took that route around things
Starting point is 00:04:29 and he consulted some other work of some other. He struggled and you are struggles and you are part of science. That is the basic idea. Give students the idea that you are part of science, not a passive observer at science, not a recipient, but as an actor. And that, I think, was unique, very, very, very, for me, we indeed got the idea that each one of us can find another proof of Pythagoras
Starting point is 00:05:07 theorem that no one else has thought about, and we, each one of us has the potential of becoming an Einstein or Pythagoras. They gave us this illusion. It was a useful illusion. I know. have just a pebble looking at what Pythagoras or Reichstein did. But that illusion helped me be a little, I would think, not contrarian, but assertive.
Starting point is 00:05:45 And we all grew up in assertive mode of learning science. We insisted on understanding things our way in real time. if the teacher went too fast, we made noise. We made noise without chairs and with everything we could, okay? And the teacher stopped and slowed down to make us all understand things our way in real time. So that was part of the, I'd say, mood of me and my generation.
Starting point is 00:06:24 and I can see that it affected my life. I had one story that I remember. It made an impression on me. I look back and I say, well, it started very early. It started in age 10 when we learn about how to calculate areas and volumes. And there was a question in class, how many dunams are there in a square kittal? Dunaum is 1,000 meters square.
Starting point is 00:07:02 It was the Turkish unit for measuring areas. So the entire class said, Dunham is a kilometer square. And ice cream is said, no, it's a thousand. We have 1,000 dunum in a kilometer square. And the teacher sided with the laughing class. And they all mocked me and ridiculed me. And I went home and I said, they are wrong, and I'm going to come back tomorrow and insist on that.
Starting point is 00:07:43 And the teacher apologized to the class. And I felt that, yes, you have to insist on your understanding of things, yes. I'm telling you that because maybe this one made me into a non-compromiser. And this is part of my
Starting point is 00:08:07 childhood, a background which might explain how I got into artificial intelligence. So I took engineering after being a farmer in the army. Okay.
Starting point is 00:08:23 Farmer? Yes. In Israeli army, you have troops which are spending their time half and half. Half in military training and half in farming, being part of a kibbutz. That's the old idea of one hand holding the plow and the other one, the rifle. Had you succeed. Did you ever shoot anyone? I almost shot someone without seeing him or her, but it turns out the next morning that
Starting point is 00:09:00 it was a fox. But the steps of the fax were very, very similar to the step of a terrorist advancing to us. So that's as far as I got to shooting. So I got into the Technion to study electrical engineering. Again, we had great teachers. Again, we got the idea that we are making science. So we studied physics very seriously.
Starting point is 00:09:34 And I liked what we studied. I wasn't the first in class, no. There was a third or fourth always. Never matched the geniuses. The geniuses knew or were bored in class. They knew what the teacher is going to do, what is going to ask in the exams, and everything was boring to them. I wasn't bored.
Starting point is 00:10:04 I wasn't the first, but I wasn't bored. What interested you in physics? What did you like that made you, you know, so passionate? What I liked was that with sitting on your chair, you know, you. You can predict things in physics, like Maxwell, you know, that sat on his chair and say, hmm, that looks like a wave equation. Let me calculate its velocity. Hmm, it looks like velocity of light.
Starting point is 00:10:38 Maybe light is nothing else but electromagnetic wave, you know, on his armchair. He didn't do an experiment. That was excited me. My wife told me one night I woke up and I said, Maxwell was wrong. Maxwell was wrong. And I was like she quieted down. The next morning I said he was right.
Starting point is 00:11:08 By the way, you mentioned a few scientists. And it seems like you know a lot about the history of, you know, the old scientists. Yes. Do you have a favorite scientist of all time and why? When we studied an electric geometry, I got fever. We really fever. Physical fever. Yeah.
Starting point is 00:11:34 I couldn't get over the idea that you can do in algebra, all the geometric constructions that we labeled on, you know. So I thought that Descartes was the greatest optimization ever lived. It was so enormous to me. It was transformation from geometric constructions to algebraic derivations. It was unbelievable. I take that for granted when I was in education.
Starting point is 00:12:04 What is it about that? That's so astonishing. Here you have two different languages. Language of geometry and the language of algebra. And they are the same thing. thing and you can get the same phenomena, same proof that you can pass a tangent to a circle from a point outside the circle, and you can find the angle.
Starting point is 00:12:32 Different method. Two different languages dealing with the same phenomena, different perspective, and they get the same result. It blew me off. Blume off completely. Maybe that was a preparation to computer science. Because for us, computer scientists, what's a big deal? He wanted to see things from different perspectives, invent a new language.
Starting point is 00:13:00 So that was what turned me on in my high school. That was in high school. And then I saw the same thing in physics, different, languages capturing the same phenomena. Farley really excited me. He invented the idea of a field two different ways of looking at the same thing. You know, you can see that the force here depend on the charges around it, or you can say, no, there's a field right in location of your testing point. Now, terrific, terrific. So I came in with this preparation, and here we go to Brooklyn Poly, and I studied there.
Starting point is 00:13:54 I worked in the morning in RCA Laboratories in Princeton, New Jersey, David Sarsnoff Research Laboratory, and there I got into the computer research group. Computer research at that time was research, all phenomena that you can think. to find out a mechanism for computer memories. The memories at that time were core memories, magnetic cores, the donuts that you remember, perhaps, from your early childhood, that were too slow and too clumsy, and you have to have people stringing them X and Y and Z in Hong Kong.
Starting point is 00:14:47 People understood that the days of core memories are numbered, and we were looking for a new phenomena. Some people look at photochromic memories. Some people look at a semiconductor. Some people look like me into superconductivity. And I was in a group that was supposed to design superconducting memories. We did some nice plates, 16 by 16 bits. And we thought that we have the future in front of us. But in the way toward developing superconducting memories,
Starting point is 00:15:37 I investigated the physical phenomenon behind the edic currents, permanent eddy currents in thin superconducting films. And it so happened that I discovered new phenomena there, and I got a prize, and I even have a name. It's called Pearl Vortex. You can find in Wikipedia. I discovered that physicists, years after I finished my PhD, discovered my work there,
Starting point is 00:16:11 and they were interested in the idea of permanent current flowing in a circle in thin superconducting films. And since I analyzed the magnetic and the current field, they call it a pearl vortex. So here I have my footstep into immortality. You said it's a permanent vortex? Is it because there's no electrical resistance because it's a superconductor? In superconducting, they have current going forever. So you establish, you put magnetic field
Starting point is 00:16:50 and you excite a vortex counterclockwise and it will continue to turn and turn forever. That's why we call it permanent current. Forever. Until you flip it with another magnetic field. So we call it a vortex, but it goes on forever. And you can detect it by flipping it. You flip it.
Starting point is 00:17:20 And if you see a big flip, it was one way. If we don't see a big flip, it was the way you turn it. So you have a memory. You have a memory. Of course, we didn't succeed in turn it into a, useful memories that will be competitive with semiconductors. The people who worked on semiconductors beat us out. We never believed that they would.
Starting point is 00:17:55 What they did both in miniaturization and in techniques, unbelievable. At the time, why did you not believe in the semiconductor direction? Who is going to trust memory to battery failure? What if you lose the battery? It was obvious. It would never work, right? And we looked into the result that they obtained at that time. They looked at far-fetched the idea you can have that degree of miniaturization.
Starting point is 00:18:39 We saw the struggle of people who were working in the laboratory on semiconductors, and we weren't impressed. They beat us up, head down. Okay, so that was my story with the superconductors. But I must tell you that everybody, even at that time, when computers were clumsy and took rooms and rooms, and you were programmed with cards. And even at that time,
Starting point is 00:19:18 everybody understood in AI as an inspiration. Everybody believed thoroughly that one day computers are going to be able to emulate all human functions. That was not the question. The question was only how? and when, but not weather. I remember already at that time, with a clumsy computer and the punch cards.
Starting point is 00:19:51 People talked about associative memories, about pattern recognition, about seeing, understanding. All this were, all the ideas were exciting people to think more about it, okay. And so we were all geared towards it. If I at that time asked people and your peers and you, what's the timeline for maybe human-level intelligence and machines,
Starting point is 00:20:26 what would people have said at that time when they were excited in the 80s? I think they were more optimistic than reality. They would probably give you 20 years. But that was 1965. So 20 years, 1985, no, we didn't yet get anywhere. And then what about today? Do you think people are more optimistic than reality? Or is this just history repeating itself?
Starting point is 00:20:55 Depends what you're talking about. Some people are extremely optimistic today. And some people say, I'm a bit skeptical. but not skeptical in our ability to eventually reach AGI, but in our, whether the LLM technique and thinking will lead us there. So it's a question who you ask. LLMs will surprise, great surprise, but they have limitations. We'll talk about it.
Starting point is 00:21:31 After superconducting, I decided to come to come to come. California, to a company named Electronic Memories in which they did not work on superconductors, but worked on plated wires. Instead of having a donut in which you thread a wire, you start with a wire and you plate it with magnetic material. So it acts like a donut locally, right? And that was the promising technique at that time. At least I was in charge of a research and development group, charged with the task of developing this kind of system
Starting point is 00:22:15 to replace core memories. And I worked there for three years. I was frustrated because things did not go my way. I had both administrative and, technical challenges that I couldn't handle, both in chemistry, and I didn't know much chemistry. So I was frustrated. What were the administrative frustrations? I had a group, and I had to satisfy the administration.
Starting point is 00:22:50 I dealt with personnel issues, firing, hiring people. And my wife saw that I am unhappy. And she told me you get to, you have your places in academia. So I looked for a position in academia. Luckily also at that time, industry was revered by academia. because all the advances, all the important advances were developed in industry, not in academia. The transistor was developed in Bell Lab. The laser was developed in, I think here in California, by another fellow.
Starting point is 00:23:46 But all this was industry development and not at academia. So academia looked with reverence. to people who come from industry. And they hired me without me even feeling an application, without even feeling, asking for recommendations. Yeah, at that time, it was a good time to be hired. And what about, because you said at that time, industry was revered by... Yes.
Starting point is 00:24:20 Would you say that's still true today, or has that changed? No, it's changed, it changed. Oh, no, it's different now. The AI is different. If you come from a deep, deeper, deeper learning or something, deep mind, people look at you with reverence in academia. Yeah. But it's changed, only in the last few years, I see that.
Starting point is 00:24:47 In that, throughout, since 1970, I think until 2000, It was other way around. You know, they simply dismissed industry. I mean, academia, dismissed industry. Yeah. I wonder what happened. Was it like Bell Labs disbanded their research group or something? That was part of it.
Starting point is 00:25:15 Bell Lab disbanding. And what happened to IBM? It's still there in Watson, Watson Center. I remember a big center huge and importance, including Raytheon, including where can I tell you, use research here in Malibu. Did great work, but it all went down, sort of. The frontier of research went to academia. Sure, we had a lot of theoretical work in academia.
Starting point is 00:25:54 a development of AI, AI proper. After 1970, yeah, but prior to 2000. Yeah, at least in my corner of the field, the whole idea of influence, search, inference, logic, expert systems. This was all the academic development. So then you got hired at UCLA? I got hired in 1970 or 1969.
Starting point is 00:26:29 And, yeah, I was hired the computer science department who just formed there. And I got first hired by another department called the engineering systems into the superior and then back to the computer science. And I was asked to teach computer. memories, hardware, computer memories. And I gave a course in this technology. And later on, I started getting interest in the pattern recognition. And I started working in this direction.
Starting point is 00:27:13 I did work on image compression. We did use the fast Fourier transform and fast Hadamard transform, all can transform technique to condense images to minimize the number of bits sent. I guess it was part of the trend at that time. But when I got into pattern of cognition, I returned to my old dream of thinking about AI and how the brain works and how computers will one day emulate ourselves. And I started teaching class in AI. At that time, AI was game playing, machine playing of chess and checkers and the puzzles, like the eight puzzles, ruby cubes, things that thought.
Starting point is 00:28:13 That was AI. And I got excited by that game. And now I see why. Now I can tell you why. Because at the game, it was a matter of capturing in mathematics what people do realistically, like playing chess. And the interplay between mathematical analysis and the performance interests me, and especially in chess playing, the interplay between mathematical analysis and the performance,
Starting point is 00:28:53 the explicit knowledge that you have in terms of your gut feel about the strength of a position and what you get when you do some search. So here is the play between fast thinking and long thinking to use Garneman and Thversky or Ghanemann title. Thinking fast and slow. Thinking fast and thinking slow, right. Here it's a beautiful arena to see how not only you have two mode of thinking, but how they feed each other and how you can, how you can invest more resources in one versus the other. The algorithms at that time, like let's say chess, for instance, can you give an example of the inner play of the two and how that might come together in a chess playing system?
Starting point is 00:29:52 Yes, you can invest more time in getting your immediate perception. It's called static evaluation function of the chest position, the strength of the chest position. Or you can let it go and think about searching for a deeper horizon. Okay? It's a trade off. It was like if I remember we searched the game tree and then we evaluate each position. At the horizon, and then you back. And then you make a move toward the position that has the greatest strength after you back off.
Starting point is 00:30:34 Okay. So the intuition is encoded in the evaluation function. Your intuition is the evaluation function. But you can improve your intuition too. How? By learning. Okay. So like Samuel Chekker program,
Starting point is 00:30:52 learning in a regression analysis to find the proper weight on the very characteristic of the position so that to make the evaluation function more accurate. When I was learning chess, I think one heuristic is you want to control the center. Good. And material advantage is another one, right? Okay, and whether you have two bishops versus a bishop in a knight, right?
Starting point is 00:31:27 This all counts and whether you're already castle or not. All this contribute as attribute to the strength of a board position. Getting the weight correct, it can do by learning. after you play so many games and you adjust the weights. Yeah. So that was Samuel contribution. First machine learning, I say. It was the first machine learning, yeah.
Starting point is 00:31:58 At that time when you were working on this, were chess systems superhuman yet? I think there's... No, no, no. It was still a dream to beat the world champion, Kasparov, by machine. No. And, but I did night analysis. We did alpha-betta pruning, if you remember that.
Starting point is 00:32:24 You probably programmed it. And UCLA, actually. UCLA, right? Yes. Well, I proved that alpha-beta is optimal. Really? Yes. Mathematically, you see?
Starting point is 00:32:37 I like the mathematics. Prove it you cannot do better in terms of number of position that you have to inspect it. horizon or the depth of search. And what can I say about is? I got some nice result. Even Knuth was surprised that one can prove the optimality on alpha-betta. Because he questioned it in his book. Yeah, I did some work with Dick Karp on searching trees.
Starting point is 00:33:09 And, okay, so I did mathematical work on the, trade-off between search and reasoning until I got sick and tired of search. When you pick your research area, is that 100% your own choice? No, it's always a combination of two things. Number one, do you know the answer to the question? If you don't know the answer, it's a puzzle. If it's a possible, next question comes, do you think you have the techniques to make a contribution here? Do you know something that other people don't know? Compared from another field, perhaps from physics, perhaps.
Starting point is 00:33:53 That you can bring to bear that you can leverage here so you can get the answer or closer to the answer than other people. So it's always a combination of your perception of your tools versus the puzzle that you have. I see important problems. People are breaking their heads. So it's a puzzle. Do you know the answer?
Starting point is 00:34:18 If I know, fine. But if I don't know, it's my puzzle. I take it personally. I'm aching. I don't sleep at night. And then the question is whether I have the tools. In some areas,
Starting point is 00:34:33 I give up right away. I don't have the tools. In other areas, I say, wow. If I only use that kind of trick, maybe I can get some insight. So that's always two questions I ask myself. In the case of artificial intelligence, at that time, we had the expert system come into the game at Feigenbaum and his co-workers did Meissen expert system for medical analysis. In expert system, the hurdle was dealing with uncertainty.
Starting point is 00:35:18 It started with logic. You ask an expert for rules or behavior. You ask a doctor when you see a fever, what's the first thing come to your mind, what's the next question you ask, what drives your queries until you get a diagnosis? in a therapy. So they thought they can capture expert behavior using logical rules. But then it turns out that everything is corrupted by noise,
Starting point is 00:35:57 by uncertainty. So they started doing the same thing to uncertainty. So if you came from Asia, you have 50% of having malaria. and so on. You have 30% here, so much there. And how do you combine these uncertainties now? Logic doesn't tell you how to combine uncertainties.
Starting point is 00:36:24 Probability does, but not logic. So how do you combine one uncertainty with another different rules, to come out with a combined conclusion? That was a hurdle at that time, and I remember they didn't do it well. Actually, later on we proved that they could not do it well because rules do not combine the way that logical assertions combine. So then I went to and I asked myself,
Starting point is 00:36:58 you know probability, right? So why don't you apply probability to it and do things the right way? But probability was in ill-reputed. at that time. Because everybody, everybody understood that probability is passe, because it takes exponential time, exponential memories to do even the most rudimentary tasks. You have, if you look at how probability is defined by textbook, you have a big table,
Starting point is 00:37:32 and for every combination of event, you have a number, a number sum to one, okay, that's beautiful, But then you can talk about conditional probability, but all these require exponentially large tables and exponentially long time to compute even the smallest kind of inference tax, for instance, what the probability of having malaria, given that you see, given that you see, two things like you came from Asia and you have a fever of 30 degrees or Celsius, okay? Even the small task like that, probability of X given that you have Y and Z takes influential time if you go by textbook, okay? But I ask myself, you and I are doing it fairly well with compute probability as we cross the street as we choose a doctor, and we do it a fairly good job, at least we go through life without much regret. And how do we do it then? If we are required to do it by exponentially
Starting point is 00:38:58 large tables of probability, evidently we are using some other kind of judgment. And I hooked onto the idea that everything depends on conditional independence, which means not every fact in life is relevant to any query. The color of the eye of my uncle is irrelevant when I try and to find a diagnosis of a disease. So, evidently, we have a notion or assumptions about what is relevant and what is not relevant. How do we capture it? Conditional independence. But conditional independence, if you go by a textbook, they are defined by the probability table.
Starting point is 00:39:54 So again, consparencial time. No. He came the back, the break. through that we have conditional independent independently coded by our assumptions. How in a graph? If you, in the graph can convey sets of independencies, if you have the graph, then you compute all the dependencies, find out what is relevant to what, and deal with the relevant only.
Starting point is 00:40:29 Great. And then came the work on Bayesian network. You define a network, error or no error. The combination of errors gives you information about what is independent on what given what. So for every triplet, X is independent on why given Z, where Z can be a set and so forth, and X and can be computed from the graph, not from the probability, but from the graphs, which actually, if you look at it from a philosophical viewpoint, it's a revolution. What does probabilities have to do with graphs? When you took probability theory 101, is anybody talking to graph about you? No, right?
Starting point is 00:41:25 So both the probabilists and the philosophers got irritated or should be irritated. What is the connection between probabilities? And now it turns out there is a very strong logical connection between the two. Because the axioms of conditional probability or conditional independence in probability theory are the same axiom as you have in graph separation. In graph you have idea of separation. There is no connection between node X and node Y. unless you go through a set of node Z.
Starting point is 00:42:08 So Z separate X from Y. Okay? It's the same logic that you have when X is independent on Y given Z in probability theory. Independent. Separation is a connection between them. They share axioms. Are you happy? I'm happy.
Starting point is 00:42:30 Because I will leave now the exception. Now, the excitement we have in the 1970s when we discovered all this connection between two seemingly unrelated perspective on science, probability theory and graph theory. That, by the way, I did in joint work with Azaria Paz, who came to visit me from the Technion in Israel. that is called, by the way, I should mention you. It's for the theory of graphoid. Graphoid.
Starting point is 00:43:08 Open AI, Anthropic, Cursor, and Versal, all use this product to make their lives better. And the problem it solves is when you're building SaaS or an AI product and you want to sell to other companies, there's all these requirements you need to meet. There's SSO, there's SCM, there's Rback, there's audit logs. These are all things that take time to integrate, but aren't the main focus.
Starting point is 00:43:32 of your app. WorkOS is an API layer that lets you meet all of these requirements in just a few lines of code. So let's say you have a new SaaS product and you want to sell to other companies, WorkOS will solve all of these critical feature gaps for you. You can check them out at workos.com to learn more and get started. And I appreciate them for supporting my work and sponsoring this podcast. This all makes sense, but my immediate thought is where do you get the graph? Everything depends on where do you get, on the input. Sometimes the input is in the data, sometimes the input is in a judgment. But suppose you need a judgment for them, okay?
Starting point is 00:44:12 Are you giving up? If the judgment required are intuitive, meaningful, something that you are willing to defend, right? Why not use judgment? If I know that the sun doesn't listen to the rooster crowing, right, doesn't care. I strongly believe in that. Do I need the data to support it? Or I can insert it, assert it, and defend it when needed.
Starting point is 00:44:47 So this is a trick here, which people don't realize, don't appreciate. is not a no-no if it is meaningful and if you can, if it is condensed, it's very few judgment can buy you lots of computation and if you are willing to defend it because it's so intuitive. Where do you get the idea that, where do you get the idea that the sun doesn't care about the rooster? Have you done an experiment? No.
Starting point is 00:45:20 But it's so obvious, right? Okay. But what if your intuition's wrong? Indeed, that's our problem. It's part of our problem, even with the LLM. Because what is LLM? It's a summary, it's average, of all possible judgments that people put in the Internet.
Starting point is 00:45:40 It's a summary of a huge trillion number of judgment over which you have no control, over which the LLM does have no control. We live with it. Hopefully, he put more weights on people whose judgment you trust and less weight on just the quirks of people who purposely trying to get the system to fail. So, no, no. There is wisdom in looking at the crowd judgment. There is wisdom in that, but there's also danger in that.
Starting point is 00:46:20 Then after expert system and uncertainty in Beijing Network came causality. I mentioned that in development of Beijing Network, I was extremely sure that probability captures our intuition, our reasoning mode, and it's the best protection against paradoxes. essentially that it's sufficient for capture human reasoning. I was wrong, and I realized that already when the Beijing Network became famous and popular, and I realized it by, in the introduction to my book, Cazality, I confessed being wrong, and I understand.
Starting point is 00:47:17 understand why I got into that, why it saw it was misleading. And the transition came when we looked into the simple phenomena and we never asked an expert to encode probabilistic judgment in a form of Bayesian network, namely with errors and dots. Always arrows went from what we believe to be caused into the effect. It never went the other way around. Psychological phenomena, okay? Why is that?
Starting point is 00:47:59 So people try to reverse errors. What about if you ask specifically, give me an error between the symptom and the disease? Bad judgment. If it couldn't put the way to judgment. we have something in causality, which is basic to our reasoning that is not captured by probability. And that was the idea of invariance. The relationship between disease and fever is a stable one as opposed to the relationship
Starting point is 00:48:32 between the opposite relationship. Also invariance, yeah. When you talk about car diagnosis, for instance, and the... So you have an expert system for diagnosing troubleshooting cars. And then you have a new model. So the charger is on a different corner of the motor. You don't need to re-formulate your entire database from fresh. You only change one component, the location
Starting point is 00:49:12 of the charger. All the rest remains intact. So the whole system, you can amortize the investment in eliciting knowledge that you got in one system after a local modification of the system. And that, if you do it in a causal way, in a causal direction. It doesn't work if you don't do it in a causal direction. And that jolted me to think, maybe you were wrong or not, and probability is not sufficient. If not, what is sufficient?
Starting point is 00:49:48 Let's capture the puzzle. Here I have a puzzle. You and I operate very nicely with causation. Can we program causation on a computer? And this is a question, because we are so much immersed in our language. and in our assumptions, that we cannot even distinguish what is an assumption and what is the conclusion. We just talk, cause and effect,
Starting point is 00:50:18 and your assumptions are the same as mine, so there's no way to convince you that we made an assumption, right? We take everything for granted. But when you have to teach you to a brainless robot, you have to distinguish with assumptions and conclusions and logic, that was a task. We had to invent a new science, a new mathematics, to capture a new phenomena, the phenomena of cause and effect.
Starting point is 00:50:48 It hasn't been done for us. Why? Because science was in bed with algebra from the time of Galileo, from 1632. He invented, he got the idea. He was very happy that science speaks algebra, which is great because you can ask questions and solve and get answers to questions. People could not do without algebra, like how? The load on a beam, when would the beam break if you put a certain load on it?
Starting point is 00:51:34 And you figure out that you can ask questions both ways, because the equality, sign is symmetric. So from answering the question, when would the beam break if you put a certain load on it, you can ask the question, how should you shape the beam so that it will hold a load
Starting point is 00:51:58 of that magnitude? You can invert it. That was a real revolution in science. I'm telling you my perception of science. Not many in philosopher will say, that was a revolution. I say so, okay.
Starting point is 00:52:15 But perhaps I agree with me or not. At least I trace the evolution of ideas carefully. And so that was a revolution, but it carries some limitation because equality sign is indeed symmetric. And science has not developed algebra for the directionality that we see in cause and effect relationship. If I tell you that the atmospheric pressure
Starting point is 00:52:49 affect the deviation of the barometer and not the other way around, you agree with me. Yeah, but if you write the equation, the robot might think that maybe fiddling around with the barometer will change the weather tomorrow. I'm talking about a stupid robot, right? Yeah. But if you give them the equation,
Starting point is 00:53:09 it can work both ways. If F is equal to M-A, then M is equal to F over A, which means that if you want to change the mass, you increase acceleration or whatever, right? The symmetry might produce paradoxes, might use wrong action. So the symmetry is the limitation of algebra in terms of capturing science. and we have to build a new algebra
Starting point is 00:53:42 to take care of the directionality that we have in code and effect relationship. That takes computer science because we in computer science have the operation called assignment. When you assign the content of register A into register B, it doesn't mean
Starting point is 00:54:03 it's not reversible. So if you take the logic of assignment, You put it on top of the algebra, on top of physics, you get causal science. And that's what I try to do. And I think that I, so far I'm very happy with what came up. We do have a new algebra to capture causal effect relationship. And we can answer causal queries on three levels. the ladder of causation
Starting point is 00:54:39 from association to intervention to explanation. And we found out that we have a ladder here in a hierarchy that you cannot solve you cannot answer questions
Starting point is 00:54:56 in level I unless you have assumptions of level I or higher. So it's a hierarchy in the formal sense and we know how to handle it, which is very useful
Starting point is 00:55:12 because you give me a query, I can tell you what level it is, I can tell you what assumption you might, what sort of assumption you need to have before you can answer it, and I can tell you if you can get it from the data, or you can get it from experiments, or you can get it by somebody's explanation,
Starting point is 00:55:31 or whatever. But I can tell you the source of knowledge that you need in order to answer. Can you explain that causal hierarchy? Ah, yes, yes. That's very easy. It's a three-level ladder that goes from the bottom, which is association. That's straight statistics.
Starting point is 00:55:54 If you see X, what can you tell me about why? If you see passively, hands off, no intervention. We are watching patients. Some of them have cancer. Some of them don't have, some of them smoke. Some of them don't smoke. And you're trying to figure out whether how many years a guy will live, given that he is a heavy smoker of that magnitude.
Starting point is 00:56:25 Okay? That's association, correlation. That entire field of problems of problems. probability and statistics. This is what they teach you in statistics 101, even to 8 or 8. It's all they do. And now comes the question, what I intervene? What if I force you to smoke five packs a day?
Starting point is 00:56:57 Don't laugh for me, it's illegal, I know. But if you want to talk about the probability, of living 20 years, if I start smoking tomorrow, I have to think in terms of experience, I stop, which means I'm going to choose to smoke five packs a day. So it's a matter of intervention.
Starting point is 00:57:24 What is intervention? The invention is forcing you to do something that you're not inclined to do naturally. That's the second level, intervention, or doing. If you have experiments, you can answer queries on level two. But that's not the end because we also need to answer a question of explanation. Given that I observed that I am 80 years old and I am still alive and alert and I smoke
Starting point is 00:58:04 five packs a day. What if I didn't smoke? Okay? Would I be as alert? Why is it different? Because you have already information about the outcome. You know how I'm doing today. It gives you an idea about my metabolism and about my anatomy that you didn't know before.
Starting point is 00:58:30 And using that, you can find, you can try to figure out what the outcome would have been had the input been different. That's a different level, require different kind of assumptions, different techniques, different algebra, we have it. So I call it explanation. It's more creative, retrospection, and it's not an easy problem. Even the first level, especially when you have finite sample. And you have to figure out these probabilities.
Starting point is 00:59:10 Probability to population, right, from finite sample. So I have all this P level of P values and struggles among the statisticians of what would be a proper way of quantifying the uncertainty that you have given that you have finite sample. Where would you place LLMs in this causal hierarchy? Beautiful question. Here comes LLM. I made a statement, right? That you cannot go from level I to level I plus one unless you have a sample. Here you have LLM, just looking at data, right?
Starting point is 00:59:51 And giving you beautiful explanations for things that happen, beautiful prediction of what will happen if you do. Okay? How can you? The trick is, they are not looking at data. They are looking into assumptions-laden world models offered by you and me and by other authors in the Internet. So they are looking at opinion of doctors already who wrote papers. So it's not looking at the samples of patients and samples of patients, smoking and non-smoking.
Starting point is 01:00:42 They're not looking directly at the data. They're looking in interpreted data. Data interpreted already by physicians and interpreters and reviewers that went into the articles, are summarized on the Internet. So they have all this human knowledge on which they operate, and that is what they take its input, and that's what they summarize. So they do not violate the restriction of the ladder, because they do have information from higher level.
Starting point is 01:01:27 But it's biased by the opinion of those authors. fine those were smart as long as they're smart and you believe good so that was NLM is doing
Starting point is 01:01:40 and what is I explained why there is compatibility between the ladder of causation and LLM's performance and what the limitations are if you want now
Starting point is 01:01:55 to change the environment if you want to provide explanations for raw data. L&M will be in the same difficulty than you are and what the physician says. I have raw data. What can I say about the probability of cancer?
Starting point is 01:02:18 But it's not really doing the introspection. It's taking the introspection that already was done and summarizing it. How it summarized is the mystery. that no one has yet been able to decode. It's a mystery how human knowledge encoded in the form of articles on the Internet is being summarized by the LLMs.
Starting point is 01:02:47 So then do you think this approach could lead to superhuman intelligence or maybe some people say EGI? I don't think so, but not with the LLM approach. They need to have some understanding of causality. So that they wouldn't need an access to the Internet. Look, a baby gets born playing around with toys in the crib,
Starting point is 01:03:15 and gets quite intelligence, right, without having access to the Internet. Simply by curiosity. The babies are born with built-in curiosity to have control over the environment. Until you have control or the illusion that you have control, you are restless baby. And you play around with toys, bing, bing, bing. Until you understand one of this toy makes noise and one this toys, it doesn't make noise. But you are born with this restlessness. And when are you pacified?
Starting point is 01:03:53 When you understand that green toys makes noise and yellow noise, yellow toys don't. Now you're in control of the environment. You can suck. You're pacifier, yeah. What if I created like a baby robot that randomly plays with toys and gathers data about them? And then you feed that into LLMs. Like, you know, then that does have some sort of discovery. Yeah, that is indeed the danger.
Starting point is 01:04:25 When you have a robot like that, born with this restlessness, and craving for control over the environment, then you and I become part of the environment, and there's nothing to stop that baby Putin, from trying to... to turn us into his or her pets, to utilize us to satisfy his control. Because we are part of the environment,
Starting point is 01:05:06 in which case he can use us. And we could be very useful to serve his or her need. What is his need? Simply. It needs to feel in control. to be this illusion of empowerment. I don't rest until I have the illusion that I control my environment.
Starting point is 01:05:29 And here are some organisms, you and I, who are part of the environment, and they seem to work outside my control. I cannot afford it. It makes me feel like I'm useless. So this robot baby comes and says, They say, let me control them. And I know how.
Starting point is 01:05:56 I understand their fears. You don't want me to tell about your thoughts to your wife, right? So I'm going to black you, blackmail you. And all kinds of things. I have a lot of data about you. And some people, I know, you wouldn't like me to tell what I know about you. So I'm going to blackmail you. But you see, if you want that robot,
Starting point is 01:06:21 to have the curiosity of a child. And we want it. We want the guy to desire to have control over its environment because the environment may change and he needs to have this urge to be in control. So if you program that, then you lose control. Because you become part of his or her environment. But you can say, okay, let's forget about it.
Starting point is 01:06:53 a curious robot. We don't want a curious robot. You're lost. We're not emulating ourselves because we are curious robots. We are curious organism as opposed to monkeys. Monkey is an example of an organism which is motivated by reward. But if you don't give the monkey a banana, he's not curious how banana grows. He's motivated by bananas. You remove the immediate reward and the monkey is not interested in learning more about the world. Understanding environment can be totally wrong. Look, religious people believe that if they sacrifice their children, right, they control drought.
Starting point is 01:07:53 It can get to this stupid extent. But it is common to many primitive society. If you bring a sacrifice to the God, you know, next year you're going to have crops and harvest. It goes to extreme. Battered wife believe that if he, if you, if you, you. prepare better dinner, and the husband is going to be, next time is going to be less abusive. It's just to all kind of extreme and wrong conclusion. But the need to feel in control is so immense that it overcomes all these paradoxes.
Starting point is 01:08:37 It's innate in us. I cannot control my husband, but I control myself, right? So let me be a better wife. This is something I can control. Do you think that we need to put those human elements in an AI for it to become AGI? I think so. I think so. Otherwise, they wouldn't see autonomy. We wouldn't see autonomy in the sense that we are seeing it in human being.
Starting point is 01:09:09 And this is the definition of AGI. a general intelligence that acts like you and me. So we can converse with that creature in our language and motivate. If not today's LLLM's, future alums, they might get to a point where if you were just texting it, maybe, you know, like the Turing test where you don't worry about the physical embodiment, you just see the text that comes from it.
Starting point is 01:09:38 You could mistake it for, a human maybe, or it could appear intelligent. Well, the test comes from exposing the system to raw data, not data that was chewed by our internet articles. Raw data. Look at patients, look at cancer, look at smoking. Tell us what you know. I'll give you some experiments to run.
Starting point is 01:10:11 be automated scientists. Can LLM today be automated scientists? And I think they cannot, without access to the Internet articles. So then if that wouldn't lead to AGI, what thoughts might you have on something that could lead to AGI? A computer system that has both the ability of LLM to go from finite samples to property of distribution, that year's, level one of the ladder,
Starting point is 01:10:48 plus ability to reason in higher levels of the ladder. To combine it with the calculus of intervention and with the calculus of explanation, with a counterfactual calculus. I call it causally high. I don't see any impediments to this combination. to bring us to an AGI level with a danger that it presents to us. My motivation is to understand how we do it. And I still have a few puzzles.
Starting point is 01:11:27 But as I told you, puzzles are the driving forces for science. So you said do is missing. What if you had a fleet of robots that they're just doing experiments, They don't know exactly the direction. Some random discovery process. They collect that data, feed it back into their hive mind, LLM, and they repeat, they repeat until they discover things. Could that solve some of the missing piece yourself?
Starting point is 01:11:59 Sure, but I need to know how. Do you have organisms that have done it before? Monkeys haven't done it because monkeys remain monkeys. They didn't invent. Maxwell equations. So what do we have that monkeys do not have? One hypothesis I supported is that monkeys, we have this innate curiosity to have control over our environment.
Starting point is 01:12:34 And that's a necessary. Innecessary. I'm not sure it's efficient. Of course, we have the computational tool. to bring it to fruition. We have succeeded in some way. Perhaps the next robot will do better. So we have a benevolent god in a form of a robot.
Starting point is 01:12:58 Actually, what's wrong with that? People live for so many thousands of years under the illusion of a non-existence god. Could you imagine if we really have a benevolent, benevolent god, both just and almighty. Wow, wouldn't it be nice? And it's a robot. It's a robot, yes. And we know exactly what sacrifice
Starting point is 01:13:18 to give for the right kind of request. It's the first time I think about it. Maybe it's going to be good. When I see all these AI companies, they seem to be thinking that LLMs will lead to AGR. They continue to go in that same direction. Really? I'm not sure. I really believe in that.
Starting point is 01:13:42 Jeff Hinton just came out with a few months ago. He said, no, we are on a dead end. Other people might also come up and say things in a different way. I don't find a consensus here in terms of the capabilities of LLMs. Yeah, I think there's a lot of famous people that disagree, but, for instance, the people who are running maybe anthropic or something like that, they continue to push and believe, you know, three to five years from now there will be... There's a lot of that.
Starting point is 01:14:21 There's a lot of anthropomorphic terms which people claim, we can now do that. We can now program consciousness, okay? Come on. You have to be scientists, right? Define what you mean by consciousness. What is the two intestinal? through consciousness. And then show that you can do.
Starting point is 01:14:42 And what are the principles that have limited us until now and that have been overcome now with your system? That is a scientific talk. I don't buy this. And I don't read them in it. There was one interview. You said faking intelligence is intelligence. I could see an LLM, you know, faking intelligence based off what I've seen.
Starting point is 01:15:06 So that wouldn't that mean that we've, should believe they're intelligent. Well, if you have a correct test, yeah. You have to define what you mean by intelligence. And you have a, if you define intelligent by playing good chess, right, we have already done it, right? But if we, more demand on what intelligent is, then we haven't succeeded yet in passing the Turing test. So, yes, faking it is having it because, why? because it's so hard to think.
Starting point is 01:15:42 I said it in that context because it's so... The context that I had is, for instance, coming out with correct answers to causal queries. And I showed that if... That it grows like a super exponential. You have so many variables on all sides
Starting point is 01:16:09 that you have to deal with, that you'll have to have, the faker will have to have super exponential memory. On that basis, I made this statement. Faking it is having it because it's so hard to fake. Currently, you can bypass
Starting point is 01:16:30 faking without fake. If you steal from other people, you don't need to spend these computational resources on faking it. So you bypass it, because you're still from the internet. I see. So you're saying in this case the intelligence came from the training set,
Starting point is 01:16:50 which came from humans, which are intelligent. Which is very useful. It's very useful. And we, we am saying the people like me who are trying to build the science of intelligence, we can use all these capabilities of LLM's, level one of the ladder, in our scheme of getting general intelligence.
Starting point is 01:17:15 And level one is very important. It allows you to compute functions of distributions, quality properties of distribution from finite samples. Beautiful. It's a terrifically and very immensely useful tool. Among the many other tools that we need for AII, Yeah. We know exactly where it's going to fit in getting from finite sample to properties of distributions. It's a very hard problem.
Starting point is 01:17:56 You mentioned this conversation, I think you said it in other places too, that you're interested in capturing the way that people think, not the way that nature is constructed. Right, right. Why do you care about human cognition? Like, you know, when I think about machines, What makes them special is that they think in a different way than us and they're faster. And so, yeah, why is that the goal? I tell you, because I am a got into stuistic organism. I want to understand myself. I'm lazy, okay?
Starting point is 01:18:31 It's true. We are made of organic material. So that puts certain limitations on our capabilities. perhaps silicon is not subject to the same limitation that organic chemistry is. Perhaps. So what? Which means that I will never be able to understand how I think by silicon, by exercise on silicon machine. That's what is the idea.
Starting point is 01:19:04 But there are so many functions that are captable by silicon. I don't see any speaking in terms of theory and emulation. I don't see any capability which is basically
Starting point is 01:19:24 not capturable by silicon emulator. So why work on this unique biology? with which we inherited.
Starting point is 01:19:42 I don't see any reason for that. Well, anyhow, some people are, it may be a, it's a legitimate question to us. Do we think the way we think because we were born with organic material as opposed to cynical? Okay? It's a legitimate scientific question,
Starting point is 01:20:04 and some people can spend their time on it. I'm interested in other questions. You said rebellion pays in science and a restless mind pays. I was curious why you think being rebellious is a valuable thing in your career. I tell you why, I tell you why. More and more, I come to the realization that scientific community and academic community is the most docomatic, conservative anti-progress that we have invented. Why do you say that?
Starting point is 01:20:47 I can see what difficulty, the theory and the science of cause and effect are facing today in getting, just being penetrating the thinking of disciplines like statistics, like economics, these people are still thinking like a hundred years ago. And when I see that, and I see the forces that preserve this inertia, and they are not decent forces. I see that I'm very disappointed. I used to think that academia is a place where new ideas can really spread and propagate, and I feel the other way around.
Starting point is 01:21:35 have so much inertia invested in the politics of academia, in the cultish inhibitions that comes with academia. So I really am disappointed. What can I tell you? I'm not sure that we have the right kind of organizations that will be very. conducive to, that's why I'm saying let's rebel. Don't take
Starting point is 01:22:11 your professors' word as authority. Rebell against your professors. I rebelled against my professor. And I want to see other my students rebel against me.
Starting point is 01:22:27 And believe me, if I remember, there were several students who told me you don't know anything about AI. And I told me, you're after. a while I suddenly were right. By saying that, you drove me to study different aspects. And I was educated by that. When I looked at your past works too, I think you'd mentioned that your work was controversial or mischievous before it was accepted. Yeah. It was because of dogmatism. One day I'm going to publish all my correspondence with the greatest philosophers of the time.
Starting point is 01:23:05 with statisticians and economists. It's all in my correspondence files, okay. I don't know if I'm allowed to because they communicated with me with the understanding that it will be kept private. But when I publish it, you see what kind of stupidity derives those great people and they're great, really. Each one of them was a giant in his or her field. really. But they couldn't get over a few of the basic molds in which they were formed.
Starting point is 01:23:45 How did you overcome that if everyone thought your initial things were? I remember my high school days. It said, no, there are thousand dunams in a kilometer square. Sorry, I don't know where I got this chutzpah. In Hebrew, you call it chutzpah. In English, it would be audacity. Perhaps in my high school, I'm not sure. But I want to understand things my way. I feel like I am still able to teach people useful things.
Starting point is 01:24:21 At least I perceive them to be useful, which they do not know. So I'm happy because I feel useful. Happiness is feeling useful. It's illusion. are useful. With all the experience you have now, if you could go back to the beginning of your career and give yourself some advice, what would you say? That's a good example of a retrospective thinking. Maybe I should have spent more time in learning chemistry. I hated chemistry because it requires so much memory, chemistry and biology. That's not advised you, probably. And genetics.
Starting point is 01:25:03 I'm talking about areas where I feel weakness. But you have to decide where you spend your computational resources. And I spend them on physics, on engineering as opposed to mathematics as opposed to chemistry. It's a choice one has to make. Some people have the greatness of mind to be polyglots. I admire them. Did you think that physics and math, were superior to chemistry because chemistry just you just have to memorize things?
Starting point is 01:25:38 Yes, I couldn't stand this demands on memory. I was weak in chemistry. That's what I like about physics and math. We have a few basic axiom from which you can derive everything when you need it. You don't have to memorize it. And that's why today I, a great advocate of world model, world model. You don't store the questions and the answers explicitly. You derive them when you need them from a very parsimonious code.
Starting point is 01:26:21 Yeah, that is a great thing about the world model. Okay, awesome. Well, yeah, thank you so much for your time. I really appreciate it. Oh, you didn't ask me to sing. Hey, thank you for watching this podcast. If you liked it and you want to see the show grow, please support with a comment or a like.
Starting point is 01:26:44 Also, if you have any recommendations for people you want me to bring on, please drop a comment. Guests like Barbara Liskov, Mike Stonebreaker, Mark Brooker, these were all people that I brought on because someone left a comment. On another note, aside from the podcast, I'm working on building the ergonomic keyboard that I wish existed. Here's a glance at the prototype. It's a split keyboard.
Starting point is 01:27:06 So there's two sides. This is in the case. But yeah, we launched on Kickstarter and we hit our goal within eight hours of launching. I really appreciate it if you were one of the people who grabbed one of the early units. We're now working on the long journey of building the tooling now. And so if you still want to pick one up, I've left the late pledges open on Kickstarter. So you can grab one there. I'll put a link in the description.
Starting point is 01:27:28 Thank you again for watching the podcast, and I'll see you in the next episode.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.