Conversations with Tyler - Philip E. Tetlock on Forecasting and Foraging as a Fox

Episode Date: April 22, 2020

Accuracy is only one of the things we want from forecasters, says Philip Tetlock, a professor at the University of Pennsylvania and co-author of Superforecasting: The Art and Science of Prediction. P...eople also look to forecasters for ideological assurance, entertainment, and to minimize regret–such as that caused by not taking a global pandemic seriously enough. The best forecasters aren't just intelligent, but fox-like integrative thinkers capable of navigating values that are conflicting or in tension. He joined Tyler to discuss whether the world as a whole is becoming harder to predict, whether Goldman Sachs traders can beat forecasters, what inferences we can draw from analyzing the speech of politicians, the importance of interdisciplinary teams, the qualities he looks for in leaders, the reasons he's skeptical machine learning will outcompete his research team, the year he thinks the ascent of the West became inevitable, how research on counterfactuals can be applied to modern debates, why people with second cultures tend to make better forecasters, how to become more fox-like, and more. Read a full transcript enhanced with helpful links, or watch the full video. Recorded March 26th, 2020 Other ways to connect Follow us on Twitter and Instagram Follow Tyler on Twitter  Follow Philip on Twitter Email us: cowenconvos@mercatus.gmu.edu Subscribe at our newsletter page to have the latest Conversations with Tyler news sent straight to your inbox. 

Transcript
Discussion (0)
Starting point is 00:00:02 Conversations with Tyler is produced by the Mercatus Center at George Mason University, bridging the gap between academic ideas and real-world problems. Learn more at Mercadist.org. And for more conversations, including videos, transcripts, and upcoming dates, visit Conversationswithtyler.com. Today I am speaking to Philip Tetlock, who quite simply is one of the greatest social scientists in the world. Up until now, we've been doing conversations with Tyler,
Starting point is 00:00:36 face-to-face. But for obvious reasons, Philip is in Philadelphia. He teaches at University of Pennsylvania, and I'm here in Arlington, Virginia. Let's just jump right into it. First question, Philip, with our forecasters, do we want accuracy, or do we want them to be a kind of portfolio to make us more aware of extreme events and possibilities? I think we want a lot of things from our forecasters, and accuracy is often not the first thing. I think that we look to forecasters for ideological reassurance. We look to forecasters for entertainment, and we look to forecasters for minimizing regret functions of various sorts, so that we would really regret not having anticipated X, Y, or Z, so we want to pump up the probabilities of those things.
Starting point is 00:01:24 But if we take, say, the coronavirus, if we had had a few more extreme nuts, who were maybe wrong most of the time, but insisting that we needed to fear the next pandemic, wouldn't we have been better off with that kind of portfolio? And thus, we don't actually want more accuracy from our forecasters. Well, in some sense, we already did have that portfolio. It was a mainstream position among epidemiologists for the last 20 years or so that you had a recipe for a disaster. David Epstein, the guy who recently wrote Range, a very interesting guy. You may have had him on your show. I don't know. But he recently quoted it. himself from a 2007 newsletter that he wrote. And he said something like the presence of a large
Starting point is 00:02:06 reservoir of SARS-like viruses among horseshoe bats combined with the culture of eating exotic meats is a time bomb. And that was in microbiology books. In the first decade of the 21st century, that was it was common knowledge. And indeed, even before SARS won, even before the first SARS outbreak. But even before SARS-1, epidemiologists were acutely aware of this. So it's not as though we didn't have it in our portfolio. We did. So those forecasters maybe weren't entertaining enough. Isn't then the margin we want to work on to make our better forecasters more entertaining and not more accurate? Yes, no? Well, the signal to noise ratio isn't going to be great. I can assure you of that because there are plenty of people who are naturally more entertaining than epidemiologists.
Starting point is 00:02:58 But maybe the whole portfolio needs to be more vivid rather than trying to fine-tune the accuracy of particular parts of it, right? Well, you have lots of people competing in the marketplace of ideas for attention, and that's a hard competition for scientists to beat. What do you think of the argument that science only exists at all, because most scientists are overconfident, that if they were rational Bayesian, they would just latch on to the opinions of the smartest and best-trained people before them, that there's only progress precisely because people are making forecasting mistakes. Right. So you hear aversion to that argument, I suppose, on Wall Street as well. Sure.
Starting point is 00:03:38 I think there's a good deal of truth to it. I certainly have been guilty of overconfidence at many junctures of my career, thinking I'm going to be able to take on things that looked impossible and often turned out to be impossible. Most projects that most scientists embark on, I think, don't succeed. it does take a certain amount of quasi-irrational persistence. As with Columbus, right? Or the founding of the United States, arguably was irrational to break away from the British Empire, right?
Starting point is 00:04:07 Which was doing pretty well back then. It was a, it seemed like a risk-seeking move. But say I set up an alternate research program, and I sought to take forecasters and, A, make them more entertaining, and B, maybe I'd give them uppers so they were more over-concounter. confident. I mean, would that do the world good? Well, you could certainly have induced the epidemiologists who are worried about the horseshoe bats in central China or other possible sources of zoonotic viruses. You could certainly have induced them to pump up their probabilities. And you would, of course, started to run into a problem of crying wolf if they'd been saying there's a 30, 40, 50 percent chance of a viral leap into human beings each year.
Starting point is 00:04:50 and it didn't happen, didn't happen, didn't happen, didn't happen. People would, you get a crying wolf effect, right? Right. How do you think about financial markets in relation to your work on super predictors? Are financial markets, in essence, super predictors to begin with? Or can super predictors on average beat financial markets? Oh, boy. Well, you know, we play with prediction markets in the work with the intelligence community.
Starting point is 00:05:20 Going all the way back to Admiral Poindexter in the original DARPA effort to launch prediction markets inside the intelligence community, I think your colleague Robin Hanson was involved with some of that work around 2000. Actually, around 2003, 2004 in that range. We've also been working with the prediction markets in parallel with forecasting tournaments, and there are pros and cons to each method of eliciting judgments. But the IC, the inteldis community doesn't let us make those markets deep and liquid the way they are on Wall Street. people are essentially competing for reputational points the way they are in forecasting tournaments. The monetary prices are either small or non-existent.
Starting point is 00:05:58 So most economists, I think, would not consider that to be a very robust test of the efficacy of prediction markets. And when prediction markets fall short, and they do typically not perform quite as well as forecasting tournaments. When they fall short, it's hardly a decisive rebuke of the market mechanism for eliciting forecasts. I mean, there are lots of very powerful institutional actors like Goldman Sachs and so forth that are continually trying to do exactly what you describe. It's an ongoing process. I don't need to, you're an economist. I don't tell you that.
Starting point is 00:06:30 Sure. I mean, there are implicit prediction markets in, say, coronavirus, say prices of airline stocks, right? And those are very liquid. They've been very thickly traded lately. If you took your 10 best super forecasters and brought them into the hedge fund people, at Goldman Sachs, and you all sat down together. Who would be teaching whom?
Starting point is 00:06:51 It's an interesting experiment. And the person, okay, my project manager from the first set of forecasting tournaments, Terry Murray founded a company, Good Judgment Incorporated, which does things like that. So that's a proprietary venture. And, you know, you probably want to talk to Terry about how successful or not successful they've been in doing that. I think they've had some success. I think it's extremely hard to do that.
Starting point is 00:07:17 It's non-trivial. I think there's a good deal of similarity in the cognitive ability, cognizal profiles of super forecasters and the kinds of people you see on the staffs of Goldman Sachs. And it would be a tight race. What about the sports betting market?
Starting point is 00:07:33 Do you think there are inefficiencies in that because they don't have enough super forecasters? I'm not an expert on sports betting. I'm better to talk to Nate Silver about that. But let me put the question more generally. There are many markets out there which predict something. Sports betting markets are simply the most obviously most explicit about prediction. So if there was something the world didn't know about prediction already, those markets should be inefficient.
Starting point is 00:07:58 Yes or no? Why would you think the answer would be yes? Well, let's say that people favored the home team too much. So too many people might bet on the New York teams, the Los Angeles teams, and then the odds would be skewed. So if there's a bias in people without super forecasting techniques, we would expect sports odds to somehow be off, or at least they would have been off before your work was published. Well, arbitrage free-dated super forecasting.
Starting point is 00:08:28 Well, but isn't arbitrage itself super forecasting, right? People are arbitraging on the basis of some set of information. So all these markets out there, do you think they are without you already, the best available super-forecasters? The term best is a term I'm just not comfortable with. I rather doubt that they're at the optimal forecasting frontier at the moment, but they're often probably fairly close. And is there room to incentivize people to out-predict the market?
Starting point is 00:08:59 Well, that's one of those paradoxes that economists are written about, right? When people predict out to many decimal places, do you think that's absurd or do you think it's useful? What's the optimal level of granularity for different categories of forecasting? I think for the kinds of things we were looking at in the IARPA original forecasting tournaments with geopolitical events, like how long the Syrian Civil War would last, or what would Russia do in the eastern Ukraine or things of that sort. Yes, it would be absurd to go to three or four decimals points. Originally, the National Intelligence Council, which synthesizes a lot of intelligence analysis, they only distinguished five degrees of uncertainty, and they didn't put numbers on it. More recently, they've moved to seven degrees of of uncertainty and they do put numerical ranges on. So, you know, somewhat likely represents a certain probability range. Now, in our work, we've explored how granular the best forecasters are doing various rounding experiments, where we round their forecasts off to the nearest tent, that kind of thing.
Starting point is 00:10:01 If it's pseudo-precision, if when they adjust, when they move from 0.6 to 0.6, for example, if on average that doesn't, you know, improve their accuracy, we would conclude that, you know, that they can't achieve that level of granularity. Our best statistical estimates are forecasters for the types of questions the intelligence community often poses can distinguish between 10 and 15 degrees of uncertainty, which is considerably more than the seven. They think they can now, a lot more than the five. They thought they used to be able to distinguish.
Starting point is 00:10:30 But how useful is it to be able to distinguish varying degrees of uncertainty is going to hinge on the kind of game you're playing. If it's poker, you might well on someone who's very adept at distinguishing things to say three decimal points. If you could take just a bit of time away from your research and play in your own tournaments, are you as good as your own best super forecasters? I don't think so. I don't think I have the patience or the temperament for doing it.
Starting point is 00:10:57 I did give it a try in the second year of the first set of forecasting tournaments back in 2012. And I monitored, I monitored the aggregates. We had an aggregation algorithm that was performing very well at the time, and it was outperforming 99.8% of the forecasters from whom the composite was derived. So if I simply had predicted what the composite said at each point in time in that tournament, I would have been a super super forecaster. I would have been better than 99.8% of the super forecasters. So even though I knew that it was unlikely that I could outperform the composite,
Starting point is 00:11:33 I did research some questions where I thought the composite was excessively aggressive, and I tried to second guess it. And the net result of my efforts, instead of finishing in the top, you know, 0.02% or whatever, I think I finished in the middle of the super forecaster pack. So that doesn't mean I'm a super forecaster. It just means that when I tried to make forecasts better than the composite, I degraded the accuracy significantly. But what do you think is the kind of patience you're lacking?
Starting point is 00:12:00 Because if I look at your career, you've been working on these databases on this topic for, what, over 30 years? That's incredible patience, right? More patients than most of your super forecasters have. shown. So is there some disaggregated notion of patients where they have it and you don't? Yeah, they have a skill set. In the most recent tournaments we've been working on with them, that this becomes even more evident that their willingness to delve into the details of really pretty obscure problems for very minimal compensation. It's quite a quite extraordinary. They are intrinsically
Starting point is 00:12:33 cognitively motivated in a way that is quite remarkable. But how am I different from that? I guess I have a little bit of attention deficit disorder and my attention tends to Rome. So I've not just worked on forecasting tournaments. I mean, I've been fairly persistent in pursuing this topic since the mid-1980s. You know, back even before Gorbachev became General Party Secretary, I was doing a little bit of this. But I've been doing a lot of other things as well on the side. So my attention tends to Rome. I'm interested in taboo tradeoffs.
Starting point is 00:13:03 I'm interested in accountability. There are various things I've studied that don't quite fall in this. number. Doesn't that make you more of a fox, though? You know something about many different areas. I could ask you about antebellum American discourse before the Civil War, and you would know who had the smart arguments and who didn't, right? Well, I would know who has arguments to take the more integratively complex forms, on the one hand, on the other hand, and then synthesis, whether you want to consider those arguments smarter or not is another matter. But yes, I suppose that's fair. I mean, I've always resonated a little more to the fox's end of the hedgehogs.
Starting point is 00:13:39 But I think when you look at the great achievements in science, they often come from hedgehogs. If you look today at the ongoing debates about coronavirus and what will happen, and I mean now the debates amongst the smart people, the people you respect, what is the mistake you see them making? The biggest mistake. Oh, you want me to be an amateur epidemiologist here. No, mistake in reasoning. You don't have to give your numerical estimate.
Starting point is 00:14:07 But procedurally, what are they not getting right? Well, is it a mistake if you're a public health expert who feels that one mistake is much worse than the other? It's much better to overestimate the threat of the virus and to underestimate it because you have to influence public opinion and public behavior? There, you're not forecasting. You're engaged in manipulation, social influence. So, I mean, this comes back to your original. question about, you know, what do we want from our forecasters? And accuracy is only one of the
Starting point is 00:14:39 things we want from them. We look to forecasters for a lot of things, to inspire confidence, to inspire fear, and so forth. If we're trying to estimate how much people cut back on their risk-taking behavior because they're afraid of the virus, are the group of people best suited to do that epidemiologists, some other social scientists, or your super-forecasters? Maybe even economists, right? We study elastikisities. Why should it be the epidemiologists? I think you'd want an interdisciplinary team. I mean, diversity is one of these words that's been reduced to a cliche, but I think we have found in our work that cognitive diversity helps.
Starting point is 00:15:19 And it helps in certain quite well-defined ways. If you want to create a composite that out predicts the vast majority of the super forecasters, a good way to do it is not only to take the most recent forecast of the best forecasters in the domain, but it's also to extremize that forecast to the degree that people who normally disagree agree with each other. And when you have convergence among diverse observers, that's a signal that the weighted average composite is probably too conservative and you should extremize. Does the diverse team have a CEO, someone in charge?
Starting point is 00:15:55 Not in this case, no. That's done purely statistically. But would you put someone in charge? And maybe the person in charge would implement that statistical algorithm, right? But who's the person you would put in charge of the team? An epidemiologist, yourself, your best super forecaster? Bill Gates? That's a matter of managerial skill.
Starting point is 00:16:14 And I would say, you know, going back to one of my old dissertation advisors at Yale 40 plus years ago, Irv Janice on Groupthink, I would pick a leader who knows how to shut up and not reveal opinions at the beginning of the meeting and knows how to listen. And which group of people do you think that best describes? I think a lot of good executives have the intuition that you get more out of a team of forecasters or problem solvers. If you elicit independent judgments initially that are uncontaminated by conformity pressure, and then you create an environment which ideas can be freely critiqued before lifting the veil of anonymity and letting people see who's taking which positions.
Starting point is 00:16:52 Do you think having machine learning and artificial intelligence has made us much better at forecasting things? right now? I mean, social events, for sure. Sure, at the micro level, but social events. Whether there'll be a recession,
Starting point is 00:17:06 how many people will die from the coronavirus, will we settle Mars? Well, IARPA just ran a forecasting tournament called hybrid forecasting competition in which they pitted algorithmic approaches
Starting point is 00:17:17 and human approaches and hybrid approaches against each other. And I should let IARPA speak for itself about how well its programs work or don't work. I don't think there's a lot of evidence to support the claim that machine intelligence is well equipped to take on the sorts of problems
Starting point is 00:17:33 that the intelligence community wanted to have answered when it runs the forecasting tournaments has been running with our research team. Things like the Syrian Civil War, Russia and Ukraine, settlement on Mars. These are events for which base rates are elusive. It's not like you're screening credit card applicants for visa. The machine intelligence is going to dominate human intelligence. Totally. Machine intelligence dominates humans in Go and chess. It may now dominate humans and poker. I don't know. What the state of the art is quite there yet, but, you know, it is,
Starting point is 00:18:07 and they're StarCraft or, you know, whatever the next thing that Dennis Sosbis is going to conquer. So, but no, I don't see evidence that those approaches work in the domains that we study with the intelligence community. So do you think the hybrid man-machine approaches are overrated? It's very domain-specific. It sounds like a great, idea, who could be against it. But the devil lurks in the details, and it doesn't deliver as automatically as you might hope. It's trench warfare here. Do you think the world as a whole is becoming easier to predict or harder to predict? And again, I mean social events. I'm not sure there's a definite trend one way or the other. You hear a lot of talk, a lot of claims that there
Starting point is 00:18:49 are, but if you look back on the 20th century, there certainly were lots of major pockets of unpredictability. it's not clear. What if I say, again, current events aside, but it seems easier to predict. There hasn't been a world war since 1945. There's been steady economic growth in most parts of the world. More peace. Isn't that easier to predict? You just predict two to four percent global economic growth,
Starting point is 00:19:16 and you pick up a fair amount of what's happened since 1950? Indeed. Well, simple extrapolation algorithms historically are hard to be. And if you, you know, we just were running a COVID-19 mini forecasting tournament right now, and they're proving to be hard to beat. The skill, of course, is when to alter the trend, whether to accelerate it or to decelerate it or to change direction. And do you have a personal intuition on that? I think humans have been repeatedly humbled in competitions against simple statistical algorithms. Going back to Paul Meal, his famous little book on clinical versus actuarial approaches to predicting
Starting point is 00:19:54 in medicine and psychiatry, I would say be humble. Now, there's some of your early research that, if I read it properly, suggests that making people accountable leads to more evasion and self-deception on their part. Are you worried that you work with pundits by trying to make them more accountable will lead to more evasion and self-deception from them? Or how do you square early Tetlock and mid-period to late Tatlock? It's actually not too difficult in that particular case. It really depends on the type of accountability.
Starting point is 00:20:29 Tournaments create a very stark, monistic type of accountability in which one thing and only one thing matters, and that is accuracy. You get no points for playing to your – it's for being an ideological cheerleader and pumping up the probabilities of things that your team wants to be true or downplaying the probabilities of the things your team doesn't want to be true. You take a reputational hit. So the incentives are very unusually tightly aligned to favor accurate. that's extremely unusual in the social world. Most forms of accountability occur in organizational settings in which there are lots of distortions at work. And the rational political response for a decision maker located in most accountability matrices
Starting point is 00:21:09 and organizations is to engage in some mixture or strategic attitude shifting toward the views of important others or, as you put it, evasion, procrastination, and so forth. But given that when a pundit is on, say, the evening news, there's not a little box at the bottom that gives the Tetlock score of that pundit, correct? So most people don't know the actual record. So given out there somewhere is a measure of how good or bad the pundit is, are you worried that in a sense you will make those pundits run further away from objective standards precisely because they do poorly by them?
Starting point is 00:21:42 I don't know if we can make them run much further away from objective standards than they already are. but that's a very interesting point about forecasting tournaments. And I look at the kinds of people who are attracted to participate in. I mean, at the very outset, I mean, I invited lots of big shots to participate in forecasting tournaments, and they turned me down. They've repeatedly turned me down at a very interesting correspondence with William Sapphire in the 1980s about forecasting tournaments we could talk a little about later. But the upshot of this is that young people who are upwardly mobile see forecasting tournaments
Starting point is 00:22:16 it's an opportunity to rise. Old people like me, you know, aging baby boomer types who occupy a relatively high status inside organizations, see forecasting tournaments as a way to lose. You know, the best, if I'm a senior analyst inside an intelligence agency, and I'm on the National Intelligence Council, and I'm an expert on China, a go-to guy for the president on China,
Starting point is 00:22:39 and some upstart R&D operation called Liarpha says, hey, we're going to run for these forecasting tournaments in which we assess how well the analytics, the community can put probabilities unless Xi Jinping is going to do next, and I'll be on a level playing field competing against 25-year-olds, and I'm 65-year-old. How am I likely to react to this proposal to, this new method of doing business? It doesn't take a lot of empathy or bureaucratic imagination to suppose I'm going to try to mix this thing.
Starting point is 00:23:07 Which nation's government in the world do you think listens to you the most? You know, I, hmm. You may not know, right? I might, actually. Can I say? You know, I suppose the most prominent political fan I have at the moment is probably one of the senior advisors to Boris Johnson, Dominic Cummings, who recently caused a bit of a stir in the UK by appointing a super forecaster.
Starting point is 00:23:41 who had written some blogs that people interpreted as misogynist or racist or fascist or eugenicist. I don't know, some mixture of all those things. But I think his name was Andrew Sibisky, who's a young man 24, 25. The kind of the kind of person, young people are attracted, as I said before, to forecasting drugs. It's a fast track toward upward mobility. You have all the high status people making vague verbiage forecasts, people like me. But do you think Cummings actually is influenced by you? as I understand what he's doing correctly or not, he thinks he knows a bunch of things that other people do not.
Starting point is 00:24:18 And that seems somewhat non-Tedlachian, right? So maybe you're part of his portfolio of ideological armor. Maybe he's actually very non-petlachian, and it's the people in Singapore who are your true fans. Well, there are some people in Singapore, too, and that's an interesting place. You should mention that. Interesting you bring that one up. But, well, Mr. Cummings and Mr. Gove both separately brought up my work at various points during the Brexit debate. And Michael Gove, at least in the UK, famously said that Britain had had enough of experts.
Starting point is 00:24:51 I don't know if you remember that, of course, a particular quote. And he was thinking of, well, he invoked a support, at least, for that position, expert political judgment, in which portions of that book, you know, compare subject matter experts to, minimal statistical baselines like, you know, extrapolation algorithms. And the answer was often no. The goal was raising the point that, you know, where these guys get off, making these confident predictions about the consequences of the Brexit and the best empirical evidence would suggest they're probably not materially more accurate than simple extrapolation algorithms.
Starting point is 00:25:25 So, you know, that that was brought up. It was brought up for a political reason. He had a political point to make. And Dominic Cummings had a political point to make as well. He fears that I think parts of the civil service are hard. hostile toward Brexit and want to undermine the Boris Johnson administration objectives. Now, I think this comes to something deeper now. It's not just about the UK. It's about the intellectual fissure that exists between social science and conservatives, that most social
Starting point is 00:25:55 scientists are liberal, and conservatives are wary of advice from social scientists. I think we may have partly paid some price for that, but it kindly the slowness of the Trump administration response to COVID-19. Yes. Now, you brought up Britain. If we look back at speeches in the British House of Commons, who is giving the most cognitively complex speeches? I can't.
Starting point is 00:26:19 You know, Bob Putnam collected those data that I reported that study in 1984. But Bob Putnam collected the original data, and he reported it in a book, beliefs of politicians published in 1970s, and I think those data were based on interviews with members of the British House of Commons. They were not speeches. They were interviews,
Starting point is 00:26:35 confidential interviews that Bob Putnam got access to and put a lot of work into obtaining. And he shared them with me and I used them as grist for a small research program. I was running in the 1980s on Cognos style and political ideology trying to tease apart the rigidity of the right versus the ideologue hypotheses. And who sounds the smartest from that period? In that period, it was a mixture of moderate, labor, moderate conservative. It was the centrist that did better. It was slightly left-shifted.
Starting point is 00:27:09 Do you think we can draw any inferences from that about politics? Should we have more faith in the people who sound smarter or not at all? Well, I think that particular measure in integrative complexity is, I think it does have some correlation with forecasting accuracy. But I think you're picking up something more than just forecasting accuracy. You're picking up what I call value pluralism. You're picking up a tendency to end up. source values that are often in conflict with each other.
Starting point is 00:27:37 So the more frequently that you, as a political thinker, confront cognitive disness between your values, your value orientations, the more pressure you are to engage in integratively complex, synthetic thinking. When your values are more lopsided, it's easier to engage in what we call simpler modes of cognitive distance reduction, like denial or bolstering,
Starting point is 00:27:56 and spreading of the alternative. So downplay one values, push up the other value, and make your life stress free. At some ideological positions at some points in history require more tolerance for dissonance, and some people are more inclined to fill those roles. And you think those politicians are also likely to be better forecasters? You know, we're not talking about huge effects sizes here. Sure.
Starting point is 00:28:19 I think that fluid intelligence is probably a more powerful predictor. But a combination of fluid intelligence and integrative complexity, I think, does boost forecasting accuracy, yes. And if you were running the CIA, those people who are pluralistic in the manner you just outlined, would you promote them more rapidly? Now, of course, the CIA, bear in mind, is by statutes supposed to be value neutral. It's not, they're supposed to be feeding impartial, apolitical advice. Just no one believes that, right? Well, that is supposed to be the division of labor here. And forecasting tournaments are, I think one reason they may be interested in forecasting tournaments is because forecasting tournaments incentivize people to do one thing and only one thing, and that's accuracy.
Starting point is 00:29:01 You don't get points for skewing your judgments toward your favored cause. You take a hit on the long term by doing that. If you were in charge of the CIA and had a free hand, how would you reform it? Oh, boy. Well, there's a long history to efforts to reform the CIA, you know, going back to 19, well, it was founded in 1947. There have been various efforts since then. People have been unhappy with the CIA for many reasons. over time, you know, Vietnam being a big one, but there are lots of other reasons that people
Starting point is 00:29:38 have expressed unhappiness. Both liberals and conservatives have the various points been unhappy with the performance of the intelligence community. In 2001, it came to a kind of a crisis point, I think, and then there was a commission to reform intelligence analysis, and WMD fiasco in Iraq, that the pressure grew even more. One reason why we're even talking right now, today is that the intelligence community was forced, essentially, by the recommendations of the reform commission to take keeping score more seriously, to take training for accuracy and monitoring of accuracy more seriously. That's when they created the Intelligence Advanced Research Projects activity, which is the R&D branch housed within the office of the director of national
Starting point is 00:30:24 intelligence. And its job is to support innovative research that will improve the quality of intelligence analysis where, you know, that can be defined in various ways. But accuracy is certainly one of one of very important component of that. But it's not just, but it's not accuracy with a liberal skew or a conservative skew. It's supposed to be just plain, just the facts, ma'am accuracy. Which questions they choose to predict, that reflects values. Whenever a bureaucracy tells me they're value-free, I start getting more suspicious very quickly, right? I never believe them.
Starting point is 00:30:55 It might be okay for them not to be value-free, but the cynical response is the correct one here. Well, indeed, the old expression, there's no view from nowhere. There is no such thing as pure value neutrality. That doesn't mean it's not something worth aspiring to. But you're right, the values off, even if you had a perfectly objective forecasting tournament system, if you had people generating the questions promoting a political agenda, you could skew the results. I think that's one of your points, right? Now, in the middle of all these discourses, we have a segment called overrated versus underrated.
Starting point is 00:31:31 and I'll toss out a few names, ideas, and you tell me if you think they're overrated or underrated. How's that? Okay. Philadelphia, the city of Philadelphia, overrated or underrated. I got endless grief when my wife and I decided to leave Berkeley and move to Philadelphia. People thought that we were borderline insane. But we left for very personal reasons.
Starting point is 00:31:52 And without going into what those were, I would say Philadelphia has been a moderately pleasant surprise. It's a city that has many, many problems. But it's not as bad as the people in Northern California thought it was. Tolstoy. You know, I haven't, I read a little bit of Tolstoy, and I've seen a number of films. And I know some of the shorthand versions, and I know that Isaiah Berlin had a hell of a time classifying Tolstoy as a hedgehog or a fox.
Starting point is 00:32:20 But I don't think I'll pass on now. John Cleese. He's commented on you. You're allowed to comment on him. Right. Well, I had a wonderful time as a kid watching Monty Python, and also Faulty Towers. So I really enjoyed his, I've enjoyed his comedy. I haven't followed him since then.
Starting point is 00:32:40 But when I was younger, I thought he was absolutely hilarious and brilliant. And I appreciate the flattering things he said about my work. The television show The Sopranos. I fell for it. James Gandalfini fell in love with his performances. And then when he, in a way, One of the last things he did is he played Leon Panetta in Zero Dark 30. And there he is, you know, eliciting forecasts from people in the CIA about whether Osama bin Laden is in that compound in a battle of that.
Starting point is 00:33:09 Is he there or isn't he effing there? I did never do him, but I think very highly of his work. The threat of terrorism. Do we overrated or underrate it in the United States? That's a very difficult question because of the tail risk aspect to it. If you look at the number of people who died from terrorism versus other causes, it would seem that the amount of money we spend on suppressing terrorism would be disproportionate. But the tail risk complicates that a lot.
Starting point is 00:33:41 What is your favorite movie? And why? I don't have a hierarchy like that. I'm sorry. Nothing comes to mind. What's the movie you've seen the greatest number of times? That you can count. I did see myself coming back just very recently.
Starting point is 00:33:55 last week in the quarantine period, to Westworld. It's not a movie. It's a series, of course. But I think that is quite, I think the first two seasons are quite brilliant. On historical counterfactuals, by what year do you think the ascent of the West was more or less inevitable? Well, I have insight information here. We did a survey of some very prominent historians. We reported it in that book on Making the West.
Starting point is 00:34:23 We part of it anyway, also in an article. American Political Science Review. I think that if you looked at just the unweighted average of judgments for the median, I think was probably around 1730, 1740. And you think after that it was not very contingent. That would have been, you see, I'm not a historian of the West. I mean, what do I really know about that? I'm simply reporting the news here. And is there any great hinge of contingency that you think about looking backwards? Like, oh my goodness, there hadn't been a reformation, or if there hadn't been a Council of Trent, or what? Well, I, you know, I'm a fan of Steve Pinker. I think the Enlightenment was a big deal.
Starting point is 00:35:06 Because it got people thinking more rationally and more in terms of science. And that had a huge spillover effects. Now, could there have been an enlightenment in China, or could the Chinese have created certain types of technologies without science? What if they hadn't scuttled their Navy around 1400? all those sorts of counterfactuals. And how necessary or contingent do you think it is that we keep on thinking in enlightenment-like terms? Is it once you're locked into it, it keeps on going?
Starting point is 00:35:37 Or is it like good government that you have to renew it every generation or two? Well, I see the work I'm doing is very much in the spirit of the enlightenment, public, transparent standards of evidence for judging subject matter expertise. I think one of the great challenges of our time is striking the right, balance between democracy and technocracy. And I think the fissures that have emerged, it's not just, you know, a conservative, very skeptical of social science, and it's apparently some parts of biological science too, which is, I think, very unfortunate. But you see this kind of some reflexive skepticism towards science on the left as well. So, I mean, how do you manage the relation
Starting point is 00:36:14 between small D Democrats and technocrats in a society in which expert guidance is increasingly crucial? What are you learning from playing the game Civilization 5, or at least watching others do so? That's a good example, actually, of why I'm not a super forecaster. I don't play Civilization 5, but Civilization 5 was one of the simulations that IARPA chose to feature in its counterfactual forecasting tournaments under the rubric of focus. And if any of your listeners are interested in signing up to be forecasters for focus, we still have one more round, round 5, and we will be recruiting people. But again, it reflects my temperament. I don't have the patience for a game like Civ 5.
Starting point is 00:36:58 Forecasting is inevitably a mixture of fluid and crystallized intelligence. And you have to invest a lot of energy into mastering a game like Civilization 5. And I suppose as people get older, they may become less likely to make those kinds of cognitive investments. I mean, it becomes more and more essential as I get older, I think, to focus on the things where I have a real comparative advantage. The best chess players are all young, right? This we know. There's clear data. Yeah, it's interesting that domains in which child prodigies emerge, music and chess and math. If we take super counterfactualists and super forecasters,
Starting point is 00:37:36 do those two groups basically overlap, or how do they differ? I think they have to be rather intimately connected, although disentangling this one is going to be really, really hard, and it's one of the things I do want to dedicate a few years of my life, to doing. Now, obviously, when you say someone a good counterfactualizer, people shrugged their shoulders and they say, well, how are you possibly going to know whether, you know, you would have gotten, undo the assassination of the Archduke in 1914, do you undo World War I, undo Hitler, you undo World War II, you make Kennedy Grouchier during the Cuban Missile Crisis, do you trigger World War III?
Starting point is 00:38:13 you've got these sorts of arguments that are essentially unresolvable. You can't rerun history. You can rerun Civilization 5, but you can't rerun history. So that makes counterfactuals a place where ideologues can retreat. They can make up the data, make up whatever facts they want to justify. No matter how bad the war in Iraq went, you can always argue that things would have been worse if Saddam Hussein had remained in power. You have these factual and counterfactual reference points that people use and
Starting point is 00:38:43 debates implicitly to make to make rhetorical points. So a lot of people, part of what attracted to me to counterfactuals was, A, how important they are in drawing any lessons from history, B, how important they are in policy arguments, and C, how unresolvable they are. Now, one of the things I think we're hoping to do in the focus program is to develop some objective metrics for identifying people and methods of generating probabilities that produce superior counterfactual forecast and simulated worlds in which you can rerun history and assess what the probability distributions
Starting point is 00:39:18 of possible worlds are. Well, it turns out you get World War I, 37% of the time, even if you undo the assassination of the Archduke and you get something like World War II, you see where we're going. So we're hoping that one result of focus will be to help us identify people and methods
Starting point is 00:39:35 that generate superior counterfactual forecast and domains where there is a ground truth. The next task will be to connect superior performance in simulated worlds to superior performance in the actual world. This is where things get tricky, of course, because in the actual world, we don't have the ground truth. So if I ask you a counterfactual question of the form, you know, if NATO hadn't expanded eastward as far as it did in 2004 under the Baltics, NATO, NATO-Russia relations would be considerably friendlier than they are now.
Starting point is 00:40:10 We've got a counterfactual there. We can't rerun history. We don't know how relations with Russia would be if NATO hadn't gone into the Baltics. Russia would have gobbled up the Baltics. Maybe Russia would be friendly or would feel less threatened. You have people who have more hawkish or more doveish mental models of Russia. And those mental models predispose them to give you certain canned, almost ideologically, reflexive answers to those counterfactuals, right?
Starting point is 00:40:34 Right. So you can, but you can measure what people's beliefs are on the counterfactual. And then you can measure people's beliefs about conditional forecasts that are logically connected to the counterfactuals in kind of a Bayesian entrance network. If I know that you think that the Russians would be every bit as nasty and snarly, even if we hadn't moved it moved into the Baltic, it might even be nastier. It's probably a fair bit that you're also likely to think it's a good idea to increase arm sales to the Ukraine, ratchet up sanctions on Putin's cronies and so forth. So we can identify the counterfactual belief correlates of more or less accurate conditional forecasting. And in that sense, you can indirectly validate or invalidate. You can render more or less plausible certain counterfactual beliefs.
Starting point is 00:41:21 That's the longer term objective of this research program. We're not to play civilization five for the sake of getting better at civilization five. Their ultimate goal is to link the sophistication of counterfactual reasoning about the past to the the subtlety and the accuracy of conditional forecasts going into the future. Another thing you should observe, by the way, if people are becoming better counterfactual reasoners, is you should observe less ideological polarization in their counterfactual beliefs. The counterfactual beliefs should become as ideologically depolarized as conditional forecasts are. Does playing world of Warcraft a lot help you become a better forecaster or playing chess?
Starting point is 00:41:57 I don't have any evidence bearing on either of those things. But, you know, there's a third variable problem there too. But you get used to a test, right? you know if you lose. There's very little self-deception. This is true. I mean, the people who do well in these sorts of things often like games like that. Venture capitalists, when they try to spot talent and others, do you interpret their behavior in terms of a super-forecaster model? Someone like Peter Thiel, he found Mark Zuckerberg, Reed Hoffman, Elon Musk. He's a kind of super-forecaster. How do you super-forecast talent and other people?
Starting point is 00:42:30 Well, it really helps to be working in an environment in which super bright people are not super rare. It's very, very hard to forecast, to identify talent when the base rate falls below one in a thousand, one, a 10,000, one, and 100,000. You're looking for a needle in a haystack. But the great advantage that venture capitalists in Silicon Valley have is that the talent pool is relatively rich. So there's quite a few flaky people that come by seeking their money, for sure. But their odds of success are significantly better than they would be if they're working from a population-based rate. And, of course, they can tolerate a lot of mistakes because just a few hits will pay for a lot of false positives. But some do much, much better than others, right?
Starting point is 00:43:16 Mike Moritz, Peter Thiel. Right. The question is, are they doing better because they have better social networks? They're doing better or they have better judgment? Do you think super forecasting as a technique also applies to super forecasting how people will do? Yes. I think everything we know from the overlap between super forecasting and intelligence and the overlap between intelligence and success in many different professional lines of work would suggest that it was almost certainly true.
Starting point is 00:43:46 Who first super forecasted your own major success? I don't know what my first major success was. Some people might tell us whether there has been a major success in my career. but who first saw it coming? I think probably my advisor at University of British Columbia, Peter Sudfeld, who supported me and believed in me when it didn't seem to be very many good reasons for doing so. It's kind of a confused Canadian kid who was unsure whether to become a lawyer in Canada or go off to Oxford or go to graduate school in the U.S.
Starting point is 00:44:19 who tipped me toward social science in the U.S. And what did he see that other people had not seen? God only knows. Well, if anyone would know, it's you, right? You've lived with it for some time. I think he probably thought I was bright enough to do well, and he probably thought that I was contrarian and weird enough that there was some possibility of doing something distinctively well, something distinctive and different in doing it well. Do you think academic advisors in general today undervalue weird students?
Starting point is 00:44:51 I think there's probably a tendency in that direction, yes, because weirdness is a big word We're weirdness takes lots of forms that you and I would, we would not be embarrassed about turning away lots of weird people. Sure. Do you think people who grow up with the second culture are better forecasters? Oh, you're thinking of my work with Carmine, Carmine. Of course. Yeah.
Starting point is 00:45:13 Gosh, you really looked into my Vita here. This is highly unusual. I think there does seem to be some advantage there. And where does that come from? How does it work? I think it has to ties into accountability, accountability to conflicting audiences and value pluralism. You have a richer internal dialogue. You learn to balance conflicting perspectives more.
Starting point is 00:45:35 You have to be a better perspective taker. Perspective taking is a very important part of super forecasting, too. How do we create more Nate Silvers and Philip Tetlocks? What should we change in the world to get more of you? I think they and I are probably very different creatures. Sure, but we want more of you both, right? What should we do? You know, one thing I think that would be useful, I mean, there is this tendency for training in universities to have become hyper professionalized and compartmentalized.
Starting point is 00:46:02 So I think it is harder for people who have weird interests that straddles a psychology and organizational science and political science and law history. People who have weird sets of interests, it's hard for them to get traction in the current career environment. And certainly in my home discipline of psychology, I mean, you had this kind of rampant publication, inflation in which, you know, your PhD students get to get a job. It'd have to have a ridiculous number of publications. And some of them in A journals, you know, sort of thing we didn't really expect 30 years ago. We thought, oh, that's a tenure case. Oh, that's a junior higher. So those kinds of pressures, I think, produce a focus, a narrowing of focus.
Starting point is 00:46:44 Now you say, well, yes, what you, you know, the great, if I said earlier, and a lot of the great advances in science come from hedgehogs, we're producing hedgehogs on an industrial scale here. And I think there's some advantage that may get a little bit more room for weirdo eclectics. And the way the departments are carved up, would you change that at all? You know, a number of academics in the past, like Jim March, at UC Irvine many decades ago tried to do something like that. And Amy Gutman, the president of the University of Pennsylvania,
Starting point is 00:47:12 is trying to do that with, you know, Pan Integrates, knowledge, which is the chaired professorships that were used to, among other things, to hire to hire a number of other people. And we are sort of floaters. We're not connected to one unit. We're multiple connections. That's a very hospitable work environment, my point of view. But I don't see a lot of places doing a Jim March, UC Irvine experiment. Well, let's integrate all social sciences together. And Amy Gutman, P.I.K. And integrates knowledge kind of program. You don't see too many of them yet. It runs against the green. I'm not even sure it'll survive it beyond Amy, I mean, the natural tendency will be for departments to want to claw back the resources.
Starting point is 00:47:50 Let's say someone comes up to you and they say, Philip, I would like to be more fox-like. I'm not enough of a fox. What actual advice would you give them to achieve that? What should they do? Or not do? Wake up earlier in the morning, exercise more. Yeah, how do you commit career suicide, become more fox-like? Maybe they're not an academic.
Starting point is 00:48:10 They're a smart business person. How do they do this? I, you know, kind of obvious things like read a little bit more outside your field. If you're a liberal, read the Wall Street Journal. If you're conservative, you read the New York Times. You expose yourself to distant points of view. Try to cultivate some interests outside your field. Try to connect them together.
Starting point is 00:48:29 I think there's an optimal distance. I mean, so for history, for example, is, you know, sounds in it quite very different from what I did. Start as an experimental psychologist, and history looks very different. But they can be connected because historical judgment is something. that psychologists study to some degree. Psychologists are interested in hindsight and counterfactuals and so forth so you can link the two. So I think there's an optimal distance.
Starting point is 00:48:51 When you go foraging as a fox, you probably don't want to forage way, way far away. You want to forage far enough away that'll be stimulating, but still possible to reconnect. What kinds of people are best at adversarial collaboration? Rare people. Those are rare people also. But how do you spot it? Really hard to do. Is it a personality trait, a cognitive trait? Gosh, that's a hard one. I mean, that was a thing that Danny Conno, I think, coined the term when he was, you know, dealing with his various critics over time.
Starting point is 00:49:25 And my wife, Barb, actually, was involved in an adversarial collaboration between the Conneman camp and the Gigerensar camp on the conjunction fallacy. It took at least two or three years of her life. it's very hard to get people to, lots of us in principle are paparians. We believe in stating our beliefs as falsifiable hypotheses, and we also, most of us believe our beliefs are probabilistic, so we're kind of somewhat basian. But that's lip service.
Starting point is 00:49:53 There's what we believe, there's our formal set of, there's our epistemological, formal self-concept, which is kind of noble, volsificationist and probabilistic. And then there's how we actually, behave when our egos are at stake in particular controversies. And those things are quite different. I think Danny Connman, who's not known as an optimist, proposed it. It sounds like an optimistic
Starting point is 00:50:17 idea, but I think he's not all that optimistic about what it can achieve on close inspection. My efforts at adversary collaboration have not been all less successful. I'd like to jumpstart a few of them again, but it's very hard to find the right dance partners. Let's try a question from the realm of the everyday and the mundane. If I go around and I look at Mexican restaurants, I'm very good at predicting which ones have excellent tacos. What are you good at predicting? I'm not good at predicting your questions. What am I good at predicting?
Starting point is 00:50:55 I think I was pretty good at anticipating the fragility of a little bit of a little bit. lot of micro-social science knowledge prior to the replication crisis erupting. No, I mean everyday life. Oh, okay. So not in my... Not social events, not social science.
Starting point is 00:51:16 What in your life are you good at predicting? When you're going to get tired and want to go to bed at night or when the dog wants to eat, what is it? Well, we have a pet-free existence. And we don't... Our lives are actually pretty, pretty simple. I guess we subscribe to that old. added, you'd be boring and stayed in your life so you can be violence and creative in your work.
Starting point is 00:51:39 We have a kind of a routine here. So it's highly predictable. And now, you know, in the quarantine days, I mean, it used to be pre-quarantine, you know, we said, oh, you know, let's, we're going to go to Europe, we're going to go here or there. There were these little points of unpredictability that spiced up life, but, you know, those do not exist right now. And two more questions to close. First, what can you tell us about your next project? Well, I think the next project is the one that I mentioned earlier. It's linking historical counterfactual reasoning with conditional forecasting.
Starting point is 00:52:14 I think it'll be the second phase of the focus research tournaments. I think that counterfactual reasoning has for too long been the last refuge of ideological scoundrels. Insofar as we can improve the standards of evidence and proof in judging counterfactual claims as well as conditional forecasts and linking the two, I think it was a potential for improving the quality of debates among interested parties. And finally, what should a super forecaster predict about the future course of your influence? Oh, to be very cautious, because we're running against the grain. We're running against the psychological grain. We're running against human nature. We're running against a sociological grain. But the Enlightenment is stable, you tell us, right? So if it keeps on accumulating and growing,
Starting point is 00:52:55 your influence should be enormous. Based on your other presuppositions. Well, there's a lot of cognitive resistance to treating one's beliefs as falsifiable probabilistic propositions. People naturally gravitate toward thinking of their beliefs as ego-defining quasi-sacred possessions. That's one major source of, that's one major obstacle. Then you have the existing status hierarchies. You have subject matter experts who are entrenched to have influence.
Starting point is 00:53:19 Why would they want to participate in exercises in which the best possible outcome is a tie? They reaffirmed that they deserve the status that they already have. So that's not, psychological, sociological resistance. You know, it's interesting. You know, the sociologists and economists have kind of different reactions to forecasting tournaments. The sociologists, the reaction is, you know, why would anyone be naive enough to think that anyone would want to have a forecasting tournament in an organization, their status
Starting point is 00:53:42 disruptive, right? Yes. And economists would say, well, if these things are so great, how come they're not everywhere? And with that, Philip Tetlock, thank you very much. It's been a pleasure. Take care. Thanks for listening to Conversations with Tyler. You can subscribe to the podcast.
Starting point is 00:54:00 in iTunes, Stitcher, or your favorite podcast app. And if you like this podcast, please consider rating it on iTunes and leaving a review. This helps other people find the show.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.