Invest Like the Best with Patrick O'Shaughnessy - Jeremiah Lowin – Machine Learning in Investing – [Invest Like the Best, EP.105]

Episode Date: September 25, 2018

My guest this week is one of my best and oldest friends, Jeremiah Lowin. Jeremiah has had a fascinating career, starting with advanced work in statistics before moving into the risk management field i...n the hedge fund world. Through his career he has studied data, risk, statistics, and machine learning—the last of which is the topic of our conversation today.  He has now left the world of finance to found a company called Prefect, which is a framework for building data infrastructure. Prefect was inspired by observing frictions between data scientists and data engineers, and solves these problems with a functional API for defining and executing data workflows. These problems, while wonky, are ones I can relate to working in quantitative investing—and others that suffer from them out there will be nodding their heads. In full and fair disclosure, both me and my family are investors in Jeremiah’s business. You won’t have to worry about that potential conflict of interest in today’s conversation, though, because our focus is on the deployment of machine learning technologies in the realm of investing. What I love about talking to Jeremiah is that he is an optimist and a skeptic. He loves working with new statistical learning technologies, but often thinks they are overhyped or entirely unsuited to the tasks they are being used for. We get into some deep detail on how tests are set up, the importance of data, and how the minimization of error is a guiding light in machine learning and perhaps all of human learning, too. Let’s dive in. For more episodes go to InvestorFieldGuide.com/podcast. Sign up for the book club, where you’ll get a full investor curriculum and then 3-4 suggestions every month at InvestorFieldGuide.com/bookclub. Follow Patrick on Twitter at @patrick_oshag Show Notes 2:06 - (First Question) – What do people need to think about when considering using machine learning tools 3:19 – Types of problems that AI is perfect for 6:09 – Walking through an actual test and understanding the terminology 11:52 – Data in training: training set, test set, validation set 13:55 – The difference between machine learning and classical academic finance modelling 16:09 – What will the future of investing look like using these technologies 19:53 – The concept of stationarity 21:31 – Why you shouldn’t take for granted label formation in tests 24:12 – Ability for a model to shrug 26:13 – Hyper parameter tuning 28:16 – Categories of types of models 30:49 – Idea of a nearest neighbor or K-Means Algorithm 34:48 – Trees as the ultimate utility player in this landscape 38:00 – Features and data sets as the driver of edge in Machine Learning 40:12 – Key considerations when working through time series 42:05 – Pitfalls he has seen when folks try to build predictive market investing models 44:36 – Getting started 46:29 – Looking back at his career, what are some of the frontier vs settled applications of machine learning he has implemented 49:49 – Does intereptability matter in all of this 52:31 – How gradient decent fits into this whole picture     Learn More For more episodes go to InvestorFieldGuide.com/podcast.  Sign up for the book club, where you’ll get a full investor curriculum and then 3-4 suggestions every month at InvestorFieldGuide.com/bookclub Follow Patrick on twitter at @patrick_oshag

Transcript
Discussion (0)
Starting point is 00:00:03 Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest like the Best. This show is an open-ended exploration of markets, ideas, methods, stories, and of strategies that will help you better invest both your time and your money. You can learn more and stay up to date at investorfield guide.com. Patrick O'Shaughnessy is the CEO of O'Shaughnessy Asset Management. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of O'Shaughnessy Asset Management. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of O'Shaughnessy Asset Management may maintain positions and the securities discussed in this podcast. My guest this week is one of my best and oldest friends, Jeremiah Lohen. Jeremiah has had a fascinating career starting with advanced work and statistics before moving into risk management in the hedge fund world.
Starting point is 00:00:59 Through his career, he has studied data, risk, stats, and machine learning, the last of which is the topic of our conversation today. He has now left the world of finance to found a company called Prefect, which is a framework for building data infrastructure. Prefect was inspired by observing frictions between data scientists and data engineers and solves these problems with a functional API for defining and executing data workflows. These problems, while wonky, are ones I can relate to working in the quantitative investing world, and others that suffer from them out there will be nodding their heads right now. In full and fair disclosure, both me and my family are investors in Jeremiah's business. You won't have to worry about that potential conflict of interest in today's conversation,
Starting point is 00:01:35 though, because our focus is on the deployment of machine learning technologies in the realm of investing. What I love about talking to Jeremiah is that he is both an optimist and a skeptic. He loves working with new statistical learning technologies, but often thinks they are overhyped or entirely unsuited to the tasks they are being used for. We get into some deep detail on how tests are set up in this world, the importance of data, and how the minimization of error is a guiding light in machine learning and perhaps all of human learning, too. Let's dive in.
Starting point is 00:02:03 Where we'll start then is really with a question about what these models or methods are useful for and what they're not useful for. And the first time we talked, you used this idea of this is just souped up linear regression in a lot of interesting ways. But maybe we'll just begin there with machine learning is an exciting set of tools. What should people think about when deciding whether or not these are even appropriate things to consider? Yeah, it's a great question. It is an exciting set of tools. I personally am so excited about them. I once just Coltirk acquitted job to go learn about this stuff, but it's very, very easy to
Starting point is 00:02:39 end up with a chainsaw when all you need it is a butter knife. And that is sort of, I don't think people realize sometimes when they've ended up with a chainsaw, but that is the danger. That's what we're looking out for. And so I was a little bit tongue in cheek when I said it's all just souped up linear regression. But if you actually look at the math that's taking place, building an AI model is easy. It's just layering a bunch of regression models on top of each other with a little bit of finessing. It's training it that's really hard and where, you know, you could spend multiple careers and
Starting point is 00:03:08 multiple lifetimes gaining expertise. But building it and putting together is really easy. And that's sort of why you end up with a much more complicated tool than you probably thought you were getting or frankly even need. So what types of problems from, you know, starting very simply is this appropriate for? Maybe you can talk about classification versus regression and things like that. Why is everyone so excited about this? What are the main problems?
Starting point is 00:03:31 It's very important, I think, to understand when we talk about AI, what are we really talking about? We tend to look at these things as if they're black boxes, and to some degree they are. But the truth is we know what's happening in that black box with some certainty. And AIs are really good at only one thing, which is discovering complex correlations in data. In human terms, you know, that's something we call experience. And so the table stakes for having an effective AI is doing something that requires experience. And there's a subtext here, which is that AIs are really dumb. It's not a popular opinion, but I think it's a true one.
Starting point is 00:04:06 AIs are very dumb. They're not thinking. They're not drawing conclusions. They're not learning from the world. They're just, they're like puppies. They can do a small number of things, and sometimes they can do them well enough to survive. And that's at odds with the popular conception of AI. You look around, you see AIs that are driving cars and winning go championships.
Starting point is 00:04:24 And how do we reconcile that with the idea that it's just a puppy running around on the floor? and it's because those AIs that are doing things that appear very sophisticated are really just seeing there and they're redeploying experience in a very strictly defined and controlled way. And so even before we start to talk about whether it's classification or regression or many, many other things, I want to just really emphasize this idea that AIs do one thing. They take in data, they infer some correlation structure, and then they output a result. And that's it. So we can apply that.
Starting point is 00:04:55 We can redirect that result in a lot of different ways. So classification is a great place. to start. Classification means, is this A or is this B? We're asking the system basically to output a one or a zero. So is this email spam? Is a very popular example of the classification system? And a lot of inputs go in, what words are being used, what's the structure of the email, it goes through an AI. It doesn't matter what it is for purposes of our conversation right now. And out comes a guess, it is spam or does not spam. So that's classification. Regression is when we actually have a numerical relationship that we want to explore. So if we want to extend our email example,
Starting point is 00:05:33 that might be the probability that it's spam. That could be interesting. It's probably not that interesting. In a finance context, beta is a great example of regression. So the relationship between a stock in the market, a stock and a risk factor, that's an example of regression. Now, we know how to calculate betas. We use linear regression. Maybe that's not an interesting enough example for the purpose of our conversation about ML. But if we're, we know how to calculate betas, we use linear regression. Maybe that's not an interesting enough example for the purpose of our conversation about ML. But if we do believe that ML is just linear regression stacked over and over and over, then I feel perfectly comfortable talking about that and using it as an example of what
Starting point is 00:06:06 AI can do just in a very simple case. It might make sense to go all the way to the bottom of this to describe some of the terminology and how these tests are structured, talk about things like overfitting, error terms, gradient descent, etc. This might get a little bit wonky. but I believe that this is really useful to understand this stuff because if you think about it as magic, I think you're making a large mistake, both as a user and a practitioner of the tools. So maybe you could just describe a generic test using the terminology of the field.
Starting point is 00:06:38 So talk about features, feature engineering, testing, training, etc. Kind of walk through the life of a test, and then I'll poke and prod on different aspects of that to get your opinion. How about we stick with our email example? So we're going to classify an email as either spam or not spam. apologize in advance to anyone out there who actually does this for a living since I'm going to sort of invent this. We're going to start with our input. So what are we actually going to put into this model to let it make a decision? And for our purposes, let's just say we take the text of the email, including its subject. So it's a bunch of words. So the first problem we have is literally
Starting point is 00:07:08 how do we get that into the computer? We can't just copy and paste it. The model needs it in a very specific way and it needs it in a predictable way that's going to be common to all the emails it sees. So we can't even give it a list of all the words because that list of words, one email could have 100, another email could have a thousand, and the model's going to have a tough time differentiating between the two. So again, there's an entire wealth of literature just on this problem of representing data into the model. And I'll wave my hands a bit and say we're going to do something which is called a bag of words. The bag of words is basically a long list of ones and zeros. And each index in that list represents a word. So Ardbark is the first index and zebra is the last one. And
Starting point is 00:07:47 And if you have a one in the first place, that means Ardvark appeared in the email. And if you have a one in the last column, that means zebra appeared in the email. And by doing that, we can represent what words appeared in any email in a way that is the same length, the same shape, no matter what email we received. So that's where we start. We need to get the data in a form that we can present it to the model. Step two, and that's called those are the features very broadly. In this example, we just have one of them. So maybe we can introduce a second feature.
Starting point is 00:08:16 How about the domain that the email came from? That could be an important piece of data. So let's put that in and we'll call that our second feature. There is a practice. You referenced that Patrick called feature engineering. And this would be if we had data that we weren't immediately sure whether it was relevant to our model or how to extract relevance from it, then we might actually sit here and brainstorm a bit how we would either mathematically or qualitatively transform that input data
Starting point is 00:08:41 to make it more amenable to our model. In finance, we actually see this a lot with data that's, logarithmically or exponentially distributed. So that, you know, I'm thinking GDP numbers or something that scales with the size of a country, that data usually you want to transform it back into a linear space before you build a model on top of it. So this is sort of the equivalent of that in the ML space. So we've got data.
Starting point is 00:09:05 We've got it in a form that's useful to us that's common to all, you know, data that we're going to observe. Now we're going to actually get it to our model. And let's imagine that we have a model now. I don't think for the purposes of vocabulary we need to deal with what's happening. in the model. Let's just say we have a model. We're now going to ask that model to produce some output. And for us, that is whether or not the email was spam or not spam. And in this case, in this example, we're going to use a supervised model. Supervised model means that we are going to tell
Starting point is 00:09:32 the model the answers for the data that it's going to train on. So we're going to show it the text of email A, and then we're going to tell it that email A was either spam or not spam. Then we're going to show the text for email B, and we're going to tell it if email B was spam or not spam. We'll do that for all the training data. And it's supervised because we are literally supervising its learning process. We are making sure that it has access to the correct answer. And those answers in ML terms are usually called labels, sometimes targets, but in this case, we'll call them label since they're labeling a discrete outcome. And the goal of this model now is given all this data we're showing it, we're saying we're starting with A, the body of email A, we're ending up with whether email A is spam or not spam.
Starting point is 00:10:10 The goal of this model is to come up with a way of predicting that label. from the standardized input that we provided it, from the features. And that is called training. And training is another place where you could spend many lifetimes learning and getting it right. And so, of course, I'm going to do some massive glossing over here. But training is about discovering a correlation structure between inputs and outputs. And when we do training, or whether I should say when we train a model, there are many, many, many places that you can screw this up and not even know. It's just you have to be so careful.
Starting point is 00:10:44 You can overfit, which means that. you didn't give it enough training data and it learned something that was apparently correct, but in reality, false. A great example of that. Frankly, I don't even know if this is a true example or not, but there's a story told about satellite imagery and an AI that was trained to look for tanks. And it gave it all these satellite images and it found all the tanks and it reported that it was doing fantastically.
Starting point is 00:11:07 And so they put it into production. And then it just fell flat on its face the first time they actually tried it in practice. And it turned out that what had happened is all the images they had shown. shown it of tanks were taken at night. So what they had really built is a model that knew if it was nighttime. And that it's, again, I don't know if it's a true story. It's a common one. It makes the point well. Yeah, it's a common one among ML researchers because it does make the point. So if not that, that exact thing will happen over and over and over. And that's, of course, a latent correlation, which you're discovering after the fact. And you may not even realize you
Starting point is 00:11:42 made a mistake in discovering it. So we have to manage this training process, extremely carefully. And so you'll hear a lot these terms referring to actually the data used in training where we talk about a training set. We talk about a test set. We talk about a validation set. Describe those three at kind of a high level. So you talked about the training a little bit, the other two. Sure. They're super important. So let's forget validation for one second. What you really need is you need test and training. Training means the data that we discussed earlier that we're supervising and showing it to the model and telling the model what the right answer is and letting it learn the correlation structure. We're going to let it train on that data. And we're going to let
Starting point is 00:12:15 it see that data potentially more than once. You know, as many times, frankly, as it needs to learn what the correlation structure is. But when training is done and we want to evaluate the training, we can't ask it to tell us what it learned with the same data that it learned from. We need to test it on new examples, new emails to tell if they're spam or not spam. So that's called the test set. And surprisingly enough, we use it to test the model. We use it to see if the model did a good job. And that's important because the model has never seen the examples in the test set before. It never had an opportunity, you know, in a sort of diabolical sense, to cheat and to cache the answers to the test emails. It never saw them before.
Starting point is 00:12:55 And so we can get what we believe is a strong opinion of whether or not the model is doing a good job. Now, you also hear about this validation set. So the validation set comes into play because what if we do everything right? We take our data. We have a training set. We train the model many times on it. then we use a test set to see if the model was any good. But then based on the outcome on the test set, we say, well, wait a second.
Starting point is 00:13:17 Maybe if we tweak the model a little bit, we can get a better result on the test set. Well, now all of a sudden, the test set has in a very implicit way become part of the training process. It's part of this feedback loop that's sending us back into the training. And so we introduced something called the validation set. And the validation set serves as your halfway test set. So you're going to train on your training set. You're going to do validation, which is to say testing with a, validation set. And then you're going to use that validation set to go back and tweak the model
Starting point is 00:13:46 further. And finally, only when you're absolutely done, when you're confident that you have a model that works, then do you use your true out-of-sample test data set. Does that make sense? It makes perfect sense. And it's a good opportunity to highlight sort of the difference between machine learning predictive algorithms, derived predictive algorithms with kind of classic, especially in the finance, academic finance context, tests where, you know, everyone sees a T-Stat and Fama Macbeth regression to prove a factor works or something like this. ML is very different. You don't have a T-Stat in ML. And it's very, I view it as extremely pragmatic. Like the goal is to get a model that gives you the best predictions with less need to understand what's going on in the model.
Starting point is 00:14:28 Do you think that that's a fair line of difference between those two worlds of kind of classical academic finance and application of machine learning? I think it's a fair description. I think that they are actually so different. And this is, this is maybe where we start to leave behind the tongue and cheek thing I said earlier, that machine learning is just stack linear regression. Because when we build a machine learning model, we actually get away from, it's not that we don't care so much what the parameters do. It's that we understand that it's very possible that we just can't make a decision about what the parameters are doing. We're, we're setting up a model in which the parameter can be influenced by so many different things and in a time series context over such a long period of time,
Starting point is 00:15:05 that spending a lot of time deciding whether one parameter is actually influential or not, that could be a waste of that time. That parameter may turn out to only even be useful when you see an email that mentions Nigeria, just for example, or princes, for that matter. It may sit dormant the rest of the time. And it's a very hard thing to capture in a world that chooses to express most statistical confidence in terms of average and standard errors. So I completely agree with what you're saying, but I think it's important to distinguish that the two models are used very differently.
Starting point is 00:15:42 A machine learning model is used for its output. An academic model has become used for its output, and I think that's sort of the joy of quants, but was originally designed to actually explore the parameters themselves and to understand, like, what is the beta? What exactly is the beta? How is it measured? Is it significant? Is it different from zero? Or rather different than one, I should say. Machine learning models are serving a different purpose.
Starting point is 00:16:05 are producing hopefully useful output. They are pragmatic. I think you actually use a great word there, pragmatic. What do you think the future of investing in finance using these techniques is going to look like? This is something I've been trying to create sort of an inventory of, you know, among a major quantitative investment firms, you know, ours included amongst traditional firms, et cetera. How prevalent are these techniques in the research setting and in the production setting? And my sense, honestly, is that they're not all that prevalent right now, certainly not in production, which strikes me as a bit odd. And given how much time you spent thinking about machine learning specifically around time series
Starting point is 00:16:40 and also having spent so much time in investing and in markets, which adapt to changing conditions so effectively, I'm just curious your take on how important, let's say, fast forward five or 10 years, these kind of nonlinear ML techniques will be in our world. I think they're going to be critical. I think you won't be able to survive without some degree of analytic understanding, if only because the competition is going to get there first. So if on the margin, if all this does is let you see an opportunity a little bit sooner, get there faster, assess the price quicker, that's enough.
Starting point is 00:17:14 We've seen markets disappear for a lot less. And there's a lot of people out there who are very resistant to that, who don't want to hear that, who are incentivized not to hear that. Can you say a little bit more about that? Sure. Well, look, if your business is picking stocks through deep fundamental, research and you own 10 at a time and there's even a hint that a computer out there is going to do exactly what you do, but can do it for 10,000 stocks. That's just an extraordinarily
Starting point is 00:17:44 disruptive thing to happen, right? That's your business. Now, that's a huge stretch. And I'm not waving my hands and saying that there's a value investor in a box, just sitting out there waiting for you to plug it in. But I sit here and for all this talk about black boxes, And granted, I am very, very much a quant. I have always found the ultimate black box to be a person who, if they had dinner last night that didn't agree with them, might make a bad call today. And I have no way of knowing that. And that's a very tough thing to deal with for me.
Starting point is 00:18:18 Whereas if I have a model, I'm pretty sure if the model's wrong, it's because I screwed it up somewhere, or it wasn't able to learn something or the data was not amenable to what I was asking it. So when it comes to talking about black boxes, we often jump to this position of, well, humans can do it and machines cannot. And I actually say, well, the fact that humans can do it says to me that machines will be able to do it or humans aren't actually doing it and it's all just a giant accident and we're all just, you know, monkeys throwing darts at a board. Now, I choose not to believe that. And therefore, my conclusion is that a machine will be able to detect the same patterns as a human. I think the problem right now, the reason that we just don't see this being rolled out across the board
Starting point is 00:18:57 at every firm is the detail. of asking that question of building a machine to solve that problem is really hard. And I'm even illustrating that right now because I'm not saying what exactly that problem is. So let's try to define it. If we were going to build a value investor in a box, what would that system be solving for? It would make good trades. Let's start there. So what is a good trade?
Starting point is 00:19:19 Well, good trade is when you buy a stock and then it makes money. Okay, well, how do we decide if it made money? Isn't it obvious? It's made money if we sold it. Well, what if we had held it for an extra week? it might have been a better trade. Or what if we'd sold it a week earlier? That might have been a better trade.
Starting point is 00:19:33 Or what if we hadn't bought it at all? We bought something else that would have made less money but reduced the sharp ratio of our portfolio. That might have been a better trade. So the question of how to assess the quality of a trade explodes extremely quickly when you actually try to get it in a fine quantitative form that you can then challenge a computer to learn. It's a good opportunity to ask about stationarity. Just talk a little bit about this concept of stationery. It's important in.
Starting point is 00:19:59 running an ML algorithm, I think one of the things you hear as an objection is, especially with shorter term, let's say, total return labels, those are just non-stationary labels. And that makes the exercise very difficult. So can you define stationarity for us? Yeah, stationarity is usually used to refer to distributions that change over time. They're non-stationary in the sense that they move. And so stock prices are a great example of a non-stationary distribution. The price of a stock, you know, take a stock like Amazon. That price certainly moves all. the time and illustrates this principle. If you want to model this price of a stock, well, if you're modeling it at $100, it's very different than when you're going to model it at $1,000, literally
Starting point is 00:20:38 just because the output of your model has to shift. So if you have a model that can output whether or not it should be within $100 plus or minus $5, and now you all of a sudden are trying to apply it to a stock that's worth $1,000, that model literally will not work. And that's a very discrete outcome of the problem of non-stationarity. But in general, statistics doesn't deal well with non-stationary distributions. And so we seek ways to work with or transform them into stationary distributions. And the easiest way to do that with stocks and what we do all the time is instead of dealing with the stock price, we deal with the stock returns. So stock returns tend to form something that at least looks like a bell curve. It's not. It's got heavier tails. But it resembles
Starting point is 00:21:19 one enough that we believe that it's a stationary distribution. It doesn't change shape. It doesn't change its behavior. We can work with it and a model trained on it on one day will probably apply in the future. I think this idea of spending so much time, humans, spending so much time on the label side is really important. And it's kind of the way you put it is knowing what question to ask. And I mean, we've certainly spent a lot of time brainstorming on what kinds of labels might be interesting. Obviously, total returns is what everyone's after. But we also know that there are certain discrete things, maybe even more stationary that result in return. So maybe you use those as labels.
Starting point is 00:21:57 Maybe it's earnings or something like this or event-based things like a dividend cut. So I just want to highlight the importance of that label formation, certainly in our world, but I guess kind of across machine learning is something not to take for granted. Absolutely. And it's almost a form of feature engineering on the output side. And frankly, I don't even know what to call that other than feature engineering. Usually you hear feature engineering with inputs being correlated to simple outputs. But you're absolutely right in five.
Starting point is 00:22:23 the output is not obvious. It's this massive cross-temporal credit assignment problem if you want a system to trade. And so I've been fortunate in my career. I've worked with, I don't think it's an exaggeration to say hundreds of different quants who are all tackling this problem in very different ways. And when I get the most excited is when I see someone who is solving a problem, which A, they can describe, you know, at all, frankly, in real sort of quantitative terms that can be then, conveyed to a computer. But B, when they're tackling a problem that I believe is even solvable, that avoids this whole, again, cross-temporal credit assignment problem and difficulty of figuring out, was it a good trade, was it not? So picking a label, whether it's, you mentioned, earnings, or, you know, some cash flow forecast or a behavior of, I don't know, an analyst, or some potential correlate of the actual thing we're interested in, which is making money, of course,
Starting point is 00:23:24 but something that we believe is predictable on average. And I come from a school that believes strongly that the market is not predictable on average. I believe that it is occasionally quite predictable. And in some sense, if you want to win in an end-to-end machine learning approach to investing, you need to not only predict the market, you need to predict when the market can be predicted. And so the easiest way to win that game is not to play at all and to go find something else that you believe is useful to you in your investment process where, again, experience is the key determinant in assessing it. That's where we want to deploy machine learning. And then use the output of that machine learning model to supplant your existing potentially very human driven research and investment process.
Starting point is 00:24:11 Can you say any more about the experience of, you said, you know, maybe a hundred different. quantitative firms tackling this problem, all trying to do it empirically. If there are any commonalities beyond just the qualitative one you just mentioned, which is, you know, you can understand the problem that they're trying to solve, and they seem to be able to understand it. Were there any other common traits across those firms or teams or the types of questions they were asking that are notable? I'll tell you a personal bias that I have, because I would be remiss to sort of single out any firm that I've worked with in particular or anything they've said even sort of anonymous. So I'll tell you something that I believe strongly, which I think I've seen echoed in in certain successful places,
Starting point is 00:24:49 which is it's the ability of a model to shrug. And I know that sounds kind of silly, but if you think back to what we were talking about before about email and spam and not spam, it's, we're forcing answers out of the model. And we're going to, to the extent that we do believe that machine learning is very pragmatic science, we're going to do something with those answers. And again, this is very much my personal belief that the market is only occasionally predictable. I just won't trust. a model that doesn't have the ability to just shrug and say, well, I don't know what the hell you want me to do here. You know, I just not have an answer. So I look for models that that have that ability before I even evaluate whether I think they've been trained appropriately
Starting point is 00:25:29 or are applicable to the problem at hand. So for example, if you're trying to tell me that you've trained a neural network to take historical returns and predict future returns, I don't even want to hear the rest of that story. I know that that's just not viable. It's not compatible with my philosophy of how markets work. And again, we could spend a whole podcast on just that one statement. Why is that? What is it about a neural net that's not compatible? But if I start to hear people talking about probabilistic models or models that have any notion of a distribution or models that can key into certain types of things or literally shrug and return sort of a third outcome and I don't know outcome, I'm certainly more interested because all of a sudden we're dealing with a mathematical
Starting point is 00:26:08 description of the world that's compatible with my philosophical understanding of how the world works. Getting back to the actual arc of setting up one of these tests, so you've described, get your features in the email example, maybe it's the domain it's coming from and the bag of words. And then you've got your labeled output. So this is supervised learning and your test, your validation, your training, et cetera. Can you talk a little bit about this idea of hyper parameter tuning against something that humans are involved in in this process that's really important and kind of how that relates to overfitting a model? Yeah, of course. So you recall, when we were talking about all that, all the stuff that you just. mentioned, we also talked about when you would have a validation set. And you would introduce it when all of a sudden your test set becomes part of a feedback loop, where based on what the model outputs and
Starting point is 00:26:52 based on the characteristics of that output, you go back and you tweak it. You try a different model. So that process, in a formal sense, that is hyperparameter optimization. And we're not optimizing the parameters of the model. So just to draw that back to something tangible, those are, for example, the betas or the things that the model is optimizing in itself to produce an answer. We are are now optimizing the model itself. So again, in a beta setting, that might mean we go from three factors to four factors, or in a neural network setting that might be going from 100 neurons to a thousand neurons, or any other configuration of the model. And we try to do this in a scientific way, which is to say we don't just try to randomly throw things at the system and see what works
Starting point is 00:27:33 the best. Rather, we try to do it in a controlled way where we can probably test a hypothesis about the optimal number of neurons or the optimal tree depth or the optimal number of layers. And this hyperparameter tuning is where the validation set really is important and really shines. Because if you think about, let's say we want to test 100 different hyperparameter settings, well, that means that we're going to look at that validation set 100 different times and make a decision about the quality of our model on that basis. And now that it's in the feedback loop, it's being used to decide which iteration of the training was best. We can't honestly use it to assess the quality of the model as a whole, completely out of sample. And so that's where
Starting point is 00:28:12 having this completely separate held out test set becomes so critical. Maybe you could describe the major categories of types of models that are out there today. So you've mentioned a few of them, something like a neural net or a tree. Maybe just, if you wouldn't mind, take a quick inventory of what those major categories are, as I guess broad as you can boil it down. And then maybe with each sort of the kind of work that they tend to be used to do, like what they're appropriate for. And I'm going to ask in a few minutes just to warn you about things like learning theory.
Starting point is 00:28:42 so that certain models are better suited to other things. But just sort of like an inventory of the major models would be fascinating. So at a very high level, we have models that are suited to classification, which we discussed earlier is things like, is this email spam or not spam? And then we have models that are suited to regression. That's things like, what is the beta? We're looking for a quantitative relationship between two things. And models tend to fall into one of the other camp. So that's sort of, if we're drawing a box that we're going to put all these models in, those are the two columns across the top. And then we're going to have sort of two rows across the side, which are supervised models and unsupervised models. We discussed supervised models earlier.
Starting point is 00:29:18 Those are where we give the model the answers during training that we want it to produce. And we're literally, again, supervising the training process to make sure it produces those answers. And then we have unsupervised models where we don't provide it answers. And that might be because we don't want to provide it answers. That may be because we don't know what the answers are. And we'll talk a little bit more about that in a moment. But at a very high level off the top of my head, I think you could you could bisect the universe two ways pretty cleanly as classifier or regression, unsupervised, or supervised. So let me see if I can try and actually put something in each of those boxes.
Starting point is 00:29:53 We'll start with linear regression or logistic regression. These are sort of the workhorses of statistics and machine learning. If you sort of peer all the way at the bottom, again, it is sort of regression all the way down. And let's tie us back to something familiar. again, just a beta regression, a cap M model. That's what we're talking about. It seems so simple at this point. It's almost quaint compared to some of the sort of grand things that we have going on.
Starting point is 00:30:15 You don't see a cap M driving a car, but it's so important. That's the basis here. So that regression model, it's obviously a regression model, not a classification model. It's right there in the name. It's also supervised. When we train it, we give it the answers that we expect it to come up with. And it just so happens that for linear regression, we don't even have to go through a long drawn-out training process, we know what the answers are through statistical techniques.
Starting point is 00:30:40 But it's very easy to extend that one step further and get something called a neural network, which is, again, often supervised, often for regression. And the minimal version of it, arguably, is a regression followed by what scientists call squashing function, to use a technical term. It's just something that adjust the output in a non-linear way. This neural nets, when you take those two steps, the linear regression followed by the squashing function, you get this incredibly powerful building block. And I don't have the time more expertise to explain to you why it's so powerful. But if you stack those simple building blocks on top of each other and build these, if you do it enough, become called deep models.
Starting point is 00:31:24 You can get this very, very, very powerful inference from what is actually just a very simple process, this linear process. And so neural networks and deep neural networks, many of them will show up in this box of regression and supervised. Some of them will move down into regression unsupervised. And so let's talk about what that means for a second. An unsupervised model is very weird if you just sort of think about what I'm saying on the surface. I'm saying I'm going to train a machine to produce an output, but I'm not going to tell it what that output is.
Starting point is 00:31:56 Sort of makes no sense. So the insight here and what makes these powerful is that even though I don't know what the output is, I know how to tell them. the machine if it's doing a good job. So an unsupervised model that's being applied to a dataset might be told, I want you to divide this data set into two groups. And here's how you're going to be able to assess if your two groups are very distinct from each other. And then the model will go out and it'll look at each data point and it doesn't know if that data point belongs to group A or group B. But it will set up its parameters and run some inference and come up with an output.
Starting point is 00:32:28 And then it will use the function that we gave at the error function to decide if it did a good job. on the quality of its output, it will update itself and hopefully it will take one step in the right direction and do a better job the next time. And so you end up with a model that even though we didn't or we weren't able to tell it, this point belongs in group A and this point belongs in group B. It actually can come up with an answer. This would be like a nearest neighbor type idea that you feed a ton of features without actually knowing what they're related to and it can almost create in like end dimensional space like clusters, like relationships, like you said, the complex correlations earlier, between.
Starting point is 00:33:03 features and then I guess that thing could be useful, right? So is that almost like a form of creating a new feature itself is like understanding that there are certain clusters? Is that a common use? Yes, that is absolutely correct. And in fact, I'm glad you said that because that let's just take one step over in these boxes we laid out. So a nearest neighbor, what's sometimes called a K-means algorithm, that is a classifier, which is unsupervised. That algorithm is super powerful. It's unsupervised. So you don't have to tell it what the answers are. It will produce these classifications of however many clusters you ask for it. And as you said, that can then become the input to a downstream model. So if I have an enormous amount of data, high dimensional
Starting point is 00:33:44 data, I don't know if it's useful, I don't know what's useful in it. I don't want to bother setting up some huge deep neural network to turn through it and draw some inference. Maybe as a pre-processing step, we run it through a K-means algorithm, we get 10 different centroids out of that, and those become interesting to me, because those just tell me that point A is like point B and point A is different than point C. That might be enough to get started. We've got a last box, which is classification that is supervised. So this is back to our email, our spam, not spam example. So something like logistic regression, which is a special form of linear regression, which at the end, again, it applies one of these little squashing functions as it
Starting point is 00:34:22 ends up predicting just zero or one. So we're not going to let it predict any number. We're asking it, is this thing in group A or in group B? Are you going to output it? zero or even output a one. We can put that in that bucket. And one thing we haven't discussed at all is tree models, decision trees, things like random forests. These are super, super interesting to me and to a lot of people, I think. Sometimes they're called the set it and forget it model of machine learning because they have very few hyperparameters. They tend to be very good. And it's sometimes difficult to come up with an argument why you shouldn't just use a random forest and nothing else. So I'd like to spend a minute here because I'm a little biased since I just spent this morning with a research partner of ours going through trees in a lot of detail.
Starting point is 00:35:07 So I guess my question is why not just use those for just about everything? What are the example our partner gave this morning? Kevin was that, you know, for something like computer vision, this might not be appropriate because there's so many interrelationships between like pixels in an image or something like this. But it does seem to me like trees, maybe you can describe why, are sort of like the ultimate utility player. in this whole landscape. Yes. So if you think about what a decision tree does, it's going to walk through a series of steps. So let's go back to our email example.
Starting point is 00:35:35 It's just a good one to keep using. We're going to decide if this email is spam or not spam, and we're going to do it using a decision tree. So at each layer or rather at each, I guess, branch of the tree, we need to use our data to decide if we're going to go down one branch or another branch of our decision tree. And if the data has one form, which is either categorical, meaning it has, has just a handful of outcomes or very marginally well distributed. That may be easy. But in some cases, that may be very hard. So if instead of emails, put email aside for a second, I shouldn't have gone down
Starting point is 00:36:09 that road. Let's talk about census data and guessing whether, let's talk about guessing whether someone will vote. You could build a model of whether someone would vote that's based on a decision tree and you would take inputs like, well, what state are they in? How old are they? What's their gender? Did they vote in the past? All this demographic data. And that demographic data should lend itself to being cleanly bisected at each branch of the decision tree. So male go down one path, female, go down another path. California go down one path. Texas, go down another path.
Starting point is 00:36:40 Age over, let's say, I don't know, 30, go down one path. Under 30, go down another path. So that data lends itself very cleanly to the decision tree model. Our email actually may not. And this is where, of course, I'm going to get a little out of my death because there may be someone out there sitting on a wonderful decision tree. classifier. But to me, at a high level, email doesn't really fit that bill. It's very hard to imagine a model that could very cleanly bisect the email universe into spam or not spam, even if we let it run for a few iterations. Computer vision is another great one where we use something called a
Starting point is 00:37:14 convolutional neural network to actually combine neighboring pixels into more meaningful data, which is, again, very, very difficult to do with a decision tree, which only, it only sort of gets to draw a line once and then go down two different paths. A convolutional neural network is in some ways doing the opposite. It's taking disparate points within the input, and it's combining them into a single observation. So it gets technical fast, but it's sort of innately clear to me that there is some data that is amenable to being split at each branch of a decision tree, and there's some data
Starting point is 00:37:47 which is just inappropriate. However, even inappropriate data, I'm not convinced the decision tree would do a terrible job. I think it would give you an answer. I think it would just lack a lot of the nuance and power that a more specialized model could give you. Taking a step back and almost kind of like a philosophical question, so you and I have talked in the past about this classic, often repeated model that in investing, there are sort of the three primary sources of edge being informational, analytical, behavioral. They go by some different names, but some flavor of those three. In machine learning parlance, we would call this like features as information edge or feature engineering. And what I've heard more and more and certainly what we found is that that is almost all the game.
Starting point is 00:38:28 And increasingly so that the data that you have is probably the most important edge that you can garner, whether that be clean data, unique data, you know, whatever it is. Do you agree with that, that general assessment that the features themselves and how you build your data sets is sort of primary? I absolutely do. I absolutely do. It's the acquisition of data is so valuable and the acquisition of metadata is even more valuable and many of the most successful practitioners of this in the investment space are successful
Starting point is 00:38:57 because they can use their own behavior, their own trades, their own research, their own errors, in fact, as inputs to their current models, which is just sort of a fascinating feedback loop, which I think has been demonstrated to be very effective in some cases. So I absolutely agree with you. If you start with bad data or malformed data or frankly cheap data or commodity data, and I don't mean commodities as an asset class, then it's very hard to imagine you're going to get anything useful out of that. If you're buying the same data set that everyone else can buy for $250, I mean, if there's more than $250 of value in there, I'd be shocked. I'd be actually shocked if there was $250 of value in there. So you need to have high quality
Starting point is 00:39:40 curated data. You need to spend time working with it. It doesn't show up in an amenable form, and that's why feature engineering is so critical. You need to make. make sure that the data is clean, not just from a qualitative sense. You know, does it have all the decimals in the right place? But is it actually amenable to the models that it's being fed to? Is it transformed from log to linear if necessary? Is it properly sanitized? Is it properly censored?
Starting point is 00:40:06 Has sort of statistical best practice been used even in the acquisition of data, forget even what we're going to do with that data? Thinking back to your time at Lowen Data Corporation and thinking about time series specifically. I'd love any thoughts on how time series changes all of this. And one of the challenges in investing is, so if you've got to split your universe up into a training set, a test set, maybe a validation set as well, that's an interesting problem because do you split your universe based on chunks of time and assume that what worked in the 80s and 90s is going to work in the 2000s? Or do you split your universe kind of across each date and therefore have like a more
Starting point is 00:40:45 fair picture of how markets have evolved. So when you're working with time series and kind of through what I just said, what would you say are the key considerations or things that you found or thought about? Time series mess everything up. That's what I learned. They're fascinating. They're fun. They're extremely useful. They take everything that we've been talking about and they sort of throw a lot of it out the window. To your specific question, I would probably argue without the benefit of seeing a specific example, that you want to split up by time. So you want to train through some year. You want to to test from that year forward. And the reason you want to do that rather than, say, training on tech stocks and applying it to, I don't know, consumers is because your model will have an
Starting point is 00:41:24 opportunity to learn from the future correlations of the data it's all. So if it sees tech stocks go through 2008, it'll have some idea of what's going to happen when consumer stocks get there as well. So that peaking, if you will, that cheating will pervade the data set if you don't really coordinate off and take really, really close care to ensure that your model can't see the future. That's, I mean, I can't tell you how many times I've seen someone yell Eureka, and it turns out that somewhere in some completely innocuous, innocent way, the computer cached an output on one day that turned out to be important, stupid example, but it can just destroy the whole dataset. So I'd love to take the chance to turn to kind of purely selfish mode and just ask you a series of
Starting point is 00:42:10 questions based on your experience, building a lot of these things yourselves, seeing others as an allocator that are trying to do the same thing. Just thinking about the future because obviously the goal is always to find really strong relationships between what you can know today and what might happen in future market conditions. And it's a tricky problem, I think, full of potential pitfalls. So maybe I'll begin by saying, what are those most common pitfalls that you've seen of firms or or yourself trying to build predictive market investing models, and they fall in a couple camps, or do you just see all different sorts of errors? It's such a wonderful question.
Starting point is 00:42:48 The first, without a doubt, is something we said in the, I don't know, the very first question you asked me, which is ending up with a chainsaw when you just needed a butter knife. So whether it's because of hype, whether it's because of external pressure, whether it's just because you need to show some stakeholder that you are doing something cutting edge, it's so easy, thanks to tools like TensorFlow, to end up with a chainsaw. And very often, you don't need the chainsaw, you just need the butter knife, especially, especially in finance, where the data is so noisy, so random, and the insights are probably shockingly easy to find if only you can peer through the noise and find them.
Starting point is 00:43:25 So that's without a doubt the number one thing, is just over-engineering, making things way more complex than they need to be, is by far the first thing. One of the second things is something else we actually touched on, which is this idea of not really knowing what you're asking and not really knowing how to set up the model. So at our new company, and you know this, Patrick, we treat the hitchhiker's guide to the galaxy as sort of a Bible. I guess I've done that for a long time, but this is my first chance to force it on other people. There's a well-known scene in the book where people wait millions of years to learn the answer to the great question of life, the universe, and everything. and they're very disappointed when the answer turns out to be 42. And it's suggested that perhaps they didn't actually know what the question was in the first place.
Starting point is 00:44:07 So that is the example. This 42 problem is what I see in many, many machine learning deployments that go wrong. Is you think you have data, you think you have labels, you go out and use an off-the-shelf deep neural net right out of some wonderful, very sexy package. And you just don't really know what did I ask this model to do? what answer was this model even capable of giving me? Like, was it capable of answering this question? That is another very, very common problem that I see. And that one is especially in finance.
Starting point is 00:44:36 Does that suggest that really the best way to get involved in this world is to start with something that you already know very well? Or at least if you had to rank things in your process where you were the most certain, like start with the most certain and see if the way that you treat that can be improved through machine learning techniques? Yes. So if I take both of my examples and turn them on their head and try to turn them into recommendations as opposed to pitfalls, what you'd want to do, Patrick, is just what you just said.
Starting point is 00:45:01 Find something you know well that works in a rather simple way or that you think has a relatively simple process. And simple here, I don't mean stupid. I mean you understand it. It's not, it doesn't require great sophistication to come up with an answer. That's a great place to see if you can deploy machine learning to see if the machine learning model's ability to gain an exercise experience way faster than you is useful. And if all, All it does is let you get to an answer sooner or slightly more accurately, then that is probably useful on the margin. And that's a great place to start.
Starting point is 00:45:36 And you don't need to sort of dive into the deep end and build this end-to-end system. Like you don't need a system that can drive a car to, I don't know, predict earnings next week. You need a system that can predict earnings next week. So where do you start? Well, what would you do as a human? What would your process be? Why don't we just pick a piece of that that seems amenable to experience gathering and apply a model there? Yeah, it seems like a really good rule.
Starting point is 00:45:57 of thumb. I always like that Benedict Evans idea. It's sort of like having unlimited interns that like for gathering information or processing it, if it's fairly simple and you just need kind of a ton of horsepower to do it, that you can enhance your intuition or your preexisting knowledge and maybe gain some sort of edge that way. Absolutely. You know what the terrifying thing about that is, is we're so geared to think that the danger from automation is to people who do things that are repetitive, when in fact, I would argue that the danger from machine learning is to people who believe that experience is why they're valuable. Yeah, that's a really interesting idea. What are, as you look back on the investing side of your career, what would you define as sort of the frontier of the application of this technology in an
Starting point is 00:46:41 investing context? Like, what was the sort of most advanced versus more settled territory? And I'd be interested to, if there is any settled toory when it comes to machine learning and investing, there may not be. But if you had to kind of separate some examples into settled and frontier, How would you think about that? You know, with the exception of a few firms that we all know, I'm not sure I'd call anything settled in terms of quant finance. The things that were always most exciting to me, at first when I was, let's say, younger, the things that were more exciting to me were the things that I now believe are completely
Starting point is 00:47:10 impossible and absurd, which was, you know, end-to-end automatic trading and crazy models that just have no basis in reality. What's become very exciting to me is systems that enhance humans' ability to make a decision. So I believe that humans, whether evolutionarily or not, are actually quite good at solving these sort of cross-temporal credit assignment problems that we've been talking about in terms of deciding if a trade was good or not. And I think it's hard. Our biases do everything they can to get in the way, but we actually are quite good at doing that. And so if we can use a machine learning system to enhance our ability to perceive those patterns and make those decisions, I think that is extremely exciting. and I have not seen any example of that where I just be like, well, that's it, that we're done.
Starting point is 00:47:55 You know, everyone can go home. That is percolating in so many different forms and so many different firms where, you know, the most successful versions of that I've seen are places where people say, listen, I depend very much on such a signal or on knowing that such a thing is happening in a qualitative sense. And you deploy a machine learning algorithm to just pick up that one thing and just become one more greener red flag that goes into a decision process. I think that's extremely exciting. What's also exciting to me is there was this rush, especially a few years ago, into the very high frequency space where humans couldn't really be involved, right?
Starting point is 00:48:27 It was all computers. And therefore, there's a very natural tendency to say, well, whoever can have the fastest and the most sophisticated and the most advanced algorithms, that person will win. And that's great and that's true. And frankly, that's not a space I've ever done much work in myself. I've become so interested in systems that now go the other way. And look at long-term investing, look at what should we call it, Warren Buffett in a box? Value investor inbox, yeah, holding periods. The ideals that we seek in a great human investor, let's just say, I love to see those
Starting point is 00:48:57 replicated in a machine not to replace the investor, but to actually enhance them, to let them deploy themselves at scale, if you will. Just to highlight the point, it is fascinating how much what you just said seems to be the most interesting part of this whole idea. So if you extend the holding period much longer to a year, three years, or to the realm of these good business at a good price type investors, I think what you see first is the confirmation of a lot of principles that these people espouse, but with a lot more nuance and granularity, maybe than there as a single human or even team of humans could possibly ever
Starting point is 00:49:28 sort through in a business and spotting of patterns which I want to ask in a second about interpretability. But it's amazing how rich the early research is in the more traditional long-term buy and hold type of investing. Completely agree. I completely agree. And that's something that can be viewed as threatening or can be viewed as wonderful. And it sort of depends on your perspective. I talk a little bit, a couple more questions here. First being about interpretability. So, you know, you mentioned neural nets and some of the more deep learning and some of the more complicated things where, or maybe many layered tree, decision trees, human beings, unfortunately, especially
Starting point is 00:50:03 if there's a lot of dimensions or features, like literally just can't interpret what's going on here. And I'm curious your philosophical take on whether or not that matters, back to this idea of pragmatism in the field of machine learning where the goal is sort of the most predictive model, period, whether or not I can understand it, but is it the most predictive? And we heard this fantastic story this morning about how Google at one point had to effectively fire a whole bunch of people because they kept overriding models in their search algorithm, in their ad click algorithm, because they were getting in the way. And it would have been way more accurate if they just let the damn model be itself, even if it didn't make any sense. And you find these kind of counterintuitive results,
Starting point is 00:50:42 and sometimes you just need to let counterintuition work. So what's your view on whether or not interpretability matters at all in all this. I guess this is the point where I'll sort of drag out my soapbox and say two things. First, I think interpretability is overrated to some degree. I'm not sure that it is always useful. And the second thing I'll say is I'm also not convinced that models are not interpretable at all. In fact, I think they can be quite interpretable. I think what you have to get away from is the idea that every single parameter is useful from a human perspective. Big neural net could have a billion parameters, let's say. Even if they were all interpretable and useful, it would be a waste of time to pretend that we could learn anything from looking at any one of them. So in some ways, I sort of
Starting point is 00:51:25 wave my hands at the whole problem of whether or not it's interpretable. The number of parameters make it irrelevant. What is interesting, though, is if you look at these things in aggregate, you actually can draw some very cool inferences out of what's happening in some of these models. And there's an amazing website which is escaping me now, but I'll send it to you, and maybe we can linked to it. I believe it is built on top of TensorFlow, and it will actually show you interactively in your browser how a neural net is processing data at each layer of the net. And what's so cool about is you can actually see at each layer a increasingly complex understanding of the underlying data. So as the data passes through each layer of the net,
Starting point is 00:52:05 you visually see the neural net's understanding, if you will, of the data. Now, is that interpretable in the same way as looking at a T statistic on something in the same? saying, well, it's not statistically significant? No, it's not, but it is very interesting from a qualitative perspective. And I think it does a lot to undermine the claim that these models are total black boxes. You can't possibly know what they're doing. I just don't think that's true. Before we started recording, I offered maybe that we would start with this idea of gradient descent and like error minimization. We didn't. But it feels kind of like an elegant place to end the conversation just because of especially you and I conversations over the years about how valuable it is to know
Starting point is 00:52:45 the negative side of things, kind of what to avoid. So maybe you could discuss this idea of gradient descent and how it fits into this whole picture. So gradient descent is at the core of the most common ways of training many types of machine learning models, many types of neural networks in particular. Some of the models that we discussed today are not trained by gradient descent. So it's important to distinguish between the two. And gradient descent is an idea. It's actually a very old idea that's sort of found new life in the machine learning world of minimizing the error of something by literally taking steps towards it. And let me try to explain what that means. If you imagine a mountain, you imagine standing on top of that mountain with a marble and dropping it,
Starting point is 00:53:29 the marble will find its way down the mountain simply by following the slope of the mountain or what in mathematical terms we call the gradient. So that marble is going to make its way down. And eventually it's going to get to a low point from which it can't continue, right? It's the lowest point. And that would be where the marble comes to rest. And we would say that that algorithm, if you will, that we've just illustrated with the marble rolling down a mountain, has stopped. So what actually happened during that algorithm? Well, the height of the mountain is our measure of the error. And at any moment, when we measure the error of our model, we also measure how a step in any direction, which is to say how tweaking the parameters of the model in any way, would affect that
Starting point is 00:54:09 error. And then the simplest form of gradient descent says, well, whichever direction would reduce our error the most, let's go that way. So if you imagine we start at the top of the mountain, we look all around ourselves, we measure the slope at every point, and we take a step in the steepest direction. That takes us down more than any other way. And now we repeat that process. And this gradient descent or slope descent is how we get down the mountain. in the most efficient way possible. Now, there's a couple dangers here, where what if we end up in a valley, but we're still many thousands of feet above sea level?
Starting point is 00:54:43 And if we had actually gone a more shallow route, we would have avoided this valley, and we would have been able to continue even farther. This is a real challenge to call local minima when you end up in such a point. And there are a number of strategies that have been layered on top of this base gradient descent or often in machine learning we use something called stochastic gradient descent to try to mitigate problems like that. So the first one is we introduce a concept of momentum. So if you are rolling a marble down a mountain, at each point in time, it's not going to look around itself and go one foot in any direction. It's going to have momentum that's going to continue to carry it in the
Starting point is 00:55:18 direction it's already gone, even if perhaps making a hard right would have been slightly steeper. So we can introduce momentum into our machine learning models. And grading descent with momentum was one of the early ways of really enhancing the power of machine learning training algorithms. So the model takes steps to minimize its error. And at each step, it continues a little bit in the direction it's been going and also in the direction of steepest descent. It finds sort of a halfway point. And then, of course, the math just layers on from there.
Starting point is 00:55:48 And right now there's a number of really incredible optimization solvers out there that that handle this in extremely elegant ways, get that marble down really quickly, if you will. But that is fundamentally when we're training an algorithm what we are doing. Well, I should say when we're training a certain class of algorithm, what we're doing. We're starting with an error surface or an error value. And we know we want to get that error as low as possible. In particular, if we have a supervised setup, we know we've misclassified a whole bunch of emails of spam. So we then figure out, well, if we tweak the parameters of the model, we'll get more right.
Starting point is 00:56:22 We'll say more that are spam that are actually spam. And that's great. So we tweak the parameters in the way that. that most improves our error, and then we do it again, and then we do it again, and then we do it again, and then we do it again. And that's, I don't know if you were looking for such technical answer, but that's how that works. Yeah, I just love it as an analogy for just like thinking about learning more generally speaking. And you've kind of hit on all the major interesting points, which are figure out how to ask good questions. Once you've done that, and, you know, as you pointed that right at the beginning, maybe that is probably the hardest part. Then it's all about getting the right kind of information that might pertain to those questions. And then this idea of gradient descent, kind of the close of all out of learning through mistakes and error minimization is just kind of an awesome idea for solving any sort of problem. And it's certainly exciting to me to learn from you today about, you know, some of the particulars around machine learning. Yes, it's hyped up term. It's a lot more complicated. It requires a lot of human intervention still. But it certainly seems like the
Starting point is 00:57:17 path forward for a lot of interesting problems. Yeah. And I love coming to chat about the stuff with you. If we had to tie it up, I'd say, you know, it's, it's all well and good to figure out how you're going to get down that mountain as fast as possible. But if you look around and realize you were on the wrong mountain in the first place, you've got a problem. Yep. Well, this has been awesome. Thanks so much again for all your time. Always love learning from you. Absolutely. Thanks for having me on. Hey, everyone. Patrick here again. To find more episodes of Investor like the best, go to investorfieldguide.com forward slash podcast. If you're a book lover, you can also sign up for my book club at investorfield guide.com forward slash book club. After you sign up, you'll receive a
Starting point is 00:57:55 full investor curriculum right away, and then three to four suggestions of new books every month. You can also follow me on Twitter at Patrick underscore Oshag, OSHAG. If you enjoy the show, please leave a quick review for us on iTunes, which will help more people discover Invest Like the Best. Thanks so much for listening.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.