Invest Like the Best with Patrick O'Shaughnessy - Michael Recce – Tim Cook’s Dashboard - [Invest Like the Best, EP.91]
Episode Date: June 12, 2018My guest this week is Michael Recce, the chief data scientist for Neuberger Berman. The topic of our conversation is the use of data in the investment process, to help cultivate what is commonly refer...red to as an information edge. I call the episode “Tim Cook’s Dashboard” because of an interesting question that Michael poses: if you armed the best apple analyst in the world with Tim Cook’s private business dashboard, what might that be worth? Effectively Michael’s goal is to recreate the equivalent of a company dashboard for many businesses, helping analysts understand the fundamental health and direction of companies a bit better than the market does, and in so doing create an actionable edge. This is a daunting task, and you will hear why. It requires both a fundamental understanding of business and of data, statistics, and methods like machine learning. In our own work, we’ve found machine learning to be useless for predicting future stock prices, but extremely useful for other things, like extracting and classifying data. This conversation can get wonky at times, but as listeners know that is the best kind of conversation, even if it requires a second, slower listen. I hope you enjoy this talk with Michael Reece. Afterwards, I highly recommend you invest the time to read a series of posts called Machine Learning for Humans, which I will link to in the show notes. It helps demystify the buzz words and explain how these new technologies are being used. For more episodes go to InvestorFieldGuide.com/podcast. Sign up for the book club, where you’ll get a full investor curriculum and then 3-4 suggestions every month at InvestorFieldGuide.com/bookclub. Follow Patrick on Twitter at @patrick_oshag Books Referenced Crossing the Chasm One Two Three Infinity Links Referenced Sam Hinkie Podcast Episode Show Notes 2:44 - (First Question) – Changes in data science through the lens of Michael’s career 5:17 – The basic overview of using data and machine learning to create an edge 6:58 – How the state of business is more than just a single data point 7:53 – How you know when you’ve pulled a real signal from the noise of data 10:49 – The advantages that data provides 13:01 – Is there still an edge in decaying data 15:34 – Building data that would predict stock prices 19:43 – Prospectors vs miners in data mining 22:18 – Knowing when your prospectors are on to truth 27:09 – Understanding machine learning 30:10 – Defining partition 32:17 – Applying the parameters of selection process to stocks 36:05 – What’s the first step people could take to use data and machine learning to improve their investment process 38:54 – Building a sustainable advantage within data science 41:35 – Predicting the uncapped positive vs what’s seemingly easier, eliminating the negative 43:58 – How do we know to stop using a signal 46:22 – The importance of asking the right question 47:09 – Categories of objective functions that are interesting to measure data against 47:42- Crossing the Chasm 48:37 – Most exciting things he’s found with data 51:17 – What investors, individual or firms, has impressed him most with their use of data 52:17 – Will everyone eventually shift to being data informed or data driven 55:33 – Wall Street’s use of data vs other industries 55:36 – Sam Hinkie Podcast Episode 57:48 – Why everyone should know how to code 58:52 – Kindest thing anyone has done for Michael 59:22 – One Two Three Infinity Learn More For more episodes go to InvestorFieldGuide.com/podcast. Sign up for the book club, where you’ll get a full investor curriculum and then 3-4 suggestions every month at InvestorFieldGuide.com/bookclub Follow Patrick on twitter at @patrick_oshag
Transcript
Discussion (0)
This podcast is sponsored by CFA Institute, the Global Association of Investment Professionals,
whose mission is to lead the investment profession by promoting the highest standards of ethics,
education, and professional excellence for the ultimate benefit of society.
CFA Institute serves a global community of investment professionals, working to build an investment
industry where investors' interests come first, financial markets function at their best,
and economies grow.
The chartered financial analyst credential is the most respected and recognized investment
management designation in the world.
The views expressed in this podcast do not necessarily represent the views of CFA Institute.
Hello and welcome, everyone.
I'm Patrick O'Shaughnessy and this is Invest Like the Best.
This show is an open-ended exploration of markets, ideas, methods, stories, and of strategies
that will help you better invest both your time and your money.
You can learn more and stay up to date at investorfieldguide.com.
Patrick O'Shaunisee is the CEO of O'Shaunicee Asset Management.
All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of O'Shaunsi asset management.
This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions.
Clients of O'Shaughnessy asset management may maintain positions and the securities discussed in this podcast.
My guest this week is Michael Reese, the chief data scientist for Newberger Berman.
The topic of our conversation is the use of data in the investment process to help cultivate what is commonly referred to as
as an information edge.
I call the episode Tim Cook's dashboard
because of an interesting question
that Michael poses.
If you arm the best Apple analyst in the world
with Tim Cook's private business dashboard,
telling him everything going on
in the Apple universe that day,
what might that be worth?
Effectively, Michael's goal is to recreate
the equivalent of a company dashboard
for many businesses,
helping analysts understand
the fundamental health and direction
of companies a bit better than the market does,
and in so doing, creating an actionable edge.
This is a daunting task,
and you will hear why.
It requires both a fundamental understanding of business and of data, statistics, and methods like machine learning.
In our own work, we found machine learning to be useless for predicting future stock prices, but extremely useful for other things like extracting and classifying data.
This conversation can get wonky at times, but as listeners know, that is the best kind of conversation, even if it requires a second, slower listen.
I hope you enjoy this talk with Michael Reese.
Afterward, I highly recommend you invest the time to read a series of posts called Machine Learning for Humans,
which I will link to in the show notes.
It helps demystify the buzzwords and explain how these new technologies are being used.
Now on to my conversation with Michael Reese.
So, Michael, we will begin by maybe using your own history in the industry as a means of describing the changes that have happened in data science
and how it has been applied to different investing processes.
So maybe give us each of your stops and maybe alongside each stop the major kind of changes and developments
and exciting things that have happened in this space.
Sure, thanks. Great to be here. So I think it's probably better to start before the investing industry because
by happenstance, I ended up doing things which turned out to be very useful. So my initial
background was in math and physics. I did graduate work in physics. I abandoned the PhD to go
work for Intel. After five years at Intel, I could tick the box. I could earn a living and didn't
really want my boss's job. I was in my late 20s. So I thought, well, what do I want to do? I want to be
an AI researcher. New enough computer science and engineering, but I didn't know anything about
biology. So I did a PhD in neuroscience. So I could learn some biology and became a professor
teaching medical students about the brain, teaching computer scientists about machine learning. Graduated
about 12 PhDs, but I missed the impact of being an industry. So I helped my students start a couple
of businesses. The first one, we were analyzing bank transactions to find white collar crime. So we're
looking for anti-money laundering and we're looking for doing some trade surveillance and things like
that. And we had 18 of the top 25 international banks as customers when we sold the business
to Walbert Pinkus in 2005. So the second business was analyzing people's online activity to figure out
what ad to show them. So in that business, about a million times a second, someone goes to a web page,
and you have a tenth of a second to decide what ad to show them based upon their history of
their clickstream activity and how much to bid for it. And so the reason why that's relevant is you'll
see that those types of data actually end up being exactly the types of data you need to look at.
And machine learning background turned out to be pretty helpful too. But for about 15 years before
being recruited from that second business by Steve Cohen to join 0.72, I was working essentially
for my student running these firms. And the second firm, we were about 1,000 people when I left.
I was running engineering. I used to say at conferences, if you have really good students,
they employ you. But then I went to work for Steve Cohen as Chief Data Scientist there. We were on
the discretionary side of the business, essentially trying to use data to help predict the direction
of earning surprise. So I was there for 15 months. The second employer was GIC, Singapore's
sovereign wealth fund, and I was Chief Data Scientist there. And there were looking across all asset
classes, not just equities and certainly not just events. And they're a long-term investor. So we get
to broaden out how to use data across lots of different types of investing. And the third of employers is
Newberger-Burman, where I work now, and I've been there about a year, also as Chief Data Scientist.
I'd like to think of it as Goldilocks, though, too hot, too cold, and just right.
So this is a huge topic. It has become one of the most important trends in investing is the move
towards rules-based or quantitative strategies, both in simple, kind of broad index terms and also
in the most sophisticated of strategies. Maybe you could describe the very basic process.
What is it that we're trying to do? What are the key inputs? What are the outputs?
just sort of a generic study would be helpful.
The idea is that there's lots of information that's in the world that's not in the market.
And a lot of these data are in things like people's navigation.
So Mark Zuckerberg was just in front of Congress.
And part of the concern was the creatives that were being shown by Russian money.
But the other concern was that advertising started to feel a little bit strange because it's so well targeted.
In the period of time that I was in the advertising space,
advertising went from just pretty much spray and prey to pretty accurate.
targeting of advertising. But if I know what ad to show you, then I know what product you're
addressed in. If I know that across all geographies and all products, I know who's winning in the
marketplace today. And I would argue that information is not price into the market. So it's really
all this information, I call it digital residue left over from inexpensive electronics
that's laying around in the world that we can sort of scoop up and to figure out who's actually
winning in the marketplace. So it can give you a couple of concrete examples. And largely it has to
do with if you think about the high level financial statement, we're treating lots of things
if you want as scalers, as single numbers like the temperature, like same store sales and things
like that. But really, the businesses are much more distributed than that. So it could be that some stores
are doing really well or some customers are actually great customers. But, you know, there's the long
tail of a little bit of commerce from lots of other customers. And what the data allow you to do is to
really understand what's happening in the business in enough detail to know what the future of the
business looks like. So maybe go into a concrete example of something beyond just same store sales.
Yeah. So one example, which has been published in blogs.
by places like second measure is when Blue Apron did an IPO. Again, if you look at Blue Apron
from its IPO documents, it looks like a company that has customers going up into the right
and has revenue going up into the right. But if you break it down into cohorts of individual
customers based upon when they join the firm, you can look at the cost of acquisition of each
of those tranches of customers based upon the advertising spend. And you can look at the long-term
spend of that cohort of customers, particularly they will be in how much revenue you expect
in the future from that set of customers. Now, if you do that analysis, you, you'll
you see that the customers are less and less sticky with each new cohort, and they're more and more
expensive to acquire. So now the business looks completely different with one level, lower detail
than it did with just customers going up and revenue going up. And I think that, again,
people have blogged about it, that information is sort of something that's available by looking
at transactional level detail on things like credit card transactions. So I think a good starting
point or null hypothesis is that there's a lot of randomness in markets. There's an enormous amount
of information and data. And one of the key steps is understanding or having a framework for
determining when you've pulled a signal from what is largely a noisy background. So talk about that
process, like the actual testing process itself. Let's say we have a generic set of data,
and we also have some outcome. It could be earning surprise. It could be future returns,
something against which we're regressing or comparing the data set. Talk about like that actual
process and maybe how that's evolved. Let me talk about the process of starting point to end point
of what would you do with the data. And then we dig down in some more detail. So there's a data
sourcing problem because lots of data in the world. Some people think their data is very valuable. Some
people don't realize the value of their data. So you have to go find the data. The data tends to be
transactional. And so think about like a credit card transaction line or a line in your bank statement or one
visit to a webpage. And the thing is that some data is more useful than others. So the weather
predicts consumer behavior. It's nice in New York today. People will go shopping. If you have
their cell phone location, you know they actually went shopping. At least they took their cell phone there.
If you have credit card transactions, you know what they spent. So not all data is equally valuable,
but you get this data, and then once you get the data, you have to figure out what the businesses that it's involved in the process.
So you look at a transaction, it could be that there's no public business in it.
It might be interested in private companies, too.
But it could be that there could be three businesses.
You could have used PayPal through Expedia to buy United Airlines flight.
And so you have to break down these transactions, know about brands associated with businesses, know about mergers and acquisitions,
in order to figure out what these things are doing in terms of helping a business.
And then you want to take that information and put it into a model of the business.
Now, I would argue that if all you do is roll that up and try to come up with a topline revenue,
why did you need the granular data after all? Essentially, you're just defeating yourself.
And so what you want to do is understand the business at the granular level.
And the way I think about that is the spreadsheet model of a discretionary investor.
So if you populate the spreadsheet model with information you're getting from the real world,
in fact, you want it to be even more granular than their typical model.
I mean, I would say to a discretionary investor, well, you know, you're a busy guy.
It's hard to update your 50 companies, 100 companies.
if you had someone updating it for you, how would you even more elaborate your model? Let's elaborate it that way.
Or another way of saying it is you're buying a public company, which means you get to see public accounting statements.
But if you bought a private company, you get a data room. What would you want to see in the data room?
You know, maybe you want to see their online sales. Maybe you want to see their overlap with Amazon.
Maybe you want to see how they're spreading to other demographics. And so all that type of information we can use the data to get.
And a key point, which we should talk about at some point, is the whole technology stack because essentially this all has to sit on top of a technology stack.
So why I describe what we want to do is building Zillow for the stock market. You want to essentially build a model. You know, Zillow is implemented sort of a form of automated valuation of property, which people used to do. Maybe it's not as good, certainly not as good as the best valuation person who would value property, but it's automatic. And so what you want to build as a first step, something which automatically values companies before you look at the price in the market. Yeah. So there are a couple interesting ways that the data is being deployed. And I would say you've described one of them, which is that you've got a, let's
you've got an Excel spreadsheet that's valuing a business, doing a DCF or something. And that's
using inputs, usually from K's and Q's. This is a better, richer way of filling in those cells
with more accurate or timely data. So that's one. The second would be the other end of the
spectrum in a purely rules-based approach where there is no fundamental insight. There is no
human really of any kind. It is merely a predictive, pure quantitative model. So how do you think about
the relative advantage that this sort of data science work gives you in those two different ends of
the spectrum. What we're interested in is actually using the data to understand how healthy the
company is. And there's lots of reasons why that's a longer term and better, deeper strategy,
because essentially the nice thing about companies is they're relatively stationary and fluctuations
in the market are classically non-stationary. But we can come back to that issue. But the point
I would just make is that let me parody for a moment, the two extremes. So typical quant thought process
is you know a little bit about a lot of companies and you have to actually be in lots of companies
to avoid tilts of any form that you don't want to predict. And so I describe them as an inch
deep and a mile wide. Discretionary folks know a lot about a very small number of companies, but they
still make money in their book. And so I described them as an inch wide and a mile deep.
And if you imagine those two axes next to each other, there's lots of different variants
that are in between. And so there's some variants which are very close to the quant approach,
which is, hey, let's take something like sentiment and analyst reports. Use natural language processing
to evaluate it. We have 10 years of history across the whole market. And that gives us a little bit
of a wiggle that we can add to our quant strategy. And that's just really easing into it from the
quant's view. And from the fundamental view, you could say, well, look, I'm going to hire a super
analyst who can go scrape the web for me. He can go find this magic parameter for me. Then I'll
manually put it in my spreadsheet. And that's easing into this process, if you want, of using more
quant-type methods in your discretionary process. But what we're interested is going right at the
middle. Build models of the business automatically with computers. And use quants, if you
want this, if you can think of it as a second generation of quant, while there's lots of
opportunity just to find mispricings, go for it. But, you know, as that disappears, what you
want is just to build computers to implement the way that a fundamental person thinks about
the business in much more detail. So a lot of interesting things to pull on there. And maybe the
first I'll pull on is this idea of disappearing. So there's a lot of discussion on any sort of
information-based advantage or edge or alpha or whatever has pretty high decay rates or fast
decay rates. Maybe talk a little bit about examples of that that you've seen. Everyone quotes credit card
data or satellite data or all these same exact like three or four examples. Any comments there on
whether or not you think there is still an edge in the informational game or not? Yeah, that's a great
question. So I get asked that question by lots of people. They'll say, well, credit card data has been
there forever. Everyone has it. So why isn't it just completely commoditized? Why would I spend
money on credit card data? And I think the key issue here is data and everyone can have data, but
information is something completely different. And what you do with that data in order to extract
information is key. So there's lots of things you want to do with that data in order to make
the information much more rich. So for example, one of the things we think about is, well,
let me back up for a second. In my internet company, we measured demographics for free for web pages.
That was our free service, and we were an advertising company. And they still are,
quantcast. But what we do is take the, you could take all the credit card spent or all the online
activity and infer the gender, age, income level, and all these other demographics of the people
And then you can figure out what's happening in this company's customer base.
So Lulu Lemon is increasing their appeal to male customers.
It used to be just a little blips in Q4 where they'd go buy something for their girlfriend.
And now you can see this is all riding on a line.
And we can see that because we've inferred through the mixture of all your spending of people in the credit card panel where their gender is.
It wasn't something that was given at the beginning.
So I think that if you take your data and you use it to build a detailed informational picture, then that hasn't been done.
People have been finding the shiny pebbles on the surface of the sand.
They haven't been actually taking the data and using it to build models.
And then that model is actually a much longer lasting form of alpha because if I could give you Tim Cook's dashboard and you could see in live real time, by product, by geography, by cohort of Tim Cook's business, either the stock must just be completely efficient or you have a huge information advantage that persists for a long time.
Yeah.
So to summarize that and make sure I understand, I think the idea is people are wrongly viewing commonly.
available or easily available data sources as arbitraged away or as no longer useful when in fact
maybe very few people are actually using them in the best way. Or they're actually processing that
data sufficiently to really glean the information out of it. That's the trick is what information
are you trying to get out of it? You know, the data sets are very large. And again, the most common thing
people are doing with the data is rolling it up to get a top line revenue number. Let's go back to this
idea of Zillow. So building Zillow for the stock market, really interesting idea. Talk me through
the process kind of from the start of how you would begin to approach a problem like that. So if
starting at the end, the idea is here's an estimate. It's rough. It's crude, but it's a rough
estimate daily, let's say, or hourly of the value of a business, of a stock. How do you get from
cold start to that kind of end estimation? Okay. So let's say you take, let's take Apple.
And if you were to plot a distribution where the y-axis is the number of transactions you see in
the credit card data and the x-axis is the dollar amount. What you would see in that distribution is
lots of bumps. There'd be a bump that corresponds to iTunes purchases, bump that corresponds to
Apple Music, a bump that corresponds to watches, handsets, iPads, computers. And so, and obviously,
some people are buying multiple things. Not all the data points are going to be nicely in these bumps.
But if you broke the data into those bumps, you could create a dashboard that looks by geography,
by cohort of when they bought their first iPhone, because their first 18 months, they're going to be
higher app spend than other times. Buy product by geography, by cohort. And essentially, this is
reconstructing Tim Cook's dashboard if you want. And so then what a valuation person would do is they'd
say, okay, what's the year-over-year growth of this segment of the business or this segment plus this
geography? And one of the things the discretionary guys do is they'll look at where the vol is. And so they
might say, well, look, all the parts of Apple's business are growing at a relatively stable rate, but
maybe iPhone X in China, or maybe what happens with iPads or what happens with their physical stores.
And so they focus on that one piece of the spreadsheet, but the other piece of the spreadsheet will
essentially roll up to the amount of revenue generated by that.
item. So another example like that is, so let's say that you think that a loyalty program at someone
like Starbucks is going to generate lots of new revenue for them. You can measure in the data
what's the conversion rate to that loyalty program. And then once you find the conversion rate,
you can say, well, when they convert, what is their change in spent? And you can then calculate
on a future basis how much of a new revenue machine you get by doing X, maybe all day breakfast
or whatever it happens to be that is a policy of a store. Maybe they roll it out in one geography
first. And so you can use the data to figure out how much of a money machine is this in that
geography and what's going to happen when it spreads to the other places. So both those examples
sound very idiosyncratic and specific to that business. And I don't know much about how Zillow works,
but my suspicion would be that the variables used are kind of cross-sectional, meaning like you can,
let's say it's like proximity to a town or an airport. Bedrooms, bathroom, swarthy. Yeah, like simple things
that would apply across the businesses that are not idiosyncratic. Yeah. So talk me through that
discrepancy. The interesting thing is everything that sounds idiosyncratic, it really ends up not being. So for example, let's say you have a chain of restaurants. The chain of restaurants is growing very quickly, and you're looking at same store sales. Well, some of those restaurants are brand new. So something's brand new, store restaurant. It has this honeymoon period where everyone wants to go. That falls down to an organic growth rate. And so with the data, you can sort of time shift them so the opening date's the same date and you can calculate what that function looks like. And so now then you can move it back and you can then figure out what it's going to look like in the future. You can look at
cannibalization when you open a new store, how many people were actually were going to the old
store and now going to the new ones, so you can calculate the maximum density. You say, well, gee,
is that specific to one restaurant? The answer is no, because all restaurants behave the same.
You know, and after you learn this kind of thing by talking to discretionary folks, then you can
implement this across all restaurants, stores, anything that involves consumer. And one of the things
I loved was when the Friday before the deal was finished with Amazon buying Whole Foods,
Jeff Bezos said, we're going to lower the prices than all the Whole Foods. Now, everyone
thought of Whole Foods as actually good food, but expensive. And so he triggers the honeymoon period
back in all of the stores, as if they just were brand new open, to generate foot traffic as if it was a brand new store.
And so, and by thinking about it in this mode, and that's why it really requires, if you want, a partnership between people who think like quants and people who actually think with valuation, in order to think about how do you actually use the computer methods to, and ultimately, we're going to be in a world where just like in chess and everything else, where the computer plus the person does better than either alone.
If I give you Tim Cook's dashboard, you don't understand Apple as a business, what do you do with it?
But if I give it to someone who's a discretionary investor who's been studying Apple for a bunch of years, I can guarantee it's a boon for them.
And they have a huge advantage long term in terms of how to predict what's happening at Apple.
What about like a survey of the landscape of, we'll call them pure quants.
So, you know, the famous names that people would be familiar would be like a Renaissance technologies or 2 Sigma or somebody like this.
How would you silo those groups, anything unique or specific about different types of pure quants today that you see in the landscape?
So one of the ways I have of describing quant shops is I think it's interesting if you look across industry.
So when I was in the electronics industry, the next processor from Intel was the next processor.
And I described it as like strip mining because everyone's working on the same project.
Everyone knows what the target is.
The whole company is working together.
If you contrast that to the drug discovery in a pharmaceutical business, which it's more like prospecting.
You have lots of individual scientists.
They all have their own lab.
They don't need to talk to each other.
And if one of them finds gold, then the whole company wins.
And I think that it's in a brand new field where you don't know the answer, or if you have
20 prospectors, 20 people working on something, you're better off letting them prospect than
telling them all what to do.
And I think it's what causes classical disruption is that someone has an idea top down
and tells everyone what to do, and they all go and work on that.
But if it's wrong, it's a, you know, all the eggs are in one basket.
And so I think that there are a bunch of firms that are still trying,
this if you want strip mining approach to try to solve the problem and if they guess right
where to dig, then it'll be great. If they don't guess right where to dig, then it'll be a
problem. But my personal view is that in a brand new field, you're better off actually hiring
very smart people and letting them work on different aspects of the problem that they think of in
order to try to figure it out because it's an unknown area. So supplying, let's say, prospectors
we'll call them, with raw tools. And maybe that's data and programmatic tools and scripts that
run or built a library over time and giving them a sandbox basically, not giving them a job,
giving them just like a general mandate, which is find edge and report back. That is superior to
like some central genius saying, you know, here's what we're going to do. The only thing I would
add to that is that what you want to do is short sprints. It's a traditional agile process.
And so what you would have is people self-organizing into groups with ideas and there always are
more ideas than there are possible. Having the whole team decide which ideas are best,
run them for a short period of time, find out what works by everything has to be proved.
And it has to be small enough chunk that you can prove the idea works. And then you do it again.
You keep cycling through that process and stimulating innovation in the team. And people love this.
This is a great work environment for people and they get to think about ideas for how to move things
forward. Because classically, it's a boil the ocean problem. There's so much data. There's so many
problems to work on. And there's a challenge to figuring out what are the best sets of ways to
leverage data into information to support an investing process.
So I love this prospecting analogy, so we're going to keep running with this one.
So I love talking about some of the philosophy behind this work and thinking about it as the search
for truth, like an actual signal being truth.
In the prospecting context, you drill a well.
If you hit oil, you've hit truth.
You know the outcome that you're after is there and it's measurable.
It's much trickier in the statistical environment and in our world.
So describe how you think about truth.
and this is where we might use terms like P values or T stats or whatever else, however you want to
frame it. How do you frame knowing whether or not you found something? And then I want to also
talk about knowing when it might be busted and broken. Here's one way of thinking about it.
If I have a clean signal on a business, let's just say if I actually can, if I can see a running
signal of the health of this business, then I can create a much larger position because my risk
is actually much smaller, right? So I'm not waiting for whispers to tell me if my investment's
going to have a problem. I just literally can see the commerce going on at this business.
So there's just a quick sidebar that I want to come back to like some of the actual statistics
behind this stuff and when you know there's a signal. But we were talking when we first met,
we were both speaking at some investment conference and we were talking about, I think, natural
language processing or some machine learning and the idea that you need to feed it observations
at the different ends of the spectrum and that very often the negative part of the spectrum is
far more useful in building an algorithm. So maybe talk about, talk about that idea. So you had spoken
in your talk about a belief you had that shorting or knowing when a company was going to fail
was actually an easier thing than actually knowing when it was going to succeed. And so I had some
observations and data that had led to similar conclusions. So I found you afterwards and that's how
we started talking. And this data was actually had come from building automated systems and I was
helping a company that was building automated systems to score and evaluate college entrance applications,
including the essay. And it turns out when you're evaluating the essay and you're using experts
to as a ground truth to sort of come up with what are good essays and what are bad essays,
it's much easier to actually have an accurate prediction for the lower half of the distribution
because the types of mistakes which occur in essays people generally agree on. And so the
machine learning system can be trained to actually detect how bad you are from the mode of
the distribution on down to worse. But it turns out that it's very, very hard to predict
the other side of the distribution.
And you can think of that is because what inspires one person, you know, as a great essay,
is not what inspires someone else.
And I think this applies to companies, too, because when someone has a great belief in a
company, where a company is going to be in what they're going to develop into, that space
of possibilities is very large.
But the things which actually are problematic in a business in terms of, you know,
loss of market share or what's happening with their customer base or so on and so forth,
what's happening with the financials, those things are actually a smaller space.
And so if you're trying to use something to learn it, you're better off.
And the other thing I would just bring into this conversation, and I think it's really
important, maybe we should talk about it more later on, is this idea of where do you apply
machine learning?
Because, again, if you're trying to predict price action, it's non-stationer, it's all over the
place, it's driven by all sorts of mood and regime and so on and so forth.
But the growth of a business and a segment of a business is like a rock.
It's very solid and very stationary.
And in fact, if, like, say, Home Depot has a really good month in the first month
the quarter, a discretionary guy is most likely to regress it to the mean based upon a two
year stack and bring it down the assumption for the other two months instead of just assuming
the other two months are also going to be good. And so that's because the growth of a piece of
of business is a very stationary thing. And so if you think about the good applications of machine
learning, there has to be, you know, cats in YouTube, it has to be something which relatively
doesn't change over time in order to truly get the benefit of machine learning. And so if you're
using machine learning to build models of a stationary thing, which is the growth of a business,
are relatively stationary.
You're much better off than if you're trying to use models to try to predict returns.
Predict price action at any instant in time.
Yes.
Like it's a disaster because of that non-stationarity problem.
So I just want to highlight that because I think people hear machine learning and they think like,
wow, people are applying these fancy algorithms to predict price.
And that really doesn't work.
Maybe someone knows how to do that.
Well, I think mostly, most of the time it's ill-posed because it's really, you're saying,
what's the mapping of these variables to that variable?
And if you can't solve it, the thing is that we can identify a cat in a YouTube video.
And we can't pick the price.
So I don't know.
There isn't a solution that I know.
And so they assume that there is a mapping from these input variables to those output variables
is already sort of misleading.
But is the commerce that I observe in the world a predictive of the growth of a business?
Of course it is.
I mean, that's the discretionary guys do, right?
So basically, if you're using the machine learning to predict something that you know is a
solvable problem, but you want to do more accurately, you want to hire a resolution
microscope, why don't you use the machine learning for that?
It's something that we know is a solvable problem and we just want to do it better.
So, yeah, I'm happy to discuss machine.
I have actually have a little, I used to teach machine learning at lots of different levels.
And so I have a sort of a business level description of how machine learning works, which I can try on you and see if it helps.
Let's do it.
Let's do it.
Let's do it.
Let's do it.
And you have five or six people.
And you have five or six people.
And they're all starting a company.
You have a CEO.
And they don't know each other, but they're going to work together.
And so you try to make decisions.
And so the CEO is going to take everyone's vote.
And so because they doesn't know everyone, he starts with everyone, everyone has one person, one vote.
And so you all make a decision about whether we should do X or Y.
And then at the end of the day, you made it a good decision or a bad decision.
So unconsciously, in the brain of the CEO, what he's saying is, oh, my God, you know, Michael,
when he says yes, he's wrong.
And so he's wrong.
So I'm going to decrease his weight.
And so next time in the vote, I'm not going to tell him necessarily, but his weight is going to be
a little bit less in this class of decisions.
And maybe this other guy's Patrick's weight is going to be slightly higher in these class
of decisions.
So these weights gradually change based upon trying to make better decisions for the firm.
Now, Michael has a team and Patrick has a team.
And so when Michael gets something wrong, he's not going to take full responsibility.
He's going to back to his team and say, which of you idiots actually told me this decision?
And so he then back propagates the correction to his weight to the person on a team who actually
was influential in him leading to the decision.
And so you can think of that as an organizational description of what back propagation is.
And the only rest of the detail you have to worry about is, you know, what are the rules for correcting the weights?
And so the bigger the weight, the bigger the change, the bigger the, there's a whole bunch of also things
which make heuristic sense about how much you change the weights when someone's right
and when they're wrong in this process.
Okay, well, so that's back propagation, 1986 neural networks.
So what's the problem with that?
The problem is that you don't know what decisions the organization's making.
But what if you could actually frame all the questions in the right way and make sure they're
framed correctly before you do this?
And you can think of that as automated parameter selection.
So in the old days, when you're trying to solve a problem, you'd have to figure out,
well, what are the features in the world?
What are the parameters?
What are the factors that I want to use into my network to try to train?
And it might be that that set of factors just doesn't partition well in the space.
Because ultimately what the model is doing is trying to use planes in the space to partition the good answers from the bad answers.
So let me try to make that more clear.
So let's say you've never drinking a bad glass of milk.
All of a sudden you go to the fridge, you pour yourself a glass of milk, it tastes horrible.
So what do you do immediately?
You look at the date.
You look at the texture.
You smell it in everything.
And so what you're doing is you're adjusting the features in your brain to try to adjust the partition functions.
So the class of milk, glasses of milk that you drink is now more restricted by moving these partition planes in this feature space.
And so that's what I mean by feature space.
Now, what deep learning does is it automatically, in a self-organizing way, learns what the best space is in order to actually optimize the partitions.
So the classic example is you have a spiral of data inside another spiral of data.
Well, you can't construct planes to partition it.
But if you map it into a space where they're actually two separate Gaussian clusters, you can easily partition it.
So if you select the parameters correctly, you end up in a space where you can partition the data.
And then the second stage of this voting process can then learn the partition functions.
So I want to go through this kind of one more time in stages because, look, this is an opportunity given your background for listeners to understand some of these terms.
And maybe we could build up from the super basics.
So from literally linear regression.
Like I think people understand that idea.
You're trying to fit a line through a set of data points with the least average space between the point in the line.
So maybe building up from there.
Well, so rather than regression, let me start with a partition, simplest partition.
So here's the way I've used it in my class, an introductory class.
Let's say you have a friend in a foreign country and you want to send them some fruit.
You put some apples and oranges in a box.
You send it off.
And your friend receives them and he's a scientist too.
But you didn't label them.
And he says, oh, this is great.
Thank you very much.
But which ones are apples and which ones are oranges.
So you say, well, you know, the ones that are more round, those are oranges.
So, you know, he's a scientist.
He gets it out.
He measures the roundness of all these objects.
And it turns out that some of the apples are quite round.
and some of the oranges are a little bit oblong,
so the distributions overlap with each other.
He says, you know, this doesn't really partition them for me.
And he say, well, you know, the apples are more red than the oranges.
And so now he goes and measures the redness of these things.
And again, they overlap a little bit.
But you say, well, just plot on one axis how round they are and the other axis how red they are.
And the higher the dimensional space, the more that the data separate.
And so now there's a diagonal line that partitions this space.
Now, you know, the simplest model, you think and think of this diagonal line is what's used to be called a threshold logic in it.
You have inputs going into two weights.
Weight one times the first input, weight two times the second input.
And if it's greater than the threshold, it actually is accurate.
Is the thing.
Right.
But if it's equal to the threshold, it's the boundary.
And so you can easily reduce that W1 times X plus W2 times Y equals a threshold into the equation for a straight line, where Y equals some combination of the weights plus T divided by one of the weights.
And so then if you just change the two weights, you're literally changing the slope in the intercept of this line.
And so the process of learning the partitioning between these two clusters is literally just changing the two weights in the threshold, which moves the partition function in two dimensional space between these two clusters.
And so that's how they work.
Now all you need to do is scale it up to 10,000 dimensions.
It's exactly the same math.
But you're now moving this hyperplane in this space to constructive partitioning.
Now, each neuron's one partitioning.
If I want to keep all clusters, I have to have multiple partitionings, multiple planes from different orientations to contain the set of points.
So I love the example of the apples and oranges and sort of parameter selection as a key part of that.
So once you've got to red and round, red and round are a key part of that process.
So maybe talk a little bit about that as it pertains to investing.
So we always talk about there's kind of only five things that matter in any quant process.
There's the data.
There's factor formation.
There's factor weighting.
There's portfolio construction and there's trading.
So early on in that stack is you got to choose the right factors and make sure you're not pulling nonsense
or building something that's going to create like a data mining issue or spurious correlation or whatever.
So how do you think about like the parameter selection part of this process?
Here's one of the ways, important ways I think I have of thinking about this.
Again, I think what happens is, and it's partially because of the tools we had, people reduce things to a scalar
or a single number like a temperature when that wasn't the right way thing to do.
So if we think about a business, people will say, well, you know, what's the consensus number
and what do I see?
And is it different from the consensus number, maybe there'll be some earnings of price.
But I don't think that's the right way to think about it.
I think the way to think about it is that there's a notional cliff or a couple notional cliffs.
And right now the business is between those values.
And if it falls below a certain year-over-year growth rate, it's a different business.
And that's a cliff.
If it goes above a year-over-year growth base, then that's another cliff.
That's another business.
And these are just like those partitions.
And so what's really going to cause a change in your investment thesis is when the business actually crosses one of those cliffs.
And so what you want to do is you want to partition what the business looks like when it's actually.
normal. And in that, you know, there's going to be lots of sort of variation in the way that
it looks. And the better you can understand that partitioning, the better you can actually know
when the change your thesis. So is that thresholds in variables like earnings growth rates or
turns out of vested capital? Yeah, that's right. All of those types of things. Exactly. So things
like, as I say, same store sales or could be whatever you happen to be looking at. And I think the
key thing is that from a machine learning point of view, if you classify things as if you use,
use the data to sort of classify the business into its regimes.
And the other thing that I think that's important is that a lot of this information is at a lower
level of granularity, again, going back to the products or the cohorts.
Because think about it one cohort of customers.
That cohort of customers, maybe it's people in advertising talk about it as a funnel.
Maybe I don't know about it first.
And then I gradually know about it and then I want to find out more.
Then maybe I'll become a customer, then maybe I become a fanatic about that product.
And then maybe I plateau at that rate of engagement for a while and then I decide to move on.
But let me give you a real example.
So 9% of the best customers at Walmart generate 50% of the revenue.
So why am I using a scaler to look at same store sales?
It's a Pareto distribution.
They're all Pareto distributions, 80-20 rule.
And so imagine on the y-axis, I have the year-of-year growth of that business.
And there's a little line of zero and there's negative and there's positive.
The upper quadrant is positive growth.
On the X-axis, I have what's happening to that Pareto distribution?
Is it flattening or is it steepening?
So if it's deepening, then now 8% of Walmart's customers generate 50% of the revenue.
And unfortunately, Walmart sometimes has been going in that direction, whereas Amazon's in the other quadrant of actually growing and spreading it and flattening its distribution of where the money's coming from, which is a healthier, which is a healthier, which is a healthier, which is a healthier, that corresponds to what stage of its life it's in.
because when you're desperately getting growth,
when you're getting growth
and only out of your loyal customers,
you're painting yourself into a corner of a room.
If you're not expanding your customer base
and you're growing,
it's just not going to be sustainable.
So something else has to happen at some point of time.
Either the growth stops or something else happens.
So I think that I see this as one level deeper
than the Thai level financial statement.
If you can use the data to model the business
at that level,
then you get a lot more insight into the stage and health of that business.
It's a good excuse to go back to the very beginning
and to say,
let's say I'm a discretionary guy. I'm a PM. I work for Citadel or something. I'm not a data
scientist. I'm roughly familiar with statistics. I'm a smart, smart guy or girl. But I want to start
using these methods to improve what I already do. What advice would you give somebody like that as a first
step towards trying to incorporate some of these ideas and methods to improve their investment process?
Well, first of all, I would say that I still think that the folks who have both experience as a
discretionary investor and have math and computer science statistics background are relatively rare. In this
new form of data investing, they're going to demand a premium and they're going to be really hard to find.
If you want a job today, reach out to me. But basically, those people are going to be relatively hard to
find. And so you're already in a fortunate position. The thing is, I think they have to break out of us.
If they're still using spreadsheets, they need to break out of it. Okay. And they're basically,
you know, I described it as like there's a one-to-one correspondence between, let's say, a Jupiter
notebook running Python and a spreadsheet. Because in the spreadsheet, you're looking at the data,
and the formulas are hidden. In the Jupiter notebook, you're looking at the formulas, and the data's
hidden. But basically, it's the same one-to-one correspondence between what's happening and the processing.
And basically, when the data gets so big that you can't look at it all, in fact, you're much better
off looking at the formulas if you want to optimize them. And so it's, people will say,
well, she, I can't program, but if you can actually build these complicated spreadsheets,
you can program. But what I would break out of that world, and the other thing I would say is
that the new computing methods are literally orders of magnitude better than the old ones.
Two lines in Python goes a long way.
And particularly if you're doing it, you know, with Lambda processing on AWS or some, you know,
there's just some opportunities in computing.
A few hundred million rows is a hard problem with, you know, old technology.
And there's an example which I often cite from the new technology.
Because in the bricks and mortar technology, the first job of the technology team is to make the trains run on time.
The second job is to keep it secure.
Just in third job is stay up with technology.
In the Internet space, if you're not up with technology, you're dead.
So superimpose that on a rapidly changing technology.
and the technology in the internet companies is much, much farther ahead.
And I can give you some very concrete examples of that.
But, you know, an example is a blog from a French AI person at Google,
and he was talking about finding the house numbers in France.
So where is the house number in the city?
Where is it in the countryside?
How do I deal with lighting and occlusion and so on to try to extract the house number,
runs a job overnight, finds from Google Street View all house numbers in France.
You tell that to an IT guy in a traditional bricks and mortar
that I want to find all house numbers in France,
and they'll pass out or something.
So I would encourage this person at Citadel
to start engaging in how to use those technology tools
and move away from the ones that they,
you know, their comfort zone
in order to be able to process larger data sets.
So it's kind of like sharpening the sword, right,
before you even get to battle.
Battle then becomes the data.
I was going to ask a question,
something like how do you build a sustainable competitive advantage
a moat in data science itself?
So it sounds like the first part to that answer is technology.
Understand how to build
the base layer and I guess at more efficient pricing and at better, more efficient speeds,
get the answers that you want. So maybe that's step one. What are the other things that those that
want to incorporate data into their systems should think about like the moat around that?
Because I think my key takeaway from this conversation and one similar to it in the past is that
everyone kind of assumes that this is becoming, data is like commoditized and everyone has its table
stakes and it's not worth anything. But it seems like it maybe it's the opposite.
that in fact certain people are building a moat around this,
and that advantage may widen, not flattened.
So how do you think about that, the sustainable advantage?
I think the sustainable advantage is going to be the depth of the little issue process of the data to get information.
And then the extent that which you use,
multiple overlapping data sources to get confirmatory views to actually develop a stronger point of view.
And so what people, you know, again, it's digital, it's residue.
So it has to be incorrect in lots of ways.
So you have to detect where it's incorrect.
the more that you bring in lots of different sources of signal.
And the other thing is that everyone's also looking at top line.
What about the, what about costs?
You know, looking at job postings or other stuff.
Job postings are a big component of op-X.
You know, they're also a leading indicator.
How many people get hired based upon background checks, you know, from people who are doing
background checks?
What the cost of the commodities are or the raw materials?
You know, what's cost of chicken going to, you know, a KFC, right, or to a Buffalo Wild Wings?
And so a lot of the early stages just looking at the revenue side of the picture.
but that's certainly not the only thing which drives the business.
The other thing I would say is that there's lots of advantage in terms of trying to figure
out the nature and way that the business is actually developing.
One of the things I would say is, so I think there's obviously lots of different trading
strategies.
It didn't really answer that question fully.
You asked earlier that we're going to use data.
And obviously trading on events is certainly one of them.
But I think another one of them is what's the end point of this company?
Is this a small company that's going to become a big company?
And that's going to have to do with how they're expanding geography-wise, how they're
expanding the nature of the way they're interacting with their customers. And I think the data is
really going to help you pick the long-term winners. One of the things I've been saying recently is
that I think that the asset management world where I live now might actually be the leader
in the space as opposed to the hedge fund world. And partially because they're less siloed in terms
of the intellectual property and partially because it's a different game. The game we're trying
to play is an incremental advantage over a large amount of assets instead of actually trying to be
the absolutely highest return for a teeny fraction of the of the investable assets. So a big part of the
discretionary hedge fund world is about, you know, it's not surprising that you were focusing on
earning surprises or it's the quarterly earnings game and building up a better estimation of that print
than the market has and then there's your alpha. The second point you made was interesting,
which is trying to figure out through data who the long-term winners are where you've got this
cap downside of the money you put in and you've got potentially like enormous upside. So if you can
somehow better estimate the non-linear outcomes, like you could have just enormous success.
So I'm a little bit confused that versus like what we were talking about earlier where it's easier
to predict the negative side of things than the kind of uncap positive. So how do you even approach that?
Let me connect the dots. So one of the graphs I've seen recently, I forget which of the research
shops that came out of was plotted the number of years that stocks has been in the S&P 500.
Yeah. And so yeah, that's right. And essentially it's going mostly monotonically down. And it used to be
40, 50 years, and now it's down to like 12 years or something. So, and then trillions of dollars a year
are moving to passive still. So essentially, you're buying things which are less likely, lower
probability to be winners because the temperature, if you want, the churn and the S&P is actually
increasing with this decrease in time. So if I can use something like my calculation, which I talked
about with Walmart, to determine who we're going to be the losers, or if I use any of the
process we're talking about of predicting who the losers are, then essentially I can still be
market weighted, I can still essentially have an index-like thing. People say, well, is this smart
beta? It's not smart beta. But have an index-like thing where I underweight some names and overweight
other names. And let's say I make 50 to 100 bips more than the index. But if you're doing this
entirely with a Zillow-driven process, maybe it only costs 10-bips to run. And so the question is,
what happens in that world? And so I think that, again, if you, it's very hard to make an AI that's
as good as the best valuation guy in property or otherwise. But making something which is
automatic, which actually makes things a lot easier to manage assets. And then, you know,
you might end up with a pendulum swing of the money that moved to passive back to active management.
Because now you have this AI-driven active process, which has, you know, relatively low
vol, essentially follows the index, but just tracks a bit above it because it's finding the negative
ones in the process. Finding the negative ones, because that's the easier problem to solve,
based upon this higher degree of granularity that the data is providing and that you're getting
through the machine learning process. So we haven't talked about statistical significance or like what the
markers are of a successful run and or the problems of data mining. So talk a little bit about the nitty
of that. Is there a specific threshold that you typically look for? And then the second part of this
question will be, let's say we find something that we feel is a very strong signal. What are the ways
that you know, is it just that it crosses back below that threshold to know that we need to hit that red
button and stop using that? What are the ways we know to stop using some signal after updates
to the data. We touched on that a little bit, well, we'll come back to it. But one of the things I
point out is that I was reading a review of Eudae Perl's new book. And so he's one of the developers
of this area of a basing inference. So correlation is not causality, but there is now a whole area
of statistical analysis, which is actually measures the causal influence. And so the data
mining thing, which is talked about all the time, you know, there are straightforward mathematical
methods which can help you avoid the risk of just finding spurious correlations. And the other
thing I would just say about the spurious correlation stuff is that it's a quant sort of classic
quant sort of problem. So when I'm talking to folks when I'm interviewing them about how you solve
a problem, well, let me give you a real sort of joke example, but real. So I was a professor
and one of the biology professors came to me and said, I tried every possible statistical
method in SPSS on my data. This particular method gave me the highest significant score. How do I justify
using it? And lots of people in the machine learning world are doing the same thing. There's a
company that literally just uses computing power to try every machine learning method on your data,
and you pick the one that works best. That's the problem of data mining, and it's just laziness
in terms of thinking. Now, if you have an idea for a signal or a way that you think something that
should improve the data, there's a very important thing that happens. When I used to try to find
money laundering, you're finding outliers in transaction. They don't transact like normal people,
but they're also in that same distribution once you separate out those outliers, you also get
all these horrible bank errors. And so if you just look at the performance of the whole thing,
you're combining the bank errors with the really good ones, you might throw it away, say, oh, it didn't work.
But if you can partition, maybe your sigma increased, you know, because you actually have both of these two groups being found by your outlier detection.
If you can partition away the ones that are bad from the ones that are good, you actually have a phenomenal signal.
And so you should start with something you have conviction that should work.
And they should stick with it, you know, not analysis paralysis, but stick with it enough to find out why it didn't work.
And so any method where you're just spray and prey against everything is a waste of time.
It seems like there's some version of the same refrain that there starts to be a drumbeat around this, which is what's becoming most important is the hypothesis is asking the question.
Answering is getting easier and easier and easier.
So it's hypothesis formation that becomes the differentiating skill.
That's right.
And be able to use the tools to answer it quickly.
Sure.
And that's what happens with any progress in science.
If I give you a higher resolution microscope, it doesn't tell you the answers.
It just allows you to do experiments faster and experiments with higher resolution.
And so that's really what these methods are doing.
What people who are adopting the new methods should see is that they should have a higher cadence of being able to ask a question, work faster, work faster, and break things.
And it's exactly that thought process around build lots of things, try them, but build the things and work on them that you believe in, not just spraying and praying.
I'm curious if there's any other categories of the objective functions here.
So we talked about returns in the future, fundamental changes in the future, sales growth, cash flow, growth, whatever.
Are there other objective functions against which you find it fun or anything?
interesting to measure data set variables. So other outcomes, outcome variables that you're
interested in or do you think people should look at? The outcome variables I'm interested in is
literally who's going to win in the marketplace. And I think that what happens in the marketplace is
at each scale, I mean, having been at startup companies, the different scales look completely
differently. There's a great book called Crossing the Chasm. Great book. And so that process of actually
understanding who's breaking out and who's winning at what scale and how you measure it, which
parameter used for determining who's winning at what scale. And that's really where the predictive power
is going to come in. And, you know, again, if, and if they're winning in a microcosm, they could
easily spread it across other geographies and so on and so forth. And really, it's understanding,
this gives me back to a discussion I've often had with the discretionary folks. I say, why do you like
the company? They say, great management. So what does that look like? Does it look like better cost
efficiency? Does it look like new products? Does it look like new geographies, different types of people being
hired? If you tell me what it looks like, I can go get data and we can see it. But if
If it just sounds like good management, I don't know how to go there.
And so that's why it's a building a bridge process with the right discretionary.
People who can think about that process in enough detail to sort of drill down and say,
what does it really look like?
What would it look like?
And, you know, where does the data come from that actually allows me to see that?
What have been the most, the couple most recent, exciting things that you've uncovered or found?
Well, I mean, I think the one I described about looking at the parade, so if think about
that parade of distribution, the Walmart.
That's fascinating.
It's a great example because the other thing is that's, it's a, it's a great example.
because the other thing is that everything's parade.
It's like the stores, you know, which stores are doing well.
It's not just which customers.
It's every aspect of the business.
And they're all parallel.
And so I think that is actually, it just speaks very strongly to why you want to have
one level more granular in your information, which is really what the data does.
So the data just gives you a one level more detail.
Again, you know, elaborate your spreadsheet with that extra level of detail and populated
with the data.
And then work with the discretionary person to figure out what it means in terms of
valuation.
Any other favorites from across your entire career, not just recent, but
favorite aha moments of finding or discovery? Well, I think the moment I was talking about with the honeymoon
period versus the organic growth rate of lots of stores, that's another one. Again, I think there are
lots of examples where people haven't realized how easy it is to get certain data. And they've literally
walk by data that's in the world and ignored it. And they're using a less useful source of information
when the other sources are available. And I have to be careful about it. Some of them, because I don't,
I want to protect people. But to get back to your question, which I like,
In fact, I ask that all the time in my interviews also as I say, well, look, if you're in data science, there should be examples of things that were surprises where you started with one thesis.
And once you figured it out, you have a completely different point of view.
And I usually ask people in an interview to tell me about one of those moments.
Because if you just do engineering and you build something, you can sort of thinking it is a semi-mindless process because you're just building this thing and you started with a design and you just finish.
But data science is not like that.
You basically start with some hypothesis and you go and you look at what the data tell you.
I'll give you an example of that from way before I was in industry.
When I was in university, people would come with questions, and we work on projects for industry.
So Sainsbury's supermarket in the UK, I'd come, I think was Sainsbury, it might have been Tesco.
They came with a dataset, and they wanted to print coupons.
And they said, look, I want you to look at what people buy and determine their life stage, you know, single, married, all the way to retired, and lifestyle, which they meant by wealth.
And then figure that out, and then we'll figure out which coupon to print.
And it turned out that we did that for them, but the two parameters which were most predictive of their shopping were different.
The first parameter was, were they immigrated to the UK from?
This was in London.
And the second parameter was how long they'd been in the country.
And those two parameters actually were more predictive of their shopping than the other parameters.
And so, you know, you have lots of examples where the data just tells you something completely differently from what your top-down point of view was.
And you want to read out of the data instead of reading into the data.
And there are many of these sort of moments of epiphany where you figure out, ah, that's what's happening.
What investors or individuals or firms have you been most impressed with as it pertains to the use of data specifically?
I have to say that when I worked at 0.72, I was very encouraged by the enthusiasm the firm had for actually getting involved in data.
And I think that other firms are much less convicted about how powerful the data is going to be.
And so I think, you know, that's impressive.
You know, obviously, 2 Sigma is an world quant are examples of firms that have been investing in lots of different data sources for a while.
I know there are several firms that have recently been working on building large teams.
And, you know, I interview lots of folks and I see where they end up.
Certainly Citadel, Millennium, doing that sort of thing.
When I first interviewed with lots of those firms, I think lots of the firms were actually still in relatively early stages.
And I don't know what's happened in the interview.
meaning few years, but, you know, presumably they're moving along, but, you know, I think it's
going to, I still think it's very early days. It's funny because a lot of what you've described
is just that. It's early days. I think the market perception is actually quite different,
that there's this like quant fatigue almost, that people feel like everything's kind of going
quant or indexy or smart beta or whatever, when in fact, maybe it's just the very early days.
So what do you think that life cycle looks like? How long will it take? Will there come a point
when there's basically no pure discretionary managers anymore.
And everybody that's competing are able to compete in the marketplace
and an increasingly competitive marketplace
is data-informed or purely data-driven.
A couple of thoughts here.
One is when I joined the internet advertising space,
in 2009, all webpages had static ads like a newspaper.
Everyone saw the same ads who went to the page.
In 2014, we crossed the point where the majority of ads you see are dynamic,
which involves thousands of companies participating in auctions in a tenth of a second.
and dynamically assigning the ad to your page.
Huge change in technology,
huge change in the companies involved in that six years.
And that's for ads,
which costs as much as a teeny less than a grain of rice.
And so on one hand,
you might think,
well,
look,
there's much more opportunity in finance to do this right.
And so,
you know,
is it going to be six years
and the whole thing's going to change?
And,
you know,
but I think that there is,
there are people who are successfully investing with other strategies.
And I think it might,
I've been debated to myself
whether it's going to be faster or slower than that,
but I don't know,
10 years would be a long time for this change to occur.
The other thing is I would, I think that people like the cell side might be affected before the
by side because essentially the information, if it's not coming directly through them,
it's going to come through someone else and they're going to get disintermediated.
And so I think the risk of being disintermediated in the cell side and missing the boat is
much, much worse than in some of the buyside shops.
And certainly, you know, and we are already seeing it, the byside companies that are winners
are also going to change because the ones who are most quickly to adapt.
the data process are actually going to be.
And, you know, the thing is, interestingly, people get this.
I do lots of client presentations also in my firm.
And one of the recent presentations, one of the investors said, which of the investing teams
use the data the most?
Literally, it's going to become a force from that direction because people are going to say,
I get this.
I get how this is an advantage.
And let me go and hunt around and listen to the story of who's actually leveraging this
type of information.
And the success of the firms might be driven by the clients getting
up to speed in figuring out, look, I see this happening. And because the thing is that sometimes
the investing firm is going to be conservative with a small C and they're going to want to keep doing
what they were doing because it's the classic disruptions, classic Clay Christensen. You just keep trying
to do what you're doing and the disruptor comes in and does something completely orthogonal to it.
And so I think that part of this dynamics is going to be driven by clients who are trying to find
the places where they're really being smart about how they're leveraging data.
I'm curious what the high level outline of that presentation is that you give to clients.
It's actually largely what we talked about today.
I mean, I talk about lots of examples of probably getting a little more detail, which I can,
about examples of the way that internal teams have leveraged the data and the way it could be leveraged in a quantum process and a little bit more detail.
But essentially to compare and contrast the way that the world was before without the data and the way the world is with the data, again, it resonates.
I mean, the clients get it.
And the clients are very enthusiastic about seeing the firm take up, take up this path.
My episode this past week was with the former GM of the Philadelphia 76ers.
And one of the points he made was that the NBA and other sports leagues are probably a decade or two behind Wall Street in terms of their use of data and analytics.
How would you stack Wall Street against other industries?
So in the spectrum from don't use any data to everything is super cutting edge, where does Wall Street fall on a relative basis versus other other industry?
I certainly think that Wall Street is in the top half of the distribution and maybe it's in the third quarter rather than the fourth quarter.
There's lots of industries. I mean, I saw this great presentation which showed how you build a house 15200 years ago with wood framing and how you build it today with wood framing and how factories looked 150, 200 years ago and how they have factories look with robots now. Clearly, the construction industry is a relatively late adapter to how to use automation and so and so forth and some technology. And so there are going to be a lot, there are lots of industries which are still really quite backward in how they use data science and technology. The incentive system in Wall Street actually will keep the firms actually.
near leading, but they still become conservative because they don't, certain types of investing,
they don't want to change their process because if it works, certainly don't change it. And
so there's a resistance to change. I think obviously the internet companies are in the forefront.
And I think one of the risks is that there's a white combinator company that does almost
everything a bricks and mortar company does. And so when is that white combinator company
is going to grow and essentially disrupt using the new cloud-based technology and replace, you name it,
you name the industry, you name it. One of the things we're interested in is doing something like
take every single sector. I don't care restaurants. How could they be using AI? So which names in that sector are hiring, posting jobs and hiring people that most look like the future? Just as an example of a signal of who's going to win, because leveraging either the computing power or the machine learning or the data science stuff, you name it, whether it's menu optimization, inventory optimization, everything you can imagine, which is going to be leveraging a lot of this machine intelligence stuff. And so,
I think it's going to permeate through the industry. And we'll almost have, when we get a little further down this path, almost an index by sector of who the leader is and rank the sectors by which sectors are actually changing most quickly to the future.
Fascinating. Well, this has been an absolute blast. I'm wondering if there's any major lessons or topics that you want to leave listeners with. And then I'll have one more closing question.
Covered a lot of ground. Yeah, we covered a lot of ground. I mean, it's been a great fun. I appreciate being here.
Now, I think I'm going to go and talk to university students, and one of the things I tell to the university students is that they shouldn't leave university unless they can code.
I don't care what they're studying.
We're entering a new world where we're going to see AI everywhere.
And coding is like the most powerful thing.
Being able to make stuff.
Being able to make stuff with just you and a computer.
And it's a very useful skill across all industries.
And so I think that there's a Will Gibson quote, which I think I'm used with you before, but which I love, which is the future is here already.
It's just not uniformly distributed.
And the key thing is the thing which excites me the most is it's not very often
where you get to see the future before it happens.
And you get to sort of position yourself, your surfboard or whatever, to try to take advantage of the wave.
And I think that this is clearly an early stage of something which is going to permeate through Wall Street and change the whole industry.
And it's just exciting to be able to see it before it actually fully happens.
Fantastic.
Well, so my closing question for everybody is that to ask what the kindest thing that anyone's ever done.
for you is. I do have the exactly the kindest thing. It's a bit of a story. So when I was,
grew up in a working class background, and I had started in school because my older brother was in
the remedial thread, they naturally put me in the remedial thread. So the first few years of my
education, I was actually in the remedial thread. So you were in these large classrooms and they
couldn't even try to teach you. So I would read lots of books. And I was reading this book by George
Gamoff called 1, 2, 3, Infinity. And I was trying to approximate pie using circumscribed and inscribed
polygons of increasing side. And I had worked something out and I went and asked my math teacher.
And he lifted me out of the remedial program and put me in the gifted program that day.
And then I arrived in the gifted program and, you know, there are all these kids in the gifted
program who had been in there for their whole career. And they were a little bit sort of
comfortable. And I was just like, wow, someone wants to teach you something. And it really got
me accelerated in terms of, I mean, there's someone teaching me something. And the more I work,
the more I got taught. And, you know, I'm on where I am today because of that teacher who,
who actually lifted me out of that program and put me into the other program.
And so it really what motivated me to go become a professor was that I want to do the same
thing and I teach working class folks because there's so many people who just, you know,
in fact, with lots of students, I can find the math teacher who convinced them they weren't good
at math and they stopped trying to learn.
And so I've done lots of work in terms of teaching just because feeling I could want to
give back to that one teacher in sixth grade who actually moved to me out of the remedial program.
Fantastic. Great answer. Very unique. Relative of all the ones I've got.
and thanks so much for this. It's been a blast.
Yeah, thanks a lot.
Hey, everyone. Patrick here again.
To find more episodes of Investor Like the Best,
go to investorfieldguide.com forward slash podcast.
If you're a book lover,
you can also sign up for my book club
at investorfieldguide.com forward slash book club.
After you sign up,
we'll receive a full investor curriculum right away
and then three to four suggestions
of new books every month.
You can also follow me on Twitter
at Patrick underscore Oshag
OSHAG. If you enjoy the show, please leave a quick review for us on iTunes, which will help
more people discover Invest Like the Best. Thanks so much for listening.
