Invest Like the Best with Patrick O'Shaughnessy - Alexandr Wang - A Primer on AI - [Invest Like the Best, EP. 272]
Episode Date: April 12, 2022My guest today is Alexandr Wang, the CEO and founder of Scale AI. Alexandr founded Scale in 2016, having been inspired to accelerate the development of AI through his work at Quora and his studies at ...MIT. Specifically, Alexandr realized there was a lack of infrastructure solutions for producing high quality data, the lifeblood for AI models. Today, Scale provides data solutions to leading AI teams at Meta, Microsoft, OpenAI, Flexport, the US Air Force, and many others. This time last year, the business was valued at over $7 billion. Our conversation is a primer on AI. We discuss the building blocks beneath successful artificial intelligence, AI’s role in both the public and private sector, and why data is the new code. We also cover the similarities and differences between AI and software from an investing perspective and what inspiration Scale takes from AWS. Please enjoy my great discussion with Alexandr Wang. For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- This episode is brought to you by Canalyst. Canalyst is the leading destination for public company data and analysis. If you're a professional equity investor and haven't talked to Canalyst recently, you should give them a shout. Learn more and try Canalyst for yourself at canalyst.com/Patrick. ----- This episode is brought to you by Lemon.io. The team at Lemon.io has built a network of Eastern European developers ready to pair with fast-growing startups. We have faced challenges hiring engineering talent for various projects and Lemon.io offered developers for one-off projects, developers for full start to finish product development, or developers that could be add-ons to the existing team. Check out lemon.io/patrick to learn more. ----- Invest Like the Best is a property of Colossus, LLC. For more episodes of Invest Like the Best, visit joincolossus.com/episodes. Past guests include Tobi Lutke, Kevin Systrom, Mike Krieger, John Collison, Kat Cole, Marc Andreessen, Matthew Ball, Bill Gurley, Anu Hariharan, Ben Thompson, and many more. Stay up to date on all our podcasts by signing up to Colossus Weekly, our quick dive every Sunday highlighting the top business and investing concepts from our podcasts and the best of what we read that week. Sign up here. Follow us on Twitter: @patrick_oshag | @JoinColossus Show Notes [00:03:04] - [First question] - The role that AI and data play in geopolitics and foreign policy [00:07:21] - The end state of a digital arms race akin to nuclear weapons [00:08:53] - Current state of things writ large and how the public and private sectors differ [00:11:33] - The flow and importance of talent when scaling AI and whether it’s more important than software [00:14:29] - His thoughts on how to communicate categories of what AI can do well and what is still a ways out [00:20:18] - The process of creating an AI model and the stages of development [00:27:16] - Principles of building a great engine for gathering data [00:29:04] - The state of technology around annotating data writ large [00:31:31] - What Scale does as a business and their product lineup [00:35:08] - The Storage and Compute equivalents in the AI space [00:37:08] - How Scale fills the gap in producing better and cleaner data [00:39:52] - What Scale will look like in 10 years if their vision is fully realized [00:41:11] - Where AI is in the S curve of acceleration and where AI and software intersect [00:44:32] - Questions to ask about how to incorporate AI and data sets in your business [00:46:23] - What worries him about the proliferation of technology that makes AI more accessible to the masses [00:48:27] - The most interesting AI model he’s ever come across and collapsing the friction between human intent and programmable outcomes [00:51:51] - The kindest thing anyone has ever done for him
Transcript
Discussion (0)
This episode of Invest Like the Best is sponsored by Canalyst.
Canalyst is the leading destination for public company data and analysis.
Founded by a former byside analyst who encountered friction sourcing, building, and updating
models, Canalyst is now used by over 400 institutions, including the largest money managers globally,
and by a number of guests on the show.
With detailed company-specific models and data on virtually every public company,
panelists clients are able to ramp up faster, update models instantly, and incorporate the highest
quality fundamental data into any workflow.
If you're a professional equity investor and haven't talked to Canalyst recently, you should give them a shout.
Learn more and try Canalyst for yourself at Canalyst.com slash Patrick.
That's C-A-N-A-L-Y-S-T-com slash Patrick.
If you're curious to hear more about Canales, stay tuned at the end of the episode,
where I talk to Canales customer, Brandon Weir from BWCP to discuss Canales in more detail.
This episode is brought to you by Lemon.io.
The team at Lemon.io has built a network of Eastern European developers ready to pair with fast
growing startups. We have faced challenges hiring engineering talent for various projects, and
Lemon.io offered developers for one-off projects, developers for full start-to-finish product
development, or developers that could be add-ons to an existing team. Check out Lemon.io-slash-Patrick
to learn more. Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest Like
Like the Best. This show is an open-ended exploration of markets, ideas, stories, and strategies
that will help you better invest both your time and your money. Invest like the Best is
part of the Colossus family of podcasts. And you can access all our podcasts, including edited transcripts,
show notes, and other resources to keep learning at join colossus.com. Patrick O'Shaughnessy is the CEO of O'Shaunusie
Asset Management. All opinions expressed by Patrick and podcast guests are solely their own opinions and do
not reflect the opinion of O'Shaunacy asset management. This podcast is for informational purposes only
and should not be relied upon as a basis for investment decisions. Clients of O'Shaunice
the asset management may maintain positions and the securities discussed in this podcast.
My guest today is Alexander Wang, the CEO and founder of Scale AI.
Alexander founded Scale in 2016, having been inspired to accelerate the development of AI through
his work at Cora and his studies at MIT.
Specifically, Alexander realized there was a lack of infrastructure solutions for producing
high-quality data, the lifeblood for AI models.
Today, Scale provides data solutions to leading AI teams at Meta, Microsoft, OpenAI,
flexport, the U.S. Air Force, and many others. This time last year, the business was valued at over $7 billion.
Our conversation is a primer on AI. We discussed the building blocks beneath successful artificial
intelligence, AI's role in both the public and private sector, and why data is the new code.
We also cover the similarities and differences between AI and software from an investing perspective
and what inspiration scale takes from AWS. Please enjoy my great discussion with Alexander Wang.
So, Alex, we're going to talk about every dimension I think of AI, artificial intelligence,
machine learning, all the things that it's going to impact.
I'd like to structure our conversation from sort of the widest angle down to the most specific,
which will probably be around your specific business and product.
But given what's going on in the world today, I think it's probably an appropriate place to start
with the role that AI, data, machine learning, et cetera, play in geopolitics and foreign policy.
We're sitting here in the U.S. and everyone's mind is on.
Ukraine, I think everyone understands that cyber is a sort of plane of conflict that exists,
but we don't know a whole lot about, to be honest, even I don't know a whole lot about.
So I'd love you to give us your overview, given you've got a sort of insider view and seed
and understanding of all this. What role does AI play in the global theater today in your view?
A few assumptions to talk through, or a few things to talk through that I think are really relevant
here. I think one is that generally speaking, deterrence has been a,
incredibly positive thing for world peace and the stable world order over the course of the past
40, 50 years. I grew up in Los Alamos, New Mexico, which is the birthplace of the atomic bomb.
The deterrence from the atomic bomb has been an incredible forcing function for a stable and
peaceful world order. And I think where we're at today in a broad sense is that there's an
entirely new set of technologies that are either in early development being developed right now
or have been recently been developed,
that are significantly shifting
what conflict looks like in general.
One of these is cybersecurity,
which you mentioned,
which is a very challenging general attack plane
because it's generally speaking
almost impossible to defend against
and the attacks are numerous in nature.
There's sort of like these shots
that can produce a lot of attack vectors.
I think AI and machine learning
is a very, very significant one.
Its ability to intersect
with many of these attack planes,
whether that is intelligence and understanding what's happening in these various theaters,
whether it's through the actual warfare component or it's through cybersecurity,
a cross-sectional technology across each of these different fronts.
And one general thing that is important to think about is that versus the prior paradigm of conflict,
where a lot of conflict would happen in literal physical conflict in wars, in theaters,
I think we're moving much more to a paradigm where hopefully 90% of conflict,
are actually resolved or battled in a purely digital context long before an actual conflict
arises. And so this broad shift, which is a thing is much more than just AI and machine learning,
but this broad shift of sort of the physical battleground to a primarily digital battleground
is probably the greatest macro shift. And then on a micro level, the ability to have great
AI technology is probably one of the greatest enablers or the greatest power contributors in terms
of which countries are going to able to be defend the most effectively, as well as to
enable the greatest level of deterrence, the broad world order. That's the most practical view
from the conflict perspective. And then I think from a geopolitical plane, I think that we're
entering this period of what's called great power competition or strategic competition,
where there is a small set of highly technically capable countries who are great power
competitors, there are strategic competitors on a number of fronts from economic competition
to ideological competition to obviously a lot of on the field competitions. And so this shift
from the paradigm of the past 20 years being counterterrorism, this shift from counterterrorism
to great power competition means that the technology itself and in particular being
exceptionally good at technology is really, really critical. And artificial intelligence is just one
of those technologies that happens to have compounding order of magnitude effect on any of the
other technologies that we could possibly develop. And so that I think contextualizes why all this
stuff matters. And I think that something that I think is incredibly important because the way that
these world events happen, and I think the Russia-Ukraine conflict is one great example of this,
is that they do not happen in predictable ways. There's not a predictable, logical way in which
events will play out over time. It is one of these great tail risks, so to speak, to the good
and well-being of humanity.
If we think about the analogs between physical kinetic conflict culminating in nuclear
deterrence, it's hard to imagine something more powerful than whatever the largest strategic
nuclear weapon is against the city or something.
Is something similar being developed in the digital sphere?
And what pops to mind is the worry that another country or another power reaches some sort
of level of general artificial intelligence to use the common term first, and that that
thing represents the equivalent of being the first and only to have nukes, and maybe it's even
harder to copy than a nuclear weapon might be. If we are shifting to digital, what do you think
the arms race equivalent is towards? What is the end state that the different powers might be trying
to achieve in that digital realm? First of all, generally speaking, for artificial intelligence,
it's a very multifaceted technology. It's more similar, you would say, to software or the internet
than it is to like any specific technology that will be built, a very particular weapon or whatnot.
And so I don't quite fully agree with this analogy that an AGI is the equivalent end state
to a nuclear weapon.
That being said, I do think that AI and AGI, the best way to think about is it enables
maybe 10x better strategic decision making and 10x better ability to actually take action
against various tactics that you want to enact.
Maybe the best analogy is like playing chess versus a human versus playing chess versus
an incredibly talented algorithm or incredibly talented computer.
the way in which it will play out strategically dominant behavior in extended periods of conflict
or extended periods of time.
How would you describe based on what you know, the current state of things amongst these
great powers?
What's the status of the United States, I guess is the first question?
How does that relate to the status or digital capabilities of other countries?
And how does public versus private sector look?
Is the U.S. private sector far better equipped than the public sector or something like this?
What's the state of things?
I'll talk about AI technology, which is the area that I understand the best.
So fundamentally, AI technology, modern deep learning was initially developed primarily in the private sector
through a lot of incredible research that was done at Google and DeepMind and OpenAI and Facebook
and these incredible firms, as well as a lot of contemporary research done in China, the clear
other major location where a lot of great artificial intelligence research has been done.
And a lot of that research, there's been this very very important.
very fast integration of that into very particular use cases in China.
Facial recognition technology, I think, is one of the primary examples
where this companies like SenseTime, Face++, et cetera,
that are primarily contractors of the Chinese government,
have built world-class computer vision algorithms for facial recognition
for use by the Chinese government in furthering a bunch of national objectives that they have.
We may have questions from humanitarian perspective, like the work being done with the Uyghurs.
That all has happened in China.
Then at the same time, I think there's starting to change.
change now, but over the past, call it five years in the United States, there's been this very
unclear relationship between the private sector and the public sector, or the private sector and the
government, around the use of AI for applications like defense and intelligence. The Google Project
Maven, Google Cloud Project Maven, conflict was probably the most visible, most clear example
of this. But as a general rule, I think that there's not the same level of partnership between
the best and class technologists in the United States, as well as the government. That is a significant
cultural problem that needs to be solved in the United States for us to really be well positioned
in the future with respect to artificial intelligence. I think broadly speaking, we're in the very
early days with all of this technology, the speed of innovation in artificial intelligence is incredibly,
incredibly fast. I think what we do over the next five years, or do over the next 10 years,
is probably 10x as important as what was done over the past five years, just given the pace
and acceleration, the technology of artificial intelligence. I think the real question,
for policymakers in the United States,
global leaders,
for technologists in the United States,
is what are we going to do
over the next five to 10 years,
and how does that compare to what other great power competitors
will do over the next five to 10 years?
And that, I think, is what will really set the stage
for the next phase of this great power competition.
In the world of software,
you always hear about these 10x engineers,
and one wonders if in this realm it's even greater than that,
is the talent at the tippy top
of this world
are going to be accessible
to great powers and governments
like how this is so important
and there's this multiplier effect of like the absolute
best talent. It seems like that's
the thing that matters. It's like competition for that
talent, not just between great powers, but
also between the top engineer, AI person
at Facebook versus going to work for the NSA
or something. What's been your sense of the flow
of talent and the importance of talent?
Do I have this roughly right that it's
even more important that it might have been in software?
There's three critical components.
to broad scaling of artificial intelligence.
One is certainly talent, which I'll get to in a second.
The second is compute and computational power,
which is impossible to ignore,
and especially as we think about a lot of potential disruptions
to the supply chain of large data centers
or large shipping manufacturing,
that's a really important one for us to consider.
And then the last one is data and data scale.
And that is obviously near and dear to my heart,
just running scale,
but it's also something I think is also impossible to ignore.
So I think those are generally speaking the three vectors that matter.
And there's been a bunch of research that shows that more or less power of these AI systems,
both quantitatively and qualitatively, scales on these three dimensions.
Then if you think about talent in particular, the interesting thing is there probably is a greater
scaling effect, but not because the technology itself is one that is more inclined for a
brilliant mind than software per se, but actually it's probably much more a function of
the power law impact of various AI systems
that is actually even greater than that of software systems.
I think that a lot of software systems,
because they have to operate within fixed paradigms,
they have more capped total impact or total utility
that they can deliver.
These AI systems,
like you think about the large language models
that have been developed recently
by Open AI and Google and others,
those algorithms have an incredible ability
to adapt to new domains.
The cap on the level of impact
or the cap on the amount of utility to be delivered,
is much, much higher.
And so what that means is that the set of people
who can develop these technologies,
Open AI, I don't know what the latest count is,
but I think their total head count is 250 people as a corporate.
That's an incredibly small number of people
to generate the broad-scale impact
that I think they will deliver
over a multi-decade timeframe.
So I think that's a very accurate point.
And this is one of the things that I think,
by the way, matters a lot,
is high-skilled immigration into the United States.
America's an amazing place.
And I think that one of the things
that is differentiated about America
a lot of people want to move to America and want to do their lives work in America.
They want to build families in America.
And I think we need to do everything we can to enable and encourage that as much as possible.
If you think about inputs or you think about the big needle movers that are actually maybe easy to change,
this is one of the ones that has, again, multi-decade consequences.
I couldn't agree more on that point.
Maddening that we don't become just the perfect beacon for all the most talented people.
The interesting analogy that I've heard before, just to wrap our minds around the sorts of things or tasks or functions or whatever that constantly improving AIML models can accomplish.
One model that was funny and interesting was like anything that an intern could do for you, you might be able to scale up through one of these models.
It's complicated enough that a person's on it now, but it's simple enough that you give it to an intern and sort of repetitive.
and I always kind of like that conception.
What's your way of thinking about how to communicate to your audience,
other businesses using your tooling,
and just people in general, like what categories of things AI can do well
and maybe what categories of things we're excited about,
but might be a very long time until AI can do well?
I think this is one of the general misconceptions about AI and machine learning,
which I think causes a great deal of FUD,
which is that the intuitive belief is that the things that are easy for humans to do,
are going to be the things that are easy for AI and machine learning to do,
which is absolutely not the case.
And the things that are easy for algorithms to do are relatively orthogonal, frankly,
to things that are easy for humans to do.
One simple example here is I think that it's going to be a very, very long time
before we have home robots that can do things like fold your laundry and do your dishes,
but a much shorter time span slash, you know,
I think this is already today,
where you can have artificial intelligence systems that are world-class copywriters
and can write better rhetoric, better words than most people ever could.
There's probably a few frameworks I would assign to this.
I think in a broad, general sense,
one way to think about the potential impact
or lower bound the potential impact of artificial intelligence
is kind of as you mentioned,
which is the ability to scale repetitive human tasks.
So take repetitive human tasks,
instead of going from zero to one,
go from one to end human work.
And I think that this is a generally amazing thing to happen
because I think that humans, for the most part, don't enjoy repetitive tasks or generally find
those relatively unpleasant and find a much more exciting to be creative and to constantly be creating.
This ability to scale human tasks from one to end is going to be this incredible, normally
economic good or economic enabler for the world, but also going to be a significant enabler
for humans to be more leverage, more happy, more creative, et cetera. So I think that's one way to
sort of contextualize the broad impact that AI can have. And there's a bunch of other nuances,
which I'm sure we'll get to.
If you think about what tasks humans are good at
versus what tasks algorithms are good at,
generally that more or less boils down to data availability,
which is that where there are large pools of digital data
an algorithm can learn from,
and those pools of digital data either have been collected in the past
or easy to collect in the future,
those are going to be the problems which algorithms can do effectively
and can learn to do effectively.
And then areas where there does not exist digital data
and is expensive to collect this digital data,
those are going to be the last things to be automated.
So a great example is, if you look at GPT3 and these large language models,
the real secret behind it is that it leverages two decades of Reddit data,
which is two decades of humans using the internet,
and basically typing language into the internet in various forms
for decades and decades and that is the pool of digital data
that it used to be able to do these incredible things
and writing long-form text.
Then if you think about this parallel that I mentioned around home robot,
there's so little data about actual capture data of,
let's someone folding a shirt or somebody folding a towel
or going around and doing chores,
the ability to actually collect and produce that level of digital data
necessary to produce algorithms that can understand that
and actually perform those tasks is an incredibly, incredibly, incredibly hard road.
This extends, by the way, to things that are like really unintuitive,
So, for example, DeepMind and Open AI very recently released algorithms, some of which are very good at
deep mind release an algorithm that's very good at competitive programming, Open AI released an algorithm
that can prove very difficult math problems or math theorems. And these are both things which are
very, very challenging for humans to do, very, very premium skill sets as far as humans go,
but there are incredible pools of digital data as well as abilities to verify or simulate the outcomes
here, which allow these algorithms to perform actually incredibly, incredibly well.
There's this very interesting process by which artificial intelligence will slowly automate
or meaningfully change what human jobs that are primarily digital in context will look like.
And then a lot of the physical work will, I think, be generally untouched for a very long time.
In many ways, if you're right, the whole idea of blue-collar work, let's call it,
or something like that being in jeopardy is maybe completely wrong.
the white collar work, the sort of knowledge work that's mostly digital in its form today
is more susceptible. As you said, this should be a good thing, right? It should open up human leverage
to higher order, more creative, more interesting tasks, non-repetitive tasks. I think as of this morning,
jobless claims are at their like all-time low or since the 1960s or something. So all this great
technology doesn't seem to have affected things that badly. But is that right that some of the jobs
that we thought might be automated first by robots or something actually might be the last
things to get automated by this model and the more digital knowledge work stuff may be first?
It's a strange thing, but I think that probably is the case. Maybe one simple mental model is if you
spend all day in a word processor or all day in Excel, the actions that you take in those
products are some of the most likely to be automated, frankly. And for the most part, I think a lot of
that work is probably not very inspiring. And so it will be this incredible enabler of leverage to think
bigger about those kinds of jobs.
Maybe it makes sense to help people understand the process of creating one of these
models in the first place.
I think the discrete steps, let's say the outcome is a model that makes a useful prediction.
Ultimately, this is all predictions.
That's sort of what's being generated by the models in the first place.
I don't know where to start, whether it's with raw data or annotation of data, and we're
starting to get into what scale now provides for companies.
But how do you think about explaining the discrete or the important stages of building one
of these models in the first place. I think just understanding that architecture will let us
dig into each piece a little bit more. Again, everything starts with the data. I often will
analogize the data for these algorithms as the ingredients that you would make a dish with or the
ingredients that you would make something they would eat with. It is incredibly, incredibly important.
We often say this thing, which is data is the new code. If you compare traditional software versus
AI software, in traditional software, the lifeblood is really the code. That's the thing that
informs the system what to do in artificial intelligence and machine learning, the lifeblood is really
the data. That is certainly like one major change. That's really important. The life cycle for most of these
algorithms is a fewfold. So first is this process of collecting large amounts of data. By collecting
it could be data that is already sitting there. There's a lot of software processes that already
collect a bunch of data. There's a lot of cameras in the world that already collect a bunch of data,
but you need to like get the raw data in the first place. Then it goes through this process of annotation,
which is the conversion of this large pools of unstructured data
to structured data that algorithms can actually learn from.
And so this could be, for example, in imagery or video
from a self-driving car, marking where the cars and pedestrians
and signs and road markings and bicyclists and whatnot are
so that an algorithm can actually learn from those things.
It could be, for example, in large snippets of text,
actually summarizing that text
so that now you can understand and learn
what it means to actually summarize text.
So whatever that translation is
from unstructured data to a structured format
that these algorithms can learn from,
then it goes through a training process.
So these algorithms basically look through
these reams and reams of data,
learn patterns, and slowly train themselves,
so to speak, to be able to do
whatever task is necessary on top of the data.
And then you launch one of these algorithms in production
and you run them on real world data
and they're constantly producing, as you mentioned, these predictions.
The very important piece is this is not a sort of like one-way process.
This is actually a loop.
And if you look at almost every algorithm that has launched out there in production,
it is not a sort of you build the algorithm and then you're done
because these algorithms are generally very brittle.
And unless you're constantly updating them and maintaining them,
they will eventually do things that you don't want them to do
or they'll eventually perform poorly.
There's this critical process by which you are constantly then replenishing them.
you're constantly going and recollecting new data, annotating it, training the algorithm,
launching that new algorithm onto production.
And if it constantly undergo this process to create very high-quality algorithms.
I want to make sure that this interesting point you made about data being the new code
really hits home for people.
Maybe even put that in a business context.
So if the IP or the moat of a software company is this code base that takes a very long time
to develop, has all sorts of dimensions to it, maybe it's microservices, maybe it's some code
monolith, it's unquestionably like an incredibly valuable asset. It's digital, but it's an
incredibly valuable asset to the company. And you're talking about, I think, a transition where it's
something different where maybe, I don't know, maybe Google's data repository or something is
this unbelievable advantage that they have because no one else has access to all of their data.
Is that kind of what you mean that ultimately maybe something like Google, their data is worth
a lot more than their code base and that that will become a trend that we see sort of across
industries. If you look at the highest performing algorithms across a variety of different domains,
image recognition and speech recognition and summarizing text and answering questions of text,
so these very different cognitive tasks, look under the hood. They actually all use the exact
same code base. That's been this very meaningful shift that's happened over the past few years
in artificial intelligence. And so we're at this point where the code has become
effectively the same and more or less a commodity, so to speak, when it comes to artificial
intelligence and machine learning. And the thing that enables the differentiation is really the
data and the data sets that are used to power these algorithms. To your point, if you think
about, one of the ways that we talk about us in a business context is if you think about what is
your strategic asset. In general, in business, your strategic assets are the things that allow
you to differentiate yourself against your competition. In a world where 99.99% of the software
in the world is traditional software, and then only 0.01% is AI software, then you care
the most about your code. Your code is what will differentiate your product versus your
competitor's product or your processes versus your competitor's processes, et cetera.
But then as more and more of the software in the world is written infused with AI, using AI, or
over time, the interfaces shift to AI interfaces, in alex-like interface, for example, as that shift
happens is you go from 99.99 to 90-10 or 80-20 or even 50-50 over time, the vector of differentiation
totally shifts to data and the data sets that you have access to. And so what means is that your
strategic differentiator, to your point, as a firm, is going to be primarily based off of what are
my existing data assets? And then what is the engine by which I'm constantly producing new
insightful differential data to power these core algorithms that are actually powering my
business. And these algorithms at the core that will power the future of business, I think,
are relatively core. I think there's definitely algorithms around automating business processes
that are going to result in significantly more profitable firms over time. There's going to be
algorithms that are based around customer recommendations and customer life cycle, which is a lot
of the algorithms that we've seen to date. Imagine TikTok recommendation algorithm, but for like
every economic interaction or every economic transaction in your life that is constantly identifying
the perfect next thing that you may want to transact with.
And that is going to exist across every firm or every industry is basically going to have to
build their version of that.
And that's going to result in like significantly more efficient trade.
The long-term impacts of that you can think of as like a general reduction in marketing
expenses or sales and marketing expenses because the algorithm just is a better job at knowing
what the user wants to do next than having to do all this marketing and all this very active
sales.
There's a lot of very real changes to, I think, the physics.
of what the best businesses will look like in, let's say, a decade or two decades or three
decades that come from artificial intelligence. And if you think about what will allow me to do
these things better than someone else, it's the quality, efficacy, and volume of the data
that is used to power these algorithms. If getting to that state of some sort of differentiated
data advantage is the goal, then that engine that you referenced becomes critical idea for any
company, a piece of infrastructure, what principles are there about building a good engine for
data gathering that may be cut across different types that you've observed? What's behind the great
engines that you've seen for gathering data? There's a few tenants. I think first, quantity is a
quality of its own. You want to just have lots and lots of data coming in. That will be a
differentiator no matter what. Not all data is created equal. And so you want, in general,
the way these algorithms work is there's sort of like needles in the haystack almost in most of your
data that end up being differentially valuable. You'll want to develop a process by which you can
actually what's called curate, but identify and understand what are the really valuable pieces of
data and general data inflow that we're getting. You'll want strategy for data diversity, or you'll
want a strategy by which you're not just going to keep collecting the same old data by which you're
going to constantly be expanding the domains or the kinds of data or the diversity of data that
you're going to be collecting over time. And that's a really interesting one. I think that one of the more
intuitive examples that I think people understand the best is that Tesla's autonomous vehicle strategy,
one core part of their strategies that they have, all of the Teslas in the world are in some
sense collection vehicles. They're all encountering very rare situations that then improve and help the
machine learn algorithm. There is a truth to that idea, which is that because not only have a greater
volume, but they also have great systems to identify these needles in haystack, that's a strategy
by obtaining an incredibly diverse data set over time that's very beneficial to the long-term.
differentiation of their machine learning stack. By the way, Tesla has other challenges as well,
so this is not a clearly dominant strategy, but there is a truth to some of these ideas.
What should we know about the state of technologies around annotating data that I know,
obviously, scale is deeply involved with the example in Ukraine around like satellite imagery?
There's this interesting jump from like raw data or raw input into some sort of refined
signal or piece of information or whatever. I don't know what the right terminology is.
What is the state of that in the world today?
Like, has that changed a lot?
Has that been pretty constant over the last few years?
Where is it going?
I think it's meaningfully changed.
It's one of these very interesting problem domains
and one that we find incredibly exciting
because it's undeniably a, quote unquote, dirty problem
or it's undeniably a very complex,
multifaceted, operational and challenging problem.
Historically speaking, this has been done
in an incredibly low-tech way
and an incredibly very manually intensive way.
for the AI industry to date. And a lot of that has been because it hasn't been a problem that
technologists have wanted to dedicate themselves to. For the most part, most people who go into
computer science, most people who go into artificial intelligence, by definition, they go into
computer science because it's a significantly more constrained problem that doesn't involve
these complexities around human operations. But I'm proud of the scale team for having really done
a lot of this work that I think will be seminal, is starting with a very manually intensive or
manually challenging process that is operationally very intensive, and then adding meaningful
amounts of automation, whether that's algorithms that will do a lot of the work before it gets to
people, or algorithms that are able to identify most of the errors that people might make, or in general,
injecting a lot of operational efficiency processing into this process to enable significantly
better outcomes and either more efficient processes or higher quality data. That is really the
core engine behind what we've built at scale that has enabled some of the biggest platforms in the
world, such as Meta or Microsoft or many of these other large tech firms to significantly scale
a lot of their machine learning efforts. One of the companies that we really look to for
inspiration, we've been deeply inspired by is what Amazon has done in two different spots in
their business. They've done it not only with logistics for e-commerce, but they've also done it
with AWS. He's taken both of these very large, complex, half-operational, half-technological
problems, and they've added significant amounts of technology and automation and operational
efficacy to result in orders of magnitude better performance than you could have accomplished in
the past without sort of this embrace of the messiness of the problem.
A good excuse to talk about literally what Scale does more than an hour in, and it got into
the business itself, a whole bunch of amazing context and less than so far. What are the building blocks?
I'm struck that if you go to the little products tab on Scales website, it looks not dissimilar
are from the AWS tabs of old where there's lots of individual functions that are part of this
value chain, if you will, of building one of these things that you can quote unquote higher scale
to do. Just talk us through the business itself, what it does for customers and how that product
lineup fits together. If you are the retail operation infrastructure equivalent or the AWS
equivalent for the AI future, what does that look like today? Where we started was really this
problem around data annotation. Because what we notice is that it was this massive bottleneck for
the overall progress of machine learning is that most firms could simply not get high enough
quality data sets for them to actually build great algorithms. Those very much the initial
problem that we set out to solve, which is how do we enable companies to get better data to fuel
better AI at a score? And we started by doing that with some of the most technologically sophisticated
firms in the world. You know, we work with, as I mentioned, meta, Microsoft,
We work with large automakers in building autonomous vehicles like Toyota and General Motors.
We work with Open AI on a lot of their cutting edge research,
work with large internet companies like Instacart or Etsy or Flexport on a large-scale machine learning problems.
And then through doing that, we've built what we really believe to be the best engine in the business
or the best engine in the industry around producing these very high-quality data sets.
And a lot of our view in expanding beyond that has been that primitive is incredibly powerful.
And then integrating that primitive with other components of broader machine learning lifecycle
that was referencing before enables us to produce really powerful products just like how for
AWS, they started with the primitives of storage and compute.
And then they sort of combine those primitives in many interesting ways to produce like
these incredibly powerful products like RDS or the,
their DNS products or the list kind of goes on and on to solve customer business problems.
Really look at the same way. We have built this core primitive around data annotation and the
production of high quality artificial intelligence data sets. We can combine that with other
primitives that we've built, such as data management products or model monitoring products or
algorithmic development products. And we combine all these primitives in producing these products that
solve really core customer problems, whether that's in the government space, working with
defense and intelligence agencies and building really high-quality artificial intelligence
algorithms to solve core national security problems, or that's in solving, working with large,
complex Fortune 500s who have incredible amounts of unstructured data, such as documents flowing into
their business, and we build automation to support a lot of that document processing, or we work
with large map builders, and we use artificial intelligence and human processing to enable this
map creation. One of the great lessons from Amazon as a business is that parallel execution is
incredibly powerful. The traditional business logic or the traditional business truism is around,
hey, you should really just focus on a few things and do those things really, really well.
And that's your sort of like strategic edge. Amazon really threw that out and became
very focused on parallel execution while focusing on what are the primitives that are going to
enable them to build a lot of products very successfully. And we take more of that approach,
which is we've built incredible primitives. We're going to make those primitives better and better and
better, and then we're combining these primitives into building best and class products for our
customers. One of the things you'll hear about Amazon Web Services specifically is that at the end
of the day, what they're really trying to drive people towards is more S3 and EC2 storage and compute
spend. Those are like the most basic building blocks and a lot of the other services that
AWS provides effectively maps back on to increasing the volume of storage and compute as a business
model. Do you think that that's true for you all too? And if so, like what are the S3 and ECT?
equivalence for scale?
We think that's true to an extent in that, again, we're almost religious in this belief
that better data results and better AI and the data is a new code.
And so we really think about as like how does scale become the long-term infrastructure
provider to powering more or less the global expansion of data to meet the needs of
AI globally.
That I think is certainly one very large driving factor.
And a lot of our goal is to build the most data-centric AI infrastructure platform out
there.
But at the same time, I think that maybe one of the things that we really,
index on as a company. And I think that some Silicon Valley companies do with this, some don't,
but we would certainly take this approach, is to really index on the customer value creation,
which I think is really important in the space of artificial intelligence and machine learning,
because frankly speaking, this has been a technology that people have been talking about for a very
long time. And then the actual value-creating use cases are few and far between. There's a lot of
distrust of AI systems. There's a lot of belief AI isn't good or AI is so bad or AI is just this
snake oil technology. That is absolutely not true. Artificial intelligence is incredibly powerful
technology, but there's been an incredible dearth of firms that are focused on how do you
index against the actual customer value and the customer outcomes. They're able to generate using
technologies. On the one hand, we probably do believe this large-scale growth is a primitive belief around
large-scale growth of powering these data sets for machine learning. We also believe on the other end of
working very hard to power as big or as meaningful of customer outcomes as possible,
and that being a primary metric on which we value ourselves.
If we zoom in on the original use case, helpful annotation of data,
what does that mean in the literal sense?
If the outcome is well-labeled, clean, large data set and the input is whatever,
like what is the gap that scale was filling?
What is the function that it was actually doing to help the company produce more
cleaner, better data?
This is sort of the state of the market that we saw when we entered it.
Even when we had started, there were a large number of very talented machine learning teams
that were solving varied problems from autonomous vehicles to building AI for you commerce,
to building sort of voice assistance systems.
There were a lot of interesting use cases of the technology.
Each time we'd go to one of these firms, we'd ask them, what are your biggest problems,
always in the top three data quality and data volume is one of them.
And you dig into it, and it's because they have this incredible challenge of motivating
their incredibly brilliant, smart ML engineers and ML scientists and working on dirty problems
around data and data quality.
And so there's gaps are technological, sometimes gaps are cultural.
I think for us, we really view a some mix of both.
But what we did is we went into all these firms.
We took incredible religion and care to build great technology to solve data quality,
data volume, bottleneck.
That was really limiting the efficacy and power of their.
machine learning models at the end of the day, and sort of rebuilt, we took, I don't know,
maybe the Stripe Playbook is what you would say it was, or the AWS Playbook before them,
we built great developer APIs and a great developer platform that enabled the machine
learning engineers and machine learning scientists to basically hidden API, get as much high-quality
data as they needed or they wanted, and behind the scenes, we would do an incredible amount
of dirty work to make this possible and actually enable them to build these great algorithms.
Maybe just to nail the point home, like a couple examples of these dirty
Dirty jobs might be the categorization of images or some components of an image or something like
that. Would that be like a good example? Yeah, great examples are like given a huge amount of imagery
or a huge amount of data off of a robot, whether that's a software and car or some other robot,
marking in all that imagery. What are the things that are really important for that robot to see?
So enabling the robots to see in the first place. Other similar quote unquote dirty jobs are
for financial services firms. They have incredible amounts of transaction data. All this transaction
is fundamentally unstructured.
So understanding who are the vendors in these transactions, what are the locations,
what are the other pieces of insight that you can pull out of these transactions, and
another ones around other forms of audio data, for example, and pulling out the relevant
intent and understanding from audio data.
So all these dirty jobs, they fundamentally start with some data format that machines can't
read today and making them effectively legible for these machine systems.
If I fast forward and think about scales future five, ten years down the line,
how would you describe the absolute best case of success?
So if the mission is accelerate the development of AI,
make it easy for companies to build more AI models effectively,
which I think is kind of the core mission of the firm.
What is the best version of that look like?
Do you think five or ten years from now?
I mean, I think we really take the view that it just comes down to the customer outcomes
that we're enabling.
So rather than thinking about what scale exactly looks like in 10 years,
I think the right thing to look at it as in 10 years are the majority of companies in the world
able to effectively use artificial intelligence in some very high impact application to solve a
business need. And is that application in some way or another powered through scale products?
And we take a very open-minded view is we're going to build lots of products over time
that can power these products in different ways. And the delivery mechanism of the algorithm,
we may not build the exact product that a Fortune 500 uses. We need to build.
may be powering some SaaS company that is integrating AI into their product that the Fortune
500 uses. We basically want to take this very long-term platform view that as long as we're there
powering decade-long massive expansion of AI and machine learning, we're going to be really pleased
and happy. If we zoom out and go to the more market side of things and put our, my investor hat on
and think about what drives enterprise value, value creation, the things that investors ultimately
care about when they're putting money into a business, they want to get a lot more money out,
The world of software has obviously been a center stage for seven, ten years now because they've tended to be very scalable, fairly high margin, incredibly fast growing businesses.
And the word that you never want to hear as an investor is deceleration in the growth world where maybe they're reaching saturation points and software is no longer a new thing.
It's a fairly mature thing.
How do you think about, you mentioned this concept of thinking about like an S curve and maybe we're for software approaching the diminishing part of that S-curve.
curve. Where is AI in that same thing? And how might these two things intersect to form lots of
new enterprise value in the future if software becomes overly saturated?
One thing to think about software for a moment, the sort of alchemy or the magic of software is
that, A, you're able to collect very large-scale datasets in a very coordinated way, be that you're
able to build simple workflow tooling on top of these data sets. Think about your traditional
CRM or frankly, the majority of SaaS tooling is workflows on top of these data sets.
that enable business value.
And then three is basically infinite scalability
of a lot of these systems.
These are some of the technological primitives
that have enabled SaaS, broadly speaking,
or software in general,
to produce a lot of value for most enterprises,
but these primitives or these forms of alchemy
have some cap.
That's where the saturation of software
that you're mentioning.
Well, then if you think about AI technology,
and you use this mental model
that I mentioned before,
which is the fundamental promise of AI technology
is that you can take repetitive tasks
that people are doing, you go from one to end with those repetitive tasks, so you can automate
the endth repetitive tasks rather than relying on humans for that. Well, if you look at the majority
of Fortune 500 business or the majority of largest enterprise in the world, there are an incredible
number of parts of their business where they spend enormous amounts of money on large teams of people
to do repetitive tasks. The alchemy that is possible there is not only the automation of
meaningful parts of that work, but also the ability to even go further than even the best trained
humans could do in many of those tasks. The sum value, that potential economic value, or this
TAM, so to speak, of AI machine learning is just absolutely astronomical. And I think that that is
at minimum 10x, probably 100x, the total business value that has been generated by SaaS systems or
software historically. I think if you think about it, you have this one S curve of the saturation
of software. And then there's this very, very early S curve that is being developing right now
around the productization and productionization of large-scale AI systems, let's say in the enterprise
or let's say across businesses. And the real question is, okay, what's the pacing of that S-curve
versus the pacing of the saturation and deceleration of the current software S-curve? And I'm an
optimist. In not too long, we're going to have a massive proliferation of AI use cases within
the enterprise that are going to be way more impactful than the use cases of software in the past.
And the way you'll see that, the business RIs generated from high-quality AI systems are going
to be 10x more than the business value generated by, let's say, deploying a CRM or deploying the RP
system. How would you advise those listening who maybe run businesses and are nodding their heads
and think, yeah, this all makes sense, this is a new competitive plane or competitive frontier,
and I want to make sure I'm not left behind,
but I'm not a data scientist.
I don't have this in my background strategically or tactically.
What would be the questions that you would have them ask themselves
to start incorporating this thinking into their business?
I think the big questions are really around understanding the data that powers for business
and understanding at a meaningful level,
what are these potential data sets within your business that can fuel a lot of this future
wave of artificial intelligence and machine learning?
And then I think it's really thinking about,
and I think this is probably in partnership with partner or in partnership with advisors,
is thinking through what are the highest value business problems that artificial intelligence
could solve that would meaningfully move physics of my business.
I kind of mentioned a few of them, but one that I think applies to nearly every business is
this problem around customer recommendations or basically building better recommendations
to empower customer life cycles.
Another one I think powers almost every business is taking some of the most expensive
repetitive process internally
and figuring out ways to automate
or make those more efficient.
And I think it's thinking through
what are these core business use cases.
One very tactical piece of advice that I would give
is one of the challenges of AI
is we're going to have this huge shortage
of human capital that is trained
in AI and machine learning and deep learning
for a very long time.
I think that's going to be big bottleneck
for many decades, frankly.
And so I think it's important that you identify
for most business owners,
that you identify partners who have some of that human capital,
have a lot of expertise and experience to help guide you through that process.
Whether that's scale or whether that's another company,
it is important to identify these business partners who can help and accelerate journeys here.
Yeah, I love the human capital angle.
I mean, we're still short traditional software developers and engineers,
and we're into that cycle a long way.
So I can imagine this is that on steroids.
What if anything worries you about the proliferation of all of this technology,
and maybe even the technology that you yourselves are developing,
that allows faster development of AI models.
I'm of the view always that technology is sort of like a morally neutral.
It can be used for good or bad.
I like to think that it's mostly used for good.
It seems that way based on outcomes over time.
But what is concerning to you about the potential or the leverage
that this might give people that shouldn't have it?
One major one that actually keeps me up in night is this thing that we talked about
the very beginning, which is AI in the context of geopolitics and great
power competition. I think that there is an incredible potential for AI to either rapidly accelerate
or result in bad outcomes for some of the conflicts for the next, call it 50 years. And so I think
that that is one very pointed specific use case of the technology that does keep me up at night.
In addition to that, I think there's these two competing curves in AI. It's interesting to
see how they'll play out. But one of the things that I think generally technology,
concern about any sort of technology
accelerating is inequality in the world.
And I think that it's not totally clear with AI
what the exact effect is.
Intrinsically, the technology benefits from scale.
The bigger and bigger your models are,
the bigger and bigger data sets are
that results in you building better AI systems.
Generally, technologies that favor scale
do result in greater levels of inequality.
And I think that's something that we really need
to watch out for.
At the same time,
that's why the democratization of the technology matters a lot.
And that's also why the building it into use cases that affect a lot of people and enable a lot of people to live better lives is also really critical.
And so, again, there's sort of these two competing curves in the technology.
One is the benefits to scale and the other is lowering of fixed costs and the democratization of the technology.
And I think we would need to be very mindful about how these curves intersect over time or what the directions look like and the relative speeds.
What is the most interesting AI model you've ever seen?
I'll nerd out a little bit here.
I'm predisposition to like these scientific applications,
but both alpha code and alpha fold
are deeply interesting machine learning algorithms.
I think that alpha code is,
it's almost like brain breaking
in terms of thinking that you have these machine learning systems,
they can solve these competitive programming problems
better than the median competitor.
And the competitors who compete in these things
are already incredibly high up
on the distribution of humans
who can do some of this work.
Beating the median competitor is probably beating 99.9% of humans
at these programming tasks.
And that's just insane.
You know, if you had asked me even five years ago,
if I thought that was going to be possible,
I would have said, no, I don't think that AI systems
are going to be able to reason and think through these complex,
creative problems well enough.
But here we are, again, the availability of the digital data,
we have systems that can perform that well.
What's an example of that?
What would be something that Alf Code is producing
that a very talented human competitor is also producing
so that we can compare them?
There's these classic programming questions, almost like the interview questions that a software
engineer would be asked in an interview process. So classic brain teesery algorithm questions
that for decades we've used in Silicon Valley to test incoming engineers for their talent,
that's something that the machine learning system alpha code can perform just as well as some of the
best engineers out there. So another way to look at this, if your touring test is a set of
programming interviews to get a job at Google, you probably have a machine learning system now
that can pass that turning test. It's a pretty shocking thing. And I think that this is more
philosophical, but I think that myself as a programmer, I think a lot of programmers take pride in their
role as sort of the master of the machine, so to speak, or the person who tells the machine
what to do and programs the machine. It's this funny role reversal, but the machines are now better
at a lot of that work than humans. What are the implications of that? As you said, it's kind of hard to wrap
your mind around. Does that start to mean that we just continue to collapse the frictions between,
I'll call it, like, human intent and programmable outcomes? It's almost like our imaginations
become the limit rather than the elegant work. Yeah, that's exactly right. This is like a broad
trend within software. But I think to your point, the beautiful end state is that we as people,
we're just going to be able to dictate into a machine learning system, what kind of software we want it to
build, and it'll build something that basically accomplishes that exactly. So we can all be
product managers effectively in our own mini fiefdoms. That future state, I think, is like reasonably
far away, maybe not infinitely far away, but reasonably far away. But we're already seeing this
where take these GitHub co-pilot systems, they meaningfully accelerate the programming task.
There's a thing that's going to happen where programming and product management, in some sense,
we're going to keep converging where differential task or the differential skill is in understanding
what to build rather than actually being able to build it in the first place.
It's sort of like the classic idea in machine learning that it's the label.
It's the outcome that you're targeting that requires all the imagination and creativity,
not the features that you're using to predict that outcome.
Questions become more valuable than answers, I guess is the other way of saying what you said,
which is just is very cool.
It unlocks the core of, I guess, what makes us human, which is pretty exciting to imagine.
I hope I get to live to see a lot of that happen.
Alex has been so much fun.
I love this topic.
I am so interested in these technologies.
I think, you know, I ask the same traditional closing question of everybody.
What is the kindest thing that anyone's ever done for you?
It's not one specific action, but I think that my first violin teacher, my name was K and Unum,
I don't know if she'll ever listen to this, but maybe she will.
Over the course of many years of working with her, I think she really instilled a few things
to me.
I think first she really created a love and joy from art and creative activity.
and she herself, which was just so incredibly joyous and very inspired and clearly loved very deeply
the art and process of the creation of great music and great performance, which was just incredibly
infectious. And I think one learning is just that passion is an incredibly infectious thing.
There's a very specific moment that I remember really quite vividly, which is there's this
period where I wasn't practicing very much and I thought I could sort of like skate by or
she wouldn't notice. And obviously it's plainly obvious to her.
At one point, she just ended the lesson, you know, we were two minutes in, and she said,
if you're not practicing at all, we should stop wasting your parents' money and just stop on lessons.
It was like this shocking thing. I think I was in sixth or seventh grade at the time, and so I was
shocked to have something so direct, be confronted with something so direct. But I think what
it really showed me, if you're going to do something, you have to do it with full passion and full
force. And if you're not going to do something with full passion, full force, then it probably
just isn't worth doing in the first place. And this has been really a driving belief behind
a lot of my life. I wrote this blog post a while ago called Hire People Who Give a Shit.
I remember reading it, yeah. Yeah, it just becomes so core to, I think, how we even higher at
scale or a lot of what we do at scales, don't half ass. You got to do things with full force
and full passion. Absolutely love the story. One follow-up question. I remember the post well and it was
excellent. How do you do that? What is the most effective way to understand not having yet
worked with somebody if there's someone that gives a shit? One of the funny things about giving a shit,
you notice it along the edges, so to speak. You notice it in maybe the depth of thought,
the depth of obsession, the attention to detail along the edges of something they really care
about. I think one great example, let's take music, for example. When you're preparing a great
performance, you will start noticing just the smallest freaking things and you will just stop yourself
and keep practicing until you're like impeccable all the mini flourishes and all the mini components.
And the way that you notice great performance is not do they get the notes right, but it's like
how like impeccable and how clearly perfect is each little detail. That's one general thing is like,
I think you notice it along the edges more so than just looking full frontal at the thing itself.
You can notice it by talking people about like, what were the little.
details that you really paid attention to or what are the things that really bothered you about
how something was being done or what were the things that took 80% of the time that other people
wouldn't notice. That's one big thing. You'll notice it's just from like the passion in people's
voice and how they care and some people express that more outwardly. Some people are maybe more
reserved, but you'll notice in just the care and sort of the attention by which they express those
things. And then the other thing I think is from my time meeting lots of people and talking
lots of people and learning about humans in general.
There are people who are generally very inspired
and there are people who are generally less inspired.
And the people who are more inspired,
you'll notice this theme in their life.
They find things that inspire them
and they find the next thing that inspires them
and they find the next thing that inspires them
and inspires their full force and will.
And I think you can notice that pattern a lot
when people just talk through their lives
and what they've done in the past.
And sometimes it's a very wide set of things.
People can be inspired by punk music
when they're a kid and then they become inspired by DCFs and investing when they're an adult.
But I think that just noticing that strength of will in these instances is really critical.
I think it's such a wonderful, interesting place to close that giving a shit manifests in the details and at the edges.
I love that concept. I think it's so true if you ask yourself the things you're proud of producing,
it's very easy to answer that question. What were the details that took all the time?
and when you can't answer that question, it's probably indicative that you did a just okay job at it or didn't care that much.
I just love it as a closing thought.
Alex, this has been a total blast.
I've learned a ton.
Thank you so much for your time.
Yeah, thank you so much.
As I mentioned when I first chatted, I'm such a fan of the podcast, so I'm excited to be on here.
Pleasure is all mine.
This episode was brought to you by Canalyst.
In this three-part miniseries, I sit down with Brandon Weir, founder and portfolio manager of BWCP,
a fundamental oriented TMT consumer hedge fund manager based in Dallas, Texas.
Here, Brandon's biggest lessons from launching his fund, their unique blended investment strategy,
and how BWCP has integrated Canales into the investment process since day one.
So, Brandon, the best place to begin is with your personal origin story and background.
What were you doing prior to BWCP and then we'll start to talk about how that experience
led to the founding of the firm?
We can take it all the way back.
I got into investing in a strange way.
I'm from Kansas City. Middle class background. My parents made me have jobs. And the only one I could
stomach at the time was mowing yards. And I started a business when I was probably 14. It turned into a
passion for investing. And the reason that was, is I noticed early on when I was able to drive,
you should drive to the nicest areas you can find. They had the best profit margins on mowing.
And I got to know some of my clients that I was mowing yards for. And they were young guys that were
obviously doing very well. And I asked them a couple of times what they did. And they told me they
worked for American Century. And I didn't know what American Century was. I'd never heard of it.
They put me on to some books. And it was simple. This was 94, 95. It was Peter Lynch, one up on
Wall Street. They said, find things you're passionate about, what products people are passionate about
and do that and make it your own. And that's what we do. We figure out how the world works.
So I started investing in 1996. I opened these accounts. And what I saw,
saw then was phones. Phones were starting. It was the beginning of the cell phone era. And so the first
three stocks I bought, which really kind of ingrained a passion for investing, were things that were
involved with cell phones. And so I go to college and I'm basically wondering if I ever going to have a job.
These are 10 bagger type stocks at the time. And it was the late 90s. And I really wanted to figure out
how to turn this into a real career, not just the day trading stuff. And I was lucky enough in
college to have a great financial professor. My finance professor helped me get a job with Goldman Sachs.
So I went and spent a couple of years in New York doing investment banking. And I was really fortunate.
My first job came with a firm called Highside Capital. Highside, this is why BWCP's roots are in
Dallas now. Highside was a Tiger Cubs spinoff. I guess it spun out of Mavericks. It was a Tiger Grand
Cub. I spent eight years there doing technology, media, and telecom. Phenomenal firm.
We were a few billion dollars.
I got an opportunity.
I was actually kind of thinking of forming this exact firm back in 2012,
but Citadel came along.
And Citadel was expanding into Dallas under the name Surveyor Capital.
And they hired me, and I ran a portfolio for Citadel for six years.
And it was a wonderful thing.
And I tell you, what Citadel does so well is risk management and portfolio construction.
And what Tiger does so well is how do you find great ideas and uncover things that people aren't
really thinking about and use time as your advantage.
What I wanted to do, and this is what BWCP is, is kind of a bringing together of those two things.
How do you use great fundamental stock picking?
But understand that the world's different.
When I started in 2004, 70% of the dollars traded were done so every day by fundamental managers.
When you fast forward that to today, it's kind of the opposite.
It's the inverse.
And today, only 30 or 40% of the dollars are traded by fundamental managers.
And so you need to be aware of what's going on.
And so we use some risk tools, some of the factor awareness, some of the things that were from Citadel,
but really just doing it around that framework of the Tiger Cup style of stock picking.
I love the fusing of those two disciplines of the rigorous portfolio construction quantitative
that Citadel obviously is famous for with the deeper fundamental insight on businesses,
which is a great bridge into the key things that you found setting up a new.
investment firm. There's a lot you need to do. It's a lot of effort. I've seen it done. I've done it
myself. I know that it's difficult. And I think companies and research is one aspect that gets
talked about less, like most of the focus on legal and systems and team and all this kind of stuff.
How was Candlest a part of that early setup? What did the company models allow for relative to your
prior experience that let BWCP get up and going faster and sort of how were they used in the early
days. Let's paint a picture. You're running a two plus billion dollar hedge fund inside of Citadel. You have a team of
six. You don't even know what back office means because every single thing is taken care of. And you have a lot of
resources. And you have a lot of young analysts that need to spend a lot of time doing the Citadel way or
whatever that is. Roll that forward. I have a couple analysts myself, small back office, friends and family
capital and my money, basically, our internal capital. And,
we've got to figure out how do we divvy up resources? Because right now, there's two things that
are our most limited thing, time and money. And we have to find across the board, where can we use
our analysts' time to do the things that are most impactful to the research process and where
can we outsource with great partners to leverage both that time and money? And that's where
Candlest came in. And we used some of our friends. We asked them about best practices, how they utilize
their limited resources up front. And Camelists came up time and time again. So what Canales does for us,
think about the amount of time that it takes to set up a model to build a proper model, to update that
model each quarter. We probably, with the number of stocks that we look at, we're saving one and a half
to two analysts a year. It sounds like a huge number, but that's probably only.
looking at, you know, maybe 100 to 200 companies, but to do that right, that's probably what
it would take from a manpower point of view. Now, if you think about the research process,
what is it about the model that's so important? Models are the output of a lot of inputs.
And the value add that we need to spend our time with our team on is what are the inputs?
What do the numbers mean? It's not necessarily the building blocks of the nuts and bolts of the
model, but it's what it really comes down to is what are your assumptions, how good are your
assumptions, and what do your assumptions mean? So whether or not you're a super long-term investor
and you're going to make a decision once every 10 years, you still need to know then what are the
cash flows over those 10 years? If you're a super short-term investor, where you want to know
everything about every margin, about every top line item, about every revenue trigger, you obviously
need to make those assumptions somewhere and put them in a model. So it's a basic building block.
and it's the foundation of everything that we do,
but it's not something that we want to spend all of our time doing the things that other people can do,
so we can then turn around and use our time more valuably.
And that has been a huge, huge help in how we allocate our resources.
How do you think about the uniqueness of company models that Kandalous provides sector to sector?
So you mentioned TMT being a focus.
I think GAAP accounting is great,
but it also isn't perfect as a way to interpret or measure or investigate
how a specific company works.
Say a little bit about like company-specific metrics,
KPIs, you know, things that matter for one business that might not matter for another
and how that factors into modeling.
So that's the other important part.
I've used services in the past that give you sort of a starting place.
There's two problems with those usually.
One, they don't have hardly anything above sales.
And to us, usually in the TMT space, in particular,
the top line is one of the more important variables.
margins have less controversy, but are obviously important, and then which brings us to free cash flow.
So when you're using a generic service, sometimes you can get the basic building blocks of what's
reported in maybe a queue with no context above for what's maybe reported on conference calls.
Canlist adds that extra layer, which is important.
But the most important part of it is the models live with us.
So everything that we change or that we adapt to the things that we need, even if it's not already
preloaded, we can add that.
And the model just simply updates.
I used to have to send models away.
And when they come back, they're stripped.
Anything that I've added, any sales that I've done, anything has been a problem.
Now, the models sit here with us.
And the reason that's important is because those models lead to our investment templates,
our investment spreadsheets.
Those investment spreadsheets are the things in which we pitch our stocks off of in our interim meetings, which then flow into our rankings.
Our rankings are what create discussions about what in the world's going on.
So you put a number into a model with Camelist, and it stays on my system and goes all the way through to price target changes.
I do it somewhere else, and it has to go send off.
The links break, it comes back, and all the work that we've done updating is gone.
So how do we focus our attention on what is the most valuable stuff that we can do with the least
amount of friction? And that's the important part. When you're using the catalyst system,
there's no friction. And it allows us to maintain our process and almost like we just added a team
member. Where do you see the most opportunity today in the coverage universe that you care about?
And what's unique about how those businesses look in terms of their financial statements?
and the levers that you think are important to underwrite and understand as an investor.
I'll take that a slightly different way, but I'll get back to the original question.
The thing that's most interesting to us, and this is kind of reshaped versus when we started,
again, we started really small and I've been fortunate up to grow to a few hundred million,
but what has changed in the last 12 to 18 months is the number of companies that have come
public, whether that's IPOs, direct listings, all of these types of things are a big tax on
the cell side, and they don't have enough time and energy to cover these things.
And so if you think about, if your goal is to find something that is less well-known,
undiscovered, something that has massive amounts of innovation, and your variant perspective
could be simply knowing more about that company and actually being able to model it on a timely
basis, that's where you have to have these types of partners.
I would say the world from two years ago has changed dramatically in the sense that there's
so many more companies. The sell side has not gotten bigger. If anything, it's getting smaller.
The resources that they have to be able to help managers is not good, but it has gone down somewhat.
So what I think is going to happen, if you think about active management and you think about
the number of ideas, we'll probably shift more of our capital into small and midcaps over time
where we have a balance between small, medium, large. I didn't think that was something that was going
happen. Two, three years ago, people would say, you know, the number of companies are just
disappearing. What is the role of active management? You can buy a few indexes. I think given the
valuation, given that the stock market just continues to grind higher, some of the best places to go
out and look are some of these more undiscovered things that are newer, both on the long and the
short side. There's a lot of stuff that came to the market that has no real meaning and no real
value. So that, I would say, has been a real big change in the number of the number of the number of
of types of companies that we're going to have to use for financial modeling. And I will tell you,
when you're looking at things that have limited history, that's where a place like Countess can
come in. They'll be able to scrape that, go back, find historical stuff. And really, what we're
trying to do to sift through all of these, we have to get some baseline level of knowledge.
And you can't do that building out your own models. It would do. It would simply be prohibitive.
In closing, thinking about other emerging managers, new managers that are hanging up their own shingle and setting up shop, what other key lessons did you learn about that process relative to your expectations? What was harder? What was easier? What did you have to do that you didn't even think about? There's a lot of lessons. Look, I will say a few things. One, starting a fund can be a very humbling experience. I had a manager and an allocator tell me early on, if I could tell all new managers one thing, it would be that the world doesn't need a
another hedge fund. What makes you different and why do you need to exist? And that's a tough fill
to swallow when you're when you're out on the road marketing. I think people really wanted Citadel.
So when we left, they love Citadel and they wanted us to replicate Citadel. I will tell you
some people have been able to do it, but replicating Citadel is a very tall thing. So the advice I would
give is find what you want to do. And what we wanted to do was slightly different. And if the thing that
what you want to do is slightly different, in our case, it was a,
bringing together of the Tiger Cove and the Highside model with the Citadel model,
be prepared to take a few steps on your own before somebody will underwrite you.
And I don't think that it's a negative thing.
I think we all get convinced that one or two people, you know,
raising $5 billion day one means that the fundraising environment is very good.
And I think we all have survivorship bias towards all of the people we look up to,
whether it be Ken Griffin or Steve Mandel or all these people who have come before us,
those funds started really, really small.
And they proved it.
And they worked themselves out for a few years.
And I think that's probably the biggest lesson.
The other one is getting off to a great start.
You can never underestimate the power that that has.
And that was something that, you know, we were very fortunate to be able to do.
And then the last thing I would say, and I've had a lot of mentors over the past.
And they've always said, well, if you want to do something,
you're going to have to find something that you love to do.
You can't do something that you don't love to do.
And I used to think that that was just something that people said that were super successful
and that was a nice tagline.
But investing is a hard business.
And what those allocators mean when they don't think everybody needs to be a hedge fund manager,
they mean that.
And so if you don't, investing can have its ups and downs.
And if you don't love what you do on a day-to-day basis and love building something,
you know, from the ground up, you're really going to have trouble.
And I think don't underestimate the amount of time it's going to take you to build it.
Don't underestimate the amount of money that it's going to cost to build and find the most talented people you can possibly find.
I think sometimes new managers fall into a trap of not wanting to hire people they think are a lot smarter than they are.
I tell everybody that we hire.
I hope that I'm the dumbest person in the room.
I want the smartest people around because that's how we'll become more successful.
Well, Brenna, I love your story. Love the fusing of styles and also love the fact that the leverage point for Canalyst existed in the early days that it let you do more faster. And I think you've made great points about sort of why repetitive processes that are nonetheless important and detail oriented matter a lot. But the one and a half to two analysts sort of stuck per couple hundred companies per year. That's a real thing. So I think we've done a great job of overviewing your story and also the ways in which it intersects with the
Candlest product. Really appreciate your time and wish you the best of luck in your brand new firm.
Thank you, Patrick. Appreciate it.
If you enjoy this episode, check out join colossus.com. There you'll find every episode of this
podcast complete with transcripts, show notes, and resources to keep learning. You can also sign
out for our newsletter, Colossus Weekly, where we condense episodes to the big ideas, quotations,
and more, as well as share the best content we find on the internet every week.
