Invest Like the Best with Patrick O'Shaughnessy - Gaurav Misra & Dwight Churchill - Building Captions - [Invest Like the Best, EP.405]
Episode Date: January 7, 2025My guests today are Dwight Churchill and Gaurav Misra, co-founders of Captions, which uses AI to generate and edit talking videos and has grown to significant scale at remarkable speed. We explore a k...ey distinction in AI: tackling bounded problems like video generation versus unbounded problems like general intelligence and what this means for building sustainable businesses. We also explore their unique data flywheel, why video generation could reach Hollywood quality within 18 months, and why building advanced AI products doesn't require huge teams. Please enjoy this discussion with Dwight and Gaurav. For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- This episode is brought to you by Ramp. Ramp’s mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Ramp is the fastest-growing FinTech company in history, and it’s backed by more of my favorite past guests (at least 16 of them!) than probably any other company I’m aware of. Go to Ramp.com/invest to sign up for free and get a $250 welcome bonus. – This episode is brought to you by AlphaSense. AlphaSense has completely transformed the research process with cutting-edge AI technology and a vast collection of top-tier, reliable business content. Imagine completing your research five to ten times faster with search that delivers the most relevant results, helping you make high-conviction decisions with confidence. Invest Like the Best listeners can get a free trial now at Alpha-Sense.com/Invest and experience firsthand how AlphaSense and Tegus help you make smarter decisions faster. – This episode is brought to you by Ridgeline. Ridgeline has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. I think this platform will become the standard for investment managers, and if you run an investing firm, I highly recommend you find time to speak with them. Head to ridgelineapps.com to learn more about the platform. ----- Invest Like the Best is a property of Colossus, LLC. For more episodes of Invest Like the Best, visit joincolossus.com/episodes. Follow us on Twitter: @patrick_oshag | @JoinColossus Editing and post-production work for this episode was provided by The Podcast Consultant (https://thepodcastconsultant.com). Show Notes: (00:00:00) Welcome to Invest Like the Best (00:07:49) The Evolution and Impact of AI (00:09:14) Challenges in Video Data and AI (00:10:36) AI in Media Generation (00:12:07) Building a Sustainable AI Business (00:14:56) The Journey of a Video AI Company (00:25:41) AI Video Editing and Creation Tools (00:29:58) Future of AI in Video and Business (00:37:51) The Future of Likeness in Video (00:39:25) Training Models on Human Data (00:41:15) Competitive Landscape and Copycats (00:44:01) The Role of Research Talent (00:46:25) Pricing AI Software (00:51:51) Investor Perspectives on AI (01:02:44) Lessons from Snap (01:07:04) The Kindest Thing Anyone Has Done for Dwight & Gaurav
Transcript
Discussion (0)
Most software companies try to maximize your time on their app to juice engagement.
Ramp does the exact opposite.
Ramp understands that no one wants to spend hours chasing receipts, reviewing expense reports,
and checking for policy violations.
So they built their tools to give that time back,
using AI to automate 85% of expense reviews with 99% accuracy.
And since Ramp saves companies 5%, it's no wonder that Shopify runs on Ramp,
Stripe runs on Ramp, and my business does too.
To see what happens when you eliminate the busy work, check out Ramp.com slash Invercored.
Hello and welcome, everyone. I'm Patrick O'Shaughnessy and this is Invest Like the Best. This show is an
open-ended exploration of markets, ideas, stories, and strategies that will help you better invest
both your time and your money. Invest like the best is part of the Colossus family of podcasts,
and you can access all our podcasts, including edited transcripts, show notes, and other resources
to keep learning at join colossus.com.
Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions,
expressed by Patrick and podcast guests are solely their own opinions and do not reflect the
opinion of positive sum. This podcast is for informational purposes only and should not be relied
upon as a basis for investment decisions. Clients of positive sum may maintain positions in the
securities discussed in this podcast. To learn more, visit psum.vc. My guests today are Dwight
Churchill and Garoff Misra, co-founders of captions, which uses AI to generate and edit talking
videos and has grown to significant scale at remarkable speed. We explore a key distinction in AI,
tackling bounded problems like video generation versus unbounded problems like general intelligence,
and what this means for building sustainable businesses. We also explore their unique data flywheel,
why video generation could reach Hollywood quality within 18 months, and why building advanced
AI products doesn't require huge teams. Please enjoy this great discussion with Dwight and Garoff.
And a key side note, the first person you'll hear is Garoff. So guys, the topic on everyone's mind,
I think, is this shift from AI as this incredible technology. That's decided. Everyone understands
how amazing this is to, okay, great, what are we going to do with it? And how can we build
enduring generational businesses with this technology at the core? You were very early in,
building a business that charged Customers very early on using this technology. Maybe you can begin
by just riffing on the lessons that you've learned so far about building an AI business that are
maybe distinctive from a normal software business or something. And also get into some of the
open questions that you yourselves have trying to evolve your business model. I just think this is
becoming the important question in the marketplace right now. And you're one of the earliest
adopters. So you're the perfect people to answer. Getting into it, like I think the first
first question behind the question that comes to mind is like, what exactly did we actually
achieve with this AI revolution? Like, what is actually the difference? Like, AI existed before
and it exists today. Obviously, there's something magical about what is there today. I think when
you get into it, you realize that it's really about the ability to train larger and larger models.
Yes, that's actually a combination of we have better hardware to do it. We have better ML
architectures. Like there's transformers, there's diffusion models. There's all these new types of
architectural unlocks that we've created. And then there's other techniques that we've created, too,
which just allow us to train larger and larger models. And turns out the larger and larger
you make these models, the more problems they can solve, the better they can be a solving
like text generation or like towards AGI or video generation or media generation in general.
I think when you realize that what actually you get to is that what really matters is the data
at the end of the day. A lot of companies are like scraping the internet and the internet is also
limited in some ways. There's only so much information on the internet even and that's growing every
day. But I think at the end of the day, beyond that, we're going to have to find what are those
sustainable sources of data that can continue to grow bigger and bigger models? And I think that's
going to be the fundamental question behind who actually ends up winning in a lot of these different
areas that AI is excelling in today. I think for us being on the video,
generation, video editing side, it comes down to like video data, which is actually much heavier,
much rarer to find, not as common as text or even audio, and potentially much more expensive
to train on as well, much more limited in terms of being created in the world. And so that tends to be
like a big challenge. One of the big things that we're thinking about is how do we actually
create a flywheel where we can ingest data on a continuous basis and a growing basis? And that data can
actually create bigger and bigger models for us from keep us at the forefront. I also want to call
out here, like there's a pretty fundamental difference between different types of AI companies that
are out there. I think if you look at a lot of the text generation companies, they're not solving
text generation. Like, we don't call it text generation. They're actually kind of solving a totally
different problem, which is intelligence. Intelligence is an unsolved problem. No one's figured that out
yet. And yes, we're achieving some levels of intelligence in these models. And there's a long way to go.
it may not end at human intelligence.
There's people in the world who are really smart.
There's people in the world who are not so smart.
They both exist.
And clearly, there's a range of intelligence as possible.
There's not one value for like you're intelligent or not.
So, yeah, is there a chance that there's the ability to go smarter than the smartest human?
It's possible.
But that's a frontier that we've never reached.
And so it's kind of solving this unsolved problem.
But I think if you think about audio generation or video generation or music generation
or these types of things, right, is I think a little bit less.
of solving an unbounded intelligence problem
and a little bit more of solving,
actually rendering a solved problem.
And video, for example, like, CGI exists.
We can make fake things.
We can make fake humans.
We can make fake sceneries and dragons.
And so this is a solved problem.
We know that there's solutions to these.
And with AI, we're actually just making it easier
to solve these problems,
not just a little bit, but like 100 times easier,
which in the end that means more accessible, larger market, more people can use these sets of
technologies. So I think that's one of the fundamental differences there is if you look at business
models for like artificial intelligence companies that are really working on AGI, then you kind of
have to think about this unbounded problem with like, okay, we put in a bunch of capital into it,
we create a model, only for that model to be beat by the next model and that model becoming
essentially useless and obsolete. And then there's the next model after that. And how long does this
go on for? Actually, we don't know. It may go on forever.
there may be like no end to this intelligence race. Whereas if you look at the media generation
companies, it actually is creating an asset. And there might be very soon a point where, oh, wow,
it's just really good. It's just perfect or close to perfect. And we've kind of solved it. And then
it's an asset. And then after that, it's just a software company. And the asset's really expensive
to create. But once it exists, it just generates value. And it doesn't lose value that easily.
So what is going to make those models better and better? I think it's going to be like,
tuning with more data, fine tuning for specific use cases, different types of things you want to
generate, different types of visuals, whatever it might be. Use cases like, oh, it's going to be
used in ads or movies or social media or something else. But there may be a point where it's like,
wow, yeah, this is pretty good. It's realistic. I think that's a pretty important thing we're
thinking about right now. How do we bootstrap that data flywheel to be able to reach that level?
What does it like to work with video data where I imagine like just the terabytes or petabytes
or however you measure it of data that you have is sort of insane?
How do you think about something that might just get as good as it can get?
I love the point that if you give a Hollywood studio or WADA or something enough money,
they can literally create any visual that you can imagine.
The friction between imagination and output is already gone.
It's just really, really expensive.
So really what you're doing is just making something cheaper.
when do you think that could be achieved?
I think it's pretty soon, honestly.
I mean, at the rate of which video models are going,
I mean, you probably remember seeing like the Will Smith spaghetti thing.
Everyone's seen this meme.
Right.
And it went from like really horrible to like, wow, this is actually good.
And I think really, really good is probably around a year, year and a half away.
I only say this because if you compare, for example, like text models to like video models,
text models are already like in the 400 billion parameter range.
People understand better how to scale LLM technology today,
just because more money's been put into it,
more times been put into it.
Like diffusion models,
still in the tens of billions.
It's still early,
not even close to the text models.
So as that grows,
there's just no doubt it's going to get better and better.
And like,
the experts kind of know that this is all possible.
It's just that very few companies in the world have the funding
and the expertise to actually go after this.
So like it just takes some time.
Like, it's not like some unsolved problem.
People know what needs to be done is just we're all getting there.
We're all moving towards it.
And we'll see those models getting better and better, especially on the video side.
I could easily see within a year and a half or so something coming pretty close to like indistinguishable,
essentially from like a real recording, maybe even sooner.
That's not like the worst case.
Yeah, I don't think people are entirely able to grasp that yet.
I think the way that that influences how they do,
their work every day, the workflows that end up being reinvented, new paradigms of all that,
which is arguably part design problem, part just product problem in general. This is pretty
close the timelines that Garav was talking about. People are experimenting today. It's extremely
early. And I think that companies' adoptions and stuff around that, we're not far off at all of
really reinventing a lot of how people end up doing their everyday work. Can you describe the stages
that you've gone through as a company, maybe we'll use like the Tesla analogy. One of the beautiful
things about their model is the cars, by virtue being driven, are gathering data all the time.
The product itself naturally generates data exhaust. And I think you've had a somewhat similar
story on the video side. And so you don't need YouTube or some massive proprietary library
of video to do what you're doing. Can you just describe, take us back to the day one of the business,
what it was to start, why you started there, and then how it's progressed since.
It has been a pretty interesting journey and we've been through some interesting twist and turns through it.
But I think if you like connect the dots end to end, it's interesting.
When we started the company, the first app that we made was captions.
We launched it.
And why did we make it?
The goal was to get content creators to create content on a video creation platform of some sort.
Not easy.
I was a snap before this and Snap had tried this many times.
They launched apps and I mean, video is kind of a commodity.
Video editors are commodities.
a lot of these companies are actually foreign,
and that's because we're just trying to minimize cost at this point,
and really difficult to compete in.
Our thought was the way we're going to crack this
is we're going to use AI to help create video somehow.
That's going to be our differentiator.
That's why people are going to come to us.
And so we saw that there was a need around speech to text.
It was a technology.
By the way, at that point, that was pretty good.
In text circles, people were like,
of course, speech to text, we understand that.
It's pretty good at this point.
But I think the average person actually didn't understand.
understand how good the tech had gotten and how accurate it was with names and like obscure terminology
and all kinds of stuff. So when we built the first product where it was just like, hey, it's just
literally put text on the videos. And by the way, this was built in like two days on a weekend,
really just band-aid together. And we put it on the app store, went to sleep. The next morning,
it was top the app store. There's no explanation. We didn't do anything to make that happen.
Somebody saw it. They posted it on something. It blew up and then woke up and I text.
Dwight and I'm like, hey, I think there's like 600 videos per minute being created on the app,
by the way. And so that was kind of like an instant success. But even in that two days of work,
we had already instrumented the app in such a way that we would be able to continue training
better and better models so that we can deliver better value to the user. So the idea was like,
the app is an AI app where people come in. They use the app. We use the data to make
the bottle better and deliver even better experiences the next time the person comes back.
That was done from day one, literally.
That was the original plan.
Now, post the launch of the app, we've added so many more features over time,
expanded the offering so much more.
And we cover now the entire space of everything from like script writing to recording to
video editing, distribution as well, and how AI can like transform each of these different
areas because there's applications and all of them.
And there's data that can be collected across all of those that can improve those models.
And that's what makes our offering really unique because all the other companies are not really
thinking about the data collection side and just generating outputs.
And that's why they have to kind of scrape the internet to make their models better.
And for us, really, it's more about growing a user base so that the data can actually power
better and better models.
And a lot of that comes through like video.
So video being funnels directly into video generation models, that gives a significant advantage.
That's potentially a possible way in which a future sort of business model could be set up.
And actually is kind of familiar, by the way.
Like it seems to me similar to the Facebook or Google business model where you have a mass consumer free product, basically.
And the data is used to power essentially like a B2B pay product.
If you think about the literal process of training, maybe you can explain it to people that are curious about.
like how this actually works. So you have raw video. A lot of it has voice in it. You can start it
obviously by translating that voice into text. But let's say you're trying to train a model.
I like how you guys referred to many people like Asora focusing on what we'll call Broll background
video or just like a landscape of video. And your focus has been on A role like a human being
on an iPhone looking video talking. How do you train a model where the output is A role like that?
like just imagine a portrait video of someone reading an ad read or something like that that's indistinguishable from an actual live video taken on an iPhone.
What is the literal training process?
What is the target of the model as it's training?
How similar or different is this to just next token prediction?
What's the mental model for next X prediction or something in a video?
Like how do you think about the literal actual training process of what's happening?
It's interesting to think about because for the models that we train, they're diffusion models.
So they actually work by starting from noise.
It starts from literal noise like static you see on TV.
At every step, based on text that's provided, it looks at the noise and it tries to like
predict a layer of clarity in that noise.
It says man wearing blue shirt.
So it starts to like draw a little bit of man wearing blue shirt out of noise.
And then every pass is taking through it, it's discovering a little bit more of the
man wearing blue shirt.
So that's the text conditioning that's helping it decide how to reach the destination of what men wearing blue shirt looks like.
So that's how the diffusion models work, which is slightly different from like how a next token prediction model like GPT works, which is kind of just as you might think about it, just predicting the next word, based on all the previous words that I've been spoken, which are considered the context.
So these models are different.
We are still earlier on in the diffusion model training path.
We're still in that 10 billion, 20 billion, 30 billion.
Meta's movie gen was, I believe, 30 billion parameters.
People haven't really scaled up.
We actually don't know how big open AISora is.
They didn't, I think, release that information.
But a lot of the work is going to go into scaling up these things.
Video obviously is really heavy.
That's what makes it different from text.
It consumes a ton of space, a ton of processing.
For us, even if we were to download, just download all of our training videos.
it will cost a million dollars to download the training videos.
That's a whole different regime than like text.
It brings different types of challenges to training these models, basically.
What does that mean in terms of the sap on resources that video models will represent relative to text models?
Like one of the big discussions in public markets and private markets is how big do the GPU farms need to get?
Are these video models to get to that point of perfection necessarily more consumptive of GPUs than text models would be?
what's your two cents on this big question of do we need to build nukes next to data centers
to train the perfect Lord of the Rings model or something?
You never know.
But honestly, I think what will save us on the video model side is actually the fact that
it is an easier problem than the text problem.
The text problem is intelligence, as we're talking about.
And the video problem is more rendering.
We already know how much rendering costs.
We already know, yeah, it's GPU intensive.
If you were to like literally CGI render or seen out, like, yeah, it will spend some time
with the GPU.
There's no doubt.
Can we be more efficient than that is possible?
It may not be the most efficient today.
Maybe there's better ways of doing it.
Maybe AI will be cheaper and faster than regular rendering.
And I think if that's the case, then that's a good thing.
But I think we know that it shouldn't be worse than that.
We should be able to solve it with fewer resources than that,
potentially or at least the same.
We generally understand where it's going to fall.
It's still early.
Just like on the training side, we're still scaling up these models,
and it's still, oh, it's 10 billion parameters, 20 billion parameters, whatever.
on the inference side, it's similar learnings happening simultaneously.
We don't need to do 100 steps of diffusion for inference, like 100 denoising steps to reach
like a clear picture.
We can distill models and have them work with a few steps of diffusion now.
I think we're definitely the most inefficient will ever be, and it's only going to get
more and more efficient.
It could be a factor of at least an order of magnitude like 10x or something.
Can you talk about the felt experience of having a business?
We won't quote how big the business is, but it's very big.
and it's grown ridiculously fast.
One of the things you hear that's a common idea now
is that a new technology like this unlocks distribution.
Distribution used to be really expensive,
late in the mature last SaaS cycle or something.
People sort of had their tools.
But when tools are just 10x better or more,
100x better, distribution for a time becomes really easy.
I think you've been beneficiaries of that unlocking of distribution.
Just talk about what that is like.
What does it like to see revenue and users
and all this stuff scale at this pace
because it seems like the revenue growth rates
of some of these AI application companies
are faster than anything we've ever seen.
And I would just love you to riff on that a little bit.
And I want to describe what it was like,
but also just reflect on anything that it teaches us.
I mean, it's definitely the most exciting
that for anybody who's working on the engineering or product side,
I feel like there's nothing more exciting than seeing direct results
if I did a thing and the next day it caused an impact, right?
Like, people cared.
There's just nothing more exciting than that. And I think we see that, which is great, which is why we've been able to build a great team and hire all this great talent, really set us up for success. But I think maybe the most interesting part of it for me is how you can almost see how as you're expanding the use case, it's actually growing the potential market. And that potential market has no competitors. As you expand the use case, you kind of see we're doing ads now or we're doing like higher quality.
video, even on that axis, and you see like entire new areas of market unlock where there's
actually no competition. And actually that's what causes the fast growth. It's actually nothing other
than just we are the only company that can do something for a period of time. And that will change.
And I think that's why it's going to be interesting to see as more and more use cases unlock,
at some point all of it's going to be unlocked, all if it's going to be having competition in it,
that would be a different time. It might be years from now. I don't know when it will be. But
At least for now, what we're seeing is this ability to expand use case. And by the way,
we really think that the use case unlocked so far is somewhere in the one to five percent range.
We've barely scratched the surface of what's possible. As that grows, we see these entire new
markets unlocks. Like, wow, this is a whole new set of people who can now do something actually
useful with this. Yes, they're completely willing to pay. They're running to us. We don't even need
to sell it. And we're the only option. It makes just growth really fast. I think that's been
probably the most exciting thing for me.
Can you level set what the platform can do today,
the major use cases,
like everyone can imagine feeding it a video,
getting a captioned video back?
That's very simple.
Can you lay out the other ones
and give us a sense for their relative popularity?
What is the revealed preference
of how people want to use a platform like captions?
When you think about us,
we actually divide the product in two areas.
So there's the traditional video editing
and video recording,
which is just as you would expect it.
It's a video editor and a video
recording software. And this is built for consumers, completely free. The play here for us is to
provide a service to a large number of people, kind of a freemium business model in a way that
they're already familiar with creating. But the goal is to actually upsell them into the AI use
cases. You actually don't need to spend all this time video editing and recording. You can just generate it.
So on the flip side of that, we offer the AI suite, which is two products, AI Creator and AI Edit.
These are exactly mirroring recording and editing.
AI Creator literally just makes videos of people talking, whether that's you or an actor that we provided
or anybody you might choose that you have the license to use.
We can make them say whatever you want, deliver whatever message you want, and we can even
create people that don't exist.
So in between that, you get a bunch of optionality of how you want.
it has to be delivered. A lot of the use cases like marketing and sales and these things are
like very close to revenue. And then there's AI Edit. Just a recorded video isn't exactly enough
to create value. You want it to be edited in some way to tell a story that you want to tell.
That's why we have AI Edit. The purpose of that is take a video in and edit it for you.
You actually don't have to worry about keyframes and animation curves and timelines and all these
concepts. Video editing is not easy. And a lot of people avoid it because they just don't want to deal
with this complexity. And our thesis is, we have a foundation model that just does the editing
for you. So you don't have to worry about actually editing anything. So that's the suite of products,
basically. The traditional versus the AI, and in the AI, we have AI creator and AI edit.
Just to clarify, like in something like edit, am I prompting it to say I want it to do this specific
thing? And then is it sort of like prompting? That's where it will go in the future. Currently,
it's more style preferences that you provide to it. So it's in the early.
early days of that. A lot of what will happen in the future is as we get more video editing data
from our traditional products, we're going to use that to funnel into our foundation model the
ability to essentially prompt with text whatever you want to say. Say things like, I don't like
these images. We want like different images with a better vibe or let's cut it down to like 30 seconds,
like 45 is too long or it sounds a little slow. We want to tell the story a little faster pace,
general prompts, what you might actually say to an actual video editor.
And the type of thing that someone who doesn't have the detail and intricate knowledge of video editing might say.
So what is like the relative breakdown of what tools people use?
The free versus the paid and the editor versus creator, like how does it shake out?
Yeah.
So today, like a vast majority of our users are paid users.
In between AI creator and AI Edit, they're both about equally popular.
There's some people that just use AI Edit.
There's some people that just use AI Creator depending on the use case.
And then there's a bunch of people who use both one after the other.
So using both lets you basically get from absolutely nothing to a fully edited video with just a couple of words typed,
which is a great first-hand experience.
Now, some people might want to just record their own video or they might be editing on somebody else's behalf or something like that.
So they might actually take a real video and pass it to AI edit to be like, okay, I already have a video, edit this for me.
And on the AI creator side, some people don't want the editing or they have very specific use case of what they're trying to do with it.
They want to just figure it on their own.
So they just do the AI creator part.
A lot of times it's like using their own likeness so they can just mass produce videos of
different types.
But it also can be like using one of our actors.
A lot of that is marketing content and things like that, things that go on social media,
but also like ads and anything that might be marketing related.
So those are sort of the relative popularity.
I would say like they're about equal.
Does it feel like you're in an arms race right now with other companies?
To an extent, yeah. I mean, I think the most interesting thing that I've seen is a lot of new companies popping up. All of them are trying to do the same thing. Like, I'll give you an example. I was a snap before this. Literally five other people have left Snap and tried to start the exact same company. Yeah, it's worrying. We should be doing that thing. It makes sense. I don't blame anybody. Like, I think it's great that they're doing it. But I think what I like about it kind of in a way the most, people are copying us. I think it's like a great sign. It means that we're doing the right things. And we kind of
avoid looking at other companies too much. Our product strategy and what we build and what we do
is really decided by our mission and vision and where we see the future being. It shouldn't be
decided by what somebody else is doing because they may not have a strategy at all. We don't know.
Their strategy might just be looking at us. So a lot of hands, we'll look at competitors only to the
extent of understanding, okay, this is what they're doing. What we really focus on is thinking about
our North Star and where do we see the future being and are we building towards that future,
not just from a technology perspective,
from a product perspective,
and a user experience perspective.
And I think that's the fun part.
I think that's so much fun.
When do we get a chance in history
to actually invent the entire stack
from the bottom to the top
all the way from the hardware level?
Like, there's bugs in the Nvidia drivers.
There's bugs in the hardware level.
Like, it's crazy.
And we get a chance to literally invent the UX.
How are people going to interact with these things?
Like, I think people are not even thinking enough about this yet.
They're just literally taking models and throwing it on UI.
I mean, like, press button, output.
What if it was more interesting?
What if you could see the steps of diffusion or you could like preview things in the middle of the diffusion process, change things according to like what you want it to generate? There's just so much that's still to be unlocked. Every function, whether there's design learning about like how the technology works or technology people learning about like how the marketing is going to work. This is going to get so much more evolved and so much more integrated and that's what we focus on. I think the arms race is ensuring that we're delivering always way in front of what our customer even needs today. Whenever we
releasing something, it gets commercialized on day zero immediately. We're not like testing it with a
bunch of people and seeing what they need and seeing if we're actually solving anything. No, no, no.
We're building this for their work. We're incredibly ingrained in how they do their work,
whether you're a large enterprise or all the way down to the free consumer. Ultimately, to Garo's
point, by inventing those design patterns and the way someone can interact with these new models,
we're literally paving the way for how people even think about doing their work.
And that's the really exciting stuff.
That is the arms race in my mind.
But that's not necessarily against another company.
What are the tradeoffs that you've had to choose one way or another as you build?
Video is a big category.
That could mean I get to make a Lord of the Rings quality movie or it could mean something
much more provincial than that.
We've actually niched down quite a bit on purpose because, as you said, like video is huge.
It's like a massive market and it's almost too many poems to solve.
I don't think if we tried to focus on everything, we would solve all of these things.
So our focus is very much on videos oriented or on communication.
These are talking videos, people saying stuff.
A lot of it tends to be marketing, sales, education.
These are like the big categories or maybe communications to some extent.
And it's about generating those types of videos.
It's about editing those types of videos.
But I think generating stock video is fine.
I think that's a great thing to solve.
But our goal isn't to create stock video.
It's actually to create a role video, telling the actual story of whatever it is you're
trying to convey.
So not just bunnies jumping around on Mars type of thing, more like telling a story,
pitching a product, or whatever that might be, something really communicative, informative.
And that's where we've seen a lot of our product market fit.
We're actually the only company training a foundation model to do this type of thing today,
to generate a role. There's a couple of technical reasons why that's the case. There's other
companies in the space, but they're not trying foundation models. So we'll see how the space evolves
in the future. I think it will actually tend more towards what we're doing. What are the surprising
hard limitations of what the models can do today or might be able to do in a year? Like,
I'm imagining or sitting at this table. There's a bunch of stuff on the table. My specific brand of
water bottle or something, I want to tell the thing to be able to hold it like this certain way.
and like I want to be able to sort of direct an object that's not the person,
but that interacts with the person.
It's something like that relatively straightforward.
Yeah, I think that will happen within six months, guaranteed, essentially.
We'll probably start seeing the first version that's coming out within months of now.
How does that work?
Are you creating like a 3D representation of this thing somehow?
What are the steps that go into the ability to create something like that?
You have to find training videos where people are already interacting with objects,
you drinking a can of coke or whatever it might be. And then you have to be able to identify those
objects and then provide them as conditioning. So for example, it might be text conditioning. So if you can
adequately describe this particular can of coke in text, that might be enough. But it may also
not be right. Like Fiji water bottle has a very particular design unless the model has seen one before,
it may not be able to precisely recreate it. And text might not be enough to describe what it looks like.
So you might imagine like image conditioning.
Here's a picture of a Fiji water bottle.
And then text that says,
Man in blue shirt holding Fiji water bottle.
And then it'll be able to figure out the rest from there.
Because it's seen bottles in general.
I understand what bottles look like.
If it sees it from one angle, it can predict what it looks like from the other.
So if you're like rotating it around and moving it around,
it'll guess essentially what it probably looks like on the other sides,
but it'll be pretty accurate because you can see the bottle from one angle.
You could imagine a world in which we provide multiple angles of the bottle
just to make it a little bit more accurate. Maybe there's something on the other side that isn't
visible in one image that you want to make sure it's like clear to the model. So those are the types of
things that are just obvious. This will be the first of what's going to happen. How do you think the
value of these things will change over time as the cost and frictions to create them false? Humans are
really good at scarcity and assigning value to scarce things. And so a beautiful video that shows a
product was valuable because it's costly to create in some sense,
How does the availability of perfect high fidelity, unbelievable quality video at a moment's notice,
how do you think that that changes the value of the video itself?
And I'm just curious of other knock on effects of what you're doing that you've thought about.
I mean, I think it's interesting.
One comparison, you can kind of draw with this.
If you think about the 2010s generally, like it was a phase of design really taking off.
Companies that Canva and Figma are created in this decade.
And not just that, but there were a lot of, like, companies.
that we're doing, make a website with a few clicks, it looks awesome, great designed websites,
just like one click away. This was an AI. There was a huge movement to just like, if you want to
sell something on the internet, if you want to have a business of any start, you need a great
design website. If your website looks like is from the 1990s, no one's going to buy anything from there.
I think that's cool again, though. Yeah, it is cool now, yeah, which is crazy how fashion moves,
right? Yeah. It all moves in cycles. That's right. Yeah. There's almost nobody that has a bad website
anymore. But that doesn't mean that having a good website is not valuable. It's still valuable.
If you don't have a good one, then you might still suffer today, even though it's like commodity, essentially, everybody should have it.
The video is more worth taking out this decade. I think we'll see more and more people adopt it. It feels like there's a lot of people adopting it today, but I think it'll be even much larger than that because the portion of creators within the video ecosystems will grow. More people will be creating it and potentially even more people consuming it.
So I actually think that the value of the video will not shrink exactly. It'll still be high-quality video.
will be high quality video, and it'll be a requirement if you want to like market, sell,
or whatever you're doing. But I do think that there's going to be other things about video
that are going to become more valuable. So, for example, if you think about likeness,
if models can just generate likenesses of people that don't exist at a whim, and they look like
great people, people you would want to represent your brand, you could even own a likeness as an IP
of your company of a person that doesn't exist and have them be the spokesperson of the company,
that sounds awesome. That sounds great. But that means that the value of the likeness is just going to
zero. The average likeness is not worth anything because anyone can make one out of nothing.
And what does that mean for the cost of likenesses in general or on the high end? I think it's
going to be determined by who's known. A likeness that is actually known by people, trusted,
understood by thousands, hundreds, thousands, millions of people is valuable now. It's much,
much more valuable all of a sudden. And by the way, that person may not have exist either. Someone might
create a completely fabricated person, post videos and stuff, become famous. Yeah, little Michaela was
way ahead of his time. Shout out Trevor McFedderick. Way ahead of the times. Yeah. So it doesn't
sound crazy in that world. I mean, I think you can grow crazy with this stuff. What are the surprising
limitations of these things? What would people be surprised that they have an especially hard time doing?
We've all seen video models struggle with people at the end of the day.
Fingers.
Yes, fingers, arms.
Drinking.
Yeah.
Olimpics.
Spaghetti.
Yeah.
I think we're kind of taking the unique angle on this generally, which is that we are training
specifically on people.
Our data is all people.
And we are specifically generating people.
We also are going to have conditioning the ability to provide like a
skeleton, for example. This is the exact animation I want to play out. This is the exact
TikTok dance I want you to do, for example. It'll just make it happen. And that actually makes
the model much more likely and better to be able to learn what human anatomy looks like and what's
normal and what's abnormal. People do have six fingers. It does happen. The model doesn't know that.
Obviously, it's not that that's the training data that's like causing it, but it may not fully realize
that if not enough training data has been given to it, that shows hands in all.
kinds of configurations and doing all kinds of things. So our goal is to solve that human generation
problem, like just actors essentially in general. The scarcity aspect, too, is that some of these are
not new problems. The corollary around movies is that a Michael Bay film, a $250 million budget or something
like that blows up half L.A. Transformers or something, I don't know. Tons of people got and see it,
blockbuster film. But all of those people are paying $25 per ticket or something. The same thing
happens for a low budget film if they can get into the box office. But the ticket price is the
exact same. I'm actually very excited about a world in which lower budget filmmakers and video
creators in general can just create more and do more complex things with not necessarily the
budget restraints. That's a massive hurdle for film creators and just creators in general.
I think it just up-levels everyone. I think the craft maybe shifts a little bit or this and that,
but those high budget films, as Garv mentioned, like technically it's general.
it's not synthetic or it's not real. Some of those things are real, I've realized, but it maybe
even creates more premium on some of those aspects. What does it feel like in the competitive
landscape to have established something so successful and important? I love the idea that companies
pass a level of maturity when someone else tries to kill them for the first time. Have you had that
experience yet? I'm really curious for like the sharper, rougher elbows part of building something
so fast. Any experiences like that that are interesting? Definitely. I mean, I think with all these
types of things, we're always, let's go with our mission and not worry about what others are doing.
But yes, a lot of people care about what we're doing. In fact, I think I would say in terms of
bigger companies, I think we're seeing an interesting evolution. Like, we kind of fall in an interesting
spot where we semi-collaborate with a lot of social networks because we're beneficial to
their growth. We create content and all social networks need content. And we have on-watermarked
content, content that is original. And this is a big problem for Instagram. If you remember, like,
When they launched reels, everything had like a TikTok watermark on it, and it was recycled TikTok,
basically. But we have a lot of that type of good content that's being generated on our platform,
by the way, like hundreds and hundreds of thousands a day. That's going to social media.
And so we end up being a valuable partner for a lot of social networks. And we've seen like
the social network landscape evolve in that sense. A lot of VCs ask the question, like,
what if Facebook copies you? What if Google copies you or something like that? And I think
what we're starting to see is like Google and Facebook are not the copying companies anymore.
They're not copying anything. They're just doing their own thing. And the copying company actually is
TikTok or by desk more generally. I don't know how this shift exactly happened. Facebook suddenly
became the good guys. I think Mark Zuckerberg is a hero now for putting all these models out,
making all this open source stuff. Suddenly his vibe is completely shifted. And then I think TikTok
has become essentially what Facebook was. Capture, kill, destroy.
everything that exists in every market that exists. Don't collaborate with anybody. And I think it'll
be interesting to see how that plays out. Obviously, there's many talks happening about a band and all this
kind of thing. We'll see, like, where all that goes. But their leadership is very well aware of our
existence. And they have tried many, many times to try to kill us. To their career, they were the
first to be aware of our existence of anybody else. What does that look like them trying to kill you?
Literally just copying the product. Blame and copying. They literally were to the extent of copying our app store
description, our website, exactly putting that in their press release word for word, copying our brand
colors, exact, precise brand colors, pretending to be asked beyond anything you would imagine.
And just kind of crazy that like a company of that size would even try these types of tactics.
But at the end of the day, like the software that they just create is just very mediocre.
And it just works because they have great distribution through TikTok.
And I think we win because we just have better product.
It seems like early days of all these models getting built that research talent was one of the most important scarce resources in extremely short supply.
Can you talk through what you've learned about that, how it's changed?
Is it still a handful of people that you really need a couple of them to be able to build the cutting edge thing?
What is the role of extreme research talent in building the models that fuel all this great product?
So the talent side is still, I think, evolving.
I don't think it's completely solved for what it's worth.
As the use cases are growing, as more and more people are realizing what's possible,
as more companies are getting started, trying to solve similar problems.
There's only going to be a more and more for shortage of talent.
Talent isn't created overnight, right?
Like it takes years and years of experience before someone can be considered experienced in an area.
And I think we will still see continued pressure on that talent side,
especially for like building generative models and foundation models and things like that.
Obviously the more VC dollars that are poured into this area, that'll have an effect.
But I do think, interestingly, it doesn't take an army to build this type of stuff.
It takes a few good people.
And it might take an army to like scale it and really make it big.
But to deliver world class results can be done with maybe a dozen people or less.
And you can beat everybody in the world with a team that small.
if you had the right ingredients in place.
A lot of the challenge just becomes like finding those people.
You can't get it wrong.
You want a very specific set of people.
You want a specific skill set.
This is all new.
So very few people have experience in it.
It's all cutting edge.
There's discoveries and inventions like happening every day, every week.
So the more you care about it,
you want to find people really close to the cutting edge
who really know what's happening today
and what are the small little wins of techniques
that will get us the edge over everybody else.
So it still is a challenge.
What our duty ends up being then is we have to give them all the resources in such to be able to do their work.
There are a lot of these folks in AI labs that can't release anything they're working on for better or worse, at least from that place's opinion.
And when you're able to bring some of the ingredients that we have, whether it be like compute data, the environment, it ends up not being that complicated in terms of recruiting negotiation and stuff.
How do you think these products will price over time? This is like always the weird question with software where the marginal delivery of it costs nothing. A lot of people have been talking about, let's look at, let's say, Accenture's Market Cap or something like that. And it's a $250 billion company, basically selling labor, very high, expensive labor, important labor. Do you think that AI applications will take labor budgets and be priced like heavily discounted labor because that's what they're doing and replacing?
or is it just going to end up pricing like all software does?
We've run these playbooks for 20 years and we kind of know how to do it.
What's your sense of how people should think about pricing AI software applications
and what its equilibrium state will be in the future?
I don't know if we completely understand it yet.
Basically, like, I think it's almost too early to tell in a way
because we aren't completely able to replace labor all the different aspects.
So we don't know what people want to be willing to pay for it yet.
in the use case graph we're like a four three four or five percent whatever something in that range
is just early and we aren't able to like fully replace certain workflows or like very
operationally heavy like processes that might exist some companies and stuff and we will get
there slowly and steadily we're moving towards that and I think we'll see what people will be willing
to pay for that so I think we'll figure it out one of the big questions there is like how does that
between consumer and B2B I think consumer pricing is evolving pretty clearly I think we're starting to
see what that looks like it seems like
it's coming down to consumer subscription. And it also seems like people are willing to pay a little
bit more than they would have otherwise. So for example, traditionally for like video related apps
on the app store, for example, web apps and Android whatever, this standard price is somewhere in like
the $799 to $1299 range. That's just considered normal. And there's a freemium business model to it.
I think what we've seen different is like for us, for example, we have been for a long time like we're
completely premium. There's no free product. You cannot even use it once for free. And that worked
just fine. People were like, okay, whatever, here's money. Let's move on. And so that wouldn't have
worked in an older world. Without the newer technologies and stuff, people would have been like,
yeah, I'm not paying for this. Right. I'm moving on to the next one. I think the other thing we're
starting to see is, can we charge $25 a month? Yes, we can. People are paying that. And so people are
clearly willing to pay much higher prices. Like, if you look at a lot of different AI companies out there,
the video generation companies and stuff, they're going across this range too and people are paying all
these prices. People are paying up to like $2,000 a month, consumer subscription. I think there's a lot more
ability to go higher on the subscription pricing than there was previously. Now, that might change.
I think a big factor of that might be like, there's just not enough competition still.
Still might be like there's only maybe one or two models in the world that are like off that quality
and people care about that quality. So you really don't have a lot of choice. Maybe if there's like a
how did these types of model floating around? Maybe the price comes down in the future.
So that's what we're seeing on the consumer side. And then on the B-to-B side, I think that's where
we will figure out a lot of it. Right. I think some of the big things that need to be solved there
is will businesses buy models that are trained on on licensed data? That's like an open question.
And yeah, they are to an extent. We'll see kind of how all that plays out. We're planning to go
much more on the fully licensed side. That's going to be one of our main differentiators because we are
uniquely positioned for that. We actually collect data at like a massive scale so we can actually
train fully licensed models. My feeling is that towards the end game, not today, but as this
area gets very saturated, so many years from now, I think things like having fully licensed
models will factor in because you'll be able to win on that very easily in like a competitive
deal and people will care about those types of things. And people might even be willing to pay more.
for that type of guarantee or just like the reps that it's licensed.
And then I think besides that, it really just comes down to like how much of the use case
we'll be able to cover.
And that's the big question.
Okay, we're at 5% today.
But is the limit 100%, is it 75%, is it 50%, where does this stop?
My guess is we can go all the way to 100%, or at least very close to.
Just because it's a solved problem, we know that this is solvable.
And I think if we can get there, I think a lot is going to change about how video workloads work.
in the world. And the pricing around labor and hot topic right now is seat licensing versus labor
or like aligning to labor costs or something. I think people are maybe rushing into the labor
argument that it actually has a very similar path or has had a very similar path. Turns out the
CFO would like that number to go down. It's not some special number or something. And if you
remove the human element to it, my guess is that that probably only puts more downward pressure on it.
It's like, great, we can do more with less. And it's like, perfect. Whatever the software is doing,
if it's writing code or if it's the automated SDR, you know, whatever it might be.
Like, there is downward pressure against those things.
I think people are getting a little excited around running towards that.
And don't get me wrong, pricing towards output and such is pretty cool.
I'm sure there is something there and there's some equilibrium we'll find.
But I do think people are rushing to it maybe a little faster than they should be
in that there actually might be more continuing alpha in the typical subscription.
Yeah, sure, maybe that's not the Salesforce seat, the classic comparable.
But there's just some market exploration that needs to have.
happen, and we probably haven't fully seen that yet. I think the entire world of investors,
VCs, growth equity investors, public investors, pretty much every single one of them is trying
to figure out how to think about AI and its implications on their companies, prospective companies,
equity valuations, all the normal important questions. How would you advise them from the other
side of the table? You've talked, I'm sure, to a lot of the great investors. You have several
of them that have invested in your company. What do you think investors understand about AI well?
What do they feel like categorically? They're not understanding as much detail as you do from the
builder's side. Give us your lay of the land of how you think investors are doing, give them a grade
or something at understanding. Yeah, if I have to grade it, maybe from a public equity side,
there are a lot of smart people out there. So I'm not going to give them to our. I don't think it's
fully being appreciated how much this is changing. We're fully like everyone's saying
and all that. But I think there's continued talk of, oh, there's all this R&D span or capital expenditure
and it's like, where is the value and stuff? And I think there's just so much attention on the
large AI labs that are effectively what Gar was talking about before is that solving intelligence.
That's a very, very different mission than the company who maybe is creating the automated
software developer, two very different worlds. And so I think paying more attention to things
outside of that is pretty important for it to really understand how this is changing inside of
their companies. I think it would be hard to find a company today, a successful company today,
that hasn't and or isn't exploring AI tool of some sort to either completely replace an activity
inside the company or, quote unquote, do more with less in another capacity. I think that's true
of every function of essentially every successful company today. And that's where you're seeing
like a lot of the adoption. So even discussing the foundation model versus some of these companies
who are just fine-tuning a model as something open source or something, there's a ton of alpha
getting these tools inside of their company. So if you're talking to someone who's, how do we do
this and roll it out as a larger enterprise, I think there are already examples. There are massive
enterprises that you go in. Someone was telling me the other day that L'Oreal, the beauty company,
you can go in and they have like an internal GPT, basically, an internal L.M of some sort.
any employee can ask any question. I don't know how much that's really being baked into their thought.
I think there's just so much attention towards these particular AI labs in the way that they're
running their businesses, which is extremely different than some of other companies,
in particular to be if they have the backing of Microsoft or something. Yeah, it's inherently just being
driven differently. And yeah, I think if you were to go into those, I think your viewpoint would
potentially change, and that could definitely inform a better understanding on how this is actually
going to change work. I love this framing of the unbounded problem nature of intelligence versus the
bounded problem nature video or some of these other things. Kind of a fascinating bifurcation.
I actually think that that applies to the text side too. Even on text, we already have created
essentially what is a tool for intelligence. It's like intelligence in a box. Intelligence you can just
apply onto something to solve a bounded problem. So whether that's coding, now think
of it in the coding context. I think as Dwight was saying, engineers are smart people. Does that
mean we need AGI to solve coding? Not necessarily. Because essentially what is doing really is just
translating. Think of how like computers evolved over time. We used to literally do the punch card thing.
Then we were writing assembly language. Who knows that anymore? Then we were doing C++, right? Exactly.
Right? Just two. Yeah. Then we were writing C plus plus and then there's these higher level languages like
Python, coming to the modern era. Scott from Cognition was the guest today. So he's building the next layer.
Perfect. Yeah. And then we're kind of just saying like, hey, the new program language is English.
That's not a crazy job. It's actually a very bounded problem. It's a problem of like inventing a new
programming language essentially. Like a program language that is even more understandable to people
because they already know it. It's the language that we already know. Intelligence is a special case.
Exactly. The general intelligence idea of, oh, we're like creating conscious.
Oh, it's like a thing that's going to exist, grow around, do things, like have his own thoughts and
have its own, like, dreams and hopes and stuff. And maybe it'll start a company at some point.
That's a whole different mission than, like, solving intelligence in the box, which essentially
all exists and it's getting better and better.
I'd love to just extend the analogy one step further to business model. Most of the commentary on
AI businesses has been, again, focused on foundation model companies that have, well, have had two
problems, huge KAPX outlays to train the models. And then huge inference business.
so often early on, really negative gross margins just to service their $20 a month subscription
product. Inference has fallen 100x in cost in the last 18 months or something crazy.
These costs are going down. But those were the two criticisms of the business model was,
oh, my God, this unbounded race, I got to spend 10x every time to build the next thing.
What am I ever going to make some money? It seems like this other category of more bounded
problems have pretty normal, great business models. Is that right? Is your sense that you
guys are just going to have really high gross margins like a normal software company. And yeah,
you have to spend money training your foundation models, but it's not $10 billion. And walk
me through the business model expectation margins, capbacks, things like that, what the J-curve looks
like in these businesses. Educate us a little bit on this second category. The way we think about it
for our business specifically is that there is a bounded cost that actually solves this problem.
That bounded cost is probably in the hundreds of millions of dollars. But it's a lot of
it actually gets us to a solution.
It gets us to something that, hey, this is actually reasonably good
at generating anything that a CGI studio might be able to do.
And that is the level that we need to be at.
Now, will that evolve?
Yes, it will need to fine-tune it.
But fine-tuning is generally cheap.
It's actually not even close to it as expensive as like training a foundation model
from scratch.
And yeah, new data will come in, which we already have a flybill we're building for.
And it's going to be massive amounts of data.
We're going to be continuously training the model
and making it aware of what's happening today
and what things people might want to generate today.
But that's just incremental fine-tuning.
It's going to be a low cost that's underlying the business.
On top of that, inference costs are going down.
So I think it's going to start looking more and more like a traditional software business.
I think what's going to happen is like initially with these tests and models existing,
whoever truly solves this problem, will have a moat for a while as long as they are ahead.
I think for us, we're also trying to build that data mode simultaneously so that we are
permanently ahead. And then once enough data is out there, enough people have raised enough money
and have tried the exact same playbook and built these models. And this could be many, many,
many years in the future. It's going to become a software race, building the workflows,
building all the traditional stuff that we know about. Pricing and packaging, like, all this stuff
is going to become really important. We've seen it all. People are going to do APIs,
then do like B2B consumer, all this stuff. There's going to be all these use cases. And I think
that's where the real competition will happen, and there's going to be winners in that.
I think our theory and strategy on this is the winners are going to be really determined by
who has the best model that's consistently outperforming everybody else.
All that comes down to like data acquisition, flywheel, essentially, and the ability to
constantly improve the model. I do think this won't be the end, though. I think new problems
will get unlocked, and we already have line aside into that, what those other problems look
like, and those problems will have their own foundation models and their own data to be collected.
And essentially, you could imagine a series of foundation models that are solving like a family
of problems across a whole set of a workflow that's broad across like video and maybe even
other types of media, different types of use cases, like film, TV, whatever you want, basically.
Maybe it's dubbing, maybe it's...
Hands.
Yeah, hands.
Post-production.
Like, I don't know, right?
Lots of different possible use cases.
So, as always, that will happen.
no doubt about that. You actually can see that these models will reach a point of maturity.
And I think on the like, what does this end up looking at a mature business end,
call it some threshold. I genuinely believe that these can look like very high margin
businesses, whether that be the deflationary behavior of GPU and just compute in general.
Like, it's incredibly early. Like we're talking about the video's latest chip, et cetera.
You're already seeing the cost come down from on H-100s as the H-200 architecture and stuff is
being rolled out.
Like, throughout history, like, these prices have never gone the other way.
It's highly deflationary as the next one rolled out because ultimately that is their business
model.
They make them more efficient.
They make them more powerful, whatever it is.
And so I think, generally speaking, that that is 100% guaranteed, at least from my
perspective, I think the interesting thing, though, is that, like, when you're earlier
stage and companies are earlier stage, and just talking about startups in general, is that
higher margin businesses actually sound to me like perfect attack vectors for another
entrepreneur. And I think you should be very wary of companies that are operating at really high
margins in the particular earlier stage. For these types of businesses, there's a ton of margin
expansion opportunities. And then that also goes for the later stage companies, though, too.
I think you're seeing it right now, the CRMs and companies, you know, operating 80, 90%
margins. It's great, typical SaaS kind of stuff. They seem like great opportunities for companies
to essentially go right after. In reinventing some of this stuff, those companies don't have the same
pricing power that they had do that they did 15, 20 years ago. At the same time, though,
the great ones are thinking about that right now and reinventing themselves. And so it does feel
like a little bit of, you know, as we discuss some of these business model changes, I think
it is a bit of a shifting ground. If you think about the future now, what is on the other side
of the mission accomplished banner of you just did all video and you can create anything you can
imagine in CGI with a hundred million dollar budget? Now you can do in captions. Then what?
What do you think you would do that?
I mean, I think if we actually achieve that within a reasonable time frame,
I think that would be just the beginning.
Because I think you could go so much beyond that.
I think these industries are massive.
Like, you could imagine a social network based on something like this.
You could imagine film and TV and stuff being dominated by these types of technologies.
You could imagine education being completely transformed.
The list essentially like endless.
This would be the starting point of a potential,
complete transformation across like multiple industries.
So I think today we're really exciting about accomplishing this particular mission,
but I think the possibilities beyond that are practically endless.
Any major lessons from your time at Snap, which strikes me as a very unique culture
and an extremely product-centric, like a good place to train for product maybe?
What lessons do you take from your time there and what lessons do you leave behind?
I mean, I think Snap had, as any company, a lot of good things and some bad things.
I think the great things that I was able to get from Snap was the ability to work with a lot of great people.
I think Snap was in a tough spot in many ways.
They were in one of the most competitive possible businesses you can exist, monopolistic by nature,
where it's really difficult to get something started and very easy to get killed.
And only the biggest one actually wins and survives.
and in that arena they were able to make a place for themselves,
mainly because of innovation.
And this just comes down to the CEO.
He was able to, out of all the random noise,
see something and understand, yes, this will work.
And nobody will see it, but I know why it will work.
And I think at the core, if it was,
he had like an understanding of the product and the customer
in a way that nobody did.
And nobody even came close to it.
There were many moments in the company's history where he was like, we're going to do this.
And everybody was like, no, we shouldn't do that.
Like, this is a bad idea.
And he'd be like, I don't care we're doing it.
We did it.
And it was the best thing we ever did.
That's the level to which his intuition was there.
Snap was famous almost for like constantly innovating.
Stories came out of Snap.
The old maps location sharing product and idea came out of there.
There's so many things that were innovating on.
I think they kind of lost a little bit on sort of the public TikTok thing.
but that actually was something that didn't fit in their like pillar vision strategy basically
because they're a private sharing platform.
Their whole purpose was low abuse.
People don't feel like you can't even reshare posts because that's a way to like embarrass
somebody by like reposting that thing to other people who are not supposed to see it.
Everything was designed around feeling good, having fun and sharing with friends,
which was really everything that people cared about at that time.
And I think they kind of missed the TikTok thing because it was the exact opposite of that.
It was actually shared to everybody.
And interestingly, it created similar dynamics where, like, sharing to everybody actually made you feel more private because there were so many people that people you know would never see it.
Somebody else would see it.
But that's the details.
I think on the downsides of Snap, one of the interesting things there is, and one of the learnings there is product market fit often doesn't have a lot to do with.
what people are doing day-to-day within the company. And once it exists, it can stay there
despite the actions of the people. So I think what ends up happening sometimes in bad cases is
like people think that the wrong actions that they're taking are the contrary and right view
because, well, the company is growing. So of course, whatever I did was the right thing.
but actually the company is growing despite the wrong thing that was going on at the time.
And so it's difficult to tell what actually is calling the company to grow and what's the good thing
and what's the bad thing.
And a lot of people walk away from these types of high product market companies thinking that
all the things that did were good things.
And there were no bad things because the company grew.
But the reality is the company is growing despite those actions.
I think identifying those was a skill that I had to like really work on building to understand
And how can we truly measure what we're launching, what we're building and understand what's a good thing and what's a bad thing.
So what I'm really grateful for from that time is the ability to work with the CEO there.
He really brought me into the circle.
Like he had a great design team that he had built.
A lot of the decision making was driven through the design team.
It was a small set of people like 10 to 12 people on that team, even when the company was many thousands and thousands of people overall post IPO.
So being a part of that team, learning from the great people on that team.
team. I evolved my design career through this process. And I think his ability to like identify
this is a person who will fit in well and will be able to learn and figure these things out.
Props to him. Definitely doing something right. The closing question I ask everyone in this show,
fun to get to do this twice today. What is the kindest thing that anyone's ever done for you?
I mean, it's hard for me to not say the kindest thing is probably my wife. And we started this company.
We were already married. We had her first.
kid, pretty hard not to call it, that it could have obviously not gone that way. I decide not to
start the company, not to do a bunch of this stuff. And yeah, enabled me to take more risk.
And yeah, yeah. Now I can't use that answer. Yeah, exactly. Yeah. I mean, I think it's a little unfair.
Besides that, I would say, likewise, by the way, for me, if I were to give you a different answer,
I think it would be just my parents making sure that I was born in the U.S. literally. Because, like,
I was only here for, like, the first couple years of my life. I was born well,
my dad was doing his PhD at Northeastern. He was studying economics. So he was there for like
four or five years, basically that I was born in the middle of that. They moved back to India
after that. But I had the U.S. citizenship. Without that, I'd still be in India. Yeah.
Simple and powerful. Yeah.
Guys, thank you so much for your time. Thank you. Thanks.
If you enjoyed this episode, check out join colossus.com. There you'll find every episode of
this podcast complete with transcripts, show notes, and resources to keep learning.
You can also sign up for our newsletter, Colossus Weekly, where we condense episodes to the big ideas, quotations, and more, as well as share the best content we find on the internet every week.
