Invest Like the Best with Patrick O'Shaughnessy - Aravind Srinivas - Building An Answer Engine - [Invest Like the Best, EP.363]

Episode Date: March 5, 2024

My guest today is Aravind Srinivas. He is the founder and CEO of Perplexity, a startup that he describes as an “answer engine” built from scratch with AI. Aravind has set out for perplexity to bec...ome the most powerful answer engine backed by up-to-date sources. He helps me pick the technology apart, describing the behind the scenes of what it takes to build Perplexity to reach its potential and compete alongside the likes of Google and OpenAI. Our conversation goes deep into programming this type of infrastructure, the competition around latency, and constructing a business model around deep learning. There is so much on this horizon, so please enjoy my conversation with Aravind Srinivas.  Listen to Founders Podcast For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- This episode is brought to you by Tegus, the only investment research platform built for fundamental investors. How hard do you work to get the insights you need to make a great investment decision? How many hours do you spend digging through public records and expert transcripts, or manually updating complex models? Investors should compete on their ability to analyze investments, not how well they aggregate data. That’s why Tegus offers a unified, end-to-end research platform that combines robust qualitative content sets, up-to-date financial data, management and culture checks, and more — all in the same easy-to-use, streamlined user experience. 95% of the top 20 global private equity firms use Tegus. Shouldn’t you? Learn more and get your free trial at tegus.com/patrick. ----- Invest Like the Best is a property of Colossus, LLC. For more episodes of Invest Like the Best, visit joincolossus.com/episodes.  Past guests include Tobi Lutke, Kevin Systrom, Mike Krieger, John Collison, Kat Cole, Marc Andreessen, Matthew Ball, Bill Gurley, Anu Hariharan, Ben Thompson, and many more. Stay up to date on all our podcasts by signing up to Colossus Weekly, our quick dive every Sunday highlighting the top business and investing concepts from our podcasts and the best of what we read that week. Sign up here. Follow us on Twitter: @patrick_oshag | @JoinColossus Editing and post-production work for this episode was provided by The Podcast Consultant (https://thepodcastconsultant.com). Show Notes: (00:00:00) Welcome to Invest Like the Best  (00:04:08) First Question - Redefining 'Great' in Search (00:07:16) The Mechanics of Perplexity (00:10:59) The Evolution of Perplexity From Wrapper to Robust Infrastructure (00:13:19) The Future of Large Language Models (LLMs) (00:22:39) Aravind's Strategy For Constructing The Business Model (00:28:16) The Process of Building an Index for a Search Engine (00:35:55) The Impact of Scaling and Data Quality on AI Models (00:40:00) Bottlenecks in AI Development (00:47:06) The Talent Landscape in AI (00:57:59) The Vision for the Future of Search (01:00:19) Importance of Execution in AI Startups

Transcript
Discussion (0)
Starting point is 00:00:02 Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest Like the Best. This show is an open-ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. Invest Like the Best is part of the Colossus family of podcasts, and you can access all our podcasts, including edited transcripts, show notes, and other resources to keep learning at join colossus.com. Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of positive sum. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions.
Starting point is 00:00:45 Clients of positive sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc. My guest today is Aravind Shrinivas. He is the founder and CEO of Perplexity, a startup that he describes as an answer engine built from scratch with AI, has set out for perplexity to become the most powerful answer engine backed by up-to-date sources. He helps me pick apart the technology, describing the behind the scenes of what it takes to build perplexity to reach its potential and compete along the likes of Google and OpenAI. Our conversation goes deep into programming this kind of infrastructure,
Starting point is 00:01:23 the competition around latency, and constructing a business model around deep learning. There is so much on the horizon, so please enjoy my conversation with Aravind Shrinivas. I thought A lot of fun place to begin would be a game I like to play with people, which is I call it the one-minute bio. I'd love to hear the one-minute summary of your life up until the founding of perplexity, just to set the context for all that we'll talk about. Yeah, I grew up in India from a pretty humble background, was focused a lot on engineering and programming right from the beginning, studied in one of the IITs, got really excited about AI because I happened to do a course on machine learning. Deep Mind published their paper on
Starting point is 00:02:08 training AIs to play Atari games. That got me to be resourceful. I got consumer gaming GPU cards from other people on my lab, traded my lab desk for them so that I could stay at the hostel and work on training these neural nets. And in general, when I look back, I've always realized I've been good at making deals. But traditionally, you would consider me a nerd, just excited about research and programming. My research there got me into Berkeley. I got to do more exciting AI research in Berkeley. That got noticed by institutions like OpenEI and DeepMind, which got me more exposure to the cutting edge. And at DeepMine, I had a pretty crappy apartment when I was an intern, so I just mostly stayed in the office. And I used to stumble upon books like how Google works there. And in
Starting point is 00:02:58 that, like I got really excited about entrepreneurship, and that got me into thinking a lot of starting a company to work on a hard problem. Little did I realize I would literally go on to work on search itself, but that was faith laws irony sort of thing, where the people who are excited about you taking on them. I worked at Open AI, and my entrepreneurial ambition exceeded what I could do as an individual contributor there. So I left and started perplexity. Initially, as a small project, to just work on searching over Twitter and LinkedIn and things like that. And as we kept seeing the progress, we just expanded our ambition to just searching over the entire internet, providing a much better experience by giving answers rather than links.
Starting point is 00:03:42 What is your theory of good deal making? Create a win-win situation. And don't be greedy. There's a show called Succession. Yeah, great show. There, Logan Roy advises. It's not a deal unless you really screw over the. other person. But in Silicon Valley, at least long-term deal-making is the best deal.
Starting point is 00:04:04 When I think about you, I picture that scene in Star Wars where Luke Skywalker is flying the little tiny plane into the Death Star and you're the plane and the Death Star is Google or something. Google's been considered one of the most unassailable moats in all the business and the product's been ubiquitous forever. We've all used it many times a day for decades now. Tell me about how you conceive of the components of a great search product. To go up against Google, you obviously have a theory of this is what great search should be, and maybe Google has gotten away from that. Just talk us through what great is to you when you think about search. When I think about the word great, I'm reminded about Steve Jobs, insanely great. You don't want to be just great. You want
Starting point is 00:04:46 to be insanely great. Look, search has always been a hack. Ten Blue Links was always a hack to get us information, but not need it anymore when we can more or less answer your question directly. So that is the great experience. And what is insanely great then? Insantly great is the AI doesn't even let you struggle to articulate a good question. This is the next part. Of course, you still have to make the first part work reliably, accurately, you know, address the long tail of mistakes.
Starting point is 00:05:21 but assume that that's going to get solved with a good amount of engineering. The next part, actually, to get into insanely great territory is make it so easy to even ask a question. Why are very few people good podcasters? Why are very few people good interviewers? Because asking good questions is not a skill that most people have. Everybody in the world is curious. Curiosity is unlimited, unbounded. But not all curiosity in every individual can be precisely.
Starting point is 00:05:51 They articulated into good, interesting questions that elicit the most from an interesting mind. And AI is an interesting mind. It is a very knowledgeable mind. It is access to basically all the world's knowledge in an instant. But it's up to you to harness the power. Now, why do we require humans to be great prompt engineers? Why do you want to sell that vision? The open AI vision of, oh, AIs are amazing.
Starting point is 00:06:18 You guys figure out how to be good at using them. But what would Steve Jobs do? Steve Jobs would be like, bring the power of these amazing AIs to the mere mortal. In fact, he's used the word mere model in many of these emails. Our job is to bring the joy of personal computing to mere mortals. Those who can see it before others, but democratize the joy. That's what we want to do for knowledge. Bring the joy of learning to everybody.
Starting point is 00:06:46 And so that's what we want to address, either helping people ask questions or staff. Start with some dumb version of the question and help the AI refine it for you and profile you enough that suggests interesting questions to ask and how the knowledge feed of questions that you just see because you ask questions about those topics before. And just every single day make you some X percent smarter. That's what we want to do. It's really cool to think about the sequencing to get there. We've had search engines.
Starting point is 00:07:20 Like you said, it's a hack to get to end. answers. You're building what I think of today as an answer engine. I typed something in, you're just giving the answer directly with great citation and all this other stuff we'll talk about. And the vision you're articulating is this question engine. Anticipate the things that I want to learn about and give them to me beforehand. And I'd love to build up towards that. So maybe starting with the answer engine, explain to us how it works. Maybe you could do this via the timeline of how you've built the product or something. But what are the components? What is happening behind the scenes when I type something into perplexity?
Starting point is 00:07:51 either a question or a search query or whatever. Look us through in some detail the actual goings on behind the scenes in terms of how the product works itself. Yeah. So when you type in a question into perplexity, the first thing that happens is it first reformulates the question. It tries to understand the question better, expands the question in terms of adding more suffixes or prefixes to it to make it more well formatted.
Starting point is 00:08:18 So it speaks to the question engine part. And then after that, it goes and pulls so many links from the web that are relevant to this reformulated question. There are so many paragraphs in each of those links. It takes only the relevant paragraphs from each of those links. And then an AI model, we typically call it large language model. It's basically a model that's been trained to predict the next word on the internet and fine-tuned for being good at summarization and chats. That AI model looks at all these chunks. of knowledge snippets that you surface from important or relevant links and takes only those parts
Starting point is 00:08:57 that are relevant to answering your query and gives you a very concise four or five sentence answer, but also with references. Every sentence has a reference to which webpage or which chunk of knowledge it took from which webpage and puts it at the top in terms of sources. That gets you with a nicely formatted rendered answer, sometimes in markdown bullets or sometimes just generic paragraphs. Sometimes it has images in it. But a great answer with references a citation so that if you want to dig deeper, you can go and visit the link. If you don't want to just read the answer and ask a follow-up, you can engage in a conversation. Both modes of usage are encouraged and allowed. What percent of users end up clicking beneath the summarized answer into a source
Starting point is 00:09:44 webpage? At least 10%. So 90% of the time, they're just satisfied with what you give them. It depends on how you look at it. You want it to be 100% of the time, people always click on a link. That's the traditional Google. And do you want to be 100% of the time where people never click on links? That's ChatsyPT. We think a sweet spot is somewhere in the middle.
Starting point is 00:10:03 People should click on links sometimes to go do their work there. Let's say you're just booking a ticket. You might actually want to go away, Expedia or something. Let's say you're deciding where to go first. You don't need to go away and read all these SEO blogs and get confused on what you want to do. you first make your decision independently with this research body that's helping you decide. And once you finish your research and you have decided, then that's when you actually have to
Starting point is 00:10:29 go out and do your actual action of booking your ticket. That way, I believe there is a nice sweet spot of one product providing you both the navigational search experience as well as the answer engine experience together. And that's what we strive to be doing. Can you talk about the relevant pieces of third-party technology that makes something like this possible and how you blend them with your own technology. So we could talk about retrieval augmented generation here. We could talk about the various lineup of OpenAI versus Claude versus Google and how you think about the underlying LLMs that you can use and tap. I'd love to talk about
Starting point is 00:11:04 how you think about that as a business person too. I hear a lot of chat GPT wrapper or something like this. Tell us about the stack that you use. And don't be afraid to get technical, super smart audience. I'm just curious how these things all tie together and work. in combination? First of all, we started off with $2 million in funding. When you have $2 million in funding, you have no job trying to build infrastructure yourself. Yep. Your only goal is to validate if you have a product that people want to use on a day-to-day basis,
Starting point is 00:11:36 or at least a weekly basis, and get enough traction and awareness among users, and then think about building infrastructure that allows you to scale from where you got to 10x, 100x more. Now, that is the level to which we were ambitioning at the start. And we decided to be a wrapper. We want to do things that allow you to get the product out as quickly as possible. So we decided to be a wrapper. We connected Bing API with GPD 3.5 API and launched perplexity.
Starting point is 00:12:09 Anybody could have done that, except once you do it, that's when the game begins. The game begins after you got some excitement. and some users are using your product even after your initial hype, and you get the sustained usage. Now that's when you say, hey, look, this thing I'd wrap together is not going to scale. Sure, Open AI is going to continue making 3.5 more scalable, and Bing has already done decades of work to make this search engine API scalable. But the orchestration layer that takes both of these together and handles so many queries
Starting point is 00:12:43 and can handle any outage in any of these APIs, requires you to build infrastructure yourself. And that's when we started building infrastructure ourselves. When 10,000 people were on the site at once, when Jack Dorsey tweeted about us, even though it's a wrapper, it went down. Because we don't have the right rate limits with Open AI. We don't have the right rate limits with Bing
Starting point is 00:13:02 or Chancheapity goes down frequently. There are all sorts of issues. AWS servers go down. That's when you start to actually build all the foundation layer of your infrastructure and back. back in to support a more scalable product. And when you start doing that, you keep encountering newer issues every time. And every time you solve the newer issues, your infrastructure keeps getting more and more
Starting point is 00:13:26 robust. And after a year, if you look back, you're like, oh, damn, this is impossible how far we have come. And we would have never imagined we could have built such a sophisticated backend when we started off with just two or three people. Can you explain from an insider's perspective and someone building an application on top of these incredible new technologies, what you think the future might look like or even what you think the ideal future would be for how many different LLM providers there are,
Starting point is 00:13:55 how specialized they get, scale the primary answer, so there's only going to be a few of them. How do you think about all this and where you think it might go? It really depends on who you're building for. If you're building for consumers, you do want to build a scalable infrastructure because you do want to ask many consumers to use your product. If you're building for the enterprise, you still want a scalable structure. Now, it really depends. Are you building for the people within that company who are using your product?
Starting point is 00:14:22 Like say, you're building an internal search engine. You only need to scale to the size of the largest organization, which is like maybe 100,000 people. And not all of them will be using your thing at one moment. You're decentralizing it. You're going to keep different servers for different companies. And you can elastically decide what's the level of throughput you need to offer. offer. But then if you're solving another enterprise's problem, where that enterprise is serving
Starting point is 00:14:48 consumers and you're helping them do that, you need to build scalable infrastructure, indirectly at least. For example, Open AI. Their APIs are used by us, other people, to serve a lot of consumers. So unless they solve that problem themselves, they're unable to help other people solve that problem. Same thing with AWS. So that's one advantage you have of actually having a first-party product that your infrastructure is helping you serve. And by doing that, by forcing yourself to solve that hard problem, whatever you build can be used by others as well. Amazon built AWS first for Amazon.
Starting point is 00:15:25 And because Amazon.com requires very robust infrastructure, that can be used by so many other people. And so many other companies emerged by building on top of AWS. Same thing happened with OpenEI. They needed robust infrastructure to serve the GPD3 developer API. and chat GPT as a product, that once they got it all right, then they can now support other companies that are building on top of them. So it really depends on what's your end goal and who you're trying to serve
Starting point is 00:15:53 and what's the scale of your ambition. Say a click more about how you view the relative merits of this, what seems like just a pure arms race that's happening. This context window gets longer, this latency gets lower, the sophistication goes up. The generations of models are dizzying. How do you build the business in a way that, that wins no matter what happens at all these companies or however many LLMs there are,
Starting point is 00:16:15 to make it forward compatible with a rapidly changing environment. I think there is no solution. You obviously need to have a good engineering team and people who are very nimble and can use and learn new things pretty quickly. One thing I would say very inspired by Jeff Bezos is the end user doesn't care what models you're using or what indexes you're using. all they want is a great product experience. None of your users are going to say, hey, Arvin, one year from now, I want your product
Starting point is 00:16:47 to be slower. Or, hey, one year from now, I want your product to be less accurate. One year from now, I want your product to be rendering the answer and these huge paragraphs. I don't want a better format. They're not going to say that they only want these three things to keep getting better, and they don't care how you achieve it. So the way we think about this particular question is whatever helps us get there. Ideally, it's something we build ourselves because usually for speed, you have to build stuff
Starting point is 00:17:15 in-house. If you rely on others, at some point, you'll hit the limits of how much you can speed up. On the other hand, if you're building the back-in-in-house, you can go even further in terms of making speed a priority for you. Let me give you one example. There are people who build on the Flutter or React Native stack so that there's one single code base for both mobile platforms, iOS and Android. But because you choose to do that, to minimize engineering overhead for you, the apps might get
Starting point is 00:17:46 slower for the end user, more unusable, more large in terms of memory that it consumes on the device, things like that, and it just makes for a worse user experience. On the other hand, if you build natively, if you directly go to Swift UI and build your app, the apps can feel a lot faster, snapier, consume less memory, and allow you to take more advantage of the native components, render more natively. Earlier, we used to render the answer on perplexity by using a render on the web and using the web view to render on the app. But then if you render more natively on Swift UI and build components for that, it feels even snappier, even better. Small, small optimizations like these, if you keep shipping them every few weeks,
Starting point is 00:18:30 the user loves it. And everyone loves a fast, snappy, accurate, reliable app. And then that increases your retention, your word of mutt and grows, you're now a better company than before. So I genuinely think backwards from what the user wants and try to do whatever it takes to get there. When I think about the history of the product, which I was a pretty early user of, the first thing that pops to my mind is that it solves this hallucination problem, which has become less of a problem, but early on, everyone just didn't know how to trust these things, and you solved that. You gave citations. You can click through the underlying web pages, etc. I'd love you to walk through what you view the major timeline product milestones have been
Starting point is 00:19:10 of perplexity dating back to its start. The one I just gave could be one example. There was this possibility, but there was a problem and you solved it. At least that was my perception as a user. What have been the major milestones as you think back on the product and how it's gotten better? I would say the first major thing we did is really make the product a lot faster. When we first launched the latency for every query was seven seconds, that we actually had to speed up the demo video to put it on Twitter so that it doesn't look embarrassing. And one of our early, friendly investors, Daniel Gross, co-invests a lot with Nat Friedman, he was one of our first testers before we even released a product. And he said, you guys should call it a submit button for a query.
Starting point is 00:19:56 It's almost like you're submitting a job and waiting on the cluster to get back. It's that slow. and now we are widely regarded as the fastest chatbot out there. Some people even come and ask me, why are you only as fast as chat chTPT? Why are you not faster? And literally they realized that chat chimpD doesn't even use the web by default. It only uses it on the browsing mode on Bing. So for us to be as fast as chat chepti already tells you that
Starting point is 00:20:22 in spite of doing more work to go pull up links from the web, read the chunks, pick the relevant ones, and use that to give you the answer with sources and log, more work on the rendering. Despite doing all the additional work, if you're managing an end-to-end latency as good as chat GPT, that shows we have even a superior back-in to them. So I'm most proud about the speed at which we can do things today compared to when we launched. The accuracy has been constantly going up. Primarily, a few things. One is we keep expanding our index and keep improving the quality of the index. From the beginning, we knew all the
Starting point is 00:20:58 mistakes that previous Google competitors did, which is obsess about the size of your index and focus less on the quality. So we decided from the beginning we would not obsess about the size. Size doesn't matter in an index, actually. What matters is the quality of your index. What kind of domains are important for AI chatbots and question answering and knowledge workers? That is what we cared about. So that decision ended up being right. The other thing that has helped us improve the accuracy was training these models to be focused on. on hallucinations. When you don't have enough information in the search snippets, try to just say, I don't know instead of making up things. LMs are conditioned to always be helpful. Always try to
Starting point is 00:21:40 serve the user's query despite what it has access to. May not be even sufficient to answer the query. So that part took some reprogramming, rewiring. You've got to go and change the weights. You can't just solve this with prompt engineering. So we have spent a lot of work on that. The other thing I'm really proud about is getting our own inference infrastructure. So when you have to move outside the Open AI models to serve your product, everybody thinks, oh, you just train a model to be as good as GPT and you're done. But reality is Open AI's mode is not just in the fact that they have trained the best models, but also that they have the most cost efficient, scalable infrastructure for serving this on a large-scale consumer product like Chan GPT. That is itself a separate layer of
Starting point is 00:22:23 mode you can build, tech mode you can build. And so we are very proud of our inference team, how fast, high throughput, low latency infrastructure we've built for serving our own LLMs. We took advantage of the open source revolution, Lama and Mistral, and took all these models, trained them to be very good at being great answer bots and serve them ourselves on GPU so that we get better margins on our product. So all these three layers, both in terms of speed through actual product backend orchestration, accuracy of the AI models and serving our own AI models. We've done a lot of work on all these things. Can you explain the early discussions that you and your team had about how to set up the business model? Because while the
Starting point is 00:23:06 10 blue links is fundamentally broken, user experience maybe from this point forward, it did play very nicely with an amazing ad-based business model. This is a different business model like chat GPT. It can use a version for free or you can pay for a better version. But I'm really really curious how you consider the various different business models. What were the business models that you almost did but didn't do? How do you think about the tradeoffs of the business model that you chose? This is all happening in real time. We're trying to figure out how to build companies and products around this new technology. I'd love to hear the early story of how you made those decisions. Honestly, I had no idea of business models when we were initially growing. Of course,
Starting point is 00:23:44 investors were all here growing, keep focused on the growth. Don't worry about making money. when you're growing, people are willing to support you to keep growing bigger and see how far it can grow before deciding what is the business model to stick to. So until we got to late hundreds of thousands of queries a day, we were not even concerned about making money. But at one point, we were really interested to know, hey, are you guys getting all this usage? Because people want to use some chat GPT alternative when it's down or some free GPT4 usage. add like 10 queries a day of GPD4 or something at one point. Are you guys just getting all this usage because of people not wanting to pay for Open AI or when Open AI is down, but you don't actually have real product market fit.
Starting point is 00:24:33 So that was a question that we were asking ourselves. And it was a valid question. So how to best answer this question? You create a subscription version of your product that has the same pricing as ChatGPT Plus. you charge it the exact same thing, $20 a month, and see how many people convert to paying users. They cannot be just paying for GPD4 because they're getting that in Open AI as well.
Starting point is 00:24:59 And they cannot just be paying for browsing alone because they're also getting that on Open AI. So despite that, why are they coming and paying for you? They come and pay for you because they like your product experience. They like what you offer. And so that needed to be known. Or else there's no point raising another round to scale this thing further up
Starting point is 00:25:17 and build them business. Because when you don't really have true product market fit and you're just a subsidy, then you shouldn't be raising more money. There's no real PMF there. And that was what we validated quickly. We put in a subscription plan. We saw how many people converted to paying users. We hardly even tried to convert them.
Starting point is 00:25:36 So despite that fact, a lot of people chose to convert and start paying and the business was growing really fast. We just said, okay, the subscription model works. It's not the ultimate model. I don't think that's the final piece in the profits that are going to be generated in the AI chatbot sector. But it's a good start. Everyone is doing that.
Starting point is 00:25:58 Open AI started it. We are doing it. Google is also trying that. Microsoft is trying that. So it's a good start where it's like something that, you know, potentially could be billion dollars in revenue for us. It's already billions of dollars in revenue for Open AI. Let's start there.
Starting point is 00:26:13 And also when you have sufficient scale of usage, try to think of what advertising. still. What does advertisement in this new medium looks like? How would it even work? How would you not compromise the quality of the answer? How would you not compromise the quality of the citations? Despite that, how can you help creators of content reach more consumers? So that's an interesting challenge to figure out.
Starting point is 00:26:35 So we will also work on that and expect the others to also keep thinking about these things. And the other business is like APIs. We have our own online LLM APIs, which is basically an, an LLM that has no knowledge cut off. So it's always live, up-to-date, real-time information, unlike the GPD APIs. And we're slowly expanding access to it and letting other people build on it. For example, Rabbit devices are using those APIs.
Starting point is 00:27:03 We're also partnering with other devices. Browsers like ARC are using those APIs. So it's small steps towards also building a developer or enterprise-focused version of the business, which can make use of all the infrastructure we built. similar to how we use opening eyes infrastructure. How does that work? So if you think about it in simple terms, opening eyes training this huge model up to a date,
Starting point is 00:27:26 and it's using information available up to that date to train the thing, that's the knowledge cut off, how do you continue to have something that valuable that is up to date, including today's dates to web pages? How does that actually get built? How do you build that? It's the same thing as a product. You ask a query, and it goes and pulls pages from our index,
Starting point is 00:27:45 and then uses signals from the web to rank it. and then gets you back the answer that uses knowledge from these snippets that it pulled up in the form of a concise paragraph. So it's whatever happens in the product, it's the same thing, except it's been black boxed to you as an API, and then you just send in your request and you get a completion. And you can make it a chat assistant too so that you can create products that enable this conversational answer engine experience. And people have built WhatsApp assistants using that. You ask a question, grab it, it can hit the API. and give you the answer.
Starting point is 00:28:19 So we can do a lot more. The reason our APIs are valuable is because nobody else offers this level of speed and accuracy for an end-to-end search plus LLM experience. Can you expand on index? You've referenced that a few times for those that haven't built one or haven't thought about this. Just explain that whole concept and the decisions that you've made. And you already mentioned quality versus size.
Starting point is 00:28:41 But just explain what it means to build an index, why it's so important, et cetera. Yeah. So what does an index mean? It's basically a copy of the web. The web has so many links, and you want a cache. You want a copy of all those links in a database, so a URL and the contents in that URL. Now, the challenge here is new links are being created every day on the web, and also existing links keep getting updated on the web as well.
Starting point is 00:29:10 New sites keep getting updated. So you've got to periodically refresh them. the URL needs to be updated in the cache with a different version of it. Similarly, you've got to keep adding new URLs to your index, which means you've got to build a crawler. And then how you store a URL, the contents in that URL, also matters. Not every page is native HTML anymore. The web is upgraded a lot, rendered in JavaScript a lot,
Starting point is 00:29:35 and every domain has custom ways to render the JavaScript. So you've got to build parsers. So you've got to build a crawler, indexer, parser, and that together makes up for a great index. Now the next step comes to retrieval, which is now that you have this index, every time you hit a query, which links do you use?
Starting point is 00:29:57 And which paragraphs in those links do you use? Now, that is the ranking problem. How do you figure out what is relevance, relevance and ranking? And once you retrieve those chunks, like the top few chunks relevant to a query that the user is asking, that's when the AI model comes in. So this is the retrieve part. Now the generate part.
Starting point is 00:30:16 That's why it's called retrieve and generate. So once you retrieve the relevant chunks from the huge index that you have, the AI model will come and read those chunks and then give you the answer. Doing this ensures that you don't have to keep training the AI model to be up to date. What you want the AI model to do is to be intelligent, to be a good reasoning model. Think about this. When you were a student, I'm sure you would have written an open book exam, open notes exam, school or high school or college, what are those exams test you for? They don't test you for
Starting point is 00:30:49 road learning. So it doesn't give an advantage to the person who has the best memory power. It gives advantage to the person who has read the concepts, can immediately query the right part of the notes, but the questions require you to think on the fly as well. That's what we want to design systems. It's very different philosophy from Open AI, where open AI wants this one model that's so intelligent, so smart, you can just ask it anything. and it's going to tell you, we rather want to build a small, efficient model that's smart, capable, can reason on facts that it's given on the fly. And this ambiguous, different individuals with different names.
Starting point is 00:31:24 Say if there's not sufficient information, not get confused about dates. When you're asking something about the future, say that it's not it happened. These sort of corner cases handle all of this with good reasoning capabilities, yet have access to all of the world's knowledge and an instant through a great index. And if you can do both of these together, end-to-end orchestrated with great latency and user experience, you're creating something extremely valuable. So that's what we want to build. If you think about the history of the business so far and every episode of what you've had to build, what stands out in your memory as the most difficult period or thing that was built? What would you least want to go back and live through again in terms of its difficulty and stress?
Starting point is 00:32:05 Well, I started working in deep learning in 2014, and we were not even doing deep learning in Python at the time. Everything was done with C++ and Kuta. I was using this framework in deep learning called CAFE, that literally where if you had to build a different architecture outside of the traditional CNNs, you had to go and write those layers in C++ and Kuda, recompile the library again, because everything's in C++ plus. it as we compiled. And after you get a compiled object, you write the new neural net using those layers, create a proto buff file, and just rewrite all the data layers again, and then launch shops. It was a nightmare. Most of the Kuda drivers would have to be reinstalled again for different GPU cards. There was no standardization. So you would probably spend hundreds of hours just
Starting point is 00:32:57 installing Kuda and installing these libraries, changing the layers, changing the libraries, that the amount of patience and willpower you needed to still do all this to succeed in your research was just crazy. But it's good. It's a good proxy to test if somebody is a good engineer or not, because usually people give up very fast. I didn't give up. And of course, life got a lot easier once Python-based symbolic languages came like Tiano from Montreal and then TensorFlow from Google. Intensur flow was a pain in the ass too because debugging it was really hard. Every time you got something wrong, you had to actually change the graph and not be able to print any intermediate things. And then PyTorch came.
Starting point is 00:33:39 This is the core AI deep learning stuff that was such a pain when we used to work with. And I would not want to go back to those days, honestly. How would you explain the feeling of the transformer coming online and what that was like to experience as an engineer? How would you explain what a transformer unlocked to a person? and that's less technical. Yeah, so what the transformer primarily did is it just made the description length of the architecture of a neural net so minimal. It's very homogenous architecture.
Starting point is 00:34:11 Until the transformer, you would have a recurrent layer, a convolutional layer, and a bunch of hidden layers. You would have to be sophisticated to know the right combinations of them. It's almost like you're cooking a meal, but you have to get the right mixtures of so many different parts that any mistakes anywhere could just cost you so much. What the transformer did is one simple model, which is two layers, attention, matmos, attention, matmos, alternating each other. That it's the same layer repeated again and again. You just had to decide three or four hyperparameters and that's it. So it just lowered the barrier to entry. You don't have to be a sophisticated
Starting point is 00:34:52 neural network expert to design architectures anymore. Instead, the work went more into getting the data right. The architecture problem was solved. You just got to literally scale it up in terms of layers, a number of hidden dimensions, but that's it. More work was spent on getting the tokenizer right, the data right, how the word is converted into the vocabulary, what parts of the internet is scraping,
Starting point is 00:35:19 quality of the data, which do you leave out, which do you train on? How do you ablate for what do you evaluate on in terms of how do you know the model is good. So that created the different set of skill sets who are more like physics PhDs, who had that rigor and experimentation, less background and ML, to come and have a big advantage right now.
Starting point is 00:35:42 And that is the core skill set of the Anthropic team. There's a company called Anthropic. They used to work at OpenEI. Their CEO, Dario Amodi, he's actually a physics guy, physics PhD. But he became incredibly, incredibly skillful leading teams like this because of his background. And he hired people like that. He hired people who were having physics background to come work with them. And they built
Starting point is 00:36:05 all GPT3 testing for Q capabilities, ablating clearly at a smaller scale, forecasting, scaling loss. They brought in this new discipline there. That's what has led to most of the breakthroughs that we see in Chachapit and all the stuff. Like nobody launches a hundred million dollar run, YOLO. You cannot do that. It's most likely going to fail. It's not like how people on Twitter talk, well, why is Google not doing this? Why are they not taking all the data that they have and launching a huge model and just destroying opening eye? Because you cannot do that.
Starting point is 00:36:36 If you just put a lot of data and model is going to be confused. It's going to look at so much that it's not going to learn any one thing properly. So there is a science towards figuring out the right data mixes at smaller scale, forecasting what will happen if you scale it up, and then rigorously launching larger and larger runs. And I believe that it's less about being a great transformer designer and more about being a great data expert and experimenter. Do you think that the transformer architecture is here to stay and will remain the dominant tool or architecture for a long time? This is a question that everybody asks in the last six years or seven years since Sederan's transformer came. Honestly,
Starting point is 00:37:20 nothing has changed. The only thing that has changed is the transformer became a mixture of experts model where there are multiple models and not just a single model, but the core self-attention model architecture has not changed. And people say there are shortcomings, the quadratic attention, complexities there, but any solution to that incurs costs from where else to. Most of the people are not aware that majority of the computation in a large transformer, like GPT3 or 4, is not even spent on the attention layer. it's actually spent on the matrix multiplies.
Starting point is 00:37:57 So if you're trying to focus more on the quadratic part, you're incurring cost of the matrix multiplies, and that's actually the bottleneck in the larger scaling. So honestly, it's very hard to make an innovation on the transformer that can have a material impact at the level of GPD4, complex cost of training those models. So I would bet more on innovations auxiliary layers, like retrieval augmented generation.
Starting point is 00:38:21 Why do you want to train a really large model? when you don't have to memorize all the facts on the internet, when you literally have to just be a good reasoning model, nobody's going to value Patrick for knowing all facts. They're going to value you for being an intelligent person, fluid intelligence. If I give you something very new that nobody else has an experience in, are you well positioned to learn that skill fast and start doing it really well? When you hire a new employee, what do you care about?
Starting point is 00:38:47 Do you care about how much they know about something? Or do you care about whether you can give them any task and they would still get up to speed and do it. Which employee would you value more? So that's the sort of intelligence that we should bake into these models and that requires you to think more on the data. What are these models training on? Can we make them train on something else? Just memorizing all the words on the internet. Can we make reasoning emerge in these models through a different way? And that might not need innovation on the transformer. That might need innovation more on what data you're throwing at these models. Similarly, another layer of innovation that's waiting to happen is the architecture,
Starting point is 00:39:22 sparse versus dense models. Clearly, a mixture of experts is working. GPD4 is a mixture of experts, mixtrall is a mixture of experts, Gemini 1.5 is a mixture of experts. So even there, it's not one model for coding, one model for reasoning and math, one model
Starting point is 00:39:39 for history. That depending on your input, it's getting routed to the right model. It's not that sparse. Every individual token is routed to a different model, but it's happening every layer. So it's still spending a lot of compute. How can we create something that's actually 100 humans in one company. So the company itself as an aggregate is so
Starting point is 00:39:58 much smarter. We're not created the equivalent item model layer. More experimentation on the sparsity and more experimentation on how we can make reasoning emerge in a different way is likely to have a lot more impact than thinking about what is the next transformer. I'm curious then to think about bottlenecks in two ways. So bottlenecks specific to perplexity and what it wants to build and your perception of what the bottlenecks are in AI writ large. If you had to answer for both, what do you think the number one bottleneck to progress is in both those cases? I would say for us perplexity, the main bottlenecks today is just getting reasoning to emerge in smaller models. If that happens, the cost per query is just going to go down tremendously.
Starting point is 00:40:43 If you don't need GPT4 for being accurate, let me give you a rough statistic. It's not actually rigorous. Let's say a model like GPD 3.5 or a mixed ral fine-tuned version of that that matches 3.5 gets 8 out of 10 queries or something like 7 out of 10 queries, no hallucinations. And GPD 4 will get 99 or 100. The accuracy rate is so much better in the long tail. Now, if I can get a 3.5 or mixed trial model to be as good as 4 at hallucinations, which is basically connected to reasoning capabilities. When you don't have enough information, just say no, deduce it, then that makes a tremendous impact. I no longer need to serve a large model anymore, and the service can run way more profitably. That sort of a skill is lacking in smaller models, and that connects to the first
Starting point is 00:41:32 point I made about making how do you train models in a different way so that the most reasoning capable model shouldn't necessarily be the largest model. That hasn't happened yet. And if that correlation breaks, I think it will have a huge impact. The other important, impactful scenario in general for the field, not just specific to perplexity, is synthetic data. What happens when all the data on the internet is saturated? We've trained on all of it, that every new data set that's being created on the internet doesn't add a lot of value to. It's very marginal. How can you make these models create the next generation of the data for themselves and recursively improve?
Starting point is 00:42:12 I'm not talking about scenarios where these models are going to go rogue and start thinking for themselves and take humanity or something. Very simple experiment where GPT5 were designed by GPD4 largely instead of human annotators. Now, this will have an impact because I spend a lot on human annotation for hallucinations. I don't have to. I can have the model do it. I can have a smarter model do it for me. So I can spend less and get data annotated faster. I can make improvements on my core models that are on production much faster because an AI can look at million queries, figure out what's wrong, annotate what is wrong, and tell my smaller AI model to train on them. And I can finish the training run in a week instead of doing it over a month.
Starting point is 00:42:58 That way my users get to feel the product better much faster to accumulate more users that way. And I get more data. The improvements on the product can be tremendously faster. So I think both of these will have a lot of impact, just not just on ours, any other startup, synthetic data and reasoning and smaller models. You talked about how in search, speed, latency, and accuracy are obviously two things that you focused on a ton.
Starting point is 00:43:23 I'd love to talk about when you think they'll become customized, more context aware of who I am relative to the next perplexity user, how that happens, and also when they become more agentic, when they can actually start doing stuff for me. Because if you think about the broken 10 links, you could argue the answer engine's broken too. The end of the day, I just want to have an idea and have an action happen. and the end of that idea that I don't have to do. How do you think those next two components of context awareness and agent behavior might start
Starting point is 00:43:54 to find their way into models like yours? First, let's start with context awareness. Maybe we call it personalization, more hyper-personalized versions of perplexity. Let's achieve very simple things. Location, gender, age will already solve a lot of personalization for you, for what's worth. People think you need to literally put all of your activity on the prompt. For what it's worth, these models get confused when you throw a lot of information at them. People advertise long context a lot, but the more you throw at these models in the context, the more confused they get
Starting point is 00:44:29 in terms of what to focus on. So personalization can be done when you know exactly what to retrieve from your past and focus more on the highest order bits like location, gender, and demographics. and create a much better experience that's more catered to you than the average user. I think we can already do this this year, and we focus on doing that. The second part, agentic versions of perplexity, we believe that's likely to happen very fast. Let's say there's three parts towards taking an action. You do your research, you make your decision, and then you take your action. We are doing the first part pretty well, allowing you to do your research.
Starting point is 00:45:09 The decision you're still exercising, you want to decide. And the final part, the action, you just want to task the AI, as if it was your executive assistant. That's what you want to do today. Now, you can go a step further and say, I don't even want to do decision. Let the AI decide everything for me. Let the AI do the research for me.
Starting point is 00:45:27 Let the AI act for me. I just want to have no agency. I just want a student chill, watch Netflix all the time. And AI is like, will work for me. Have meals show up for me. So I think the second part seems more dystopian. I don't want that to happen, though if people want that to happen, it likely happen. I think the first part we can work towards that once we have models that are better
Starting point is 00:45:50 reasoners than GPT4, today I can confidently claim that GPT4 is not there yet to be a good action bot. That is the biggest reason why the GPD plugin store failed because it cannot handle all these different APIs calls together at once and therefore it didn't work. when you want an action bot to work, it needs to chain a lot of decisions together and handle corner cases. Now, why do we still need executive assistance? The reason you need them is because sometimes you're scheduling something with somebody and they might not have availability for the availability you have and you might want to move some things around because you might want
Starting point is 00:46:29 that to happen that week itself. These kind of thinking and corner cases, you want to be able to handle And you don't want the AI to keep coming. Hey, Patrick, that guy's busy. What about this? You want your assistant to think and act on your behalf so that you're able to focus on other things. Now, we don't have the AIs that can do this today. GPD4 cannot do this today. Maybe 4.5 can.
Starting point is 00:46:51 Maybe 5 can. I don't know. When that model is available, definitely we will also add more agentic experiences in our product. We are not capable of training those models today. We don't have the budget. We don't have the compute and the talent to do that. I think somebody has to show the proof of existence of such models, and then we can figure out how to get there.
Starting point is 00:47:11 And until then, we have our jobs right in front of us to just reduce hallucinations and improve the research part. You mentioned the word talent there, which feels like such an important topic to cover. What is the talent in AI right now as a leader of a business that obviously wants to and needs to recruit awesome talent? This is certainly the most actively fast-moving, exciting area in technology today. It's attracting lots of really talented people, but I'm sure you're a supply constraint. There's not enough great AI talent out there.
Starting point is 00:47:41 So yeah, just describe the talent where what's it like. How do you participate in it? Anything you can share would be really interesting. Yeah, I would say that if you want to compete on pre-training, a large model, like GPT4, Anthropic, Claude, Mistral, Lama, Gemini. That's it. Maybe Elon's XAI. A lot of potential, but not done anything major there yet. But that's it. That's over. Everybody who has some jobs at doing large scale training, a lot of rigorous scaling law analysis, data experimentation, are working in these six companies today. Competing for building a competitor and a number seven, you just not only have to raise a lot of
Starting point is 00:48:28 money and give these people a big cluster, you also have to hire the talent away from them, zero sum at this point. And people don't want to leave because when you don't have anything, when they have peers to work with, and when they already have great experimentation stack and existing models to bootstrap from, for somebody to leave, it's a lot of work. You have to offer such amazing incentives and immediate availability of compute. And we're not talking of small compute clusters here. I tried to hire somebody from META, very senior researcher. And you know what they said? They said, come back to me when you have 10,000 H-100s. And you know what? 10,000 H-1-1-billions of dollars over a five to 10-year period. Why would I have it? And then how do I create it?
Starting point is 00:49:14 Also, by the way, it's not just about having the money to buy these clusters. You need to make it available today. And there is a supply chain problem. Most of the GPUs are getting booked out one year in advance or two years in advance. So even if you have the money, you have to wait. By the time you waited and got the money and booked the cluster and got it, the guys that are working here have already made the next generation model. And they're like, look, the world has changed. I'm already in the next generation. I'll come when the next version of the model is finished training. This time you'll come back to me when you have 20,000-8-100. So it's a game that is being played by seven people right now, six people. Largely four.
Starting point is 00:49:55 four or five, I would say. My hope is that it gets commoditized. It's not a hope in vain. There is some good reasons to why it could happen. AWS, Azure, GCP, all have incentive to commoditize this and be the biggest winners of all these models, actually, more than Open AI or Anthropic or Mistral. So the cloud service providers, and Nvidia, of course,
Starting point is 00:50:22 all of them have incentive to commoditize these models. and how so many other businesses make use of these models and deliver a lot of profits than just a few people eating all the profits. That way they get to win the most. So my hope is that since the cloud service providers are the ones bankrolling all these companies directly or indirectly, they will put it on their clouds for enterprises, and then enterprises can take them and post-trained that. I'm talking about a different kind of training.
Starting point is 00:50:53 The first type of training I've talked about is pre-training. But pre-training alone is not enough. You cannot take a model and just put it into a product and do nothing. It'll still fail at a lot of consumer use cases, customer use cases. You have to post-train them and address the long tail of issues you get on serving a product. Now, there you actually have a huge advantage if you have a lot of users because you have the data fly wheel. So if you have a lot of users and establish the data fly wheel and establish all the tooling and evals to constantly make use of your existing data sets to improve your product
Starting point is 00:51:28 and build that machine, that flywheel machine and accumulate enough users and a brand, you have tremendous advantage to create a lot of value. And we are focused on that. And that talent doesn't need to be as sophisticated or scars as the pre-training talent. This talent, you can hire people who want to get into AI from other industries, like crypto or like e-commerce and teach them. And there are so many resources out there. They are fast learners.
Starting point is 00:51:59 They even can learn themselves, and they can add a lot of value that way. Proplexity is one of the companies doing that. I'm sure there'll be many more companies there. If you were to oversimplify it and ask the investor community, what is stopping them from funding more AI application companies, not infrastructure, not LLMs, companies more like yours, I think they would say, well, we're worried.
Starting point is 00:52:21 about defensibility. We don't know how these companies control their margins and control churn, et cetera, over time if they're very reliant on underlying infrastructure. I'm not saying this is my opinion. I'm just saying a common take. What would you say to the people that want to build AI applications as they think about their business model, their defensibility, their sustainable differentiation? What advice would you give to other entrepreneurs that want to be successful building apps using AI? I would say that this is going to be an issue until you're a monopoly in your sector. Honestly, I've thought about this too, and I've expressed my frustration to a lot of people.
Starting point is 00:53:00 I'm getting asked the same question that I was asked when I had 10x fewer users. When this is ever going to stop? Will it stop when I have another 10x more? And the answer that guy gave a pretty successful billionaire entrepreneur was, no, you'll still be asked, you'll continue to be asked. so get used to this and you'll only stop getting asked when you are the number one. And then what you'll be asked is, oh, look at these smaller guys trying to do the same thing. When are you going to crush them?
Starting point is 00:53:30 So it's always going to be a thing. I would say the best answer is always, this is a hard market. There are big players, but what we are offering is this. Look at what users are saying. Trust the users, that's a signal. it is not unfair for investors to feel like they might be making a mistake. Look at what happened in the SORA text video release that Open AI did. Until then, the runway ML and PECO were every investor wanted a piece in those companies.
Starting point is 00:54:01 Now like, oh, what should I do? Because until then, their mindset was perplexity is such a dumb investment to make because they are directly in the text interface. and OpenAI is not focused on alternate modalities like voice or video, and therefore I'll go and invest and start doing that. And now does that reasoning apply anymore? I would rather do Enterprise Chad GPD as an investment because Open AI doesn't care about enterprise,
Starting point is 00:54:28 and then they come and release chat GPD for enterprise. Everything has competition, and I think at one point you just got to realize that for one company to be doing so many projects at once like Open AI is doing today, definitely not all of them are going to succeed. Even if they do succeed, their success doesn't mean your failure. They are trying to create as much value for themselves. Same thing with Microsoft, same thing with Google. For you, the only thing you can focus on is fast execution and be the best in your sector.
Starting point is 00:54:58 If those two things are not true anymore, if someone else is kicking your ass there, then you have to be worried. But that is true, anything you do. You're always going to have competition on anything that can create billions of dollars in revenue. Because for these guys, billion dollars in revenue is what they care about. If you're focused on building something that is only going to be hundreds of millions of dollars in revenue, you are just focusing on building a billion dollar company. Then you don't take as much venture funding. Do it yourself.
Starting point is 00:55:24 Try to be bootstrap. Like how the Mid-Journey guy is doing it. There's a great quote that I've seen you post, which is the successful warrior is the average man with laser-like focus, which is a great Bruce Lee quote. It sums up everything you just said. There's one more quote, by the way. I don't fear the man who practiced 10,000 kicks once. I fear the man who practiced one kick 10,000 times. And that's what you're trying to do.
Starting point is 00:55:45 If you have only one thing to protect, you go out of your way to be the best at it. If you apply that quote to your own experience, the laser-like focus piece, what have you not done in order to stay focused? What have been some things that you maybe would love to have tried or gone and done and maybe you'll do them in the future, but things that you've actively said no to in order to
Starting point is 00:56:06 stay on that laser-like focus? Well, we never did image generation. We have image generation as part of the answer. Somebody wants to enrich the answer, but not as a chat experience where someone can just type in a prompt and get an image. And then we've never done free-form chat. Everybody said, why don't you just support both modes like Bing does? Dude, I'm just trying to build a search product here.
Starting point is 00:56:33 Everyone's like, AI is not meant for information, man. Hallucination is the feature. You should build products where hallucination is a feature, not a bug. When you're building AI chatbots for search, hallucination is a bug. So you're doomed. You should take advantage of the hallucination being a feature. You should build something like character AI. You should build a GPT store.
Starting point is 00:56:54 You should have a travel perplexity, like health or like shopping. You should have a store, how it looks on character AI. You should go click on one agent and talk to it. and you should allow people to create their own agents to. And you should work on agents. You should allow people to book restaurants. You should go enterprise. You should allow internal search.
Starting point is 00:57:12 Honestly, I give a lot of credit to one of my co-founder's Johnny Ho. He's actually running our product division, basically. And he was a competitive programmer and a very successful quantitative algorithmic trader, too, before starting perplexity. I would give him even more credit at saying no to things than even myself. Because sometimes I'm also tempted. I have that founder experimentation energy. But the good thing is I'm not very arrogant.
Starting point is 00:57:44 So when somebody that's done a lot more thinking about it than me says, no, we should not do this, tend to trust their advice. And of course, there are sometimes I'm not trust, despite the saying, we have to go and do this thing. and it's happened. But most of the time, I do listen to the Steve Jobs code of, I'm as proud of saying the things I said no to
Starting point is 00:58:07 as I am of the things that I chose to do. If you think about the future now and where this all might go, I'd love you to paint the biggest possible picture for us in search specifically, not for AI at large, could go any direction. But for you and what you want to build
Starting point is 00:58:26 or in 2030 or whatever, pick your date. What gets you the most excited? What potential future states get you the most excited? I would say that disrupting search categories where we are currently clicking on a lot of links and having a lot of commercial intent there would be insanely amazing. And I think it's possible to do that.
Starting point is 00:58:49 We will be working hard on that. So right now, perplexity is associated with this amazing knowledge assistant that you use for fact checks, trivia, learning about things, digging deeper, but more like a research buddy, that experience needs to expand to the average consumer search categories, shopping and travel and insurance and legal and medicine, where there's huge amount of dollars of advertising thrown at Google for these categories, that Google has zero incentive to make these categories good on Gemini, even if they want to. And that's what we want to go and disrupt.
Starting point is 00:59:24 What gets you the most worried? I would be lying if I said I wasn't worried of opening. I was trying to do similar things to us. Bravado and all that is great, but let's be honest. I'm definitely worried about Chad GPT, trying to go more in the direction of search. And the only thing that we have going for us is our speed and accuracy in our U.S. That is very much better than what they have, at least regarded by many users that way, even if it's not my opinion.
Starting point is 00:59:51 If they chose to focus more, it's going to be more competition there. In that case, the differentiation is going to come from us executing even faster and better. Because unlike a big company, if you call them a big company, they're actually pretty fast. So we have to be even faster. So that is one thing I do think about. Unlike what most people say, I'm not really worried about Google. Not because they cannot execute on this. They actually are way better engineers and researchers than us.
Starting point is 01:00:17 It is their own business model. Yeah, it's a counter positioning thing. Yeah. You've just raised around from some well-known investors and individuals. So I'm sure you have talked to lots of investors that invest in this sort of stuff. If you think about all of those experiences, what do you think investors understand best and least about your kind of company right now? I think what they don't understand, at least a lot of them, is how hard it is to actually create a product like this. Their mental models are all in the vertical SaaS era.
Starting point is 01:00:51 where they found companies that went more verticalized and found customer lock-in effects and then succeeded as a business. What they failed to understand is verticalization might actually be the wrong strategy in AI because one generic bot that can do many things is a lot more valuable to the end user. When people think of AIs,
Starting point is 01:01:16 they think of the most generic way of interacting, which is natural language-powered. Whereas when you are going, going verticalized, there are always going to be certain query categories you cannot handle. But you cannot instruct a human user to only interact in a constrained way. You have to design the product that way. It takes a lot of product design and verticalization on the product layer to still let humans interact in the most free-form way, yet have a constraint experience.
Starting point is 01:01:44 Nobody has succeeded at this. And further, people don't realize that incumbents in the verticals can just sprinkle an AI chatbot within that app and have all the other layers like product depth and other support like databases, existing vendors on that one single platform already going for them. Most of the things are handled. So that is one thing that took me a lot of time to explain to people in the beginning. Only one particular investor said this to me without even me having to explain this. Mark and Reason.
Starting point is 01:02:15 Mark and Drickson talked to me in January, 2023, and he said, all I'll tell you is when Google came out, there was so much of investor frenzy in investing in Google for a vertical. You don't know any of those companies today because they all shut down, but they got a lot of funding. So don't do that mistake. Everybody's going to tell you to make perplexity a vertical so that in their mind, they're investing in something safe. But if you do that, you're doomed. You better go all the way and go all in,
Starting point is 01:02:54 or you just don't do this company. I was like, damn, finally one guy said exactly what I was thinking and happens to be the guy who pioneered the browser. So I got a lot of courage from his advice today. From the outside, it's so interesting to watch a company like yours just as a user and it's been a blast to use it and see it progress. Is there anything else that we haven't talked about that you think would be the most surprising to the non-builders out there, whether that's investors or users or people that aren't actively building these products themselves? Is there anything surprising about the state of things today that you think we should talk about that we haven't?
Starting point is 01:03:33 This is not surprising if you put more thought into this, but a lot of people are worried about competition and modes and sharing they don't get destroyed by opening. AI. Some amount of thought there is very useful. I'm not discarding that at all. But if that is this only way you make your decisions on what to do for your company or your product, you're likely to fail because these guys will do everything under the sun if it's actually valuable worth doing. So as a startup, your only job is to figure out if there's something you can do that has not been done yet already and that can deliver value to the world. That is truth. Startups are all about finding a truth vector. If it is truth and if you can generate value, there's no reason existing players don't want to do the same thing too. They will also try. The way you establish
Starting point is 01:04:26 your modes continue to execute faster, make sure that it's not easy for the other companies to do because it takes a lot of non-trivial work and keep going. On the other hand, if you're like, hey, I want to build a sales AI co-pilot because Open AI is not going to go stat vertical and Google doesn't care, Microsoft doesn't care. The Salesforce would do it. It's not hard. These are things that HubSpot would do it. So just don't be so naive in the way you think about strategy and modes.
Starting point is 01:04:55 And don't spend so much time in the beginning of the company thinking about strategy. Try to iterate. We made a lot of mistakes. We built sex with SQL. It's the dumbest idea you can work on, honestly. Because nobody even writes SQL. 80% of the SQL that actually makes money for Snowflake or Databricks is not even being written. It's just power BI generated or already written queries that are constantly periodically running.
Starting point is 01:05:19 When you start thinking about these, charging someone based on consumption rather than writing new queries, Texas SQL is such a bad idea because once the same sequel is being written, you're not making any money out of it. So usually when you try to think on white paper a great idea or a whiteboard, it doesn't happen. Very few people that build companies of that nature. And I would say probably all the existing big players were all built with iteration and trying out things and stumbling upon something awesome and then building the strategy around it. Build the execution muscle first. Don't try to be a great strategist right away. Build something.
Starting point is 01:05:57 Make sure it has some traction. Get the muscle that you can keep iterating. And then you deserve the right to strategize. This is actually been said by this other guy called Frank Slutman as well. the Snowflake CEO, he has this whole line in his book, Amped Up, where you only deserve the right to strategize once you have earned the track record of execution. And I strongly believe in that. Really, really interesting closing thought. I'm so appreciative of you letting us behind the scenes here into building one of these things very actively. I'm sure it's been stressful
Starting point is 01:06:30 and all-consuming and very fun and very interesting. I ask everyone that I interviewed the same traditional closing question. What is the kindest thing that anyone's ever done for you? We were going through the SVB incident. You remember? The bank collapse. Yes. I was supposed to be on vacation that weekend and I was not. Yeah, we were going through that and I was very stressed and a lot of people checked on me during the time.
Starting point is 01:06:53 And Nat Friedman just came and said, I'll give you the money. Don't worry. I'll take care of your payroll. That was very nice of him to do that at that time. Him and Daniel obviously have done an amazing job of supporting this ecosystem. Pretty cool. Yeah, exactly. I had no idea because I was actually in Redmond at the time doing some company visit and
Starting point is 01:07:12 this whole thing was going on. I was in the airport. Before I could get on to the flight, I tried to wire the money out to my personal account, in fact, because I didn't even have another account. And I checked with my lawyer, this is okay. He said, this is very emergency. Just get the money out somewhere, man. It doesn't matter.
Starting point is 01:07:27 And Nat and Dan will like, don't worry. Even if that money goes out, we'll fund you. Amazing. Well, thank you so much for your time and for a great conversation. Thank you, Pat. If you enjoy this episode, check out join colossus.com. There you'll find every episode of this podcast complete with transcripts, show notes, and resources to keep learning. You can also sign up for our newsletter, Colossus Weekly, where we condense episodes to the big ideas, quotations, and more, as well as share the best content we find on the internet every week.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.