Invest Like the Best with Patrick O'Shaughnessy - Gavin Baker - AI, Semiconductors, and the Robotic Frontier - [Invest Like the Best, EP.385]

Episode Date: August 27, 2024

My guest this week is Gavin Baker. Gavin is the managing partner and CIO of Atreides Management, and he has been on the show many times before. He is one of my favorite investors to talk to and this m...ay be my favorite conversation with him. Gavin first started covering Nvidia as an investor at the turn of the millennium, making him the perfect guest to discuss all things AI and investing. There is so much detail in this discussion and I’m incredibly grateful to Gavin for sharing his wisdom with us again. Please enjoy this fantastic conversation with Gavin Baker. For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- This episode is brought to you by Ramp. Ramp’s mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Ramp is the fastest-growing FinTech company in history and it’s backed by more of my favorite past guests (at least 16 of them!) than probably any other company I’m aware of. It’s also notable that many best-in-class businesses use Ramp—companies like Airbnb, Anduril, and Shopify, as well as investors like Sequoia Capital and Vista Equity. They use Ramp to manage their spending, automate tedious financial processes, and reinvest saved dollars and hours into growth. At Colossus and Positive Sum, we use Ramp for exactly the same reason. Go to Ramp.com/invest to sign up for free and get a $250 welcome bonus. ----- This episode is brought to you by Tegus, where we're changing the game in investment research. Step away from outdated, inefficient methods and into the future with our platform, proudly hosting over 100,000 transcripts – with over 25,000 transcripts added just this year alone. Our platform grows eight times faster and adds twice as much monthly content as our competitors, putting us at the forefront of the industry. Plus, with 75% of private market transcripts available exclusively on Tegus, we offer insights you simply can't find elsewhere. See the difference a vast, quality-driven transcript library makes. Unlock your free trial at tegus.com/patrick. ----- Invest Like the Best is a property of Colossus, LLC. For more episodes of Invest Like the Best, visit joincolossus.com/episodes.  Past guests include Tobi Lutke, Kevin Systrom, Mike Krieger, John Collison, Kat Cole, Marc Andreessen, Matthew Ball, Bill Gurley, Anu Hariharan, Ben Thompson, and many more. Stay up to date on all our podcasts by signing up to Colossus Weekly, our quick dive every Sunday highlighting the top business and investing concepts from our podcasts and the best of what we read that week. Sign up here. Follow us on Twitter: @patrick_oshag | @JoinColossus Editing and post-production work for this episode was provided by The Podcast Consultant (https://thepodcastconsultant.com). Show Notes: (00:00:00) Welcome to Invest Like the Best (00:04:42) The Magnificent Seven and Tech Competition (00:06:29) Generative AI and Scaling Laws (00:08:36) Challenges in AI Infrastructure (00:15:02) The Future of AI and Data Centers (00:17:51) Efficiency in AI Models (00:35:14) Synthetic Data and AI Training (00:42:37) Inference and the Role of Smartphones (00:48:35) Investment Implications in AI (00:49:09) Opportunities for New Companies (00:51:20) Challenges at the Application Layer (00:52:25) AI's Impact on Advertising (00:53:40) AI ROI Debate (00:54:39) SaaS Metrics and AI Disruption (00:55:59) AI-First Application Companies (01:00:50) The Future of Robotics (01:14:01) Leadership in Tech Giants (01:24:05) The Evolution of Investing

Transcript
Discussion (0)
Starting point is 00:00:02 Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest Like the Best. This show is an open-ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. Invest Like the Best is part of the Colossus family of podcasts, and you can access all our podcasts, including edited transcripts, show notes, and other resources to keep learning at join colossus.com. Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of positive sum. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions.
Starting point is 00:00:45 Clients of positive sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc. My guest this week is Gavin Baker. Gavin is the managing partner and CIO of a tradies management, and he has been on the show many times before. He is one of my favorite investors to talk to, and this may be my favorite conversation with him. Gavin first started covering NVIDIA as an investor at the turn of the century, making him the perfect guest to discuss all things AI and investing. There's so much detail in this conversation, and I'm incredibly grateful to Gavin for sharing his wisdom with us again.
Starting point is 00:01:23 Please enjoy this fantastic conversation with Gavin Baker. All right, Gavin, you and I have actually not done this in years, even though we do it offline much more frequently than that. So I've really excited to do it again on the record. I have a list of 50 things I want to talk you about, so we'll see how many we get to. But a really fun opening framing was something I saw you put out into the world, which was back in 1960, the Magnificent Seven were led by Yule Brenner, and ultimately only three of the seven survived the shootout. We've got a new Mag 7 today, and we could probably spend the whole time talking about them. We won't, but I thought it would be a fun opening moment, just to hear you riff on why these massive companies might actually be in some form of business
Starting point is 00:02:09 shootout now with a lot of what's happening in the world of technology. So I think these companies were all in their own discrete swim lanes for a long time, competitive swim lanes. The only place where they really overlapped was cloud computing, where you had Google, Amazon, and Microsoft all competing, but that was a very stable oligopoly. Google cut prices aggressively, something like 2014, 2015, Amazon matched, and that actually materially impaired Amazon's revenue growth and fed into all of these fears that cloud computing was going to be a commodity,
Starting point is 00:02:46 which was a real fear, hotly debated topic, looks deeply ridiculous now. I think that set the stage for like, hey, we're effectively going to agree between the three of us on some markup on cost, but then try and differentiate in other ways. Amazon had e-commerce. Facebook had advertising that was higher up the funnel from Google. Google had search. Apple obviously had the device in the OS and the app store, minor competition from Android.
Starting point is 00:03:14 Netflix is doing streaming video. And again, a little bit of competition with Google there maybe. But they're pretty distinct competitive sets. And then Microsoft, while they had cloud, they also had this massive enterprise software business really focused on productivity. And just with gin AI, with LLMs, to the name, I mean, Gen. A.I. means generative AI, but the G and GPT means general purpose. It's such a general technology that they're all of a sudden in the same swim lane and they all feel like it's existential. Mark Zuckerberg, Satya, and Sundar, just told you in different ways. We are not even thinking about ROI.
Starting point is 00:03:54 And the reason they said that is because the people who actually control these companies, you know, the founders, there's either super voting stock or significant influence in the case of Microsoft, believe they're in a race to create a digital god. And if you create that first digital god, we can debate whether it's tens of trillions or hundreds of trillions in value. And we can debate whether or not that's ridiculous. But that is what they believe. and they believe that if they lose that race, losing the race is an existential threat to the company.
Starting point is 00:04:31 So Larry Page has evidently said internally Google many times, I am willing to go bankrupt rather than lose this race. So everybody's really focused on this R.O.I equation, but the people making the decisions are not because they so strongly believe that scaling laws will continue. And there's a big debate over whether these emergent properties are just in context learning, et cetera, et cetera. But they believe scaling laws are going to continue. The models are going to get better and more capable, better at reasoning. And because they have that belief, they're going to spin until I think there is irrefutable evidence that scaling laws are slowing. And the only way you get irrefutable evidence, why has progress slowed down since GPT4 came up?
Starting point is 00:05:20 because there hasn't been a new generation of Nvidia GPUs. It's what you need to enable that next real step function change and capability. It is actually going to be interesting. So everyone, the Blackwell delay plays into this. It's really, really hard to create what's called the coherent trading cluster of 10 to thousands of GPUs. And coherent just means each GPU, we can say, knows what the other is thinking. Technically, it's more like they have a shared memory space.
Starting point is 00:05:48 and that cluster has to be coherent to train. And the biggest coherent cluster in the world until very recently was 32,000. So you had 32,000 H-100s, probably only using like at most 15,000, 16,000 of those H-100s at once because of efficiency problems. But X-AI decided they were to build 100,000 GPU cluster. And I would say with Elon's unique physical engineering mind, which we've seen play out at SpaceX and Tesla, And now in NeurLink, where he figured out how to miniaturize everything, along with really capable teams, I think he re-architected the XAI data center from first principles in Memphis.
Starting point is 00:06:30 It's very different than other data centers. And because of that, they're able to get enough density that they can effectively make a 100,000 GPU cluster coherent, even with Hopper, without next generation networking technologies from Nvidia, Broadcom, and others. and they've started trading on that, and that means that I think we're going to see, probably, you know, scaling long as old, the first GPT, four and a half class model,
Starting point is 00:06:56 sometime, I don't know, in the next six, nine months, I don't know the timing. And then after that, you'll have Blackwell, and then you'll go up to 300,000 GPUs in a cluster because you have next generation networking technologies that make that easier, and that is going to be a massive step function change. and then that's where you can get maybe a GPD5, 5 and a half, 6th class model.
Starting point is 00:07:19 And then the reason it's existential, it's just like if Gid AI is so generalizable and you have this ASI, you're like, what's the value of content in a world where AI can make content that's infinitely better than any human? What's the value of Netflix in a world where I can say, I want to watch a mashup of Star Trek and Star Wars tonight? Search might just go away and be replaced by agents. It's important to have a lot of humility about, this, Chase Coleman, who runs Tiger, is an exceptional investor, had a really interesting statistic,
Starting point is 00:07:50 which was chat GPT came out in 2022 and was to AI as Netscape Navigator, was to the internet in 1994. Only less than 1% of current global internet market cap was founded, founded in the two years after Netscape Navigator came out. And nobody could imagine the companies that were going to be founded. So it's just the biggest companies were founded many years later or several years later, five or six years later. So just we're very early, important to have a lot of humility, but as long as scaling laws continue, and the only way they can be disproven is if you have a new generation GPU that has the traditional, whatever it is, three to five X performance improvement, and then you get better networking such that you can link three to five X more of them together.
Starting point is 00:08:38 and then you get that 10x, the order of magnitude improvement and compute, and you don't see a massive improvement in model quality, then scaling laws will stop and that'll be a catastrophe for the entire data center infrastructure. But the people close to this all believe that scaling laws are going to continue. Last night, my daughter on the way home from dinner said to me, hey, can you ask chat GPT some question? And I said, do you have chat TPT on your iPad? She said, no, not yet, you have to get it for me. I said, well, just Google it. She goes, what's Google? Wow. How old is she? She goes, Google's not a real thing. That's what she said. They just used chat TBT for everything. My son uses it for everything. She's eight, he's 10.
Starting point is 00:09:22 Wow. It's just fascinating to think about, wait, what? What did you just say? And they use it all day. They use Dolly all day every day. Search in their mind is chat chvety, which is just a fascinating thing for kids that could use both if they wanted to. And rather than a search engine, they just want the answer engine. And I just find that completely fascinating. Absolutely. What does it mean for humanity? The open internet has been really, really good. And yes, we have these walls, content gardens now, but there's still like a very robust, vibrant open internet where when you Google, you do get divergent opinions, generally, you know, even though we're all in these well-known filter bubbles. But if you just get one answer, I should think we're at an important moment. It's very rare that
Starting point is 00:10:08 To me, investors can really, really contribute to the world. We contribute an aggregate in an indirect way, in a massive way, because we fund all these new technologies that are solving cancer, indirectly funding AI. But I think it's rare that you can have a more direct impact as opposed to kind of a diffuse impact. I think it is supremely important for humans that we do not end up in a world where there is just one dominant model.
Starting point is 00:10:38 That is the most dystopian future I can imagine. Because that model then, if kids all over the world, follow your kids' behavior, whatever values that model has will be imbued to the rest of humanity. And I think we are at a time in history where, for a lot of reasons, the idea of objective truth is under attack. For a lot of younger people, their feelings are facts. the idea of science and objective truth is really under attack. And I think it's really important to have AIs that are dedicated to the idea of objective truth, no matter how unpopular that may be.
Starting point is 00:11:18 And I think the best way to accomplish that is to have them compete. You only have one dominant AI. That's very scary. And forget 1984. Hello, who knows what kind of future. But whereas if you have three, four, five of these, that's very different. They'll compete. They'll have different value systems. And I think we can, as investors, by lowering and improving the cost of AI infrastructure, whether it's breakthrough networking, storage, technology, software that improves the utilization rate of GPUs, and directly funding. I think some of these labs that are competing with Google and meta, I think it's very important
Starting point is 00:12:01 for the world. Can we talk a little bit about the world of data center, and semiconductors because you are one of the most seasoned investors in both of these spaces. I think you started your career covering semis 25 years ago or something, and you just know more about it than just about anyone else I've talked to. And I often find myself asking you or calling you with a question about these two areas. And most of the attention has been on the models and less on the world of semiconductors. Everyone knows Nvidia, obviously, but they just sort of assume, like, yeah, it makes these amazing chips that power the rest of it. And there's just so much
Starting point is 00:12:35 else that's going on. You've hinted at it with like the new clusters that are being built. But just give us like a stated from your perspective how this has evolved and what aspects of it are most interesting to you, whether that's the data center or the individual chip or the systems or however you want to approach it. I just think you have probably the most interesting and valuable perspective of anyone I know on these two topics and have been thinking about them for a long time. And now it's like the main event. It's for sure the main event. I started covering Nvidia in January of the year 2000. I would say watching, the early days of
Starting point is 00:13:08 Nvidia has a public company, Tesla has a public company, and being a substantial investor in both, have been by far the greatest privileges of my career and the most exciting thing I've ever done as an investor. And Elon and Jitson are for sure the two best CEOs I've ever seen with Lisa Sue right behind them at AMD.
Starting point is 00:13:32 And I just said that she took AMD from company five times levered. literally four years behind Intel to essentially utterly dominating Intel in every way. I think they had 20 days of cash when she took over. Everybody talks about Satya, and he's very impressive. He took over a monopoly that had been extremely poorly run with recurring revenue at high margins. But anyways, I have been at Simis for a long time. It is my first love.
Starting point is 00:13:55 Where are we? So first, you know, I think it's been touched on a lot of your podcasts. The number one thing we have all heard is tech investors. One reason that tech has been so amazing for 30 or 40 years is that, software has zero marginal costs. These companies all have extremely high gross margins, recurring revenues. And AI is the exact opposite. AI has extremely high marginal costs, because scaling laws literally mean that the only way you get an improvement in quality is by spending a lot more. Full stop. That is what scaling laws mean. If you believe in scaling
Starting point is 00:14:31 laws, you believe AI will have very high marginal costs. Now, because of a variety of things that we're going to talk about, those marginal costs go down really, really quickly. But they're still really, really high at the leading edge for models, especially for training and for inference, but much less so for inference. Inference and training are two totally different markets. So what this means that it has really high marginal cost is that infrastructure, efficiency, and excellence, I think, is going to emerge as the single most important success factor, particularly for the model companies themselves. And this has been measured today in something called MFU, model flops utilization, and that generally runs around 35 to 40%. And that's literally
Starting point is 00:15:17 the percentage of compute, theoretical compute, flops, that you're actually applying to trading. So of the theoretical Flops, most companies that have been published, and people stop publishing this because it's so competitive. But the technical papers for GPT3,
Starting point is 00:15:35 I think what was called Google's Lambda and Nvidia's Megatron all showed their MFU and it was between high 20s and high 30s for all of them. Google had the highest.
Starting point is 00:15:45 If you have a higher MFU, it means that you can choose for the same amount of money you were spending. You have the same amount of GPUs and the similar amount of power, presumably. You can choose between faster time to market.
Starting point is 00:15:57 If you run a 50% MFU and your competitors running 40, for an equivalent amount of trading flops, you can be in market 25% faster. You can choose between better quality. You might just do the trading route as long as possible, or you can make the model a lower cost in a variety of ways, most of which relate to quantization, and we can get into that, but it's very technical.
Starting point is 00:16:19 It's almost like if you have a 25% higher quality model because you have a higher MFU and you build in quantization, then you can actually get almost a 50% reduction in inference costs because if you can quantize one level lower than competitors, it's a profound advantage. And so I think MFU, if I were to pick one metric to evaluating lab success, because now there's been five GPD four class models from it's Google, OpenAI, Anthropic, XAI, and meta.
Starting point is 00:16:49 So there's five GPD. Does Mr. L have one that good? Maybe right under, arguably just as impressive because it gets really good results in all the evals with a much smaller parameter count. By the way, these models are commodities today, but I am suspicious once we get to Scaly Laws continue in GPT 7 or 8. Literally costs $500 billion to trade. I don't think they're going to stay commodities. Scale is the most powerful barrier to injury, and that's a lot of scale. So MFU is the most important metric because it gives you all of these advantages
Starting point is 00:17:25 and ways to differentiate yourself amongst five people, six people who've trained these GPD four class models. But I think there's something better than MFU, and I actually came up with it this morning, and it adds to MFU and decomposes MFU. I would think of it as maybe like a unified AI efficiency equation. So MFU, I would decompose into two things. The first thing is something called Mammoth, maximum achievable matrix multiplication flops,
Starting point is 00:17:57 Mammoth, Matt Moll. And what this measures is software efficiency. And this goes to Kuda. And this guy, Stan Beckman came up with it on X. What he did is each chip has a theoretical maximum performance, which is just easily calculatable, you know, flops. And then he looked at what can you get in practice? It's Nvidia GPUs run it, per his testing, 83%.
Starting point is 00:18:21 And I'm sure that there are some labs that are running them closer to 90%. But this is just a single GPU. One reason it took AMD so long to break into this market beyond the fact that it almost always had semiconductors, it almost always takes you until your third generation chip to really, really hit it. That's what it took with Google, TPUs. Instinctin by 300 is a good chip. because their Rockin open source software was terrible,
Starting point is 00:18:47 they didn't have the internal capabilities to really improve it. No one in the community took time to improve it. So the instinct that my 300 first came out, it ran at something like 25 or 30% mammoth. Now, per stance tests, it's up to 60%, but because it has more flops, it's actually slightly ahead of the Nvidia GPUs. And this was actually really,
Starting point is 00:19:11 I thought it was brilliant, as a way of conceptualizing and quantifying the CUDA advantage. They can run an 83, and AMD is running at 25, and it answers the question. MI 250 was actually a pretty good chip. Nobody used it. Why? Its mammoth was probably 10%. So first, there's mammoth, and that is how good is the software for your chip.
Starting point is 00:19:33 And then how well do you, as someone who's using that software, optimize it? Netflix used to say that we know how to use AWS. more efficiently than Amazon, that a lot of people have said some of these labs know how to use Kuta more efficiently than at VINIA, that's a big secret sauce. But they're running at 95, right there, that's 10%. If Nvidia GPUs are running an 83% efficiency
Starting point is 00:19:59 on a per GPU basis, why are we down at 35, 40% for MFU? And the reason is something I would call SFU, system flops efficiency. And this really captures networking, storage, and memory. In each one of those, and we could decompose SFU into each of those, but I think SFU is like a helpful way to think about it. That's not all.
Starting point is 00:20:24 So you need to multiply mammoth times SFU, system flop sufficiency, and then you need to multiply that by the percentage of time that is spent in a checkpoint per rating run. And these things all compound. And they compound out to some wild things. And just in case, should I explain checkpointing or no? Because as a reference earlier, trading cluster has to be coherent to function.
Starting point is 00:20:51 And that means each GPU needs to be aware of what every other GPU is thinking. If anyone GPU fails, you lose everything from the last time you saved the model, which is called the checkpoint. The GPUs fail all the time. not just GPUs, but GPUs melt. If you read the Lava 3 technical paper, I mean, it's like the list of reasons that GPUs fail is like astonishing. An optical link goes down, a switch goes down.
Starting point is 00:21:19 There are so many points of failure in the chain of storage, networking, and memory that feeds each GPU, that there are innumerable ways for them to fail. The GPU without storage, memory, and networking is worth us. And if you want to really, really understand, this, something I highly recommend to everyone, build a gaming PC. Because every computer, whether it's the iPhone, whether it's a laptop, whether it's a data center, has the same, I would call four fundamental elements. Three are memory, storage, and compute, and then the networking that connects all those things.
Starting point is 00:21:58 And when you assemble a gaming PC, you literally have to plug the PCI Express into the GPU and then into the CPU. and then you have to connect slot into DRAM so that it'll connect. And then the DRAM has to connect to the flash storage. And DRAM is obviously memory. I highly recommend this. But I think the way to think of a data center, R.D. computer, is just imagine that it's a restaurant. In this restaurant, the primary unit of compute and at an AI server, it is the GPU,
Starting point is 00:22:29 is the head chef. A head chef can do nothing without food and ingredients. in utensils. I would conceptualize storage as like the delivery truck that brings food to the restaurant. And then the way storage connects to the rest of the server is today always over PCI Express, and that's a networking technology. And that is literally the guy who moves the food from the food truck, it's the restaurant's refrigerator. And we'll call the restaurant's refrigerator. This is an imperfect analogy, maybe the memory. And then you have to move. the data from the memory, actually in this case, into the CPUs, into the big pool of DRAM,
Starting point is 00:23:14 the CPU can do its job, and then you move it to the GPU's memory, and then the GPU can finally do its job. Maybe the GPU memory is like the stove, and what flows between the stove, and the soushaft is the connection between the CPU and the GPU. And we can go into this, but just unless that chef has a stove, cook, utensils at food, he can do nothing. And the big problem in the data center is over the last five years in particular, these numbers are going to be directionally accurate. GPUs have gotten 50 times faster and the rest of the data center has only gotten four to five times faster. And that is why MFU is so low because the GPU is sitting around waiting for all those things
Starting point is 00:23:58 to do their job and it's doing nothing most of the time. So I do think it is sensible, really sensible, to invest in next generation networking, storage, and memory technologies, particularly in networking. Because if we're going to get to a million GPU cluster, that'll be a million Rubens, which is the generation after Blackwell, we're going to need profound breakthroughs in every step of that thing, every single thing, all new stoves, all new refrigerators, robotic sous chefs, new utensils, everything. Otherwise, it's going to be wasted, and MFU is going to be 3 to 5%. But anyway, it's coming back to checkpointing, because there's so many points of failure
Starting point is 00:24:43 in that chain, the head chef in data centers, being the GPU, goes up and flames a lot. They literally melt, and that every other component breaks a lot. The cluster is 32,000 of these metaphorical kitchens, where each step, if it fails, brings down the whole cluster. So because of this, people checkpoint frequently, and that means save the model. Now, if you through better networking topologies or better cooling technologies such as GPs don't melt,
Starting point is 00:25:15 if you have a lower failure rate, if you have a more reliable cluster, you need to checkpoint less. So we've had mammoth, SFU, then checkpointing frequency. And that gives us, and like let's just say, there's one company that runs at 90% mammoth,
Starting point is 00:25:31 and then it runs at 50% SFU. Now you're at 45% utilization, and they have to checkpoint. I'm just going to make something up a third of the time. You're down to 30% MFU. You have a different company that can run it close to 100% Mammoth because they're Kuta Masters, and they run it, let's call it, 60% SFU.
Starting point is 00:25:56 Now you're at 60, the other, you're at 45, so that's already a massive difference. It's a 33% difference. And that if you have to checkpoint only 10% of the time, you're at 54%, your competitor's down at 30, that's not even all of it. The last thing that we need to multiply it by is PUE, which is power utilization efficiency. The power has a cost, and it's going vertical for these big clusters. There's only basically three places in the United States today where you can get a gigawatt of power
Starting point is 00:26:28 to a single data center that's reliable. old enough, that's nuclear. I think it's like eight cents per kilowatt hour on average in the United States. What they're going to be charging for that gigawatt is like Tidex that. But your PUE then really matters because your cost of electricity really matters. So things you do to optimize SFU might actually increase your PUE in a really negative way. And so I think if you're, I just thought of this morning, maybe it's super obvious. and every lab already is doing this. But I just think if you leak it all into an equation,
Starting point is 00:27:06 and then you do dollars that it takes to get it, you can really make trade-offs. Oh, wow. If I spend 2x more on networking, I get much higher SFU, which is great, but then I have a much higher PUE, which is bad. And what we ultimately want is not just actual ex-aflops per second, but we want exa-flops per second per dollar of capex per watt of electricity.
Starting point is 00:27:39 And the equation I just described, which I'm going to write out at least for myself, captures all of that. And everybody's going to be making different design decisions and data center architecture and these design decisions you make are going to be immensely important. And I think you're going to see, particularly with these 100,000 clusters, some of these companies are going to have greater than a 100% advantage at exaflops per dollar of CAPX per watt consumed. And this is when you're going to really separate these labs, because that is the difference between GPT 7 or 8 costing a trillion dollars or 500 billion.
Starting point is 00:28:25 The next class costing $200 billion or $400 billion. And then for inference, it's the same. It's the exact same. Not only is it cost to serve, but it directly affects user experience in terms of tokens per second, which we know from Google is one of the most important things for search UX. So all of this is also going to apply an inference. Inference is a lot easier because it really, really just comes down to memory bandwidth the non-chef memory. This reminds me so much of the Vacloss-Smill history of energy stuff,
Starting point is 00:29:00 where you would get a new, he called him prime movers, some source of energy, fossil fuels, wind, whatever. And then you get this long period of the gains all coming from the efficiency. So if you're spinning a turbine or something, in the early days of coal, the turbine only captured 10 or 15% of the available energy from coal itself. And now we're at like 98 or something like that. And it basically sounds like that's the exact same thing that you're describing here. The GPU is the coal. And those will keep getting better, which is cool in technology on like coal. But it sounds like that's basically the story, which is so interesting. That is like turning sunlight into usable energy via different fossil fuel mechanisms and motors.
Starting point is 00:29:42 And this is turning sunlight into compute, literally. And now that sunlight can come in the form of actual sunlight. It can come in the form of artificial sunlight. That's nuclear. It can come you come in the form of stored sunlight, that's fossil fuels. But this is the efficiency at which you turn sunlight into compute. And the dollars you pay for that ratio. Yeah. By the way, it is just interesting. It's one reason I was never that excited about wind because turbines were really efficient.
Starting point is 00:30:12 Whereas you could look 10 or 15 years ago, solar was terribly inefficient. I actually think it's awesome. You know, it's a little sad to me. You know, like all these kids are really worried about global warming, and you have all these things about 20-year-olds. I don't want to bring children into a warming world. Global warming is a big problem. It is a solved problem.
Starting point is 00:30:30 And it is solved because photovoltaic cell efficiency is compounding just under 10% a year. And battery efficiency is compounding 200 bits less than that. And if you compound that out over the long term, not even the long term, the world is going to run on sunlight directly. Possil fuels are just going to go away. It's because of economics, and it's because sunlight and storage is going to be cheaper than every other way of providing power. It's just not going to be a problem for the next generation now. Maybe we hit some tipping point, and it's irreversible, and this and that.
Starting point is 00:31:06 Emissions are going to collapse in my lifetime, literally collapse. Obviously, like pre-industrial age humans, we're actually massive polluters. because the pollution per unit of firewood is really high, and you needed fires, otherwise you'd get eaten by a saber-toothed tiger. Now, there weren't that many of them. It is possible that we're going to be back to, like, neolithic levels of emissions in my lifetime, assuming I can live a little bit longer because of AI. So that's an awesome and encouraging thought.
Starting point is 00:31:43 That's good, and I just wish somebody was shouting that from the rooftops. Well, here we are. I'd love to keep going up the stack here because every single level is interesting. So the semiconductors themselves and the potential innovations there are interesting to me. The data piece is really interesting to me all the way up to the application layer. And I'm just curious what you think about all of this stuff. So one of the things that we're assuming in all this is that we're going to have more data to train these things on at GPT, 5, 6, and 7. And I think there are some interesting discussions about what available data there will be or how we'll get more data or how we'll create.
Starting point is 00:32:18 synthetic data and whether or not that will work to create a better model. Like, I'm interested in what you think about all of this stuff that is required if we're going to have the big arms race that you were describing earlier. So I do think this was a real bare case, maybe nine months ago. But I think it was hitched out in the Claude 3.5 technical paper. It may be a little more explicit in the Invidian Nemotron technical paper. You know, it's awesome. Invidia, whatever, it feels like there's going to be a rate limiting factor for AI.
Starting point is 00:32:46 They've solved it. and that's what Nebatron did. But I do think for reasons that no one understands, and no one understands these models. No one understands how they were, why they were, why scaling laws. There's all sorts of theories, and maybe we're getting a little better at understanding them, but no one understands them.
Starting point is 00:33:02 So no one understands, point one, but it does look like synthetic data works. No one understands why, but it looks like it works. Now, again, will it continue working? I don't know. Nobody knows. No one knows. people who have seen Kevin Scott just did a podcast, and he basically said, look, I've seen
Starting point is 00:33:24 some early checkpoints of GPT5, and scaling laws are continuing. I think in a lot of ways, that's the best indication we have that they're scaling. And I actually think largely in respect to XAI, I do think Open AIs, the combination of what XAI is doing and Blackwell delay, the Blackwell delay means that if you're waiting for Blackwell and you're trying to get 100,000 cluster, XAI is going to have. have arguably a one-year advantage, which is untenable to these other labs. So they're all now frantically working on stating up their own 100,000 cluster, but they don't have Elon designing the data center, redesigning the data center from first principles.
Starting point is 00:34:00 Data center architecture was always like kind of nice to have. Now it is must have. It's existential. I think synthetic data does look like it's going to work. This goes to where will the value for these models come from? Why is meta so comfortable open sourcing? You may ultimately see everyone open source. You know, X-A-I has open-source, Grockwad.
Starting point is 00:34:23 But the reason is the value may not come from the model. Now, look, if you're running that equation I described and you have a one to 200% advantage on exaflops per KappaX dollar per watt and scaling loss hold, you're going to have such a massive advantage of model quality that you're never going to open source it. You can still effectively kind of steal models, distill them. If you have that computer advantage at Scaly-Lawl's hold,
Starting point is 00:34:46 wow, you're going to have something very valuable. The value clearly comes from distribution and unique data. Meta is open-sourced Lava. They haven't open-sourced all their data, and that data is just going to be for their version of Waiba, so it will for sure be better. Google, they have YouTube, and then all sorts of other data sources
Starting point is 00:35:09 that they've developed when they tried to boil the ocean to create the knowledge graph, which are those little knowledge panels that kind of appear when you do searches. So between YouTube and the work they did for the knowledge graph and Google Maps, they have crazy data. They don't care because they can use that data to monetize their model. XAI, whether they open source or not, they will always have, for sure, access to X data in a way no one else does because XO's 25% of XAI. And then I would think over time, XAI will kind of be an intelligence layer that cuts across Elon's ecosystem of companies.
Starting point is 00:35:43 By the way, we should talk about AI and robotics, which I actually think may be the biggest disruption in our lifetime, comparable to like artificial superintelligence and these digital gods. If these models, unless someone develops a compounded advantage of mammoth, SFU, checkpointing frequency, and PUE, such that they have a dramatic difference at exaflops per dollar of capax per watt, I think they're probably all going to converge to run. roughly the same intelligence. And the reality is, given the way that I think at least Google and better are thinking, even if someone else is way ahead of them and more efficient, they will
Starting point is 00:36:22 try to solve the problem with money. Oh, wow, it only costs them $300 billion, no problem. We're happy to spend a trillion. Although ultimately, economics will apply. These may impact stock prices. Like, and I think it's not inconceivable. Some of these magnificent seven, forget buybacks and dividends. They may eliminate their dividends, stop buybacks, and start issuing stock to fund this. We could be in a really wild world in a lot of ways. But these intelligences will probably converge on kind of a similar what I'll call IQ.
Starting point is 00:36:52 And then it's just who has the most differentiated real-time data about the world. It's these unique data sources coupled with every time you rate an answer from one of these models, you're helping it improve. And so if you can couple you. unique data with internet scale distribution, then you're going to have a winning formula. And there's only a few companies that have that. You know, it's XAI. It's Google.
Starting point is 00:37:23 It's Microsoft, although Internet scale may be. And Open AI gets that through Microsoft, probably Amazon and Anthropic gets it through that. And then, you know, obviously meta. And then Apple is the big wild card. And we should talk about that because I think one of the biggest things is where is the inference going to happen? and compute tends to cycle in between centralization and the decentralization. And we're at the end of like a long period of centralization in the cloud, where a lot of compute right in the cloud and these big data centers
Starting point is 00:37:54 is just because you could get much higher efficiency in these big data centers. Now, clearly for trading, that is going to happen in giant data centers that will be in the cloud, although something that I think is underappreciated about these AI data centers and why we shouldn't worry about them over the long-term crushing power. demand is you can put them anywhere. You put them in Wyoming. They don't need to be near a big city. I think we'll eventually see giant data centers in shale gas fields with power plants on those fields long away from any humans. That will probably be an intermediate term solution to the lack of nuclear power in America. The U.S. we do. We have all the best models. And the U.S.
Starting point is 00:38:34 wants these models trained in America for better or worse. Obviously, Mr. Hollis, French, I don't know where their models were trained. cycles of centralization and decentralization based on where can you get the lowest cost compute at the highest utilization, kind of a variant of the equation I described earlier. And I do think the subcomponents of that equation, not only are they maybe helpful to labs, but they're really helpful to investors. If you can improve SFU by 20%, and it all comes down to millimeters of silicon. for only a few more millimeters of silicon
Starting point is 00:39:08 or maybe even less millimeters of silicon, they're like, oh, my God, those millimeters of silicon are the ultimate cost. Wow, you have like a massively winning formula. And so you can just almost go look at each step of that long chain and where do you have really high cost for square millimeter of silicon? And then that is an opportunity to really optimize.
Starting point is 00:39:31 And whether it's at the infrastructure software, the data center hardware, or the semiconductor layer. It's all there. But I think you're going to see inference increasingly done on phones. And this is clearly Apple's play. It is one reason that they are in such a advantaged position. All these other companies are trapped in this prisoner's dilemma.
Starting point is 00:39:57 I am sure they would like, given how much it costs. They're all economic animals. and they could just like reach an agreement. It's, hey, you know what? Nobody is going to stand up a Blackwell cluster until 2006. Like, they probably all cite it. But it is literally a classic prisoner's dilemma, and that would be a Nash equilibrium.
Starting point is 00:40:18 But we're not going to reach a Nash equilibrium in the race to create digital God with the states are existential. So there is prisoners to blubimba. Apple's not. By the same way, Google spins collectively, cubulably spent, I don't know, hundreds of billions of search between KAPX, OPEX, all this stuff. And Apple monetizes it almost as well as Google because they have this distribution chokehold, toll booth in iOS. They're clearly going to do the same thing with Apple intelligence.
Starting point is 00:40:47 Then Google will obviously do that with their phone to Android. And so you're going to have a small model running on your phone that for simple questions, think of it as like a 100 IQ model that with like really, really sophisticated knowledge. They'll probably be two of them. They'll check each other. And a lot of inference will then happen at the edge. And if inference is happening at the edge, then I do think we're going to see superphones
Starting point is 00:41:14 because the advantage to you as a human. What bounds inference at most times is memory, and you will be willing to pay, today iPhones are sold based on how much flash storage they have, what you will most care about is the amount of DRAM you have on your phone, because that will determine the perimeter count of the model you can run local. And that is, therefore, the quality of your own local intelligence that has access to all of your data in a privacy safe way.
Starting point is 00:41:44 And when you ask it for something, if it can be done on a compute, it will. And I think from networking like a truism is not when you can, switch when you must. But I think the next thing for inference, it will be local when you can, cloud when you must. So if you can inference locally on your phone, you will always do that because it's free. And cloud inference definitely costs money because you're burning GPU hours. Now, it's very efficient, but the inference on your phone is effectively free. And then increasingly our competitiveness as humans, you know, it's very funny. A lot of really smart people are a little skeptical about AI.
Starting point is 00:42:23 I get it. It's like you have a 125 IQ or 135 or whatever it is. AI, particularly your domain, isn't it all impressive to you. And it's like, okay, it's a little bit for search for me learning about other domains. But man, a lot of humans don't have 120 IQs. And what people are missing is that there's a lot of humans with like, I don't know, I have no idea. 100 IQ, the most with 100 IQ.
Starting point is 00:42:47 Yeah. So let's just say it's 100 IQs. And they can use AI. And all of a sudden, they're like a 115. And it's like, holy shit, it's incredible. And that's just going to keep going until these ASIs have IQ of 1,000. And it's like, why I'm as a human, am I bothering to do anything other than work with an AI to create art? Anyways, if I can have like a $3,000 iPhone that has four times the DRAM on it that the $1,000 iPhone has,
Starting point is 00:43:17 it may be more storage so that local model could do rag. and this is my model, it likes me. I pick the voice that talks to me in, it knows me, it's my friend, it's my agent. All it wants is thumbs up for me in that process of improving these models. I think they will ultimately be, there will be a way that they can be RLH staff on an individual basis on the fly, which we're going to need a lot of breakthroughs to make that happen. And it's going to be my model that likes me and knows me and knows what I want, and then this is where you get this agent future
Starting point is 00:43:54 that everybody's talking about where it's like I have an agent on my phone and if I have an agent on my phone because I have a super phone that has an IQ of 115 or 120 and someone else's agent only has an IQ of 100, I'm going to be profoundly advantaged as a human.
Starting point is 00:44:10 And then that continues all the way up because I think eventually Apple will monetize this and the way Apple will monetize this clearly is when the odd device LLM isn't smart enough. They're going to send us to the cloud. They're going to use something called a router,
Starting point is 00:44:25 which is what every application company uses. You just want to route to the best model for the query per dollar. Apple will charge companies to be in their router, and then, oh, if you want to be routed to,
Starting point is 00:44:40 eventually you're going to have to pay a bigger and bigger toll, and everybody will pay that toll the same way Google paid that toll. But if I, as a human, I can have a smarter model locally, and then I can opt in into maybe Apple will just make it easy. They'll say, okay, you can have cloud superintelligence or you can have cloud intelligence.
Starting point is 00:44:59 And I pay $60 a month for cloud super intelligence. Now, or whatever it is, you know, $1,000 a month, $10,000 a month. What would you pay for that? And so I have 20 points of IQ on my phone relative to a lot of people. And then I'm paying $10,000 a month for cloud super intelligence. our hyper-intelligence, super-intelligence is just $1,000 a month, and then regular intelligence is $20 a month.
Starting point is 00:45:25 Like, I'm going to be really advantaged as a human. And that seems dystopian to me, but it's also kind of hard for me to not see that happening. And many, many investment conclusions flow from this in the same way that they flow from, like, hey, okay, mammoth is at 90%, okay, so that's probably not a good place to invest. SFU is at 30%? Wow, that's a great place to invest. PUE is at 1.8, and it could theoretically be at 1.3.
Starting point is 00:45:55 That's a great place to invest in that equation. You kind of want to invest at the most inefficient chain. A lot of investment implications flow from where does inference happen. And then a lot of implications for humanity flow from that as well. I'd love to talk a little bit about the common or part of this whole equation. We've talked about like Mount Olympus and like the game of kings fighting each other for dominance. We haven't talked it all about. What about just like a normal new company that's being started to take advantage of these super intelligences at the application layer or to do something else?
Starting point is 00:46:27 Then we'll get to robotics after that. But if I force you to just be a early stage series A, series B investor in the world of startups, and they don't have this kind of resources that all these Magnificent Seven and others have, how would you think about the aspects of companies and opportunities that would get most attractive in the world where GPTX keeps scaling and getting better. First thing, I would just say it's always what are people doing? What I am doing is really focusing on companies that improve that equation, elements of that equation I described earlier. Because I think that's the choke point. That's the highest likelihood of success. And in tech, whatever you find a constraint, if you can invest in something that alleviates that constraint,
Starting point is 00:47:08 you usually do well. And right now I profoundly believe the constraint is it SFU? checkpointing, which goes to reliability, and then PUE. So that is where I am targeting my dollars. I think the application layer is really, really hard. There are people who are killing it. I think you have Sarah, who I've never met, but I admire on your podcast. And she's like absolutely, from my perspective, crushing it at the application layer. But just, geez, investing there today feels really hard to me.
Starting point is 00:47:39 And I just think anyone who has a lot of conviction about that needs to be reminded of that 1% Chase Coleman's debt. Yeah, I was just thinking that. All these VCs had really strong views. Like, there are all these companies that were funded in 95, 96, 97. And they seemed like they had got in the net. But it just takes time. I guess they did go into that for the VC. They were public and they were able to cash out.
Starting point is 00:48:06 But they actually didn't go into the net with the fullness of time. So I just think it's important to have a lot of humility at the application layer. And maybe it's just that I am so comfortable with this infrastructure layer, that is where I have concentrated to date. There's really only been what application company that I've been excited about. We've seen a lot of them. Maybe I'm like too aware of the Chase Coleman stat. And I'm too, and I'm being too careful at how I approach this application layer.
Starting point is 00:48:39 But I'm going to miss out on a lot of CMGI was incredible, you know, what were some other besides Yahoo, what were Lycos, you know, there are all these companies that we don't even think of today. But they're incredible venture outcomes. Maybe I do need to be a little more mindful of that, but they're clearly people, whether it's benchmark, whether it's serious firm, who are succeeding. And what I would say in the application there, I heard this from Dishria, but it really resonated with me.
Starting point is 00:49:07 And I think it also goes to this whole, silly AI-R-I debate. Not only is it an ROI debate about creating a digital god, but I think there's a super clear ROI, and you can actually do math to show that. And a lot of people are conflating CAPX and OPEX in ways that are just not helpful. Meta went down like 80%, partially because they were overspending on the Metaverse, but partially because Apple took away their ability to target through IDFA.
Starting point is 00:49:34 Meta is a 5X, and their revenue growth is re-accelerating. Why is it reaccelerated? Because they spent a vast amount of money on AI to figure out how to target. It's like maybe the meta return alone justifies all of the spending. They call that meta advantage, but meta advantage is just part of it. That's just where as an advertiser, you can let meta do the targeting. And then Google has performance max. And all of that is, and that's where there used to be ads were created.
Starting point is 00:50:01 We're going to come back to applications, but ads were created. We're going to pay some humans to figure out that we need. need to show this creative to like white, 40-year-old guy in Boston and this type to this type of demographic category in this city. And we're going to show them at this time of day. Oh, we're going to show them after the local sports team is won and it's sunny because that's actually the best time to see an ad. But now AI can do all of that. And we're going to have a million creatives and it's going to be optimized on the fly. And that's what AI is enabling Google and meta and other firms to do. And probably just on that. Just that justifies it. Yeah.
Starting point is 00:50:36 been a massive ROI on AI. What are we talking about? The other thing that's really funny to be about this whole AI-R-I debate is, okay, we're going to have this abstruse debate about ROI. These companies are all public. There is something called return on invested capital. And RIC has gone up for all of these companies since they ramped CAP-X. What are we talking about?
Starting point is 00:51:00 If you're an AI-R-Y skeptic, why has R-R-Y got up at these companies? Well, yeah, KAPX is up. No pat is up more. Why is that? Because they're doing exactly what you would expect to happen in World of AI. They're trading off human labor against GPU hours. That's why their R.OIC has gone up, and these GPUs are really efficient. Let's have an AI-R-O-I debate.
Starting point is 00:51:26 When the RIC at the big AI spenders starts to go down, until then, it's just the height of intellectual, ridiculousness. I mean, ridiculous. What Vichrya said that I thought was so powerful for the application layer, there's all these SaaS metrics. After your first year, you want to be at million to be on pace. After your second year, you want to be at $5 million. After your third year, you want to be at like $10. And if you're above that with reasonable cash burn, you can probably run these curves. You're going to be a successful SaaS company. And all those curves, which everybody knew is one reason SaaS multiples got so
Starting point is 00:52:03 inflated because it became almost like quantitative. Oh, wow, this company is at 15 million in year three? That means we can pencil them in for a billion dollars in year eight. A, that led to multiples expanding ridiculously. B, it led to the industry being overfunded such that the degree of competition in each vertical went up such that the curves no longer held. And this is one reason why if you made a lot of SaaS investments, particularly in 2021, you're in a world of paid. And then on top of that here AI comes along and it fundamentally changes the paradigm for application software because what application software fundamentally does is makes humans more efficient. And today we're in a state with AI where it's making humans more efficient, but you don't
Starting point is 00:52:44 need that big of a wrapper around it. That's where we're going to come to the mystery comment. And then ultimately, if it's replacing humans in your selling application software on a per seat model and those seats start to go down, you have a problem. These companies are super aware of it. they were going fast in one direction, and now they have to contend, almost in every vertical,
Starting point is 00:53:03 with this next generation of AI-first application companies, which are just these often very thin wrappers around GPT, you'll call it an LLM wrapper. And what Fisheries said to me that was so funny, it's like all these AI companies are blowing these traditional SaaS metrics out of the water. And all this isn't being counted in those AI ROI papers. Just so there's all these companies that are going like zero to 30 million in nine months. And even though AI is really high marginal cost, they're being more cash flow efficient
Starting point is 00:53:36 relative to software company. So in almost every vertical, there's multiple companies that are AI first, that are just really, I think what Vichry said was they're just paper thread wrappers around pick your LLM of choice. Then yeah, they use a router to find the best LLM. But they're like magic to their customers. they aren't going after software budgets, they're going after labor budgets.
Starting point is 00:54:02 It's very intellectually interesting to me to watch. How do you get conviction that one of these companies can use what is initially magical to their customer to build defensibility in, all while allowing for the fact that the big models are going to continue to compound at a really high rate. And this goes to things like, Maybe you just find a really good sales motion that works.
Starting point is 00:54:28 Maybe you find a really easy integration point. Maybe you build some defensibility around that integration. Maybe you make it really easy for a small business to do rag on a really unsophisticated system. How are you building differentiation into that wrapper? And hopefully you're not fine-tuning because that does extensively fine-tuning because that locks you into a model generation. How are you creating a compound AI system? such that you're serving each query at the lowest cost possible, and you're starting with a small model and treating it to a big model.
Starting point is 00:55:01 There's a lot of things you can do, and all of those are really, really important. But I just think at that application layer, there's just, it seems like for any category, any category it feels like they can be imagined. There's multiple startups that have exploded that are defying all traditional SaaS metrics, and it's not clear to me in this point. point how much defensibility each one of them has. Some of them are going to have amidst defensibility around one of the axes I described, but I think a lot of them may not, and I think that's what makes it so hard and why I've been approaching that space so carefully.
Starting point is 00:55:40 You know the amazing thing? I think the framing around labor, I mean, just take me as a stupid, simple, tiny example. I have a team of three people building one of these, we'll call them, like a really light LLM wrapper for doing research on private companies. And, And if you just think about how often I use this thing in a way that I normally would have used an analyst, it's multiple times a day. And it's returning work as good. So there's so much undifferentiated heavy lifting in every white collar labor market, I guess, is like one major learning. 100%. If you can have the answers instantly and for effectively free, it's crazy how much you use these things. We're just getting started. Yeah. And I have the same thing with like
Starting point is 00:56:22 an LLM at my firm. It was actually very interesting. We had an intern who was really talented. And he said, listen, I think this internal AI tool is amazing. And I use it every day and more every day. It's gotten much better this summer. We work on it for a long time. And then I was like, show me how you use it. Because I'd been using it a certain way.
Starting point is 00:56:40 And then this literally, you know, this 21-year-old kid was like, well, this is what I do. And now all of a sudden, my usage of this tool is two hours a day, three hours a day, four hours a day. And I do think Jack Welch had this term called scutwork, just like hard. unpleasant work that had to be done really well in a white-collar setting. And I think it's all of these little rappers, or mini-wrappers, whatever we're going to call them, are going to replace a lot of that Scott work. And initially, it's going to be in combination with humans, with science fiction Neil Asher's world, there's something called a high man, which is a human AI hybrid. And it's where they're linked through something like a neuralink to like a computer server that is,
Starting point is 00:57:24 supported by like an exostructure, robotic exostructure that they walk around on. There's like fusion chess, whatever we would call it. We'll have that for a while. And that goes to, you know, 100 IQ humans performing like 120 IQ humans and then forming like 130 IQ humans, then 130 IQ humans performing like 160 IQ humans. But then eventually it feels like as long as scale we must continue, which is a big F, it's just going to be the AIs. Maybe we could talk about robotics.
Starting point is 00:57:51 I had a really interesting conversation, call it April. of this year with an investor that has been investing in lots of these same things for long periods of time privately and publicly and has big positions and lots of the companies that we've been talking about today. His observation to me was the big underestimation that's happening over, let's say, five years, is the role that robotics and robots will have combined with all of this technology we've spent all today talking about. And I would love to hear you riff on that because in the near term, it feels like a little bit, quite a bit of frothiness, like some crazy funding rounds for these companies that you don't really know what the robots are being designed to do, sort of general
Starting point is 00:58:31 purpose, humanoid type robots. There's all sorts of interesting, more specialized ones that are cool too. But what do you think about all this? It does seem kind of under-discussed relative to just all the foundation model and semiconductor stuff. I agree. I think it may end up being a bigger near-term disruption than what we were just discussing, the automation of a lot of white-collar labor. I think the first robot, the first robot that's really going to impact the world, is every Tesla car with what they call their AI4 hardware. Because from my perspective, there's a publicly sourced miles between disengagement. And you have to remember for Tesla, Tesla is going to get the same miles between disengagement.
Starting point is 00:59:11 Like if you built a new city on Mars, and it was populated by entirely different looking cars and streets and everything, you could drop a Tesla in that. that city and it would have the same miles between disengagement than it gets in any other city. Where something like Waymo is geo-finsed, we're really only using it in cities that have like nice grids and good weather, et cetera, et cetera, et cetera. It is clear to me looking at the crowdsource data
Starting point is 00:59:39 and miles between disengagement with different versions of FSD that when they cut over to 123, which is effectively all deep learning, I think eliminated almost all human code, something dramatic changed in the rate of progress. And then when they cut over to 12.5, which runs best on the AI4, which used to be called the HW4, it's just like the local computer of the Tesla, it is now rolling out to AI3, was another step function. And these, going to that same scaling law, those step function improvements were made with a fraction of the compute
Starting point is 01:00:21 that Tesla is now installing publicly in their data center at the gigafactory in Austin. They've actually filed, sometimes I wish they, as an investor, I wish they couldn't file so many of these patents, but they've filed some really innovative patents for data center cooling related to what they're doing with what looks,
Starting point is 01:00:41 it's been publicly said is going to be over 50,000 H-100s or H-200s. F-Sunds. FSD is now on the same scaling law, and arguably on a faster scaling law, because they have a lot of catch-up to do, that GPTs have been on. So I think 12.5 is like GPT3, and it can consistently drive me most places with no interventions. I'm a seasonal driver. I really only drive by Tesla in the summer, and actually my wife, Becky, tends to do most of the driving because she likes it more than me. So we kind of get like a seasonal look. It's almost like every May we check in.
Starting point is 01:01:20 There was just always continuous progress. This year, it's like when we turned on 12.3, it was like all of the progress over the last 10 years was in that one release. From the first time I had that Tesla, and then we had that again when I went 12.3 to 12.5. and they're at probably like a GP2 level of compute. I think they're going to go really fast to GPT4.5 compute, which means you're going to get, using these orders of magnitude, you're going to get like a 100x improvement really fast. So I think there's all these people who have been skeptical.
Starting point is 01:01:57 They're all in for abject humiliation. They just are. And then unlike GPT2, only Tesla has access to a visual trading data set that is based on miles driven, we can argue whether it's 100x, a thousand X, 10,000 X, bigger than the second biggest trading data set, which is Waymo. So it's like people, oh, how are they going to make money? Well, it's like in this case, in the world of self-driving, from my perspective, it's like they own YouTube, they own all of meta's properties and the open internet and X.
Starting point is 01:02:36 and then other people are like trying to do it using Yahoo. Yeah, using Yahoo. Like, good luck. Like, who's going to win? Now, obviously that could change. Important to have humility. There may be an algorithmic breakthrough that reduces the importance of that trading data set.
Starting point is 01:02:54 And for sure, Waymo is going to try and brute force it. It just throw whatever amount of dollars they need to get the data to compete. And they have a different approach using LIDAR, Tesla. it doesn't. We'll see, like, I don't think it's a foregone conclusion. Nothing about the future is certain. But just if I look at how amazing 12.5 is on AI4 hardware, and think about the tiny amount of compute that that was trained on. And the mega cluster that they are standing up in Austin using known techniques, we're going to skip, I think 12.5 at GPT2, we're going to skip really quickly to GPT4. And then look, you know, I'm sure Waymo will reinforce it.
Starting point is 01:03:36 there may be algorithmic breakthroughs such that there are other people, we'll see. But then the other big thing is just using it to LLM for FSD. One of the best follows on X is Dr. Jim Fann who's Nvidia's head of robotics, and he's had a lot of posts about how
Starting point is 01:03:52 there's a fascinating exchange between him and Elon on X. It is amazing the extent to which AI happens on X. The Jack's team at Google and the Pye Torch team on Meta got into this bitter fight and it went to Mammoth, which framework was better for Mammoth. Eventually, the heads of each lab had to step in publicly on X and make peace. But like, wow, you know, you learned so much just following that.
Starting point is 01:04:21 And like every AI researcher is active on X. AI happens on X. And it's such a great forum for using it. But Jim Fahan had this fascinating exchange with Elon where Jim Vann talked about how LLMs could massively improve. FSD, and Elon replied, yes, the only two data sources that will scale infinitely are synthetic data and real world video. And I thought that was interesting. And then that goes to, I think, maybe the biggest risk in which this view that I just described of Tesla's autonomous future
Starting point is 01:04:53 is wrong. It's just if synthetic video data can be used in the same way that synthetic data It can be. We know that synthetic written data works. We don't know if synthetic video data works. Nobody knows. And obviously, there's a very high bar for regulator. You know, I think it's something like, I figure, whatever it is, like 50,000 or 100,000 people die in car crashes every year globally. It might even be a million. Obviously, we could take that down dramatically using AI, but we're much less willing to tolerate traffic fatal accidents from AIs than humans. That is what it is. So, you know, it's going to be heavily regulated. But Dr. Jim fan. that the reason LLMs were going to be able to really help with FSD is because of the following, this is the way my, relative to some of the people working on these problems, my comparatively low IQ brain conceptualizes it. Anything that's been trained on real world data just knows what to do, what a really good human driver would do in that real world situation. If there's a novel situation, it may not know what to do. And that's where, from my perspective, the LLM can really help because one of the emergent properties at GPT4, and we can debate whether or not it actually
Starting point is 01:06:06 is an emergent property or just in context learning, it has what's called a world model. And that means, I'm sure you know this, but if you ask GPT3, okay, what happens? If you stand a champagne bottle upside down and put like a basketball covered in soap on top of it, GP3, no idea. GPT4 will often get questions like that right. I should actually see if it gets that exact question right. A three-year-old human will say, that's going to fall,
Starting point is 01:06:34 the champagne bottle's going to shatter. It's really hard for GPT3, and that goes, you know, this jagged frontier that people talk about. So if you put a really speed-optimized small LLM in locally on each Tesla, there might be just enough reasoning capability to unlock another step function
Starting point is 01:06:57 in FSD capability. And now look, Wabo will have that too, has lots of other people, but they won't have Tesla Vision all of the proprietary data set. Just think this is going to be a reality in a way that is abjectly humiliating to everyone who is an FSD skeptic
Starting point is 01:07:15 in the next 12 to 18 months, maybe in the next six months. And I haven't never been willing to make a prediction like that before. So then you take that, and the same thing goes for humanoid robots. Google showed this with research called Tinsor RT2. We're dropping an LLM into a humanoid robot with a world model that understood what things were and what to do just made it so much easier.
Starting point is 01:07:43 Instead of training that humanoid robot how to pick up a tennis ball and a basketball and a football and how each one is different, it could reason. And so this is why putting LLMs into these humanoid. robots, I think is going to be so transformational for the world and make a lot of blue-collar labor optional. I do think politicians and political systems are utterly unprepared for what may be coming. The one thing I would say that Elon Adjitzen profoundly agree on publicly is that humanoid robots are the future, not these specialized robots. And the reason is just, of course, a specialized robot can be better than a humanoid robot at any given task, but the humanoid robot
Starting point is 01:08:23 can do any task that a human can. The world is optimized for humans, and there are massive scale efficiencies in manufacturing. And so because you can make, it's almost like humanoid robots are going to be to the field of robotics as GPT was to AI. GPT was a generalizable type of AI, and these humanoid robots are going to be a generalizable form of robotics.
Starting point is 01:08:50 And because of that, they're going to be manufactured at such a scale that they have a cost advantage. And then the physical world is going to start to be optimized for then. And that's why they're going to win. So good luck to all these non-humanoid startup robot companies. I hope you get Lycos or CMG or MySpace type venture outcome. But I don't think any of them are going to be Google. And in the same way that so much of GPT advantages incumbents, whether that's meta,
Starting point is 01:09:21 Google, X, X, X, A, Microsoft, these humanoid robots, the reason it advantage is incumbents is because they have the raw ingredients of data, compute, and capital, which is what you need to effectively monetize these. That's why their ROC is going up even as a RAPX. I do think that incumbent manufacturers who have expertise in battery design, actuators, motors, with big data sets, are going to be advantaged. I'm reasonably bullish on Optimus. That's just such a giant market. There's going to be so many competitors. And whether it evolves, FSD, I could see a world where there's just two or three companies. Maybe it's Tesla, Waybo, and some open-sourced variant. Or it could be synthetic reward data works.
Starting point is 01:10:12 LLBs really improved the efficiency of that specialized visual algorithm. And so there's like thousands of them. I think that's unlikely, but it's possible. I think robotics may end up advantaging incumbents in the same way, I think FSD and LLM, GPT have advantage, generally advantaged to incumbents. And then startups are willing to take like a new AI-first approach. But I do think robotics is going to change the world. It's super exciting. Like, I can't wait to have my own personal robot. Or 10. Yeah, or 10. Like, sign me up. Like, I just want one everywhere. What time is it? It just tells me it has, you have an AI that's load. it into my phone is also running on that and it likes me and it's friendly to me and it lasted
Starting point is 01:10:53 my jokes. It's signed me up. Yeah. Yeah. I'd love to ask a couple of closing questions that are big picture, big arching questions. The first is responding to something you said earlier, which was having watched Jensen and Elon and Lisa at AMD operate for so long and rating them as these like exceptional CEOs, it just seems like such a small handful of leaders at these companies can have such a massive impact on the world.
Starting point is 01:11:20 And so their character and way of operating is really important for all of us. Of course, this will recycle and we'll get new ones and up-and-coming ones or whatever. But what is it about that group of three and maybe throw into the recipe a few others that you've learned a lot from? What is it about those three that match them so well with the modern world and way of company building? I mean, like the crazy stuff like Jensen having 40 direct reports or Elon working on seven companies at once.
Starting point is 01:11:43 These are unusual human beings. Maybe just riff a little bit on why those three and sort of the nature of leadership in this era of technology. I guess I would say, above beyond having clearly unusual IQs and ranges of domain knowledge, you know, and I do think it's hilarious that Lisa and Jensen are cousins. It's like crazy. I just want to bet on everyone in that family going forward who's like a first cousin. Died me up to bet on those genes. What a crazy coincidence.
Starting point is 01:12:16 Beyond inherent genetic advantages, I actually think the modern American corporation, in the same way, like almost any large organization with its hierarchies, is set up to reward people who are charismatic and political. And because those people are charismatic and political, they're not always really smart. And because they're charismatic and political, they get big egos because people like them, they get used to having their way.
Starting point is 01:12:46 The same way every, I think, child under four is like a functioning sociopath. Most people who become multi-billionaires are really powerful politicians. The physical part of the brain that deals with empathy shrinks changes you as a human. It is a long way of saying that I think the way that modern corporations have been evolved to be run, you end up with a lot of not CEOs who are not, CEOs, who aren't, Excellent. But objectively terrible CEOs who were amazing at rising to corporate ranks, flattery, whatever it was to took to get there. And then by the time they got there, their ego had gotten so big, their sense of empathy had gotten so small. They're no longer
Starting point is 01:13:32 capable of functioning effectively. And often with people like that, you find there's like one or two people behind them, there's like a key leader who runs a division or whatever, who's kind of arisen with that CEO through each division or a group of people. Those are the people with the good judgment and skills to make high quality decisions. So Elon and Jensen in particular are nothing like that. Not only are they the front person, but they are always working on the most critical problems at the company. During my year off, I went to see Jensen. It was a different conversation because I wasn't working. And he just said, and maybe this has changed, but the time he said, I have no fixed schedule.
Starting point is 01:14:21 I have no standing meetings. I just find out what is the most important problem at the company. And I go and I sit my desk down in that area and I pull the best resources to work on that problem and I love it. So he's working on the most important problem. That is also what Elon does at each of his companies. Whatever is the most important problem is what he is working on. When the Raptor engine was in the critical path for Starship, I'm going to get the details wrong,
Starting point is 01:14:54 but I think there was like a standing 1 a.m. A. Tuesday morning or Monday morning meeting for Raptor engineering, only the like 12 or 18 or whatever the number is. Smartest people were allowed to be there. Everyone at the company wants to be there. So I think that is one, something that ties them together, working on the problem. Second is loving to hear bad news. My dad was a bankruptcy attorney.
Starting point is 01:15:19 He always said the number one thing. All bankrupt companies had in common was the CEO who didn't like to hear bad news. Just say it's Elon and Jensen love hearing bad news. At those companies, if there's bad news, it must immediately go to them. And that's very differentiating. no hierarchy. Wherever in the company, the problem is, that is who Jensen and Elon what to work with, the subject matter expert. Whether they're like 23 or 50, there's no hierarchy. You know, it's like the same way, like there's this apocryphal story when, I think it's true. When J.P. Borgon
Starting point is 01:15:51 was by Bear Stearns, the best modeler in the company was 24 years old. So they set up a desk for him side by side with Jamie. And Jamie would say, change that, change this, this, you know, 24-year-old This kid was like the guy because he was the best. You know, Jamie Diamond, another exceptional CEO. He didn't ask for that guy's boss or boss's boss or whatever. He's like, this guy's the best. I want to work with him. And the last thing that I think really, really ties together, particularly Jensen and Elon
Starting point is 01:16:20 is a mission orientation. Jensen started out about like, hey, let's make photo realistic video games in virtual worlds because that'll be exciting. Remove what someone called reality privilege and humans will eventually be able to be in the virtual world. the way the metaverse is still coming. I think it's just clear that it will happen first with augmented reality and then with BCI. Just VR goggles, count me a skeptic, never going to work, augmented reality glasses, ambient computing, and the BCI is kind of the in-state. It's smart
Starting point is 01:16:52 meta-bath company controls CTRL Labs and is working on this because I think that may actually be the next truly disruptive form factor. We're going to have superphones. The phones are going to going to stay as the compute layer for extended reality. But BCI may be the next true way we interact with computers after phones, the next true compute platform and maybe the ideal way for AI going back to that high-man thing. But that was a compelling vision, photo-realistic graphics, a lot of technical people are loved to play video games. So I think Jitson hired a lot of really smart people based on that. And then Jensen in a lot of ways is one of maybe the single most, I don't know, He's very important in the history of AI because very early, 10 years ago, you started
Starting point is 01:17:37 hanging out with people like Jeffrey Hinton and Yal-Lakud. I remember him saying those names to me and talking about the famous ImageNet competition when ResNet 51, and this means, effectively, I would paraphrase it, but we've turned intelligence into an engineering problem. The way we're going to solve intelligence is just by gluing more and more GPUs together. He saw that. it dedicated all of it being it to that. And then it became about intelligence.
Starting point is 01:18:05 That's a mission. And then for Elon, all of his companies are really mission-oriented. PayPal was about making local at the time commerce frictionless and reducing the tax from all this crazy fintech systems and payment flows, reducing that tax on the world. And I think that tax has been massively reduced, not necessarily because of PayPal, but because of PayPal and a lot of other companies like them. And then Tesla was really about,
Starting point is 01:18:30 making the world sustainable. And I think between battery storage and pulling EVs forward, Tesla, the world was always going to run on sunlight just because of economics. I mean, forget emissions or concern about the environment. Economics were going to dictate an emission-free world, apart from rockets, because you cannot get literally out of the Earth's gravity well without chemical propellant just from a physics of energy density perspective. Tesla's done a lot to, I think, make the world more sustainable. And it was just so striking to me when I first started meeting with Tesla executives back in 2011, 2012. They were at our so mission-oriented about making the world more sustainable and environmentally friendly. SpaceX, too. It's incredible if you talk to all of them.
Starting point is 01:19:13 I mean, it's just wild. Yeah, the bartender at SpaceX. If you ask the bartender at the Tiki bar, she's like, I know each engineer's drink, have it waiting for them at the end of a long day, not that they're having drinks that off. I'm like, maybe that makes me a step closer to Mars. But everybody, each company and all these companies, they can say the mission, and they all say it with this messianic zeal. I think you get better employees. And then XAI is about being dedicated to objective truth and new scientific advances. And if you have these missions and you're competing, a lot of these, the world's best minds
Starting point is 01:19:49 have spent 20 years trying to make people slightly more likely to click on one link than another. And if you have these missions, you get better employees. So I think that mission focus is a really, the exceptional teams, and then those exceptional teams want to work with people like Elon and Jensen, because they know if you're 25 years old and exceptional, you go to one of those companies, and you happen to be in a critical path of a critical problem,
Starting point is 01:20:12 you're going to work directly with them, and there's probably no other company where that is true. It has more than, I don't know, pick a number, 5,000 employees, 1,000 employees, I don't know. So I think that mission orientation, there's no need for hierarchy, you get really exceptional employees. I cannot tell you how exceptional the teams at Elon and Jensen's companies are. Yeah, I mean, I've met a lot of the SpaceX people for a bunch of different reasons,
Starting point is 01:20:40 and all of them have this same zeal, intensity, mission focus. It's absolutely remarkable. And a lot of them left and then came back very quickly because there was no other environment they could find quite like it. Maybe like the most fun place to close our conversation today, I could literally do this with you for seven straight hours. I got to like one third of my questions and your passion for markets and technology is just so palpable and amazing. My last question is about our business. How do you think the investing process, investing firms, edge alpha will evolve against this backdrop that we spent the last two hours discussing, whether it's the internal tool you talked about or the one we're building. I'm sure everyone else is building something. it just seems like, holy cow. And this has, I guess, been the history of our businesses that keeps getting more competitive.
Starting point is 01:21:28 But how do you envision it five, ten years from now as such a passionate practitioner? Public equity, so we both play video games. The meta of any given competitive video game is always changing based on how the game designers balance it. So if it's a PVP game, you might go from a sniper meta to shotgun meta. Sports are similar, but there's a fundamental truth,
Starting point is 01:21:48 which is you want to have the highest KD ratio in a shooter-based PVP. the more PV environment, you want to have the highest or the fastest clear of the most difficult apex in game activity. And the meta of how to accomplish all those is changing. Sports is the same. The NBA evolves from a mid-range jumper meta when Jordan was in it to a three-point meta, and it'll evolve back again, and this is competitiveness in the meta, but there's still an underlying truth in the NBA, which is you win by scoring more points than your competitor. The meta of investing, both public and private, I think, is always changing.
Starting point is 01:22:22 It changes much faster in publics. There have been all these huge meta shifts since I started in the business. ROIC was a revolutionary concept in the mid-90s. Everybody knows about ROIC. REOC was just a big improvement on ROE, which in some ways, focusing on ROE might argue, is Buffett's greatest contribution to the field of investing, more than the four filters, et cetera, et cetera.
Starting point is 01:22:47 But RICO was a big meta-shift. reg FD and how that changed company communication. That was a meta shift and one that I think was awesome. And from my perspective, Reg FD means there's almost no reason to talk to companies because I had kind of a unique experience. Like I was at Fidelity for 18 years. It's very funny to me when hedge funds are small asset managers, and I define small as, let's say it, under 500 billion,
Starting point is 01:23:12 say, oh, we have an access advantage with public companies. No, you don't. No, you don't. If you add up every hedge fund's access, all of it, it's a fraction of what firms like Fidelity Capital did T.R.O. have. It is many orders of magnitude. And what was awesome to me, I had the year off, and I just realized they never say anything that is not in a transcript. And I went back and read all the transcripts and things that I thought were like big insights through transcripts. Human beings, we have a big I-O problem, hence BCI's. And the eye for me of reading the I being,
Starting point is 01:23:48 the input. My input speed for reading is probably something like five to ten times faster than humans can talk. And my input when listening is bound by the output speed of my partner. So extremely low bandwidth form of communication, literally, in terms of bits or bytes of information per second. Like I think that and having transcripts of everything, that was a big meta shift. Credit card data was a big meta shift. Expert transcripts have been a big meta shift. I think all limbs are going to be the biggest meta shift. And I actually think they may be hardest. for purely quantitative investors, because I think there's going to be a period of five to ten years.
Starting point is 01:24:24 It's very clear that Renaissance, Bridgewater, Tussigma, they probably realize the importance of data and compute for a lot of other people. They're the only people who can compete with these tech companies who can throw up 20 or 30 million or 40 or 50 million at like a leading AI researcher per year. And they cornered a lot of data sources. They have data and no one else has,
Starting point is 01:24:46 and they're throwing more compute at it than anyone. but I am optimistic that human fundamental investors, the tool I'm using is called IntelPro, by the way, if anyone's curious. I think right now the only alpha left in the market for fundamental investors at the edge of probability because these quantitative investors, why is there stop being, why is there no dispersion anymore in value strategies? Well, because value strategies, the alpha from them, was based on human emotion and people being embarrassed to buy stuff and afraid to buy stuff and taking career risk. Well, algorithms, they don't have any of those.
Starting point is 01:25:19 And that's why there's not as much alpha. They're just really simple value strategies that worked amazingly well until, oh, 20 years ago when quantitative investing took off and they squeezed the alpha out of value strategies. That's not to say that value as a factor cannot perform really well and be the best before a factor. But there's just not a lot of dispersion within the value factor because of algorithms. I am hopeful that there is a five to 10 year period where if you're a fundamental investor like me who can get a slight edge on maybe future probability states over what's discounted in the market
Starting point is 01:25:56 through deep domain knowledge and try not to have any biases that I can combine what I have with this tool, with years of data. I'm just so grateful we have a vector database with years of data. I think it is going to enable me and other fundamentally based portfolio managers to probably benefit from a lot of the things that these firms like Renaissance have been benefiting for. And this is a long way of saying what LLMs do, what AI does, is it means the human language is the programming language. And now because of these LLMs are very soon, I'm going to be able to program effectively from an effective investment effectiveness at the same level as one of these.
Starting point is 01:26:39 50 million dollar year people are close enough. And then combine that with my own unique set of data and domain knowledge in my mind. And I'm hopeful that there's like a five to 10 year period here where fundamental investors share of the alpha in the market goes up at the expense of quantitative investors. That's my hope. We'll see if that happens. I don't know if it's going to happen. going back to that high man fusion chest, the Tart Chess, like, I hope I have a five to 10 year
Starting point is 01:27:11 here, you know, where I can prosecute that. Very quickly in the world of Venture, I think Venture's going to evolve outside of pure Series A specialists. Venture is going to evolve. Series C-it-up is going to evolve the way of private equity. Mainstream private equity. There's no sourcing advantage. There's no pricing advantage.
Starting point is 01:27:31 In fact, there's a pricing disadvantage. Because the highest bidder wins. You're just the highest bidder. Everything's an auction. literally every deal that these big firms do is an auction. So where they compete and differentiate is an operational value ad. And I think that is where series C and up growth equity is heading. And it's not just, it's true operational value ad.
Starting point is 01:27:54 You know, I've spoken about this. You know, like I think Valor does this. Valor does it. Really do it. You know, really hard operational problems. It's not just helping with HR, helping with PR. It's not the LP window dressing operational support teams. It's real operational support.
Starting point is 01:28:12 And I think that is where the world of growth equity is heading and evolving. But that's almost because of, in some ways, the combination of the rise of crossover funds and megafuns that would have happened absent LLMs. I do think LLMs are just going to make the knowledge part of venture so much more democratized. I think about IQ, EQ, and then there's KQ, knowledge quotient. and now it's going to really help to have a really deep differentiated domain knowledge database,
Starting point is 01:28:40 but you're going to have to work hard with an LLM to take your IQ from whatever it is up 30 points because people who you had between your combination IQ and KQ, you had 30 points on them, now they're at your level unless you use an LLM. But I just think it will really for a while place the emphasis on JQ judgment quotient.
Starting point is 01:28:58 Almost maybe for A's and B's, the most important skill will simply be assessing team quality. And it may be that's that five to 10 year window where VCs can still differentiate in the same way fundamental investors, maybe I hope will be able to take some alpha share from quantitative investors because of LLMs. But just it will be all about JQ at the C, the A, and the B. And some combination of JQ and EQ is this person exceptional. But I don't know. It's just a hypothesis. Fascinating stuff. Gavin, you're one of my favorite investors to talk to. I have loved today's conversation. I love how specific in detail that will
Starting point is 01:29:33 because you're also one of the most passionate investors that I've ever met about your craft. This has been such a total pleasure. Thanks for your time. Thank you, Patrick. This is awesome. If you enjoy this episode, check out join colossus.com. There you'll find every episode of this podcast complete with transcripts, show notes, and resources to keep learning. You can also sign up for our newsletter, Colossus Weekly, where we condense episodes to the big ideas, quotations, and more, as well as share the best content we find on the internet every week.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.