Invest Like the Best with Patrick O'Shaughnessy - Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Episode Date: August 25, 2026

My guest today is Neil Movva, founder of Sail. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background ...for hours or days at a time rather than answering a human in real time.  In that world, latency matters less and cost matters more, and Neil has built the whole company around driving the cost of a token as low as it can possibly go. What makes this conversation special is that it is one of the most detailed tours I have ever done through the full stack of intelligence, the software, the chips, and the power, and how all three connect.  Along the way we cover the trade-off between speed and cost that lives inside every GPU, his scavenger strategy for buying the chips and power nobody else wants, his contrarian view on Nvidia, and why the premium the frontier labs charge for being three to six months ahead may not last.  Please enjoy my conversation with Neil Movva. For the full show notes, transcript, and links to mentioned content, check out the episode page ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠here⁠⁠⁠⁠⁠.  ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at ⁠colossus.com/subscribe⁠. ----- ⁠Ramp’s⁠ mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ ⁠ramp.com/invest⁠⁠ to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, ⁠Vanta⁠ continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to ⁠vanta.com/invest⁠.  ----- WorkOS⁠ is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- ⁠Ridgeline⁠ has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ridgeline.ai⁠. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:20) Neil Movva (00:03:22) Building a Token Factory (00:05:32) The Rise of Long-Running Agents (00:08:47) Deep Research and Cybersecurity (00:15:12) The Full Stack of Intelligence (00:20:03) Throughput Versus Latency (00:24:58) The Future of AI Chips (00:33:19) Why Transformers Work (00:36:43) The Future of Data (00:44:05) The Market for AI Chips (00:47:56) Is the AI Boom Different? (00:51:08) Reinventing the Data Center (00:56:43) Scavenging Power (01:01:04) Where Compute Is Most Inefficient (01:07:02) Open Versus Closed Models (01:10:37) A Trillion Tokens a Day (01:12:42) The Contrarian Case on NVIDIA (01:14:38) Advice for AI Hardware Founders

Transcript
Discussion (0)
Starting point is 00:00:00 RAMP is the only platform built to make your finance team leaner, faster, and better, saving businesses 5% annually on average so you can stay focused on growth. RAMP customers grow revenue 3.2 times faster than the average American business. Visa, Vercell, Cursor, Stripe, Notion, 11Lab, Shopify, and 70,000 other businesses all now run on Ramp. Mine does too, and so should yours. Learn more at ramp.com slash invest. Open AI, cursor, Anthropic, Perplexity, and Vercell all have something in common.
Starting point is 00:00:29 They all use WorkOS. To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, SCIM, Rback, and audit logs. Instead of spending months building these mission-critical capabilities yourself, you can just use WorkOS APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on WorkOS. WorkOS is the fastest way to become enterprise-ready and stay focused on what matters most your product. Visit WorkOS.com to get started.
Starting point is 00:00:56 Felix by Rogo is a personal finance agent that turns a single prompt into finished client-ready work using your firm's own templates, context, and standards. Send Felix an email like, take these comments and turn them for me, or update my tracker with the context of these emails, and Felix sends back finished PowerPoint decks, Excel models, and sourced research. Felix works the way your team already does, delivering work quickly and accurately around the clock. Learn more at rogo.a.ai slash Felix. Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is a very much. invest like the best. This show is an open-ended exploration of markets, ideas, stories, and
Starting point is 00:01:35 strategies that will help you better invest both your time and your money. If you enjoy these conversations and want to go deeper, check out Colossus, our quarterly publication with in-depth profiles of the people shaping business and investing. You can find Colossus, along with all of our podcasts at colossus.com. Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of positive sum. This podcast is for informational purposes only and should not be relied upon as a basis
Starting point is 00:02:06 for investment decisions. Clients of positive sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.c. My guest today is Neil Mova, the founder of Sale research. Sale is building what Neil calls a token factory, an inference company designed for a specific kind of future,
Starting point is 00:02:27 one where AI agents run in the background for hours or days at a time. time rather than answering a human in real time. In that world, latency matters less and cost matters much more. And Neil has built the entire company around driving the cost of a token as low as it can possibly go. What makes this conversation special is it's one of the most detailed tours I've ever done through the full stack of intelligence, the software, the chips, the power, and how the three connect. Along the way, we cover the tradeoff between speed and cost that lives inside of every GPU, his scavenger strategy for buying the chips and power no one else wants,
Starting point is 00:02:57 his contrarian view on NVIDIA, and why the premium the Frontier Labs charge for being three to six months ahead may not last. Please enjoy my conversation with Neil Mova. I think it's important early in these conversations to just say the thing, literally what you're building and what it does today. So maybe just orient us there
Starting point is 00:03:16 with a brief description, like literally what the system is that you're building and why it should exist. Sale research is a token factory. We have an API where anyone can send us requests where they can use large language models, open source, language models for any tasks they want. We will serve those tokens to them at a price that is unbeatable in the market. We also support their ability to build agents on top of this.
Starting point is 00:03:37 We host what we call sailboxes, which are long-running agent virtual machines hosted in the cloud that are designed for agents that run for hours, days, or weeks. So I should think about you as a peer company to others that serve different kinds of inference. You're serving one specific kind of inference, and your goal is to be the absolute cheapest provider, an enabler of a certain kind of use of intelligence. Exactly. The theme of our company is abundance. We want to deliver this new commodity of intelligence to as many people as possible at a cost that is sustainable for almost every industry. We think that whenever you make something 10 times cheaper, it's a new product category. We aspire to do that for tokens. We think it's so profound that the machine can think,
Starting point is 00:04:15 and now our job is to make as many machines as possible in the world work towards thinking. So if you think about the theme of the day being token costs, is token cost the right way to think about this? Is there some other way you'd put it? To start with absolutely token cost. Today my North Stars, I want to have the lowest cost per token in the industry and do that by a mile. I don't think tokens are the final unit of work or intelligence, but they are what we use today. After tokens, you start to move more towards more outcomes, which is like a vague direction.
Starting point is 00:04:43 You can imagine, for example, today when you consume tokens through an agent, you don't actually control how many tokens the agent reasons for. It can reason for a certain amount of time or it can call a certain number of tools. And increasingly, I think we will have agents do some unit of work. take as many shots on goal as they can. And however many tokens they used to get there is going to be a dependent variable, depending on the task. So you think about like agents that self-administer a token budget as opposed to a company sending a budget for how many tokens engineers can spend per month. Why is there an opportunity that you can tackle?
Starting point is 00:05:16 It seems like the entire world is oriented around more, better, faster, cheaper tokens right now. It seems like the world is trying to solve this problem very aggressively. What was the unique opening that you saw that maybe the market's not being efficient in its attempt to tackle this? So I think there's two things that are tailwinds for our company. One is got to be the rise of open source. I had to talk with that first. I think we're starting to see an increasing number of our customers and the broader market care about owning intelligence. They want to have control sovereignty over the thing that they depend on.
Starting point is 00:05:47 That created a much more robust market for our customized models or even just like these vanilla open source models that no one can ever take away from you. You always have the weights, you always have the right to deploy them however you like. In that world, there's been a reasonably robust market for the past couple of years, serving these models at large scale. The challenge is all those companies, you could take your pick, Base 10 fireworks together, they all focus on low latency inference. And they were pulled in that direction by one very important customer, cursor. I think that that was the right choice about a year ago.
Starting point is 00:06:17 And as a six months ago, it started to look like maybe low latency wasn't the only thing you wanted from an agent. You wanted more persistence, more long horizon tasks. And now it's, to me, very obvious that the future of magnetic inference is long horizon tasks. You're going to run the machine for hours or days at a time. It doesn't matter if it's been set tokens at 100 tokens per second. Maybe 10 is just fine if that comes with corresponding advantages and efficiency. Why are you so confident in that?
Starting point is 00:06:43 To me, it seems like I want everything as fast as possible. When you're waiting on it, you absolutely deserve the fastest sense are possible. Yeah. My trick is I don't want you to be waiting on it. I want it to be proactive, I want it to be in the background. One way to say it is like the best latency is no latency at all. When you wake up in the morning, the work's already been done overnight. You didn't even have to ask for it.
Starting point is 00:07:00 That's the dream. We're not quite there yet. More importantly, I think the more you're in the loop as you prompt agents and wait for a response. In fact, you're the bottleneck in having the agent do more or less work. What we'd like is the agent to operate on more human time skills. You don't manage your colleagues every five minutes. You ask them to do a high-level task and you come back and check in maybe every day, but more likely once a week.
Starting point is 00:07:20 And that, to me, is the future of human agent collaboration, more like human time skills. Say more about the early indications that this is happening and therefore you should be building this company. So the first and most important thing is the idea of test time compute scaling. The idea that you can give an agent more time and it will give you a better answer. So that was theorized about two years ago now. But it wasn't really something that we could actually bet on until, I would say, late last year with Opus 4 or 5. Opus 4.5 was the first agent that was at all suitable for longer horizon.
Starting point is 00:07:49 It was pretty mediocre when it first came out, but you look at the more recent models and what we've done on open source as well. And you see that agents are capable of running for an hour at a time. I wouldn't say it's days, but definitely an hour is quite suitable today. Seeing that like average task length get longer and longer, it doesn't take many points to have you draw out the exponential and see that agents are worth running for longer periods. What do you think will be the market share of long running agents in three years or something like this? I love this market because it's unbounded. There's no human in the loop so you can consume as many tokens as you like in the background versus human attention span. If you tell me to consume 10x as many tokens at Codex or at
Starting point is 00:08:26 Quad Code, I'm actually not sure if I can anymore. I'm already in the loop and locked in coding for most of the day that I'm at the laptop. What is unbounded is how many tokens can be consumed in the background or proactively. Long term, I think, you know, we're going to end this year at maybe 50-50 background and real-time workloads, but I see this going to 90-10 in favor of background. What are your favorite examples of something that gets accomplished much better as a background task than as a human-in-the-loop task? Most deep research, most questions where you want to have a definitive answer over not 100 sources, not 1,000 sources, but 10,000 sources or more. If you want to build an authoritative index of information, like, for example, one of our customers' parallel
Starting point is 00:09:06 web systems seeks to do, they want to build an index over the whole internet, and they want to monitor the internet in real time for changes. That is the kind of crazy exabyte scale task. that you need a very different kind of intelligence or scale of intelligence to achieve. Deep research is a top category for us. And then increasingly we see cybersecurity following this direction. If you think about, yes, there's so much code you can generate, but there's exponentially more ways to break that same code
Starting point is 00:09:30 than is to generate that code. There are some great customers out there who are working very hard to find agents that can break any piece of software and proactively patch them. When Fable first came out, for example, or Mithosk first came out, basically there was this push in the cybersecurity community
Starting point is 00:09:44 to run Fable against every line of code we've ever written and look for bugs in 20 different ways, meaning you're looking for memory errors, you're looking for business logic errors, and looking for network vulnerabilities, all these things. And these are all actually things that you would write specialized agents for. You wouldn't just have Fable look at the source code once, you'd have it actually set up environments where you can pen test these applications. At some point, people started to make this joke that security has become proof of work. When you want secure software, it's really a question of how many dollars did you spend on Anthropics APIs trying to break into your software.
Starting point is 00:10:17 That is the best indication for how secure it is, because that's the best tool in the world. And increasingly we found that open source models, well, the frontier of intelligence here is quite jagged. It's not the case that Fable finds a superset of all bugs in software. You would find some bugs with a very small model that you don't find with a large model. You'd find some bugs with Haiku that you would find with Fable and vice versa. So it encouraged this very diverse approach to sampling and trying to build cybersecurity agents
Starting point is 00:10:41 that break software autonomously such that you can patch them. If you were to get speculative and imaginative about the sorts of things that very cheap, very long-running agents can enable, we talked about some very practical examples, deep research, cybersecurity, et cetera, if you get a little bit dreamier about the use cases, new product category that this sort of inference will unlock, I guess the question is just like, so what? If you're maximally successful, dream a little bit about what that might enable. I think for individual users, what I'm excited about most is this idea of proactive, intelligent agents. You can imagine a Siri that is running in the background all the time to understand
Starting point is 00:11:19 all the emails you received in a day, all the text messages you receive in a day, and has a much more encyclopedic view of your life and how to be helpful in that life. Right now, there's still point solutions. You end up doing a lot of prompting. Siri is not very proactive. That's something that we can fix with abundant inference. If you trust the machine enough that it's reliable and also trustworthy is in private. You might even imagine the machine can understand how you interact with it and proactively surface your next action. Whenever you open your phone, can we build a good model of what you're going to do next? My estimation is yes, we totally can. And the key to that is incredibly cheap intelligence. You have to be willing to spend tokens without any promise of return.
Starting point is 00:11:55 That is the unlock. The long lens view to take on this is that we have a form of intelligence that can tackle any verifiable problem. Any verifiable problem means most software. It means a of formal math proofs and similar. And it could also mean scientific discovery. These are all relatively verifiable problems. And all those things currently have a dollar cost attached to them, essentially, that's a hidden one. It's like how many tokens could you possibly harness to make this work. We have started to bring you within view a dollar cost for these long horizon tasks that is reasonable. It's not millions, it's thousands. And maybe it could be hundreds or even tens of in the near future to have a definitive answer to any scientific question, to any research problem.
Starting point is 00:12:39 So if we dream about that future, we then become limited just by the questions that people can ask, basically? Pretty much. The questions we can ask, the models are on the cusp of basically taking even a high-level question and chasing it down. Every possible follow-up, you can have the model essentially take that on its own. And the question is, what is your token budget? And we will solve the token budget problem. What about non-verifiable tasks? Those are, basically the entire category of human taste into that category. We have not solved human taste yet, and I don't know that it fundamentally can be. I'm excited to be surprised here, but we are focused on very quantitative problems. We leave the quality of writing, we leave the beauty of art to people.
Starting point is 00:13:18 I think for individual users, what I'm excited about most is this idea of proactive, intelligent agents. You can imagine a Siri that is running in the background all the time to understand all the emails you received in a day, all the text messages you receive in a day, and has a much more encyclopedia view of your life and how to be helpful in that life. Right now, there's still point solutions. You're doing a lot of prompting. Siri's not very proactive. That's something we can fix with abundant inference. Vanta automates security and compliance for over 16,000 fast-moving companies like Ramp, Cursor, and Harvey, keeping an audit ready around the clock. It's the number one agentic trust platform, and it now helps companies like yours watch for the risks that show up
Starting point is 00:13:57 between audits, across your vendors, your AI tools, and your whole environment. Every new tool your team signs up for, every vendor that turns on AI features is an opportunity for something to go wrong, and most security programs weren't built for AI's pace of growth. The Vanta agent works like a 24-7 GRC engineer in the background, finding issues, drafting fixes for you, and cutting vendor assessment time by up to 50%. Whether you're a fast-growing startup or a global enterprise, Vanta helps you earn and prove trust. Invest like the best listeners get a special offer for $1,000 off at Vanta.com slash invest. Ridgeline is the first end-to-end system of record with embedded AI for investment management firms,
Starting point is 00:14:39 running portfolio accounting, reconciliation, reporting, trading, and compliance on one unified platform. Firms are moving off legacy technology and onto Ridgeline because of how far ahead Ridgeline's AI features are compared to anything else in investment management software, which is why I believe that firms that come out ahead in the AI era will be the ones running on Ridgeline's unified platform. If you're serious about your firm's AI strategy, Ridgeline should be placed. part of that conversation. You can request a demo at ridgeline.a.ai. All right, now let's talk about the very clever stack of solutions that you hope to build, ultimately to have this giant token factory, supplier of extremely low-cost intelligence.
Starting point is 00:15:21 I think you think about this in terms of software, hardware, and power. Talk through what your master plan is to approach this challenge that's so different from what others are thinking about doing. We always have to start with software. Where is the opportunity on today's data centers to improve efficiency. And the first thing we did was we tried to build the entire LLM software stack around peak GPU efficiency, meaning we're using a video GPUs. We wanted to squeeze out more tokens from the same chip than anyone else in the world. And that starts with the lowest level of programming, kernels. It's actually my background. I spend my whole professional life working on GPUs and kernels. InVi was my first job while I was in college. I got to see how the
Starting point is 00:15:57 TensorFlowCores got to earn their right to be on the chip. What does that mean? Like, what is the TensorFlowCore? TensorCore is a specialized unit on the GPU. that accelerates matrix multiplication. Simple as that. There's been a long history of how we evolve that hits a core over time that we'll get into. And why is matrix multiplication so important? I cannot say that there is a divine truth of the inverse
Starting point is 00:16:15 that explains why matrix multiplies seem to be the atomic unit of computation. But one way I've heard it described to me is, well, it's a really succinct way to mix two blocks of numbers together and have them interact in some interesting way. That's as much as I can say about it. It is really convenient that linear algebra turns out to be a very compact representation
Starting point is 00:16:31 of arbitrary relationships and data. So, Nvidia, great graphics company, had market share dominance in GPUs and gaming graphics for quite some time. And then starting in like the mid-2010s, they started to actually start these like Skunkworks projects to make the graphics processor more suitable for machine learning tasks that they were tracking. I remember actually reading some of the lab notebooks of some of my managers when I was at Nvidia. They would visit these small ML conferences like ICML or NUREPs at the time. They would just take note of these papers like, oh, this deep learning thing seems to be catching on. And what's really interesting is that these grad students are using gaming and video GPUs in order to train their large models. We should double-click on this and figure out what's going on here.
Starting point is 00:17:12 By 2015-2016, at least Jensen had the conviction to kind of double down on, hey, this usage of our chips is only going to grow. Let's start allocating more and more precious silicon dye area to this capability that seems to be emerging. Let's put the first version of TensorFlow cores on the chip. So we're talking about taking this gaming chip, which is designed for painting pixels on a screen. and adapting it to do Michigan multiples. It was early. You would be competing against the graphics teams, essentially,
Starting point is 00:17:40 when you ask for more silicon area and any chip company, there's always competition for that. It is something that the designers guard so carefully. You don't ever want to invest in the wrong technology because that's opportunity cost that you could have allocated it to some other functionality. We fought tooth and nail and got just a tiny bit of diarrhea, maybe like 5, 10%, something like that,
Starting point is 00:17:58 for the first generation of these chips to get some amount of acceleration for basic convolutions, which were the fundamental operation for computer vision models of the day. And then we had a software team that was trying to squeeze all the performance we could out of the chip. And I think on that software team, which is where I work, that's what actually taught me the most about the ethos that NVIDIA has around this term called speed of light. They always chase the speed of light for any piece of hardware that they make. It is so ingrained in every engineer's mind that if the machine can do it,
Starting point is 00:18:27 we're going to push the machine to the frontier until it does what we think is going. The speed of light is the edge of what's possible. The speed of light is the edge of what's possible. Exactly. If we think the chip can run at this frequency and produce this many multiplies per cycle, we're going to get there. We're going to break every bottleneck and get to that peak level of performance. To this day, I tell all my engineers, we're chasing 100% speed of light.
Starting point is 00:18:46 I don't care about relative numbers versus the competition. I only care about absolute numbers. What are we able to do on the chip? How do we achieve that? Before we leave that chapter of your time at Nvidia, anything else beyond that cultural touch point that changed the way you think about things or that stood out the most about how the business ran back then or its culture? I have a ton of stories about Nvidia.
Starting point is 00:19:02 I can tell you a few of them. One of my favorites is that on the tenure side, a lot of people I worked with in Vyndi-2015, 2016 are still there today. That company has incredible retention, and these are the best engineers. On the Silicon side, at least I've worked with in my whole career. They're extremely, extremely motivated and passionate. They believed in parallel computing as a concept
Starting point is 00:19:20 through its various incarnations and have loved seeing the chip evolve. This is their life's work, and they're extremely competent in that direction. They're also a very frugal company. NVIDIA and all of Silicon Valley companies, after 2008, they had some cutbacks and like perks. So no free lunch, for example. Invidia took it one step further. There was no free milk in the fridge.
Starting point is 00:19:38 So if you wanted to drink coffee at NVIDIA and you wanted some milk, you'd actually have to chip in a dollar every month to the milk club, and the milk club would stock Costco milk in the fridge. And I remember that distinctly. We don't do that at sale. It's a frugality that permeates the company.
Starting point is 00:19:53 So coming out of this time there, you get this experience of what it's like to develop more efficient, usage of the underlying hardware through software. So link that to today's environment. The GPU is fundamentally a throughput machine. The GPU's happiest when you give it a lot of work to do and let it chew through that work at peak utilization of its compute units. But that's actually not the way that we've taken AI in the last couple of years. We've pushed AI to be an interactive chatbot tool is the most common form of AI usage today. In that world, you care a lot
Starting point is 00:20:21 about actually spitting answers out to the person of the keyboard as quickly as possible, to your point about don't make the user weight, I want things as fast as possible. That's actually quite interesting for the GPU. It's very difficult to put the GPU in its happy path of being fully compute utilized when you're trying to spit out tokens quickly. There's a fundamental tradeoff on the GPU between being throughput oriented or latency optimized. And everyone has chosen latency optimization because the shape of usage was chatbot oriented. I believe that's the most profound change we're going to see in the next year. We're going to move away from chatbots to more proactive or background agents. In that world, it makes a lot more sense to build a stack around throughput.
Starting point is 00:20:56 Can you explain technically why the tradeoff between throughput and latency is unbreakable? Why can't we have both from the same hardware? It's quite foundational in almost every system that you could ever possibly look at. There's always a tradeoff between getting a small amount of data through the system as quickly as possible and leaving a lot of buffer room for that, or trying to run wide and slow. Narrow and fast or wide and slow is like a classic tradeoff in all computer science. But for GPU specifically, I think there's one thing to focus on, which is there's this concept of like batching on the GPU.
Starting point is 00:21:26 We want to group many users work together into a batch that we can run all at once on the GPU. That's the parallel processing of the GPU. We'd like to have a lot of parallel work to do. The thing is, though, you're doing net more work when you run a large batch of compute together. So you might be filling all the units, but every step along the way as you carry a batch of work through the GPU, there's more work to be done.
Starting point is 00:21:50 individual token or any individual user's request in that batch, it's going to spend a longer time on the GPU being carried with other people's traffic. Maybe the way to say it is, if you want to get downtown an SF, you can take the bus or you can take a private transit. And the private transit is going to have its own direct path as the crow flies or using exactly the roads that you want from point to point B. A bus, it's going to have to serve many more people and it has to fundamentally do something that works for everyone. So it takes a slower path and it stops and waits for other people to get on and off. I think the bus versus car analogy is pretty accurate. It's a great analogy. And step one for what you're trying to do is create the
Starting point is 00:22:25 best possible bus on top of Nvidia GPUs. Like that's step one of your optimization. That's exactly right. It means we explore things like different parallels and schemes. And maybe that's another example I can give you is with invidia GPUs. One of the things that they've really innovated on and a great job with is the NVLink interconnect between GPUs. And in fact, that NVLink system is so good that if you have a large matrix multiply that you want to perform faster, you can actually cut that matrix multiply in half and shard it across two or more up to eight, let's say, Nvidia GPUs, and have them all work on pieces of that larger matrix multiply and have them connect to their results together at the end, reduce their results back together at the end. This is a great, great way to cut
Starting point is 00:23:05 the minimum latency of an operation. Each GPU is now doing one-eighth as much work, let's say, and therefore it can finish faster, but not eight times faster. It's sublinear scaling. You'll use eight times more hardware, but you won't get eight times the speed. You might get like four to five X the speed. You're not going to get strong scaling. This is because of communication overhead. It's because every GPU is going to be a little less efficient working on a smaller tile of work than a larger tile of work. It's the only way to speed up if you want the minimum latency possible.
Starting point is 00:23:32 You can do that. But it is not the choice I would make. For example, I would prefer to use a different parallelism scheme like expert parallelism or pipeline parallelism. We may do interesting things to overlap and hide the communication latency in a way that you would have less ability to do that for low latency service. So is right way to think about NVLink as a technology which improves latency performance? Yes. And only latency performance. Which will segue into the next segment of what we do differently as a company.
Starting point is 00:23:58 But yes, NVLink is mandatory for low latency inference. So, Nvidia is excellent at low latency inference. And I'm telling you that we don't really care that much about low latency inference. So where does that leave us? I'm not holding my breath for other companies broadly to figure out NVLink quickly. It's challenging technology to figure out. It's hard to scale. It's hard to productionize.
Starting point is 00:24:16 If I do have some other vendors chip and it is good at the foundational compute components, it can still be matrix multiplies really well. It just can't communicate there's results across its peers quickly. Well, maybe there's room for that other chip in my stack as a really, really good compute per dollar option. That's what I actually optimize for in most cases is how many flops does this chip? have and how much is it going to cost me per hour to own and operate. There are other ships that definitely rank higher than Nvidia on flops per dollar, but they may not have as much interconnect. So it's my job to figure out what parallelism scheme am I going to use that's going
Starting point is 00:24:49 to make this chip suitable for inference. It's not going to be tens of parallelism. Invidia is basically mandatory for that, but other techniques may work well for me. So before we leave the latency part of the story, can you comment on companies like Cerebris or others that can perform incredibly fast operations. I'm curious like what you think about those approach of those companies, what might happen in the future? What is your prediction for the future of very low latency-focused hardware? Cerebrus, Grock, and a couple others that are coming out of stealth now, I think have made a very interesting bet on not just building another GPU, but actually building a different kind of accelerator that focuses on a different memory hierarchy. They want to maximize the amount of
Starting point is 00:25:30 S-RAM on the chip and use that as very, very fast memory for weights and KV cache. So S-Ram versus DRAM, there's two ways to make memory for a chip. One is to integrate the memory on the logic die itself, meaning you tell TSM, I want this many megabytes of storage on my chip, and there's a way to build that. TSM has a standard cell library you can use, and you can just print out a bunch of cells of S-RAM. The problem with S-RAM is it takes a lot of area on the silicon dye. So if you want to build a large dye, like let's say the NVIDIA-Block.
Starting point is 00:26:00 well at 800 millimeter square. If you made that whole DiasRam, it would be in the maybe like single digit gigabytes, it feels like. It's not a crazy amount of data storage. Compare that to if you're willing to take a different process technology entirely, so not TSM anymore, but now Micron, SK-Hinix, Samsung. They build DRAM, which is a whole different way to build memory that's more focused on capacitors than transistor cells. The standard way to build SRAM is what's called the 6T transistor cell. It's a stable transistor arrangement that allows you to write a bit to it and then it holds that state in that bit, regardless of whether you keep applying, well, you had to apply some power, but it's holding that bit without any sort of like active management. It's static. Now, dynamic RAM,
Starting point is 00:26:42 DRAM, it's dynamic because what you do to write some data is you write a charge onto a capacitor. And as soon as you write that charge into that capacitor, the charge is dissipating. The dynamic part of DRAM is that you must every 50 milliseconds or so refresh every bit you've written. So you're constantly juggling billions of balls in the air, essentially billions of bits. Had to be managed by a memory controller, which is reading and refreshing every bit on the DRAM. Now, the benefit of that is you can get much, much higher density. And it's a whole different process technology. There's a ton of different trade-offs, hence where we split the DRAM manufacturing into an entirely different company,
Starting point is 00:27:14 like Micron, SK-Hinex, and Samsung, these are the best companies in the world to do this. They build DRM. And if you take DRM from those companies and you stack it into many layers, and you kind of print them or solder them around the main logic. die that you get from Nvidia, you can now get hundreds of gigabytes, like Blackwell has 288 gigabytes of HBM capacity around the logic die, and the logic die itself maybe only has like 500 megabytes of S-Ram. So it's possibly multiple orders of magnitude, three orders of magnitude difference in density for DRAM versus S-RAM. So let's go back to CREBRIS. What are they doing?
Starting point is 00:27:48 They see this problem. There's not really an obvious way to increase S-RAM density on the chip, But the thing with S-RAM is because it's so physically close to the logic gates that actually do the computation, the arithmetic logic units are right next to the S-R-R-Ram that they're going to pull from. The compute units that are doing the major cell applies can pull data from S-Ram at mind-boggling speeds. Scribes quits petabytes per second, 21 petabytes per second for their waiver scale engine 3. Compare that to HBM on an NVIDIDIDIDAN blackwell is 10 terabytes per second or so in that range. So once again, many orders of magnitude difference. more capacity, but proportionally less bandwidth, essentially. What Cerevers does is they say that we're going to take as many of these dyes as we can.
Starting point is 00:28:30 We're not going to limit ourselves to the 800 millimeter radical limit, the TSM, 800 square millimeter limit that TSM imposes on us. We're going to take the entire wafer and have every dye connect to every other die over scribe lines, and we're just going to try to get as much S-RAM as we can on the whole wafer. And we can get to like, let's say, 50 gigabytes of S-Ram per wafer. and then we're going to stack many wafers together in a pipeline or similar. And now we can have up to a terabyte of memory, very, very fast memory. You do all that work just to get to the ability to read data from SRAM at 21 petabytes per second per wafer.
Starting point is 00:29:04 Therefore, you can now serve these language models at extremely high tokens per second because you can move the entire parameter count of a large model like Kimi. You can move all that data in and off the logic cores in about a millisecond or something like that. So there you go. you have a path to a thousand tokens per second. So what is your prediction for like that segment of the market? What happens to them is some hybrid. We had to pair the cerebrous chip where it's very strong.
Starting point is 00:29:30 It's very, very good at fast access to memory with something that has more capacity for memory. It's true that you can take a one trillion parameter model like Kimmy and fit it on a large number of cerebris diet waifers. But you can't do something about the KV cash very easily. The KV cash is something that grows as people, use the model more. And that is always dynamic. You don't even know how much KV Cash you're going to need. It depends on how many users you have and how many users you want to serve. Can you explain
Starting point is 00:29:56 KV Cash? KV Cash, when you ever use a language model, every token you send through the language model stays in the context window of the language model for as long as you're having a conversation. We talk for 100,000 tokens. The 100,000th and 1th token is still in the conversation behind us and the model is referencing all the past conversation history in order to make better predictions about what the next thing we're going to say is. That KV Cash, it's a bunch of memory. You have to store a representation for every token that we send through the language model, and it frequently gets to be larger than the weights of the model themselves.
Starting point is 00:30:30 You have this crystallized knowledge in the model weights, and you have the dynamic knowledge of the exact conversation we're having in the KV Cash is the way I like to think about. And this is why sometimes people would observe deep in a conversation, things start to degrade because there's some sort of technical problem? So the KV cache is quite interesting in that regard. The KB Cash is an exact representation of everything that came before. We store all the information that we've seen in the conversation.
Starting point is 00:30:53 However, during training, the model did not get trained primarily on very long context conversations. It got trained primarily on, let's say, 8,000 token conversations or 16,000 token conversations. So if you take the model to 200,000 tokens, there was some training that happened at that context length, but it's not the models like core strength. And so there's always been a challenge for the frontier labs to figure out how do we make the model exactly as intelligent at 10,000 tokens as we expected to be at 200,000 tokens. And it's going to be a perennial battle for us. We've had one million context windows as a concept for years now. Anthropical was, I think, the first to hit
Starting point is 00:31:25 that one million context window length. I still, you know, use slash compact in my quad code well before one million context length. I don't think it's actually great to hit the full length. These extremely fast, extremely low latency approaches ultimately are limited by this factor. Yes. You can do whatever you want for the weights. It's very possible to have much unbeatable performance on weight storage. However, the KV Cash is going to be a big thorn on your side. So three years from now, five years from now, what role do you think these kinds of chips play? Like what sort of market chair do they have in the heterogeneous chip market?
Starting point is 00:31:56 Like Cerebris and GROC and maybe a couple others, you should think of them as accelerators. What they are really good at is being used in conjunction with a more traditional GPU-like device that critically has this off-chip memory built in. You want off-chip memory for capacity and on-ship memory for speed. We want to hybridize these two things. So if you take transformers in the limit, you take a transformer to a million context length. What ends up happening is you have this compute-bound stage, which is the actual matrix multiplies for the what we call the MLP,
Starting point is 00:32:26 which is where most of the models knowledge, world knowledge, is encoded. And then you have the attention layer, which is where we're dynamically adapting to the current conversation. Attention in the limit is usually memory-bound. The MLP in the limit is compute-bound at large enough batch size. I would say the original sin of Transformers is that you've taken this extremely fundamentally memory bound layer and juxtaposed it right next to a compute bound layer. It is very difficult to have a single chip that is good at both compute operations and memory operations. The GPU is quite balanced in this regard, but you had to choose one or the other.
Starting point is 00:32:58 Srebris has a very fast memory access for something like a matrix multiply, and it's really good to host the MLP, the weights, essentially, on the Syrebus chip. But the GPU has the capacity to scale to really long context lengths. You would like to put the attention possibly on the GPU and the MLP on the Cerebris chip. And I believe this is what's happening with Nvidia and GROC. Can you riff for a minute just on Transformers? For people that, again, aren't deeply familiar with what this innovation was in 2017, what its strengths and weaknesses are and whether or not you think it will remain the dominant architecture
Starting point is 00:33:30 or a dominant architecture for the future of AI? What it did, it allowed us to learn on unsupervised data really effectively. Because transformers, what they're all about at the end of the day, is taking any sequence, any arbitrary sequence of data and trying to find patterns in that data. And they critically, the attention operation, which is the headline component of transformers, it allows the model to dynamically adapt to what it thinks is the most relevant component of the sequence. Every step you take through a transformer, you are essentially like re-waiting the input that
Starting point is 00:34:04 you looked at before and figuring out which is most relevant for your next prediction. It's extremely amenable to learning arbitrary sequence data. And the most interesting sequences of data that we produce on a regular basis is language. And that's how we got to dominance in the language regime. To zoom out even further, I think what Transformers really did well is that they scaled. Transformers make no such human prior. Transformers just say, well, there's going to be a pattern in the sequence of data. And if there is a pattern, I'm going to find it. I'm going to throw more and more parameters at this problem until it works. Transformers benefit from a lot of the computer vision work too. For example, one of the challenges in computer vision was we had a hard
Starting point is 00:34:40 time going from hundreds of thousands of parameters, which you get for linear models like support vector machines or other legacy machine learning models. Those had thousands of parameters. Then we got to deep learning and got to tens of millions of parameters with computer vision. The biggest models were 150 million parameters was a huge model for computer vision. And now we routinely talk about trillions of parameters. And transformers are the link to go from millions to trillions of parameters. So if I think about the important units of scaling being data and compute, does it stand to reason that you think transformers will just stick around because that's the thing that we're good at getting more of those two things?
Starting point is 00:35:14 Transformers are such great sponges. You increase the compute available to a transformer by 10x, and you'll get some log improvement somewhere. And so far, the scaling laws really work. They're really quite beautiful. And to the point about what do transformers do really well, they extend to almost any data set you can throw at them. They're extremely powerful general learners.
Starting point is 00:35:31 And I think what's especially useful about transformers over other techniques that we've tried to replace attention is transformers represent any pairwise relationship that you want. Any token in the sequence can attend to any other token in the sequence. So if there's any relationship that's in the sequence at all, you're going to find it with a transformer. Now, it may be the case that you don't need all to all modeling. You don't need every token to look at every other token. But if you need to, transformers give you that option. And until we know a better way to prune that space down, a better way to kind of have information modeling be more selective. Attention is a very, very good operation.
Starting point is 00:36:09 This is another trick that we learned in the computer vision days. One of the old Carpathie sayings is that if you have a new data set that you want to train a model for, your first goal should be to overparameterize the model and try to overfit the data that you have, to prove that there is a relationship that you can model or memorize, that your learning algorithm works, that you can instill knowledge into the model. Once you can overfit, then you can compress. And the compression is how you get generalization. You don't want to actually memorize the data that you have in front of you.
Starting point is 00:36:35 You want to generalize, and therefore, once you overfit the dataset, then you can kind of work backwards and try to find the general patterns that fit into the smallest parameter account possible. What's your prediction for the future of data and riff on the importance of data in this whole story? I like the phrase that intranet was a one-time subsidy on data. We got it for free. It's extremely high-quality. There are about 30 trillion tokens of high-quality tax. 300 trillion tokens if you take a wider view on what qualifies as a good text.
Starting point is 00:37:01 And we've basically looked at it all already. Models have seen the entire internet many times over at this point. And there is not a whole lot more to be done on human data from the internet. The next phase of data in my mind is model self-improvement through RL environment, gyms, basically. Now, in fact, we don't even benefit from getting more random user interactions with AI. It used to be that the new type of data that we cared about a lot was the interaction data from people using chat chvety and giving chatubt signals on what they liked and didn't like.
Starting point is 00:37:30 I like the argument now that the median model that we serve is so much more advanced than like a random human giving feedback that the signal you get from random human preference, unconditioned human preference, is not actually worth anything anymore. You want expert human preference at this point. The model has outgrown every day job. Yeah, every day show. Exactly. So the future of data to me is giving the model a hard, verifiable task and letting it run in this gym where it's kind of isolated and it just has a problem and it can make progress on it and get measurement of whether it made progress on that problem or not. You can imagine coding problems are in this category, math problems or also in this category. More and more, we can just give the agent a computer essentially and have it act like it's a human
Starting point is 00:38:10 worker and give it feedback on whether it's making progress towards the target outcome. That environment becomes the data. I think this is not a super differentiated take, but it's been really, really productive from what I've seen so far. And you think that just goes on for a really long period of time? Or is that another like, if I think about the internet, is this one big block? Like, this is another big block that we'll have its day in the sun
Starting point is 00:38:30 and we'll kind of get it all. And then we'll have to move on to something else. I think it's actually more profound than that. Basically, the idea is that if you want artificial general intelligence, the best way to get there is to just keep stacking specialized intelligences until you have no more gaps to fill. The only thing you need to make sure you do to make this work is you must make sure that you'd have to have.
Starting point is 00:38:50 is verifiable. You need to give the model a self-grading system. If you have that, you have the recipe for self-improvement on any task you like. I think you've seen this held up by the way Frontier Labs spend. They used to spend that much on data. Now they spend a lot more on RL environments. These environments absolutely capture that relationship of recursive self-improvement on a verifiable task. Coming back to your initial task of making existing hardware more efficient by being more in control of what's going on at the hardware level through software. Keep going on what you've done so far and what you want to do. And then we're going to jump to hardware and then jump to energy finally.
Starting point is 00:39:28 I mentioned kernels. It's surprising. People think kernels are done. There are great people like Trudeau who write excellent kernels. And they form the bedrock of all of our modern deep learning is built on flash attention. Modern transformers are built on flash attention. But if you deviate from the happy path at all, if there's a new model that comes out that has a slightly different way to embed
Starting point is 00:39:46 positional information, like the change of the rope system. Suddenly, the kernel that we had is not suitable for this new model, and we may have to make a patch to this kernel. I wouldn't say we're in the phase where we had to invent new kernels from scratch, but having the ability to quickly modify existing GPU kernel, sorry, a kernel, and by the way is a general term for any program you run on the GPU. Historically, kernels tend to be put into a library where every kernel has a very, very scoped purpose. Typically, you have a kernel for a matrix multiply. You have another kernel for even something as simple as addition. You want to add two tensors together, that's another kernel. And then increasingly, we've started to fuse those kernels together. So if I do a matrix
Starting point is 00:40:24 will apply, and then I want to add it to another matrix that I also multiplied, maybe those two become one kernel, and I just fuse the operations where, sort of writing the data out to DRAM, and then reading it back in just to do the addition, maybe I can just do this easily. Why are humans still doing this? Like, it seems like the sort of thing that AIs would be exceptionally good engineering, more efficient kernels. Maybe that's where we're going, and we're just not quite there yet, but if we aren't there yet, is that where we're going? If we're not there yet, why humans still doing this? Why is TreeDow so well known? You know, it's a name I know. I don't want to speak for Tree, but what he taught me was you shouldn't write kernels by hand
Starting point is 00:40:58 anymore necessarily. I like to say we write kernels in the whiteboard. We go to the whiteboard, we describe what we think the machine should be doing. Then we succinctly describe that in the natural language to a model. And then the model is able to do the execution of, Here is my input and output. Here is the strategy of how we want to dispatch this work onto the GPU. I'm going to go implement. So we're doing the conceptual design. Exactly. I'm not sure exactly why models are not superab at doing this.
Starting point is 00:41:23 I don't think this is like R-Mote or anything like that. I'm sure in six months' time, we'll have much better models on kernel engineering. And I'm sure the labs would tell you that they already do a lot of their kernel engineering in a fully automated way. So software as an edge, but I think about software as maximally near speed of light, efficient usage of an underlying piece of hardware, is going to trend towards not being an advantage for a company like yours over time. That's right. The rising tide of something like Mythos or GPD 5.6 soul that lives all boats. It really does.
Starting point is 00:41:50 I actually don't think there's a point in specializing to say we work on making the model better for just kernel engineering. I think that's actually not the most meaningful subset of coding in general, kernel engineering in particular. Maybe there's some privilege information you inject into the prompt. That's like a useful way to steer the model to be better at writing kernels. but broadly speaking, yes, we're all downstream of the frontier in terms of this capability. I always love this from the history of energy.
Starting point is 00:42:16 There's always this pendulum between the raw source, let's say coal. And then if there's a certain amount of energy available inside of a hunk of coal, what percent of it we can harness and use? A big part of the history of energy was getting that number from 10 percent to 95 percent or whatever. If I just think about it at Blackwell or something, and Blackwell is the piece of coal, what percent do you think we're at? How efficiently can we use an interesting piece today? There's a lot of different ways to analyze that.
Starting point is 00:42:43 In some level, we are really efficient at optimizing the performance when the GPU is doing the thing that it's most happy doing, which is a large dimension matrix will apply. That operation runs at 70, 80% of peak utilization, and it's limited not by software, but by power. The way NVIDIQO's peak flops is a little optimistic. You never hit that because of power throttling. Because of heat.
Starting point is 00:43:04 Exactly thermals. In practice, you don't spend the majority, your time in a transformer in that happy path where you're doing a large batch matrix hold apply. Our job is to basically build the engine around the chip such that we are feeding the GPU these large batches of work at all times. One of the most profound transitions we've had in the GPU world in the last year has been this moving of you don't program one GPU at time anymore. You should think about the whole rack and maybe you should think about the whole cluster at the entire data center at a time. With Nvidia, they've started shipping not just a single
Starting point is 00:43:33 GPU or a single motherboard, but actually the whole rack system is something that they prescribe. They call it NVL 72. Their latest chip, the Grace Blackwell 300, that ships as a rack of 72 units, and it's an open race to figure out who can program the whole rack scale computer as efficiently as possible. And my belief is that that shape of compute is the future of efficiency and speed. In fact, the video is a great job of if you want the lowest possible latency, you should be using that chip. And if you want the highest possible throughput, you should probably also be using that chip as a right now. And it's all comes down to like, this is a very new paradigm of programming. One of the things you hear is that the market for the best chips, Blackwell's, let's say, is like a drug market or something right now.
Starting point is 00:44:11 There's all sorts of fascinating things happening to get as many of them as possible because everyone's so short. It would be to react to that analogy. Is that what it feels like? Then also to talk about what the market is like for like not the bleeding edge chips. If I am willing to accept a slightly or moderately inferior chip, what's that market like? Let us into that world. Media has a long-term view on all their chips. They see this immense demand for the black oil chips. They can do what other suppliers have done in the past,
Starting point is 00:44:38 which is crank prices. Made the market supply and demand curves well correct. They'll intersect at some point, and everyone will be technically happier. But Nvidia sees if they just let the most deep pockets buy all the chips that maybe hurts them in the long term if the customer ends up accruing a lot of more power. They understand the compute is power today.
Starting point is 00:44:56 They're quite strategic about how they allocate compute. That's the first thought. The second thought is that relationships matter a lot. Nobody wants to have a huge order of a chip rental come in from this new startup that says, oh, yeah, I'm going to rent 10,000 black wells for three years or five years. This startup has only been operating for a month, typically. Who knows that they're good for the money? The way you convince someone to give you access to compute is quite challenging these days and requires some pretty great relationships or just incredible financial backing to make this happen on the Nvidia side. And it's all because the scarcity is so high and demand is just off the charts.
Starting point is 00:45:32 Now, for other chips, I wouldn't even call them inferior. I like to say there's no bad chips. There's only bad pricing. I will make any chip work at the right price. That's one of the ethoses of the company. Let's talk about AMD. AMD, I think great chips overall. The challenges that people don't understand how to program them very well.
Starting point is 00:45:49 I've been talking to you about how we have such a great kernel team. We're so serious about squeezing the performance out of the hardware. InVIDIA is pretty good at doing that for their own chips, frankly. There's some alpha that we can squeeze out, but there's a lot more to be done other chips because the vendor does a little bit less work than Amidia does to make the best kernels out of the box. Or even better, there's alpha and just other people have this perception that AMD is not as good as Invidia. That's music to my ears. I'm very happy for them to sleep on this chip and for me to buy as much as I can. Now, I think that that's not actually super true anymore.
Starting point is 00:46:17 I think AMD is actually somewhat popular amongst some large buyers. I think publicly meta and opening I have bought a ton of AMD chips. We're increasingly seeing that all the AMD supply is also being allocated. where there's a long tail of other companies that are popping up. Net new companies are great, such as ACHD or Sombinova or D-Matrix. All these companies are popping up. I think the main challenge for them is scale. Can they actually get enough way for allocation from TSMC to pump out chips to make it into the market?
Starting point is 00:46:45 If there's a new chip on the market, I'd like to know about it as quickly as possible and evaluate whether we can buy a good fraction of that supply. And so it's fundamentally an arbitrage for you. Like if you can be much better at eking out performance from chips that have received less attention, you can then resell that at a margin and it can be a great business. Exactly. And I think that it's not the case that everyone else has a skill issue that they can't make these chips work as well. I think we're quite competent in this. I think we're probably one of the best teams in the world to use multiple silicon architectures and be quite aggressive in chasing down performance in unlikely
Starting point is 00:47:16 places. But yeah, I think it's the speed at which we're willing to kind of build our stack around a new chip. We don't have a huge amount of incumbency around while our data center providers are only stuck with this class of chip, and it's going to be a huge pain for us to deploy these new chips. We have some very creative Dea Center partners who are willing to move very quickly, and there's a new class of those that we can talk about. Most importantly, we don't shy away from the challenge. That's frankly a big part of this, is just saying, yes, we love TPUs. We're going to make TPUs work. Yes, we love Traneum. We're going to make Traneum work. And if it doesn't work easily, we're going to find a way to fit it in with the heterogeneous
Starting point is 00:47:49 serving system. It will have a place. Every chip has a comparative advantage. We had to find that advantage and then squeeze it in that direction. Just as an interlude, before we get to hardware, data centers, energy, et cetera, which will be really fun part of the conversation, I'd love me to talk about your perception of the investor classes worry. Like, you look at memory stocks, or my current favorite is you look at the chart that plots the percent of the S&P 500 that's semiconductors. Historically, it was like 2, 3, 4 percent. Now it's 19, 20, 21 percent. And it just sort of looks like if you're a student of market history, you get all these things through time that have sort of reached some crazy near-term peak and then collapsed back to long-term
Starting point is 00:48:29 norms. That has all investors worried. A lot of people made a lot of money in Micron and S.K. Hynix and companies like this, but everyone feels like long-term like computes a commodity. It will not represent a quarter or fifth of the entire market capitalization of the world. They're scared and that's the setup. Everyone acknowledges that there's a huge shortage. Everyone sort of feels like, ah, we'll figure it out and these things will revert back down to their normal place in capital markets. I'm curious what you think about that narrative. I'm less of a student of history as more of a member of history.
Starting point is 00:48:59 I was born in 1997 and my mom worked at Intel, and run up to the year 2000 and the dot-com boom and crash. I remember the time where Cisco was the most valuable company in the world, and in Intel was close behind. And mostly draw parallels to that period of history from 25 years ago to today. And I think the main difference is that a lot of the investment in networking equipment historically was speculative. We anticipated this future demand for users that never came.
Starting point is 00:49:25 And I think what's interesting about token consumption or AI consumption broadly is that it's no longer speculative. People buy tokens because they're immediately valuable to them. You don't hoard tokens, you use them immediately. This is also even different from what we had two years ago, where there was a supply crunch for Hopper generation chips in 2023, 2024. In that period, it was all training-oriented spend, and training is inherently speculative. Now it's everyone's instituting caps on how much. you can spend on cloud code. It's a very, very different world to be talking about inference spend and predicting inference spend to go up. I do think inference spent monotonically increases.
Starting point is 00:49:59 There's no speculation on inference. Coming back now to your take on hardware, the unit level is interesting to me. Talk about chips, talk about racks, talking about clusters. I'd love to talk about data centers. You said you've had some interesting partners doing some cool things. Talk us about the present and future of data centers as you see it, because of this seems like obviously a critical thing for being able to serve all this inference is lots of innovation in this part of the world. And obviously, you're focused on it. I think one of the themes in our conversation has come back to what is training versus inference. Like, what is the difference between these two categories? And what was different about two years ago being training oriented and today being
Starting point is 00:50:35 inference oriented. And I think the most conservative players in the entire AS DAC stack have got to be the infra players, whether at data centers or even more conservatives, TSMC, that ship in for people. Data centers, they ever built for training. Training is a superset workload over inference. You can make any training cluster work for inference, but maybe not vice versa. The difference there is networking. How much do you invest in bandwidth between chips and how large of a cluster do you need? There's actually a dis-economy of scale to data centers in some way. It's way more expensive and difficult to build 100,000 GPUs in one data center than it is to build 10,000, than it is to build 1,000. Now we just talk about how many megawatts or gigawatts do you have. And basically there's no way
Starting point is 00:51:14 to build a gigawatt data center in the United States easily anymore. Even 100 megawatts is increasingly hard. It's basically impossible unless you're a very special set of customers. 10 megawatts is probably on the edge of what's possible today. And one megawatt, I would argue, use plentiful. So there's this incredible floor on the market where you can find lots of aggregate power, but it will not be concentrated. And that was not interesting to anyone who's building training data centers because you just assume, he'll be in one spot for, no one wants to deal with cross data center training. The market has some lag in it. I think that the market still assumes that we had to go shake down those 100 megawatt and 10 megawatt data centers, wherever we can bind them.
Starting point is 00:51:53 It's still the attitude I hear from a lot of data center developers. But increasingly, we're seeing a few new thinkers realize that inference is going to be suitable for these distributed one megawatt data centers. We're quite in agreement with that. And we are very happy to buy small pools of compute across the United States and use that as our inference fleet. Give us a sense of literal physical size of one megawatt versus 10. Yeah. Well, so this got really want. with the advent of liquid cooling.
Starting point is 00:52:20 Now you can pack insane levels of power density into a single physical rack, a megawatt of compute. You'd imagine this massive data hall, like a huge warehouse basically. And now you can actually pack that into around like eight racks for the compute. Each rack is about the size of a refrigerator. You can just imagine eight of them lined up.
Starting point is 00:52:38 That's a megawatt. Your view would be that the future that you want to help build is a whole bunch of different chips that can be used together that you can buy. you're a buyer to eke out the most per chip, and that those chips can then be coupled in very small data centers to just do inference. Those two steps of a whole bunch of random compute, some of which is cheaper than it should be, your ability to eke more out of it, and then small units of expression in the data center equals way cheaper intelligence. I certainly think so. Yes,
Starting point is 00:53:10 there's a lot of ways to access cheaper flops if you're able to be creative with what you take. And so one of the ways that I describe what we do is we will buy any chip anywhere in the world for any duration of time. That is a level of flexibility and liquidity that no one else has right now. We're very aggressive about putting our money where our mouth is, and we will take any capacity and find a way to make it work in our fleet. And that is a big part of our advantage today. Long term, we had to create more of that advantage by investing in these data centers that other people are going to be skeptical of. because what's going to happen when you set this army of 1,000 small data centers versus the one big gigawatt data center? Few things. You're not going to have power redundancy quite often. You're not going to have backup diesel generators on site. Those are all very expensive. We cut all that overhead.
Starting point is 00:53:52 We're not going to have redundant networking in a lot of cases. We're going to put these in facilities where we have good access to power, a single source of power. And we're going to trench one line of fiber to these data centers. But we're not going to have three lines of fiber with redundancy and failover in SLAs. It's just going to go down sometimes. I won't be surprised if some of them get down to like 95 percent. Uptime. Which is bad. Very bad. That's fatal, atrocious for anyone else. Couldn't survive in a big gig, wasn't. You'd have basically zero buyers for a data center that is 95% uptime. I'm that first buyer that I will buy 95% up time.
Starting point is 00:54:22 And the reason for that is because of this background engine thing. If things running in the background, you don't care? Partially, it's actually two things. One is that we have a really robust control plane that is going to be fine handling any single failure in any single data center, as long as it's not correlated with other data centers, and I can just move the workload somewhere else. I'm cool with that. happen at some rate and I am basically linearly happy with a data center that's 95% uptime
Starting point is 00:54:45 versus 98% versus 99%. It's just linearly good or bad for me. You do need that async piece that I mentioned. We serve these long horizon agents because what happens when a request fails is that I'm going to have to go find a new GPU to put that request on. And that means that for that single turn of the agent's work, it's working for an hour but then hits a roadblock because its GPU got pulled away. In that moment in time, that agent is going to experience maybe like an extra minute or two or three, maybe even 10 of latency. My argument is that my customers don't care because their agent was running for hours. They're sleeping. It doesn't matter.
Starting point is 00:55:20 It doesn't matter if like a single turn occasionally becomes a little bit longer. We tell our customers, look, our average throughput is going to be very competitive, but our P99, our 99% tallytile latency, it's not going to be controlled. It cannot be. And in return, I'll give you unbeatable economics. And I think that's the right fit for background agents. Talk about power as a category. What the innovation that you're seeing is, where do you think it goes from here? What are you seeing that's interesting, innovative?
Starting point is 00:55:43 Where do you think this goes? Okay, so I said I want 95% uptime on my data centers. Could I even take 80% up time at the right price? Probably. And what does that mean? Well, I'm a son of California. I love solar and wind. I think solar and wind power is way undertapped in the United States.
Starting point is 00:55:58 And the challenge has always been this intramency. You would even consider solar and wind unsuitable for data centers because you have a persistent base load and an intermittent power source what are you going to do? Well, I think we're actually not that far from solving that problem. I am totally capable of tolerating a outage for my data center that's measured in even days or weeks, which is like the worst-case nightmare scenario for a data center is that we're going to have a long-term outage because the wind isn't below and the clouds are in the sky. Fog is hanging over the valley for some time. That's worst-case scenario. It's in fact highly
Starting point is 00:56:27 predictable, and I can just call in capacity in some other place of the world whenever that happens. I'll just model the weather and figure out when my data center is going to be offline, move my work with somewhere else, and it's fine. The trick is that it's going to give me better access to power that no one else is going to touch because it is so annoying to deal with that kind of outage. And if my chips are cheap enough, they're probably not going to be in video racks.
Starting point is 00:56:48 But when my chips are cheap enough, I don't mind the capital cost of having idle chips. I've heard you describe this entire system as like a scavenger strategy. Yeah, that's right. Unpack that analogy a little bit. Well, first we scavenge chips, and then we scavenge power for those chips.
Starting point is 00:57:03 The idea is in both cases, I do not want to be bidding against Anthropic or up an AI for a complete capacity. I'm not going to win against them and I don't want to. I want to be more creative and use the supply that they don't find legible today. And over time, I amass enough aggregate supply. I'm never going to get concentrated supply. I will only get aggregate supply. And over time, I build my aggregate factory that is unbeatable in economics.
Starting point is 00:57:24 We are building a factory. We're trying to build the best steel factory in the world. But it will come through mini-mills, not through large, monolithic steel plants. And if I imagine the different versions of this, how vertically integrated you can be. The extreme would be you own everything, so that it's a very capital-intensive business. You own the power source, you build the data centers, you design your own chips, you control the software that eke's the most out of those chips, and you sell the end-finished token to your user. Your user is me, and you just own the whole stack. But you can imagine
Starting point is 00:57:57 many other permutations of the business where draw the line anywhere. You could be incredibly capital-lite, own nothing and just be like the coordination plane across all this stuff, the virtual scavenger. How do you think about that question of which type of these businesses to be? You know, there's actually two parts of me to receive that question. One is the CEO of a company that needs to work every day and grow sustainable and as quickly as it possibly can. The other is the founder. And the founder is much more imaginative and loves this stuff.
Starting point is 00:58:25 The founder and me wants to do everything. This is my entire life. I spent my entire life thinking about chips, power, energy. All I care about is this stuff. So, of course, I want to be maximally ambitious. I don't want to stop ever. I will never stop until I have built the most efficient system from soup to nuts. You're doing real life factorial, basically.
Starting point is 00:58:41 Very much so. For intelligence. Very much so. So that's like the emotional from the hard answer. On the CEO side, I think we have to be more pragmatic. The capital we're looking at for owning everything is insane. Software has high leverage, so we have to start with software. But ultimately, do we own power generation, or can we get great power purchase agreements with utilities?
Starting point is 00:58:59 I'm more inclined to pursue letting other people specialize. in the things that they're historic we get at and then see if we can get to this scale. I think of it as like, I want to get to the scale where I earn the right to take this under our wing. I absolutely think that there's efficiencies to be gained everywhere in the stack. If you can break the assumption
Starting point is 00:59:16 that people I would be buying from, they made assumptions about who their customers would be, and I maybe break those assumptions. It's a pretty optimistic view. I think it's only possible because we're actually trying on to write the largest market for compute in the history of computing. We're actually gonna build
Starting point is 00:59:32 so many millions, trillions of dollars of investment into inference. Because of that focus, it makes sense to build a lot of things that are custom for inference. And it's my job to seek all the places where that's possible. And then as it become obvious to me and my partners, I look at my partners to build custom things for me. And if they can't do it for me, I will do it myself. If you had to just zoom out on this entire system, software, hardware, energy, etc., and stack rank the places that you think that we are the most inefficient today at producing useful, intelligent tokens. What does that list look like? I think compute scaling is actually very efficient. You give me more flops and I will use more flops. I would say we're actually
Starting point is 01:00:12 fairly judicious already with our use of flops. If you look at a modern MEOE model, there are very few models that are more than 10% dense, meaning 10% of the possible number of experts you can activate are activated. And I think the frontier models are close for like 1%. Fairly sparse already. I don't think that we're wasting too much on the MOE side. People have been working with MMOs for quite some time, they're pretty good at squeezing M-OEs. Where we are not good is attention, and its use of memory specifically. The KV cache is quite uncompressed right now. I think if you look at the entropy in a KV cache, it's not earning its keep. Like we're storing many kilobytes of data in the KV cache per token. And that's probably off by an order of magnitude or two.
Starting point is 01:00:52 I don't know what the frontier labs do, but Deepseek certainly publishes really interesting work to compress that further and they're making good progress. And I think the fact that they're able to make an order of magnitude in progress here every year or so signals that there's a lot more room to go. If you zoom out further, I think that we actually don't marshal our compute effectively at all. We have all this compute in the world. Invity is pumping out five million blackwell chips this year. Where are they all going? Are they all being used at all the time? I certainly doubt it. I think that at some level we just need better orchestration of compute across the world. This is very difficult to do because a lot of the compute disappears into private pools or compute that will never see the light of
Starting point is 01:01:26 day. And those GPUs sit very sadly idle. It pains me physically to see that those GPS, are just silicon and power going into that, and it's just sitting idle. And I want to fix that. How we organize and orchestrate the world's compute as a shared resource and pack it more efficiently. I would estimate that we all make fun of XAI for having, you know, some challenges with total flop utilization on its clusters. The reality for the rest of the world is it's far worse.
Starting point is 01:01:50 A ton of GPUs just sit in warehouses or sit in private pools allocated to a specific customer, just don't get utilized. You're attacking the efficiency of that very directly. That's way more effective, yeah. about FABs, what do you think is the future of FABs themselves? I think everyone is wondering, will the memory companies, will TSM, will Intel, and others, how will they expand capacity, basically? Will we do it here in the U.S.? RIFOn fabrication of chips themselves? Like, if we could just snap our fingers and have a hundred times the chips in the stock today, we probably have way
Starting point is 01:02:24 cheaper tokens. That seems like an important part of the universe to hear your view on. Everything grows in balance with each other. If we snap our fingers and double all those things, you might fix a TSM bottleneck, you're just going to run into another bottleneck. You make 20% more chips, then you have another bottleneck immediately. I will say, though, it is interesting what they consider to be a must deliver, like an invariant that their customers, me, are always going to want, versus what I think as like a more fluid relationship. I think that if the fab exposes more of their tradeoffs, to me, I'm able to make more intelligent decisions about what I can do. One of the most interesting examples here is that any fab has a lot of spread in their
Starting point is 01:03:05 worst chip that comes out of the production line and the best chip that comes out of the production line. There's a lot of variance in how chips are made. And the question is, if you have a company like TSM, they work very, very hard to tighten what we call these process corners. We want to keep the worst chip as close in characterization to the best chip, and they've got a great lens to make that possible. but that means that they are adding a lot of controls in the process that maybe I don't need. Maybe I'm actually willing to find a place for that worst chip.
Starting point is 01:03:30 You don't need to tighten the process control as much, which takes more time and cost. Maybe I'm willing to take a lot more rejects. And I think for us, it's like a more holistic optimization around cost of the dies, and then the cost of power and places we can put them. My whole goal is to so dramatically expand the supply of power across the United States that I have a home for a lot of ships that otherwise would not have earned their place in a data center. Can we talk about how you design the system of your own business? What lessons have you learned? You talked about some interesting in video lessons. Bring me into the culture
Starting point is 01:04:03 and how you structure a team and a business where this is the North Star. There's a lot of, in the limit thinking, we don't worry about the immediate nature of when we start working on a model, the efficiency is not going to be very good. We think about where we could end up in like a month, six months or a year's time. We don't accept the state of the machines we work on is fixed. Even something like the Blackwell chip, if we think that there's some bottleneck that is holding us back from achieving this performance, it's very important to me that we understand and characterize that very well and write it down so we can both A tell Indvita about it or friends and also to basically keep this in mind for future chips that we buy. We want to learn things
Starting point is 01:04:42 that are invariant for us or the company long term and kind of fold that into future decisions that we make. We're very collaborative. I think one of the most important traits that we look for are people who either, who are both good students and great teachers. A lot of our people on the team were TAs in college and loved the experience of sharing knowledge in this way. We do whiteboard sessions all the time. The collegial environment where everyone has something to teach and something to learn is extremely important for us. What are the attributes of people that you would want to hire that you think will be resilient to the work environment three years from now when more stuff is handled by machines. Curiosity. A hundred percent curiosity. The one thing I cannot teach is love for
Starting point is 01:05:22 performance. Love for digging into every microsecond. The machine is working and understanding what's happening on the machine at that time. That to me is the most important trait for a performance engineer. It's what I look for. I don't look for lots of AI experience. I don't look for Kuda experience at all. That's actually a huge red herring. I mean, Kuda as a concept or GPs as a concept, have evolved so much in the last five years. There's no point asking for 10 years of experience. I want to teach that. but I cannot teach the love for performance engineering. That is what I seek. Can you give your assessment of the major labs one by one,
Starting point is 01:05:54 but also then the relationship of closed source as a category to open source, what you think is happening and will happen? In a line, I would say the labs pay an immense premium to be three to six months ahead of everything else. I think that's probably still worth it. I think it makes perfect sense for an anthropic to do what they do. There's a sensitive topic around distillation,
Starting point is 01:06:13 which I think is a very core piece of the relationship between closed and open frontier. I'd like to offer an alternative view on that, which is there is the sense that distillation is theft, that you are taking something from the frontier models when you distill on their outputs. And in fact, even if that's not your intent, even if you don't ever try to scrape data from Anthropic,
Starting point is 01:06:32 one thing I'll offer is that an increasingly large percentage of the artifacts we put out on the internet are AI generated. You just look at GitHub alone. What percentage of repos created in the last year did we think we're created by cloud code? Do we consider that to be distillation? because it's probably all we need. I would not be surprised if you could train a Fable Class model
Starting point is 01:06:49 only on the outputs of code you consider good on GitHub that's open source. And certainly if we take the position that users own the outputs of their interaction with AI and they choose to put that up on GitHub, which a lot of them do, we're going to have latent distillation for a long time. It seems fundamentally possible for me. I don't think it's fundamentally possible to prevent the diffusion of information or model capabilities. It will happen.
Starting point is 01:07:11 The question is just how fast. And so then the question becomes, do scale, and improvement laws hold forever or for a really long period of time. And if they do, then there's value to being 306 months ahead and that will just last as long as it lasts and they can charge a huge premium for those tokens relative to a very cheap open source token. Is that the right way to think about it? I think it's possible. I don't know that the premium for being 3 to 6 months ahead is going to last that long. I mean, if you look at like enterprise deployments, they don't move at 3 to 6 month speed. A lot of enterprises are probably still handling 4.6 or opus 4.7. They don't
Starting point is 01:07:45 adopt the bleeding edge rapidly. There's a lot of questions that people have around rolling out any change at all. We're just so early in scratching the surface that I don't think there's any way to call a winner in this race. Certainly, I don't even think this is a race that can be decided ever. It's a continual process. Finally, I don't think open source ever goes away. If there's a vacuum because, you know, one leader steps out, a new leader will step in. There's too much incentive and too much tailwinds, too. It gets easier every day to train a frontier class model. So your hope of what the future looks like is what balance between closed and open, what balance between model companies doing everything because they have the advantage of owning the stack or whatever. Anthropic can do that.
Starting point is 01:08:23 It's like the new Google will just do that or something. What do you hope the future looks like? I want abundant tokens and diverse harnesses. I want everyone to build their own harness. Every company, every user, even. Make the agent your own. We're not that far away from that level of customization and capability. I want people to own their intelligence. And I want that intelligence to be customized, probably not through weight fine-tuning, but probably through more in-context learning. That's a more technical detail, but the underlying input to this abundance future is about cheap tokens. My job is to make the tokens as cheap as humanly possible. I will achieve that, and I will do it through every layer in the stack available to me. I love the supply-side levers.
Starting point is 01:09:01 I will use every source of power, and I will use every piece of land in the United States that's suitable for this. In return, people will have the incentive to explore what it's like to have abundant intelligence. We still treat the agent as a person that is expensive to consult, and you should ask them when you have a hard question. That's not the way to think about intelligence. It's incredible that the machine can think, and we should try to get that into as many hands as as many people as possible. You sit in such a unique seat and you have such a unique perspective on what you're trying to do to make this feature a reality. What do you think are your most divergent views of the world versus your friends who are really well informed and interested in this
Starting point is 01:09:37 stuff? What ideas of yours make your friends look at you like you have three heads? That's most of the ideas on chips, I would say. You know, when I talk about building custom chips and they ask me, oh, so what's different? It's about sidestepping the HBM the HBM shortage and focusing on more extreme offload to other forms of memory, such as Flash. I'm quite passionate about that idea. Everyone on my team knows that I keep banging the drama around like, what would we have to change about the model architecture to make offloading KV Cash to Flash work at a much greater level? That's in the community of like inference people. We have some divergent views on what you can do if you design a system around.
Starting point is 01:10:10 serving at one to 10 tokens per second, which is our whole North Star. More broadly, I think, there is this larger sense around how do people consume a trillion tokens per day? That's the world we want to create the capability for them to do that. What's the trillion tokens? Like ground us and how much that is? A trillion tokens. Well, okay, at opening eye pricing, that's at least $5 million at the very least for $5.5.6. Yeah, I think the dollars is probably the most probably the most. Yeah, it's a good metric, yeah. It's millions of dollars. Yeah, so what's the world in which we consume what currently costs $5 million per person per day.
Starting point is 01:10:41 We were asking for at least three to six orders magnitude improvement in cost per token. Get that under 5,000, you probably have some customers. And in fact, I would argue that for some size of model, we are approaching a trillion tokens being measured in tens of thousands of dollars. And that's something that you could imagine running for a single job. Are you all worried that just like the average person just can't and won't do that? It doesn't do that now with their own brain. and there actually isn't that much demand
Starting point is 01:11:08 for intelligence in the world? I never will believe in that. There's always demand for intelligence in the world. The on-ramps to that intelligence are our challenge as a product community. I'm not a product person, so I cannot say I had the best vision. You want to enable those people.
Starting point is 01:11:21 I want to enable those people. I want them to never be held back by the sense that my free-tier users, I can't afford to give them this many tokens. And I hear that from my customers all the time. We want to fix that. What about the inverse question, not what you think is craziest,
Starting point is 01:11:32 but what consensus thing you think is wrong? One of the things I keep coming back to is this question of NVIDIA. I am bullish on NVIDIA in the short term, and NVIDIA, you should never bet against them. They're always going to reinvent themselves. Fundamentally, I think one thing that surprises people is when I tell them that, hey, if you look at Hopper and a Blackwell to Rubin and you compare like for like, what is the performance per watt of BFload 16 multiply, it hasn't improved all that much. Or you take that one step further and go to TSMC.
Starting point is 01:12:00 If you look at TSM 5 nanometer versus 4 versus 3 versus 2, the performance per watt on these chips doesn't change like a dramatic amount. The consequence of this is people lose their minds over geopolitics like what would happen if we lost access to DSMC for any reason. My contrarian take is that it wouldn't be that bad. Supply would take a shock for sure, but the best processes that we have in the West, like Intel, not that far behind, at worst, like maybe 2X, worst performance per watt. The gap is just far smaller than you would make it out to be if you follow like the chipboard dialogue. What else is happening in the AI world that is not in your path, meaning it's not like a component of this whole system that you would end up doing something in that interests you most?
Starting point is 01:12:40 Well, we're fully downstream of models. So the model people get to decide how to design their architectures. I don't have any input to open AI or anthropic, but I can only pray that they go in the direction that is amenable to me, or I have to like do my best to predict where I think they're going to go and build my serving architecture accordingly, both software and hardware choices. They have, I think, the most interesting game in some ways to play. Once again, this is a game actually like the profundity of the machine thinking and how consequential it is to decide to use something like sparse attention versus dense attention or how consequential it is to like use a different data type. We were training in B-FLit 16, but now we can train
Starting point is 01:13:15 in FP8 or FP4, lower precision data types. That is just an arbitrary choice it feels like, but it has profound implications for what chips I can use and how I should build my hardware, think about the future of compute. If you had 100 entrepreneurs in a room, all of whom wanted to create some new compute startup, and let's say they were specifically wanted to make hardware, chips or systems or racks or whatever, what advice would you give them on how to orient their companies or like the type of company, not the specific choice they're making on a tech bed or something like this? Because it seems like we're going to try everything, and that will be great for the world. Some stuff will work. But if you had to give them advice on how to orient their business to be
Starting point is 01:13:50 successful in this coming world, what advice would you give them? It's all about the bottlenecks on supply chain. So you need to first convince me or convince an investor that you understand the like three to five bottlenecks that dictate modern chip supply. There's TSM weight for capacity, there's HPM capacity, and there's advanced packaging. Maybe a fourth one would be power. Where will you get the power? How will you build these racks? You should have a great answer to each of those four bottlenecks and how you're going to work around them because it's all arbitrage at the end of the day. You're building a chip because you think that NVIDIA has made some choices that are difficult for the change, which is true. Invidi makes a lot of choices that are difficult to
Starting point is 01:14:25 for them to change. They're not perfect. They're just really well balanced. You want to be spiky. You want to pick something and say, I think they've underpriced the impact of how short we're going to be on HBM. We're going to push really hard in this other direction instead, which I do think is probably the thing to attack most. Why? There's no easy way to bring on a lot more fabs of memory. So it's just going to be a while until we have. Yeah. Then the boys and Boise don't love huge Cappex for cyclical. They've been burned on that many times. Conceivably, like, because of that shortage, the world is going to route around it by making everything else in the system more efficient? I think they're going to make everything else more expensive.
Starting point is 01:15:00 I think that iPhones will cut their memory. iPhones are going to go up in price. We're just going to deal with it. Why doesn't Nvidia go all the way to the end and sell tokens, do you think? Invidia is really smart about this. They don't compete with their customers. Invidia takes a long view on everything. Why don't they even start with the NeoCloud?
Starting point is 01:15:16 Why don't they just sell compete at the back door? Well, Nvidia is really good. Jensen is really good at making his friends billionaires. He's made Correve, a billion dollar company, many billion dollar company. And there's no need for him to destroy that goodwill. He wants to create a diverse community of neoclods and inference providers who are all jockeying to create demand for Envidio, such that if any one of them decides to, I don't know, vertically integrate or go with AMD or any other option, he's got three more people hungry to fill that position. It's great to have competition amongst his buyers. My favorite closing question for everyone is, what is the kindest thing that anyone has ever done for you?
Starting point is 01:15:50 My immediate first thought is like all the mentors that I've had over the years. It's a rare person who takes a lot of time out of their schedule and makes it like their personal interest, essentially, to make sure that you understand something or teach you something or ingrained some value in you that they think that you're on the cusp of understanding, but just push you over the line for understanding. A lot of the people in Nvidia that I mentioned earlier who instill that love of performance engineering in me, but also my professors in college who I remember like my advisor in sophomore year, I was a very impatient student as I show up at his office hours and say, I want to build AI trips. I know what I want to do. Why am I wasting time taking all these other basic classes and networking and operating systems? And he just laid out basically like the whole stack and showed me the beauty of understanding every piece in the puzzle.
Starting point is 01:16:33 He took my entire path of like trying to focus on one piece of the system and said that it's so rare that someone can actually understand the entire stack from the gate level silicon all the way to building a great internet scale service. You should aspire to be someone who over the course of your lifetime achieves that level of understanding. It is such a rare crate. That level of expertise is so noble to chase. That stays with me quite a bit.
Starting point is 01:16:57 Amazing conversation. Thanks so much for your time. Thank you so much for having me. If you enjoyed this episode, visit colossus.com. You'll find every episode of this podcast complete with hand-edited transcripts. You can also subscribe to Colossus, our quarterly print, digital, and private audio publication featuring in-depth profiles of the founders, investors, and companies that we admire most. Learn more at colossus.com slash subscribe.
Starting point is 01:17:20 You know how small advantages compound over time. That's true in investing and just as true in how you run your company. Your spending system is your capital allocation strategy. Ramp makes it smarter by default. Better data, better decisions, better economics over time. See how at ramp.com slash invest. As your business grows, Vanta scales with you, automating compliance and giving you a single source of truth for security and risk.
Starting point is 01:18:03 Learn more at vanta.com slash invest. The best AI and software companies from open AI to cursor to perplexity, use WorkOS to become enterprise ready overnight, not in months. Visit WorkOS.com to skip the unglamorous infrastructure work and focus on your product. Ridgline is redefining asset management technology as a true partner, not just a software vendor. They've helped firms 5X in scale, enabling faster growth, smarter operations, and a competitive edge. Visit Ridgeline apps.com to see what they can unlock for your firm. Every investment firm is unique and generic AI doesn't understand your process.
Starting point is 01:18:36 Rogo does. It's an AI platform built specifically for Wall Street. to your data, understanding your process, and producing real outputs. Check them out at rogo.a.i slash invest.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.