Odd Lots - How to Build the Ultimate GPU Cloud to Power AI

Episode Date: July 20, 2023

Artificial Intelligence is all the rage right now and most of the investor excitement has so far been focused on the companies providing the hardware and computing power to actually run this new techn...ology. So how does it all work and what does it actually take to run these complex models? On this episode, we speak with Brannin McBee, co-founder of CoreWeave, which provides cloud computing services based on GPUs, the type of chips pioneered by Nvidia and which have now become immensely popular for generative AI. He walks us through the infrastructure involved in powering AI, how difficult it is to get chips right now, who has them, and how the landscape might change in the future.See omnystudio.com/listener for privacy information.

Transcript
Discussion (0)
Starting point is 00:00:10 Hello and welcome to another episode of the Odd Lots podcast. I'm Joe Wisenthal. And I'm Tracy Alley. Tracy, have you looked at Nvidia stock chart lately? And by lately, I don't mean like over the last two years. I mean like just like over the last like two weeks or two months. I don't need to look at it because everyone keeps talking about it. So I know what I'm pretty happy about could I just say?
Starting point is 00:00:31 You know, we did that episode like two months ago. Yes. With Stacey Razgan. And we were like, what's up with the Nvidia? Like, you know, I know it's at the center of the, I. AI chips boom and whatever. And then like we did that episode and it came out. And then a week later, like they just like knocked it out of the park.
Starting point is 00:00:47 The stock took off. Yeah. So, you know. We were early. We were at least like, you know, a good like two weeks early on that. Yeah. Two weeks. I'll take it. I'll take it.
Starting point is 00:00:58 So clearly something that, you know, we talked about this with Stacy. Like, you know, something that Invidia has is like everyone's trying to buy it. Everyone's trying to get it. But then it raises the next question of like, okay, but what? is that market like how do you buy a chip? Yeah, how do you buy a chip? And then I guess what do you actually do with it once you have it? Because my impression is that for a lot of these AI applications, the way you use the chips, the way you set up the data centers is very, very different to what we've seen in the past. And I think also what NVIDIA is doing now is kind of different. But maybe we can
Starting point is 00:01:34 get into this with our guests. My impression is they're trying to create a sort of like holistic approach for customers where they provide not just the hardware, but also some services to go along with it. Yes, right. And like all the software and Stacey talked about that with the Kuda ecosystem and how dominant that is. But right, like what do you do with it? Like how do you get one? If like what, you know, what would we do, Tracy, if a big pallet of Nvidia chips wound up here. Joe, you want to know a secret? Yeah. My basement is filled with H-100 chips. Just got a pile of them. It came with the house. It was on that ship that was stuck off the Chesapeake. Instead of getting your couch, you got it. I just caught a palette of H100.
Starting point is 00:02:16 We're manifesting that into reality. So anyway, like how this world works. Essentially, like, the trading and dealing of these, like, the hottest commodity in the world, right? Which is these advanced chips from AI and how that works and who can get one. I still think it's like a sort of mystery that we need to delve further into this question. I agree. And there is also, there's a lot of excitement around it right now for the obvious reasons of everyone's really into generative AI and Invidia stock is exploding as we already talked about. But we're also seeing a lot of previous, I guess, consumers of chips like the crypto miners, start to pivot into the space. And I'd be curious to see what they're doing in it as well and how much of that is just, you know, desperation versus a real business opportunity. And the video game market. Yeah.
Starting point is 00:03:09 Oh, totally. I forgot about video games. Which was like the other thing. It's like for years, I thought of Nvidia is the video game company. Yeah. Because they had their logo on Xbox. And how realistic is that pivot? What proportion of those types of chips can be used for AI now?
Starting point is 00:03:24 Well, I am very excited. We do have, I believe, the perfect guest. We're going to be speaking with Brandon McBee. He is the chief strategy officer and co-founder of CoreWeave, which is a specialized cloud services provider that's basically. providing this sort of like high volume compute to AI type companies. They recently raised over $400 million have been in this space for a little while. So Brandon, thank you so much for coming on OddLodz.
Starting point is 00:03:51 Thanks for the opportunity, guys. Really excited to chat with you all today. So let's just let me sort. If Tracy and I, like, I don't know why they would do this, but if like some VC was like, you know, we want you to do odd Lodge GPT, we want you to like do a core base large language model off of all the work. you've done. We want you to compete with Open AI. And they gave us like, I don't know, some like, you know, $100 million raise. They said, go start to do your startup. Could I call Nvidia and buy
Starting point is 00:04:19 chips? Would I be able to like get in the door there? Gosh, I mean, you're, I think you and everyone else is asking that question and you're going to have a huge problem doing that right now. It's mostly just around how much in demand this infrastructure became, right? I mean, you could argue it's one of the most critical pieces of information technology resources on the planet right now. And suddenly everyone needs it. And I like to contextualize it in that, you know, the pace of software adoption for AI is like one of the fastest adoption curves we've ever seen. Right. Like you're hitting these milestones faster than any other software platform previously. And now all of a sudden you're asking infrastructure build to keep up with that, right,
Starting point is 00:05:05 space that traditionally takes more time. And it's, it's created this massive supply demand and balance just on in-place infrastructure today. And not only infrastructure is available to purchase. And it's, it's an issue that is going to be ongoing for a bit as well, we think. So can I ask the basic question, which is core weave? What do you do exactly? Joe mentioned the capital raise, which I think has you valued at something like $2 billion. So congrats. But what exactly are you doing here? Yeah, thank you. So Quereve is a specialized cloud service provider that is focused on highly parallelizable workloads. So we build and operate the world's most performant GPU infrastructure at scale and predominantly serve three sectors. That's the artificial
Starting point is 00:05:53 intelligence sector, the media and entertainment sector, and the computational chemistry sector. So we specialize in building this infrastructure at super compute scale. It's like quite literally you know, it's 16,000 GPU fabric, and we can get into all the details and how complex that is. But we build that so that entities can come in and train these next generation foundation machine learning models on. And we found ourselves in a spot where we can do that better than literally anyone else in the market and do it on a timeline that's faster. Or I think the only entity with H100 available to clients at scale globally today. So you have an actual basement full of H100 chips.
Starting point is 00:06:36 Can you talk to us, you know, when you say infrastructure, we help clients build out the infrastructure, help us conceptualize this. Yeah, what does the infrastructure for this type of AI actually look like? And how does it differ to infrastructure for other types of large-scale technology projects? Yeah, totally. So, you know, I think during the last Nvidia quarterly earnings called Jensen put this a really great way in the Q&A section. He said that we were at the first year of a decade-long modernization of the data center or like making the data center intelligent, right? You can kind of,
Starting point is 00:07:16 you can suggest that the last generation or the 2010's data center was comprised of CPU compute, storage, and these things that didn't really work together that intelligently. And the way that Nvidia has positioned itself is to make it a smart data center. That's like smart routing of data, packets, of different pieces of infrastructure in there that's all focused on how do you expand the throughput and communicability of, and in between pieces of infrastructure. Right. It's this amazingly different approach to data center deployments. And so the way that we're building it and we're working with Nvidia infrastructure,
Starting point is 00:07:55 we design everything to a DGX reference spec, and a DGX's, like, how do you draw the most performance out of Nvidia infrastructure is possible with all the ancillary components associated with it. So all this stuff is going into what's qualified as a Tier 3 or Tier 4 data center. We co-locate within these things. We're not quite building in a basement. In our past history, we certainly had time doing that. But this is within just amazing collocation sites that are operated by our partners, such as Switch. Right. So a tier three, a tier four site is something that's qualified based on its ability to serve workloads with an extremely high uptime.
Starting point is 00:08:38 So we're talking like 99.999 percent uptime rate. And that's guaranteed by its power redundancy, its internet redundancy, and its security. And then ultimately, like, it's connectivity to the internet backbone. Right. So as it's like as a first step, you're housed within these. data centers that are just critical parts of the internet infrastructure. And then from there, you start building out the servers within there. And I can go into that detail. So you mentioned, actually, I want to just get sort of defined some terms. Can you just real quickly before we move on, Tier 3, Tier 4? What do you mean by this? Yeah.
Starting point is 00:09:19 So Tier 3, Tier 4, this all goes back to like the quality of the data center that you're in. It's all about the reliability and uptime that you should be able to achieve out of that data. center. It's another way to qualify the services around it. It's like power. You get redundant power, right? Like multiple power services in case one goes offline, there's another one. You get, you know, redundant cooling. You get redundant internet connectivity. It's all these services that, like, have extra fail safes that allow for you to operate at the highest uptime and security level possible. Is higher tier better? Like tier three, four? Is that better than tier one and tier two?
Starting point is 00:09:56 That's correct. Okay. So, quick. follow-up question then, you know, we're interested in like, okay, where the rubber hits the road, the scarcity is here. Let's say Tracy miraculously opens her basement, and there really is, like, you know, all these pallets of these Nvidia chips there. Is there capacity at the data centers right now? She's like, you know what? We want to co-locate with you. You guys have great power. You're pretty well connected to the internet. You have, like, good security guards so that's operated 24-7 and we want to set something up. Like, is there space there? Yeah, it's a fantastic question.
Starting point is 00:10:29 It's an issue that didn't really pop up until really in the last eight weeks or so. It's a sector that's been. It's happening that fast, Joe. Eight weeks. Okay. So the two-week lead time on InVideo is very important, Joe. That should be the table. You're right.
Starting point is 00:10:46 You're right. Wow. Wait, what happened, so wait, what happened 16, what described 16 weeks ago versus eight weeks ago? Sure. Even last year, right? So this is a space, the data center space, collocation space, that's been fairly chronically underinvested in because the hyperscalers just built out their own data centers instead.
Starting point is 00:11:07 But what's happened is the infrastructure changed. The type of compute that we're putting in these data centers, it's different than the last generation, right? So we're predominantly focused on GPU compute instead of CPU compute. And GPU compute, it's about four times more power dense than CPU compute. And that throws the data center planning into chaos, right? Because ultimately, let's say you have a 10,000 square foot room in the data center, right? And you have a certain amount of power. It was called 100 units of power that go into that 10,000 square feet. Well,
Starting point is 00:11:41 because I'm four times more power dense, it means that now I take those 100 units of power, but I only require about 25% of that data center footprint. Or in other words, 2,500 square feet within that 10,000 square foot footprint. So that then leads to, like, not only is the space in the data center being used inefficiently. Now, because you theoretically have to run more power into the data center to use that full 10,000 square feet due to the power density delta. But now you have cooling issues, right? Because you design that footprint to be able to cool 10,000 square feet spread out across
Starting point is 00:12:18 that entire area. But now you're dropping all the power. Sorry, I just want to back up because this is really. extremely interesting. So I don't, I just want to get this detail right. Just, sorry, just to, and then move on, but the, let's say, given an X amount of power at 100 units of power, what you're saying is that with this next generation of compute, it now only gets, that's now only sufficient for a quarter of the data center. In other words, that to power that space, that space, and that to then power the whole space, you really would need like 4x the power. That's accurate.
Starting point is 00:12:53 And the complication really arises out of the cooling that's required from that. Right. So if you imagine you could cool a 10,000 square foot space and you design for that, that's one thing. But now if you have to cool in a much more dense area, that's a different type of cooling requirement. And so that's led to this issue where there's only a certain subset of tier three and four data centers across the U.S.
Starting point is 00:13:18 that can are currently designed for or can quickly be designed and changed to be able to accommodate this new power density issue. So now not only, like, if you had all those H-100s in your basement, you might not have a place to plug them into. And that's become a pretty big problem for the industry very quickly and truly has only arisen in the last eight weeks or so. And it's going to persist for a few quarters. So you were describing the difference between CPU and GPU, how do you actually connect these newer types of or these different types of chips together? Because I imagine, you know, old data centers, I guess you just have a bunch of like Ethernet cables or something like that.
Starting point is 00:14:03 But for this type of processing power, do you need something different? That's exactly correct, Tracy. So what we, so the legacy, the generalized compute data centers are really what the hypers look like, you know, Amazon, Google, Microsoft, Oracle. they predominantly use something that's called Ethernet to connect all the servers together. And the reason you use that was you don't really need to have high data throughput to connect all these servers together, right? They just need to be able to send some messages back and forth.
Starting point is 00:14:32 They talk to each other about what they're working on. But they're not, you know, necessarily doing highly collaborative tasks that require moving lots of data in between each other. That's changed. So today, what people are focused on and need to build are these effectively supercomputers, right? And so we refer to the connectivity between them, the network between them as a fabric. Right. It's called a network fabric.
Starting point is 00:14:58 So if we're building something to help train like the next generation GPT model, typically clients are coming to us saying, hey, I need a 16,000 GPU fabric of H100. So there's about eight GPUs that go into each server, and then you have to run this connectivity between each one of those servers. But it's now done in a different way, to your point. So we're using a Nvidia technology called Infiniband, which has the highest data throughput to connect each of these devices together. And taking this 16,000 GPU cluster as an example,
Starting point is 00:15:38 there's two crazy numbers in here. One is that there are 48,000 discrete connections that need to be made, right? Like plugging one thing in from one computer to another computer, but there's lots of switches and routers that are between there. But you need to do that 48,000 times. And it takes over 500 miles of fiber optic cabling to do that successfully across the 16,000 GPU cluster. And now again, you're doing that within a small space with a ton of power density, with a ton of cooling. And it's just a completely different way to build this infrastructure.
Starting point is 00:16:16 And it's just because the requirements have changed, right? Like we've moved into this, like, this area where we are, you know, designing next generation AI models. And it requires a completely different type of compute. And it's just, it's caught the whole sector by surprise. So much so that, you know, it's, it's really challenging to go procure it at the hyperscalers today because they didn't specialize in building it. And that's, you know, where CoreWeve comes in is we only focus on building this type of compute for clients. It's our specialty. We hire all of our engineering around it. All of our research goes into it. And it's, you know, it's been a fantastic spot to be. But our goal at the end of the day is just to be able to get this infrastructure into
Starting point is 00:16:55 the hands of end consumers so that they can build the amazing AI companies that everyone's looking forward to using and incorporating into, you know, enterprises and software companies. You know, you mentioned these special or purpose-built connections that NVIDIA is making, and this kind of leads nicely into my next question, which is what exactly is your relationship with NVIDIA? And in order to provide this type of service, you know, vast amounts of processing power that is well suited to a particular type of technology, in this case, AI, do you have to have a really good relationship with NVIDIA to make that work? Like, do you have to have special access to H-100s and other chips? It's a great question. And I'll try to offer it from NVIDIA's perspective.
Starting point is 00:18:02 And it goes a little bit back to the answer I just provided as well in that I would think from Nvidia's seat, what's most important is empowering end users of their compute to be able to access their compute in the most performant variant possible. At, at, scale and to be able to access it quickly, right? Like a new generation comes out, they want to be able to get their hands on it, right? And we've built core weave around hitting every single one of those checkboxes, right? We build it at DGX reference spec. We build it at scale.
Starting point is 00:18:32 And we bring it online on a timeline that's, you know, within months of a next generation chip set launch as opposed to, you know, the more traditional legacy hyper scalers that take quarters at a time. So us being in a position to do that has enabled us fantastic. access within Nvidia. And we have a history of consistently executing on exactly what we say we'll do, right? We underpromise and over-delivered as a business. And I think that's just put us in this place where Nvidia has the confidence in allocating infrastructure to us because they know it's going to come online. They know it's going to get to consumers faster than than anyone else in the market,
Starting point is 00:19:14 and they know it's going to be delivered in its most performant configuration that exists. You know, I was thinking as I listen to some of these answers, I keep having like these like imagines like, oh, you know, there's probably like some random industrial company that's like traded like, you know, on the like S&P 400 that makes some cooling fluid whose like sales are going to be up 10x. So I'm like Googling while we're talking like, oh, what is a company that makes cooling fluid? Or like, who is some company that's like really good at making these like infinite advance because it just like invest in HVAC.
Starting point is 00:19:46 Right. Yeah. Like what are the, anyway. Right. But like right. Like you know there's going to be some charts that are like these like, yeah, or tertiary play that are like 30 X up. But you know, I want to get a sense from you of so it's really changed a lot.
Starting point is 00:20:00 And I kind, you know, in the last several months could we see it from in video results and what you're describing? Like how big is the market getting? And the way I think, you know, I know like with AI, there's training and they sort of build the model and then there's inference. the inference is how they spit out the results. Can you talk a little bit about what you're seeing in terms of the growth of both of those aspects of AI, which is bigger and which is growing faster? And how do they compare to like the size of the installed compute base that already exists?
Starting point is 00:20:31 Oh, absolutely. So this is one of my favorite topics because it's just mind blowing the scale that's going to be needed to support AI and the scale of this infrastructure. So, okay. So today, most of the funding that's going into the AI space is to, for funding to train next generation foundation models, right? So when a company's raising a bunch of money, at the end of the day, most of that money is going into cloud compute to go train this next generation file model, to build that intellectual property. So if they have this model, they can go bring it into the inference market. And what I would say is we're having a supply demand issue, like a chip access crunch in the training phase where in reality, the scale of the inference market is where all the demand truly is going
Starting point is 00:21:21 to sit. So what I'd offer to help contextualize that is let's take, you know, there's some well-known models in the market today. Let's say there's a pre- an end market trained model. And it took about, let's say, 10,000 A-100 or so to train. A-100 is the last generation in GPU, but it still applies in terms of relative scale here. So that company that used 10,800 to train their model, our understanding is they're going to need about a million GPUs within one to two years of launch to support the entire inference demand. So you can train the model on 10,000 of these chips, 10,000 of these systems, whatever there.
Starting point is 00:22:06 And then if they're actually going to be in the market and sell something or provide some service to make it work. to make it worthwhile, they're going to need a million? A million. And I think that's just within first two years of launch, Joe. Like, we're talking about something that's going to continue growing afterwards. And so what does a million GPUs mean, obviously, right? So, you know, a couple, I think it was like end of last year, all the hyperscalers combined, right? Amazon, Google, Microsoft, Oracle, you can throw a core even there.
Starting point is 00:22:36 There was about, you know, 500,000 GPUs globally, right, available across those platforms. I'd say at the end of this year, it'll be closer to a million or so. But that's suggesting then that one AI company with one model could consume the entire global footprint of GPUs. And now you start to think, wait, aren't there a bunch of other companies training these models in market right now? And I would say, yes, there are. So it can imply that there are in the short term, the demand of several million GPUs just to support the, inference market. And there's just nowhere near enough globally of this infrastructure. And it's going to be a big challenge for the market as we exit this training phase and move into the
Starting point is 00:23:25 productization or really just the commercialization of these models, like how do you generate revenue off them? And it's it's something that I don't think many people truly understand, just the amount of scale and construction that needs to take place. And now you put that in the same framework of the data centers that we were talking about, right? So there's this lack of data center space. There's lack of chipset supply. Like it's going to be an issue for for years that we see. So when it comes to scale, you know, you keep mentioning the hyper scaler scalers, which is a great term. But people like Amazon, Google, I guess, Microsoft, IBM, et cetera, how quickly, or what is your impression of how quickly they are able to ramp up in this space?
Starting point is 00:24:10 like how fast could they react to some of the trends that you've been outlining? Yeah. So I can offer what I'm seeing today. You know, the H-100s started to be distributed globally to all of us, right? Like all the entities that have these, you know, kind of upper-tier relationships with NVIDIA back in March. Right. So we started getting them this infrastructure online in April, really scaling in May. And, you know, we have builds going on at 10 data centers across the U.S.
Starting point is 00:24:40 right now, and we're delivering it to clients. The guidance that we're seeing from the hyperscalers is that they're not going to begin delivering scale access to the H-100 chip set until late Q3, maybe mid-Q-4, and some of them are even beginning to guide into Q1. And it's all driven by the fact that this is just a different type of compute that they're building relative to last generation, right? You're no longer just running Ethernet, to your point between all these. devices. You're not just plugging in CPU blades. You're having to deal with like totally different
Starting point is 00:25:15 data center power density and cooling requirements. You're having to build super computers instead with 500 miles of fiber and all these connections. It's just it's a completely different way to build the cloud and it's it's taking them some time to catch up because you have to retrain entire organizations to do this. So, you know, as of now, I'd say the direct answer is three quarters after a chip set launch. But it's seeming it might take longer. And I think that's all going to contribute to this just kind of slower ability to scale infrastructure than what's being dictated by the adoption rate of AI software. And it's going to lead to this supply demand imbalance. That will just last for a while.
Starting point is 00:25:57 You know, you keep mentioning, or we both keep mentioning, the H100 for obvious reasons. But do you look at other chips or what would happen to, you know, know, your own business, if for instance a new chip was developed that could do the same thing or better than an Nvidia H-100. Like, for instance, I hear a lot of excitement about some of the stuff that AMD is developing, and I'm not a chips expert, except maybe when it comes to Fritos or Lays. But, like, how big a difference would that make to you if we suddenly got a different chip manufacturer gain prominence in AI?
Starting point is 00:26:40 Sure. So I'd offer kind of two broad responses. One, typically when you train a model, you're going to use the same chips for inference on that model as well, right? So GPT4, for example, I was trained on A100s. They're predominantly going to use A100s going for. You might fit in some kind of newer generation hyper-efficient chips into there, but it's not like you need a, quote, a GP with more.
Starting point is 00:27:08 V-RAM on it, right? Like, you're going to need your 40-gig or your 80-grid gig RAM chip because that's the size of the model that you trained, right? You're not going to need like next multiple generations. You're not going to like really be able to adopt them to change the efficiency of serving that model. So what we view is that a chip's lifespan is like its first two to three years is spent training models. And then its next four to five years is spent doing inference for those models that it trained. And then within there as well, you do this thing called fine-tuning, which is updating the model with new information, right? Like, how do you keep a model like up-to-date with what's happened on a Twitter or what's happened on in the media? Right? You have to keep
Starting point is 00:27:53 retraining it, right? And you'll use those same chips to do that. But it's your question on other chipsets. And this is something that we have a particularly interesting view into because we have, like, you know, call it 650 AI clients. Right. And we're having to be a question. And we're in conversations with them daily to ensure that we're meeting their scaling demands. So it gives us a look into six to 12 months into the future what type of infrastructure they expect to need. And it's overwhelmingly people still want access to Nvidia chips. And the reason for this is something that dates back, I think it's nearly 15 years,
Starting point is 00:28:30 when Nvidia and Jensen made the decision to open source Kuda and to make this software set accessible to the machine learning community. And today, if you go to GitHub and you search a machine learning project, they just all reference Kuda drivers. And he's established this utter dominance of ecosystem around his compute within the ML space, really similar to like the X86 instruction set for CPU versus Arm. Right. Like X86 is used predominantly.
Starting point is 00:29:04 Arm has been trying to find its way into the space for a while now. it's just really struggled because all the engineers and developers are used to X86, similar to how all the engineers and developers in the AI space are used to using Kuda. So it's something that, like, obviously, AMD is highly incentivized to find a way into the sector, but they just don't have the ecosystem. And it's a huge moat to deal with. And, you know, kudos to Nvidia for establishing themselves and having the patience to stick with it and to continue to support that community over the last 15 years,
Starting point is 00:29:39 and it's really paying off for them in spades today. You know, if the demand comes for that infrastructure at some point, it's, you know, we can run other pieces of infrastructure within our data center. But I also find that Nvidia has such an advantage on the competition with not only its GPUs, but all of its components that support the GPUs like the Infineband fabric. that it's going to be a really difficult company to displace from the market in terms of the best standard for AI infrastructure. Can I ask you a question? And I'm going to, I want to ask this politely because it's not intended to be accusatory or anything like that. So I don't want you to hear it aside. But like when you're like talking about like hyperscalers and you're like, you know, Amazon, Google, Microsoft and, you know, kind of core wave. And it's like, okay, those are trillion dollar companies. and you're a $2 billion company.
Starting point is 00:30:34 Like, why, like, I still don't think I, like, wrap my head around, like, and I know, like, they're all, like, in, they're all talking about AI, et cetera. Like, can you still just, like, explain to me a little bit? Like, why aren't they just going to, frankly, like, steamroll you or be able to, let's put it this way, be able to, okay, maybe it'll take a few quarters to reevaluate things. But, like, you know, eventually this just becomes this sort of de facto offering from these big companies that have these huge cloud budgets that must be orders of magnitude larger than yours. Yeah, yeah, I would really love to be able to have access to their cost of capital.
Starting point is 00:31:10 That's for sure. So the way, look, it's, the way I talk about this is we don't have a silver bullet necessarily, right? I can't point to like a super secret piece of technology that we put inside of our servers or anything along those lines. But the way I like to broadly contextualize it is, is reference. another sector. And it's that like Ford should be able to produce a model Y, right? Like they have the budget, they have the people, they have the decades of expertise. But in order to ask them to produce a model Y, you would have to ask them to foundationally change the way that they produce a vehicle,
Starting point is 00:31:50 all the way from research to servicing. And that entire mechanism, like, it's a giant organization. Now you have to go ask that huge organization of people to change the way that they go about producing things. And I get that. But just to push back a little bit. And this is like a theme that comes up in various flavors on odd lots a lot, which is that like companies have internal. It's really hard to replicate sort of like tacit knowledge within a corporation. And we see that with companies that make semiconductor equipment. We see that with companies that make airplanes.
Starting point is 00:32:23 We see that with real estate developers that know how to turn an office. building into a condo. And so I think this is like a deep point. But, you know, they are offering AI stuff. Like, I can look at Google right now. Like, there's cloud AI. Like, and there's Asia AI. And they all have their announcements. So I'm still trying to understand, like, what is it that you're offering that all the hypers, they all have, they all say they have AI offerings. So what is the difference between sort of like what you have and what they say is like their, you know, AI compute platforms. Absolutely. And this will really depend on how much technical detail he'd like for me to get into, but broadly through infrastructure differentiation,
Starting point is 00:33:01 like literally using different components to build our cloud, and through software differentiation, we use different pieces of software to operate and optimize our cloud. We're able to deliver a product that's about 40 to 60 percent more efficient on a workload-adjusted basis than what you find across any of the hypers. So in other words, if you were to take the same workload or like go do the same process at a hyperscaler on the exact same GPU compute versus CoreWeave, we're going to be 40 to 60 percent more efficient at doing that because of the way that we've configured everything relative to the hypers. And it comes back to this analogy between like why Ford can't produce a model Y.
Starting point is 00:33:44 Again, like they can't. These are trillion dollar companies we're talking about. To your point, they have the budget, they have the personnel. And they certainly have the motivation to do so. But it's not just one singular thing they have to change. It's a completely different way to building their business that they would have to orchestrate. And it's what's the analogy is however many miles it takes to turn an aircraft carrier. Right.
Starting point is 00:34:07 Like it's going to take them a while to do that. And I think if they do get there at some point, which, you know, I don't disagree with you, they're certainly motivated to. It's going to have taken them some time, literally years, to get there. And they're going to look really similar to us. And meanwhile, I've dominated market share. And I've really established my product and market. And I'll continue differentiating myself on the software side of business as well.
Starting point is 00:34:32 Since we're on the topic of adaptation, can I ask about your own evolution as a company? Because I think I've read that you started out in Ethereum mining. And at one point, I'm pretty sure crypto mining was a substantial, if not the biggest portion. of your business, but you have clearly adapted or pivoted into this AI space. So what has that been like? And can you maybe describe some of the trends that you've seen over your history? Yes, absolutely. And you're right.
Starting point is 00:35:22 We did start within the cryptocurrency space back in 2017 or so. And that was spawned out of just, frankly, curiosity from a group of former commodity traders. So myself, my two crew of founders, we ran hedge funds, we ran family offices. So we traded in these energy markets. We were always attracted to supply demand mechanics. But what attracted us within cryptocurrency was there's this arbitrage opportunity that was a permissionless revenue stream, right? Like I knew the cost of power.
Starting point is 00:35:52 I knew what the hardware could generate in terms of revenue with using a power input, thus it's effectively an arbitrage. Right. So we explored that. We had some of that infrastructure operating. literally in our basements, as you said, Britain. We, and then that, like, quickly turned into scaling across warehouses. And at some point in 20, I think it was 2018, maybe late, 2018, we were the largest
Starting point is 00:36:20 Ethereum miner in North America. We were operating over 50,000 GPUs. We represented over 1% of the Ethereum network. But during that whole time, we just kept coming back to the idea that. that there's no moat, there's no advantage that we could create for ourselves relative to our competitors, right? Like, sure, you could maybe focus on power price and just kind of chase the cheapest power. But that just felt like chasing to the bottom of the bucket, right? You know, I think an area we could have gone into is producing your own chips, right?
Starting point is 00:36:52 Because if you produce the own chips and you run the mining equipment before anyone else has access to it, then you have an advantage for that period. But, you know, we weren't going to go design and fab her own chips. So what we kept coming back to was this GPU compute, man, what if we could do other things? What if we could develop uncorrelated optionality into multiple high growth markets? And those markets are where we predominantly sit today with an artificial intelligence, media and entertainment and computational chemistry. And the original thesis was, well, whenever our compute isn't being allocated into those sectors,
Starting point is 00:37:30 we'll just have it mining cryptocurrency and we'll build out this fantastic company that has 100% utilization rate across the infrastructure because it could switch immediately from being released from an AI workload into going back into the Ethereum network. And we did get a brief glimpse of being able to operate that way
Starting point is 00:37:48 in 2021 as we had our cloud live and we had AI clients in place. But Ethereum mining effectively ended during the merge in Q3 of, of 2022. But I'd say the other thing that we never appreciated was the utter complexity of running a CSP, forgetting about the software side of the business, which in and of itself, we spent about four years developing the software to build a modern cloud to do infrastructure orchestration and actually be a cloud service provider. The components themselves that the sector broadly used
Starting point is 00:38:28 for crypto mining were these retail grade GPUs, right? The kind of things that you plug in your desktop to go play. Right, the video games. They were like selling them on Stockex. Yes, yes, at bet. It was crazy during that period
Starting point is 00:38:43 to get your hands on that infrastructure for crypto mining. And all the video gamers hated the crypto people, right? Because they're like, I want to play this game and they would line up, what is it, GameStop and like the geek wire shop and all that or whatever it is. And they like couldn't get it because you got it.
Starting point is 00:38:58 Not you, but you know, the crypto people were getting access to the chips first and getting more value out of them so that you could bit them up. We were certainly part of the problem. And that's absolutely correct. But, you know, what we found ultimately is like those chips, that's not what you run enterprise grade workloads on. That's not what's supporting, you know, the largest AI companies in the world. And starting in 2019, we stopped buying any of those chips and only focused on purchasing enterprise grade GPU, chip sets that, you know, Nvidia has probably about 12 different skews that they offer, including A100 and H100 chips and really oriented our business around it.
Starting point is 00:39:39 So it's a, I don't expect to see much repurposing of this kind of older retail grade GPU equipment that was used for crypto mining. Because in crypto mining, you want to buy the cheapest chip that can do the thing for it, right, that can participate in crypto mining. But there's a huge difference in price. between a retail plug it into your computer so you can play video games chip and an enterprise grade you can run it 24-7 there's not going to be downtime you're going to have a low failure rate like there's there's a large technology difference and there's a large pricing difference between those
Starting point is 00:40:11 and the crypto miners you only needed the retail grade chip because you know if it went down for two percent five percent of the time for a failure rate that's not a big deal but the the tolerance the uptime tolerance for these enterprise-grade workloads is measured on the thousands of a percent, and it's a different type of infrastructure. So we don't expect to see the components really being reused, if at all. And then the other variable, going back to the very beginning of our conversation, are the data centers in which these are housed. So, Joe, to your point, earlier, we sit within Tier 3, Tier 4 data centers,
Starting point is 00:40:49 and that's basically the broad industry standard for being able to serve these kinds of of workloads. The crypto miners sat within tier zero, tier one data centers. And these things are highly interruptible. They do like really interesting things like helping load balance the power markets in places like ERCOT, right? Like they'll shut down when power prices go too high and it load balances the grid. But enterprise AI workloads don't have a tolerance for that. Their tolerance, again, is measured on the thousands of a percentage in terms of uptime. So not only does the infrastructure not work from crypto mining, but the data centers that they built within don't work either the way that they're currently configured. Now, they could potentially
Starting point is 00:41:37 convert their sites into Tier 3 and Tier 4 data centers. I'll tell you that in and of itself, that is an extremely challenging task. And it takes a lot of proprietary knowledge and industry expertise to do so. It's not just throwing a few fans in a room and a few air conditioning units. It's a, it's honestly, it feels like walking to a spaceship. Tracy, this is, this is an episode. I don't know about you, Tracy. There's like six another, like six follow-on episodes. It's like, how do you know, seriously, like the whole like data center market and the coolant and all, you know, the electricity.
Starting point is 00:42:11 Like there's so many different rabbit holes you could go down just like with the infrastructure you're talking about. For sure. And I think the estimates that I've seen on repurposing crypto GPUs, I think I've seen like, I've seen like five to 15 percent. So to Brandon's point, but I'm sure, I'm sure there will be people out there who try. You've got to try, right? Because what if it works? Right.
Starting point is 00:42:35 Right. If you can make that work, that's amazing. But we're just, you know, coming as an entity that was an extremely large operator of that infrastructure and has built, you know, one of the largest cloud service providers for AI workloads. I can tell you it's going to be really, really hard to do it. We've had exposure in both of those places. And at the end of the day, they're just very, very different businesses, both from the type of engineering and developers that you employ,
Starting point is 00:43:01 to the infrastructure, to the data centers that you sit within. So can I just go back, you know, just sort of like big picture? And I guess it sort of goes back to like, who gets access to what? Who gets access to chips? And I imagine that, you know, not only do you need a lot of money to, like, build a relationship with like, in video, you also probably need like a, you know, expectation you're going to be back the next year and back the next year and back the next year and back the next year and that you actually like have a relationship and so forth. But I have to imagine like planning is really tough.
Starting point is 00:43:32 And when like, you know, you have this sort of like AI machine language, whatever, like industry. And then something like jet jet GPT comes out and like suddenly everyone like, oh, I need to like have AI access. Talk to us about like this sort of like challenge of just sort of like planning the build. when it can move that fast and like everyone is just sort of guessing how big this market is going be in two to three years. Oh my gosh. It's it's been utterly insane, right? Like the, you know, back to last year, you know, the supply chain and ability to get your hands on components, you know, you would call your OEM. The OEM is the original equipment manufacturer. Like those are the super microes, the gigabytes of the world who actually, you know, build the notes, build the servers.
Starting point is 00:44:16 And you're buying through them. And then, they buy the GPUs from Nvidia, and build all the components together, right? So if you called them and said, hey, I need this many nodes to be delivered, they'll say, great, we'll start assembling. It takes us, you know, a week to two weeks to get the parts in assembling, and then it's another week for them to ship them to you, and then it takes us two to three weeks to plug them in, put them online, get them going, right? Now, that's completely changed.
Starting point is 00:44:43 As you know, like, all the supply chain has gotten thrown off, so much so that, you know, Nvidia is fully allocated, like they've fully sold out their infrastructure through the end of the year. You can't call them, you can't call the OEM and just say you need more compute chips. Like that's not possible. It's so much so that, you know, when clients are coming to us today and they're asking for, you know, like a 4,000 GPU cluster to be built for them, it's, we're telling them Q1. And increasingly it's moving towards Q2 at this point because Q1 is starting to get booked up right now. So it's something that a lot of time has been added to it. And then there's other supply chain variables within there as well.
Starting point is 00:45:24 We had a client earlier this year that we were in negotiations with them on the contract and we really wanted to perform well on timing for it. So we knew because of our orientation within the supply chain that there were some critical components that needed to be ordered ahead of time so that it would reduce our time to bringing the infrastructure online. And at that point, it was the power supply units and the fans for the nodes that the OEMs were putting together. And if we hadn't have done that, it would have been another, I think, eight weeks on top of the build process, just because not only components would have been there at the same time. So you're navigating this, you know,
Starting point is 00:46:07 within other kind of global supply chain disruptions and inflation and all these other things that are going on right now. And it's just an insanely complex. task that I think, you know, the generation of software developers and founders that we're working with today were used to being able to go to a cloud service provider and just getting whatever infrastructure they needed, right? You go to your hyperscalers and say, all right, I need this. And it was just there and available. And that just doesn't exist today because of the pace of demand growth that we've been on and just the lack of this infrastructure's availability. and it's just caught everyone by surprise.
Starting point is 00:46:48 Again, you're asking infrastructure to keep pace with the fastest adoption of a new piece of software that's ever occurred. Brandon McBee, Corleave, thank you so much. That was a great conversation. Like I said, I always sort of measure the quality of a conversation. Like, do I get seven ideas? How many additional episodes come out of it?
Starting point is 00:47:08 Like, that is a pretty good proxy for a good conversation. Do you get like eight ideas for future episodes? We got a bunch there. So thank you so much. for coming on the podcast. Always happy to chat with you guys, and thank you for the invite. Thanks, Brandon. Tracy, I want to find that company that makes the coolant for the data.
Starting point is 00:47:36 No, seriously, for the data centers that allows them to pack more computed, more energy into this space because it's like it feels like they're probably going to make a fortune in the next few years. Joe, I think you just want to talk to an HVAC contractor that's like installing your conditioners. Can we talk to an, just some random, like, I love the idea. Maybe it would have been such a funny thought like these like really advanced. And data centers, like, oh, do we have like a local air conditioning guy who can like come in? No, but I imagine. Actually, that would have been a good question for Brandon, wouldn't it?
Starting point is 00:48:04 Like the labor constraints in building and adapting some of those data centers. But there was so much in there. One of the things that I was thinking about was the point about how, well, okay, if you train a model on one type of chip, you're going to keep using that type of chip. And I guess it's kind of obvious, but it does suggest that there's some stickiness there. Like if you start out using an NVIDIA H100, you're going to keep using them. And in fact, you're going to consume even more because the processing power required, the compute required for the inference is higher than for the actual initial training. I knew that that was the case because Stacey said so as well. But I did not realize quite the scale of like how much more like, okay, like if you train a model and then we try to take it to market productize it as a business person.
Starting point is 00:48:55 I would say if we try to product type, like how much more computing power we would need for the inference aspect. And meanwhile, we have to keep training it all the time to keep it up with fresh data and stuff like that. Yeah, totally. And the other thing that I was thinking about, and again, Stacey mentioned this in our discussion with him as well, but this idea of invidia building a kind of large ecosystem around the hardware. So you have the open source software, Kuda, which we talked about a little bit. And then you have these sort of high touch partnerships with companies like CoreWeave where they're trying to make it as easy as possible for you to use their chips and set them up in a way that works for you. It feels almost like what Bitmain used to do.
Starting point is 00:49:44 Do you remember that? No. Maybe they're still doing it. Anyway, but it does feel like they're trying to build this like ecosystem moat around the chip technology. Yeah, no, it's absolutely true. And, you know, I really do take that point that Brandon made about, like, every company has a sort of like knowledge that cannot be written down on a piece of paper, which is a Dan Wong point that we've been talking about for years. And so it's like, to your point, you know, if like you have to like use different types of connectors and different types of power and all these stuff, like the ease with which any sort of traditional cloud provider or data center provider can, you know, sort of switch to it. It's like a, you know, it's like a, you know, it's like a, you know, it's. It's not trivial even with lots of money. No. But I'm coming away from that conversation thinking, like, the big question here is how quickly can those other hypers adapt?
Starting point is 00:50:34 Yeah. And, like, how big a moat can Nvidia build around this business? And then, I mean, the other question I have is, like, what if none of these companies make any money building AI models? Like, I still don't think, like, that's been proven. And so you can have this, like, huge boom and, like, we got to build an AI model. we're going to build like, you know, outlawed GPT for like data stuff and whatever. But it all is somewhat predicated on these companies being successful and making a lot of money. And if they're not, and if it turns out that like the monetization of AI products is trickier than expected,
Starting point is 00:51:06 then that also raises a question about like how long this like last. I'm sorry, Joe. So you're saying that tech companies should make money. Is that it? Are you sure? That's right. That's a, it's real post-ZERP thinking of me. I know. All right. Shall we leave it there? Let's leave it there.
Starting point is 00:51:23 This has been another episode of the Oddlots podcast. I'm Tracy Allaway. You can follow me on Twitter at Tracy Allaway. And I'm Joe Wisenthall. You can follow me on Twitter at The Stallwart. Follow our guest, Brandon McBee. He's at Brannon. Follow our producers, Carmen Rodriguez at Carmen Armin and Dashel Bennett at Dashbot. And check out all of the Bloomberg podcasts under the handle at Podcasts. And for more Odd Lots content, go to Bloomberg.com slash OddLots, where we have transcript. a blog and a newsletter that comes out each Friday. And check out our Discord. We have an AI channel and a semiconductor channel in there. So people talk about these topics 24-7. Maybe they'll be talking about them in both of those rooms when this comes out.
Starting point is 00:52:05 Discord.g.g. slash oddlots. And if you enjoy Odd Lots, if you appreciate conversations like the one we just had with Brandon McBeeigh, then please leave us a positive review on your favorite podcast platform. Thanks for listening. Thank you.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.