In The Arena by TechArena - Nebius on Why AI Infrastructure Is More Than GPUs

Episode Date: July 23, 2026

In this episode of Data Insights, Allyson Klein and Jeniece Wnorowski welcome Hitesh Kumar, GPU Cluster Design Architect at Nebius, for a discussion on the realities of designing and deploying AI infr...astructure at scale. The conversation explores how AI infrastructure has evolved beyond GPUs into a full-system challenge involving power, cooling, networking, storage, and interconnect technologies.

Transcript
Discussion (0)
Starting point is 00:00:00 Welcome to Tech Arena, featuring authentic discussions between tech's leading innovators and our host, Alison Klein. Now, let's step into the arena. Welcome in the arena. My name is Allison Klein. Today is a Data Insights episode, which means Janice Narowski is joining me. Janice, how are you doing? Hi, Allison. I'm doing great. How are you? I'm doing great. We have the most exciting guest with us today. I am so excited for the topic today. Janice, tell me who you brought with you. Couldn't agree more. It is all about understanding neocloud.
Starting point is 00:00:39 That's all the rage. And today we have Hitesh Kouar with Nibius. Hitesh, welcome to the show. I'm super excited for Hitesh. Hatesh is a GPU cluster design architect at Nibius. And one of the things we're going to discuss today is just go a little bit deeper around AI infrastructure, how it's evolving, and how it's really going beyond
Starting point is 00:00:59 kind of just networking storage, interconnecting, the overall full system architecture. So again, Hitesh, we're excited to have you here today. Thank you. Thank you. Yeah. So, Vatash, I'm so excited to talk to you. I've been thinking about this all the morning. Can you just start with a little bit? Denise introduced your role, but can you just give some contacts for how your role fits into Nebius's broader charter? And for those who might be a little less familiar with Nebius, which is hard to believe, given how much of a tear you guys have been on. But I'm sure there's some of our audience that I haven't heard a lot about Nibius. What is the company aspiring to in terms of building out
Starting point is 00:01:37 AI infrastructure? Sure. So NABUS is primarily an infrastructure provider, right, and in a fairly general sense now. So it's not just cloud GP resources, it's everything that you need around that. And they're quite well vertically integrated in the industry, I'd say, like a few other players, including the hyperscalers, achieve that. So they can provide you with all the cloud resource you need, your GPUs and storage and whatever, but also they are able to do, you know, dedicate clusters or even provide tokens directly through our token factory. So very well will vertically integrate it and we do all that through all of our own server,
Starting point is 00:02:15 rack designs, cluster designs. And my role in all that is being a GPU cluster architect. So let's say from when a site is chosen and power supply is arranged and all that, up to the point where logistics and delivery and deployment start, that gap in the middle is where my team comes in. So we have to logically plan out the clusters and then put them onto the floor plan and then figure out what connects to what and so on. Quite a broad set of tasks to do. That is a broad set of tasks. But you're also focused on designing GPU clusters at a time when the demand for AI is just so great and computers accelerating rapidly. How would you describe what's
Starting point is 00:02:57 changed in the past couple of years in terms of how these systems are designed and deployed. We can start with the obvious one, I guess. I mean, power and cooling density, right? About a year and a half ago, I was in Norway deploying Nvidia H-200 GPU servers in a mid-sized data center there, so about, I think, it was 30 or 40 megawatts. And you'd have rows of racks and you have three or four nodes per a rack. And that's about 24 to 32 GPUs in what's like essentially a cabinet. Nowadays, we've moved up to like 72 or even more GPUs per rack and way more racks within the data center. The GPUs themselves, the servers themselves have maybe one and a half times, maybe a bit more the power draw that they used to. So that sort of
Starting point is 00:03:45 explosion and density over the last four years has made it quite challenging but also quite exciting to be deploying these things. These deployments and buildings and so on lag a bit behind. is it tough to retrofit or build new buildings within two years that are able to suit these much more powerful and much more dense clusters. That's been the biggest shock definitely so far. Yeah, a lot of considerations we've had to make in the sites that we select and how we design them, how we make these clusters fit. And it's not just liquid cooling.
Starting point is 00:04:16 That's one of the new things. It's also keeping air cooling going for other generation, seeing how we can push the limits of that and design that sensibly as well. But yeah, that's the most obvious biggest change. change as power and calling density. Now, there's a lot of attention of the GPs themselves and there has been for a number of years, but I mean, if you follow this talk work lately, you can see that everyone's waking up that there's a lot more that goes in to building AI data centers.
Starting point is 00:04:40 And your writing often focuses on broader system, networking interconnects, and really the role of data movement within AI workflows. Why is it important to think about AI infrastructure as this full system rather than just individual components that are delivery white compute. You're right. GPUs do get a lot of attention, mainly because when you look at the financials
Starting point is 00:05:04 or the analysts who look at the markets, things are often measured in either number of GPUs or megawatts of compute, sometimes measured in dollars as well. But it's always never measured in the amount of networking you need or the amount of calling that you need. That's something that comes a little bit later when you have to consider the practical side of deployments, right?
Starting point is 00:05:23 All that supporting infrastructure to get the GPUs to a point where you can actually use them, run something meaningful, full on them, and keep them stable. All of that is part of my role. First of all, it's always power and cooling. So that's the first thing you've got to check is once you have that number in mind, you want to deploy 10,000 GPUs, let's say, can you get the power into the site and distribute it and keep all your components cool.
Starting point is 00:05:47 Then it's networking and storage, the management plane for the GPUs. Can I orchestrate them? Can I run things and monitor how they're going? and store all my data, connect everything with a high-speed fabric, so all the GPs that talk to each other. Once you have all that planned out, finally then you can start worrying about the 10,000 GPs that you heard in the headline of the article.
Starting point is 00:06:08 And we often even abstract away GPUs quite a lot of the time, so we get told that we're deploying these GPUs in a cluster, and we say, okay, this is what the GPU rack will look like. These days it might be NVIDIA's NVL-72s that are quite popular. So we'll just get that given to us in like a, template, we'll put all those into places that we kind of need, start working things around that and then do this kind of cyclical thing where this sort of fits, now we're going to change a little bit of this and refits it all again. But in that whole process, GPUs are
Starting point is 00:06:38 quite often abstracted away into something that we always either come back to later or figure out right at the start. So yeah, my job as a GPU cluster architect focuses, a lot of things that aren't GPUs. Okay, because I was going to be about my next question, Hittash. From your perspective, where do you see storage kind of fitting into the broader system design, especially as data volumes grow beyond our wildest dreams, what workloads are becoming more distributed in your opinion? And what teams are kind of underestimating the value of storage in the overall infrastructure? Storage has always been really important. More so now. I mean, we're seeing that from the market side as well, right? I believe people are signing longer and longer term contracts to get hold.
Starting point is 00:07:24 of storage now as well. Storage bandwidths are increasing, storage volumes are increasing as well. For AI workloads, for training, it's always been obvious, is to store those data sets and then get the parts of the data you need to your GPUs fast enough so that you don't store the training on your data feeding. For inference now, actually, it's getting really interesting to it, because models are getting so big and context sizes are increasing so long, so the amount of conversation you can have with whatever model you're talking to, for example, because they're getting, getting so large, we need to start moving the whole context of that conversation, how that's stored inside memory, needs to start getting moved off the actual GPU and sometimes even off
Starting point is 00:08:05 the host itself. So you'll see this from Jensen Huang's keynote where he was talking about ICMS or STX or whatever NVIDIA will decide to call it in a couple of weeks. That is essentially be using storage as a buffer or an intermediate tier between what your GPUs will have very close to them and what they'll need to read soon or within a certain amount of time. So, yeah, storage will become more and more important. And the demand for storage is just increasing like crazy. And I don't really see any change in that. I think storage will, yeah, will continue being just as important. How do you think it advances in interconnected technologies and cluster architectures are changing what's possible in AI training in inference today.
Starting point is 00:08:51 I know a lot of this work is built upon high-performance compute clusters and stuff like that, but I know that we're really pushing the limits of scale. How do you see that changing both from a standpoint of what's going on in the industry and what operators like Nebius are bringing to the table? So interconnects are getting more and more dense too. We're seeing in the short term, we're moving from what we now use is plugable optics, which are single transceivers, and those will connect one or two cables, which will then use optical fiber to connect between servers.
Starting point is 00:09:23 We're moving from smaller plugables to much larger plugables. And then the next one or two years, I mean, as we've seen recently at, I think, OFC it was the conference where Arista announced an XPO design, so that's, I think, extreme plugable optics or whatever it stands for. They're getting way more dense, and now you'll have up to, I think, 12, 16 or more ports per a single plugable optic, and even those would be liquid-cooled. So now the thing that you put into your server to connect your optical fibers will be liquid-cooled too. That's getting way more and more dense.
Starting point is 00:09:54 And then in the next, say, five or seven or eight years, I think, we're looking at CPO and NPO. So that's what NVIDIA has been talking about a lot recently, where you'll have the optical inches themselves within the server so you can directly attach your fibres to the ports on your server. and that's going to be coming in soon where you'll get even more dense. So our networking density is increasing a lot, and that brings with all sorts of challenges around cooling those,
Starting point is 00:10:20 cabling, those rags that are now extremely dense, and then a new supply chain for all those parts that we have to work towards. So a lot of people in the industry are really having to balance between sticking with what the mature supply chain is able to do, which is providing these plug-able optics, but then also preparing or sort of experimenting with the newer stuff as it comes along so they don't miss out on that wave. I would imagine NEPA is evaluating the same.
Starting point is 00:10:46 I wouldn't know for sure, but we definitely need a lot of optics too. So it is surely on our minds. Yeah. And a lot of organizations are really trained to build and access to GPU infrastructure, as you've talked about. But what are some of the common kind of misconceptions or oversimplifications you see when people think about just adding more GPUs? So yeah, you could just purchase and add more.
Starting point is 00:11:10 GPUs, but then everything else around gets a lot more complicated very quickly. One of the most surprising things that I think people often find out the hard way is that the failure rates of the components in your cluster, so as a cluster is running, doing its job or whatever, occasionally different parts will fail just because they randomly do. The rate at which different parts fail, now because you have so many more GPUs, so many more optics and all the other parts, the rate at which they fail grows linearly or sub-linearly, which is manageable. You can add more staff for your maintenance, have a bigger budget for that. But the amount that affects your workload now just explodes way faster.
Starting point is 00:11:49 I think meta have probably the best publicly available data set and study on this, where they published the findings from like this 16,000 GPU cluster they worked on, which I think was A100s or H-100s, like Nvidia GPUs. In short, they experienced a, failure every three hours on average. And just thinking about that, if you're not metta and you don't have the software and the monitoring systems and the staff to be able to manage all of that, how do you then run anything meaningful on a cluster that large if you're failing every three hours, right? At that point, it's the most difficult to keep pausing and restarting.
Starting point is 00:12:25 So your software has to get really fault tolerant, which, you know, we're kind of getting there. Even some open source stuff is doing that. Obviously, your staff and your operations and everything and the amount of spares you hold, the logistics that you do, all of that has to scale. Now to match that increased error rate. And this mostly affects like the really large workload,
Starting point is 00:12:44 so like training, where you want thousands of GPUs working together to improve a model progressively. For things like inference, you can break your cluster into small parts and hide it all behind a CDN, a distribution network, right? But for training or really large workloads,
Starting point is 00:12:59 it gets pretty tough to manage that. That's one thing that really surprises people. I would say another thing, thing that I found out quite recently actually is the scale of the power that these clusters now draw. So when you do add, let's say, 10,000 more GPUs, after a while you get to a point where the power draw of that site is now so significant, you might actually start affecting the grid if you're connected to it, right? So imagine your training run just finishes a checkpoint where you now want to pause and take a recording of the models to save it in case something fails later.
Starting point is 00:13:30 At that point where you do that, a lot of different components, the data transfer, the GPS compute themselves, networking, all of that will just like pause. And so you're utilising your power usage from your cluster, even say your rack goes from a 110 kilowatts down to about 30, 40 or 50 or something. That is a huge drop, actually, for a lot of these power systems. So as you do increase to adding more and more GPUs, you also need to consider components of capacitors or batteries or anything intermediate between. new in your power sources, ways to mitigate these power drops, whether you want to do that through software even. This is something that also catches a lot of people out, which not many even talk about. I'll stop there.
Starting point is 00:14:12 Those are the main two things I'd say. You'd go a lot. And those are like things that I would think that we could actually do follow-up interviews on those as just important interviews, because there's such pressing challenges. I do want to say, I think for anybody who reads Tech Arena, I think they know that we're Nebis fan girls. We write about you guys. you're positioning yourselves as an AI-native cloud provider.
Starting point is 00:14:32 How does building infrastructure specifically for AI workbloods differ in your mind from just building out traditional hyperscaler cloud design? So for traditional cloud design, let's say, web hosting databases, or if you want to do scientific computers, so supercomputers for things like weather forecasting or so on. Traditionally, what we think of as really big deployments, you do want the density of your compute, so your CPUs, you want to get as many cores as you can, or you want as much memory as you can, certain footprint, whether that's a powerful footprint or a space.
Starting point is 00:15:08 Networking and storage haven't really been that important in that era. Storage has been so, but mostly from a point of, I would say, archive and rather than having warm data access like you would for an AI workload. And networking, of course, has just exploded in bandwidth in the last, let's say, five or so years. because of the needs of AI to be shuffling so much data around. Whereas traditionally, you wouldn't really need to shuffle that much data around between the actual servers themselves. You'd want, obviously, for scientific computing, you'd want lower latency and stuff. You'd have some requirements, but not nearly as much as we do now.
Starting point is 00:15:41 So all of that has exploded when it comes to AI work close, just because of the amount of data we need to move around between servers. I think that really is what's triggered that change. I mean, that's a side effect, of course, of model size and the amount of context these language. models have. So yeah, networking and storage have they been really huge. Of course, like before I said power and cooling, we already went over that. There's a lot of adjustments you need to make for those two. But yeah, I would say those are the two other biggest points that we've seen. Nice. Amazing.
Starting point is 00:16:14 Yeah, thanks, Detesh. That's actually really insightful. I agree with Allison. We should come back to some of these topics that you're hitting on. For teams that are beginning to think about scaling their AI infrastructure and workloads, how should they approach decisions? around the whole infrastructure, whether it's to build, buy, or even partner. Yeah, very good point. So if you are scaling from a sort of a starting position, if you're going from zero to one rather than one to 100, at that point, you really don't want to be thinking about buying, I would say, your own infrastructure, because the startup costs and the overheads on that
Starting point is 00:16:44 are quite high. So you can have a few smaller GPUs, let's say, on your own site. Maybe you'll have some racks for them, that one, and a proper power supply. But when you start to scale beyond that, when you want to get these enterprise class nodes in your own site, you then need to start thinking about proper power supply cooling, of course, as we said. But you need staff to be able to run and maintain those. Then you need the software for that. You need licenses and insurance. And that's the whole minefield that we don't want to step into.
Starting point is 00:17:12 But the costs really scale. The overheads are quite high and the startup costs are quite a lot. So at the start, it's probably as to sick with cloud resources. once you do scale beyond a point where your own financial models, say it makes sense to have your own servers, then you want to look at some sort of co-location agreement where you can have a few racks and a well-built and maintain data center. And then when you are really pushing the limits of that, then it makes sense to get your own site and then you have all the infrastructure to support having site. But it's a progressive thing. And until you really, really think that you have enough guaranteed demand to scale beyond a certain point,
Starting point is 00:17:49 as best to stick to the cloud and then slowly bring things on site. When you're thinking about the next few years, we can imagine models getting more and more powerful enterprises, really driving broad proliferation of adoption, both from a standpoint of, you know, fight tuning models for their utilization, but also just a huge wave of inference. How do you see AI cluster design evolving to meet this growing demand?
Starting point is 00:18:16 And what do you think we're going to be talking about in a couple of years about challenges within the underlying infrastructure. The first thing I would say is efficiency will definitely improve, and that's the dollars per flop of performance, or dollars per performance, you could say. That would massively improve, and that would improve through the whole cluster. So every part of your stack will get more and more efficient and you'll get much more bad for your buck when it comes to running workloads on these
Starting point is 00:18:41 clusters. The one I find most interesting, though, is how I think networking designs for these will change. So I think when we say workflows become more distributed, we should clarify, I think that instead the clusters will become more sparsely connected. And workload will still be quite tightly run on those. But then you'll have lots of small islands, let's say, within your data center, within your whole multi-site fabric, whatever it might be. You'll have these lots of small islands of very tightly connected to GPUs. And then between those islands, you'll have more sparse connections. We're already seeing this now with how
Starting point is 00:19:18 Nvidia and the industry is moving towards these tightly connected rag-scale designs where you have the NVL 7-2, which is like the premier product of Nvidia, the NVN that stands for NVL, right? NVL, rather. So it's all connected with a very high bandwidth fabric within that rack. All the GPUs talk to each other really fast, and that's your sort of island. And between those racks, you have a scale-out fabric, which is, you know, let's say an order of magnitude slower.
Starting point is 00:19:43 that then means you have these tightly connected islands with sparsely connected fabric between them and then as you scale beyond that now we're seeing things like scale across being mentioned that then means that you have multiple data calls within a site or even multiple sites that could be as much as kilometers apart these are then weakly connected too so you can see how we have this sort of fractal pattern where you got tightly connected sparse and then you zoom out more and that's also tightly connected and you zoom out more that's also sparse we're seeing the same thing with Google's new TPUs. Huawei's ascend accelerators too. They're following the same sort of pattern of being quite a high bandwidth in a small region and then more sparse later. I think that will be
Starting point is 00:20:23 a predominant pattern that will be seeing again and again, which are already seeing some of. I think that would be the most interesting thing for me at least. Yeah. I've got a lot of interesting things from you in this conversation, Hadesh. And I know our audience and our listeners would love to learn more. Where should they go? So I would say the best general purpose source might still be semi-analysis, although hopefully one day there'll be more competition around that. After that, probably Google and Microsoft have very good posts and blogs that you can learn from. Some places have pretty good technical documentation too. I would say just throwing some examples, they're Corning, Naddot, both have really good blogs and technical documentation.
Starting point is 00:21:05 InVidia's own posts and documentation are pretty decent too. If you want to hear what I have to say, you can check out my LinkedIn, so just my name, Hittesh Kumar, or my substack at six rack units. I try to post topics like these there, too. I'm going to be a reader, Natasha, and thank you so much for being part of the tech arena. It's been a real pleasure. I would love to have you back on the show to probe deeper into some of the topics that you brought up. And Janice, thank you so much for another great Data Insights episode.
Starting point is 00:21:34 We keep bringing the heat with our guests, so thanks so much for the collaboration. You're welcome. Thank you. Thank you, Alison. Thank you, Hatesh. Thank you, for joining Tech Arena. Subscribe and engage at our website, Techorina.aI. All content is copyright by Techarena.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.