In The Arena by TechArena - Nebius on Why AI Infrastructure Is More Than GPUs
Episode Date: July 23, 2026In this episode of Data Insights, Allyson Klein and Jeniece Wnorowski welcome Hitesh Kumar, GPU Cluster Design Architect at Nebius, for a discussion on the realities of designing and deploying AI infr...astructure at scale. The conversation explores how AI infrastructure has evolved beyond GPUs into a full-system challenge involving power, cooling, networking, storage, and interconnect technologies.
Transcript
Discussion (0)
Welcome to Tech Arena, featuring authentic discussions between tech's leading innovators and our host, Alison Klein.
Now, let's step into the arena.
Welcome in the arena. My name is Allison Klein. Today is a Data Insights episode, which means
Janice Narowski is joining me. Janice, how are you doing? Hi, Allison. I'm doing great. How are you?
I'm doing great. We have the most exciting guest with us today. I am so excited for the topic today.
Janice, tell me who you brought with you.
Couldn't agree more.
It is all about understanding neocloud.
That's all the rage.
And today we have Hitesh Kouar with Nibius.
Hitesh, welcome to the show.
I'm super excited for Hitesh.
Hatesh is a GPU cluster design architect at Nibius.
And one of the things we're going to discuss today
is just go a little bit deeper around AI infrastructure,
how it's evolving, and how it's really going beyond
kind of just networking storage, interconnecting,
the overall full system architecture. So again, Hitesh, we're excited to have you here today.
Thank you. Thank you. Yeah. So, Vatash, I'm so excited to talk to you. I've been thinking
about this all the morning. Can you just start with a little bit? Denise introduced your role,
but can you just give some contacts for how your role fits into Nebius's broader charter?
And for those who might be a little less familiar with Nebius, which is hard to believe,
given how much of a tear you guys have been on. But I'm sure there's some of our audience that
I haven't heard a lot about Nibius. What is the company aspiring to in terms of building out
AI infrastructure?
Sure. So NABUS is primarily an infrastructure provider, right, and in a fairly general sense now.
So it's not just cloud GP resources, it's everything that you need around that.
And they're quite well vertically integrated in the industry, I'd say, like a few other players,
including the hyperscalers, achieve that. So they can provide you with all the cloud resource you need,
your GPUs and storage and whatever, but also they are able to do, you know,
dedicate clusters or even provide tokens directly through our token factory.
So very well will vertically integrate it and we do all that through all of our own server,
rack designs, cluster designs.
And my role in all that is being a GPU cluster architect.
So let's say from when a site is chosen and power supply is arranged and all that,
up to the point where logistics and delivery and deployment start, that gap in the middle is
where my team comes in. So we have to logically plan out the clusters and then put them onto the
floor plan and then figure out what connects to what and so on. Quite a broad set of tasks to do.
That is a broad set of tasks. But you're also focused on designing GPU clusters at a time when
the demand for AI is just so great and computers accelerating rapidly. How would you describe what's
changed in the past couple of years in terms of how these systems are designed and deployed.
We can start with the obvious one, I guess. I mean, power and cooling density, right?
About a year and a half ago, I was in Norway deploying Nvidia H-200 GPU servers in a mid-sized
data center there, so about, I think, it was 30 or 40 megawatts. And you'd have rows of racks
and you have three or four nodes per a rack. And that's about 24 to 32 GPUs in what's like
essentially a cabinet. Nowadays, we've moved up to like 72 or even more GPUs per rack and
way more racks within the data center. The GPUs themselves, the servers themselves have maybe
one and a half times, maybe a bit more the power draw that they used to. So that sort of
explosion and density over the last four years has made it quite challenging but also quite
exciting to be deploying these things. These deployments and buildings and so on lag a bit behind.
is it tough to retrofit or build new buildings within two years
that are able to suit these much more powerful and much more dense clusters.
That's been the biggest shock definitely so far.
Yeah, a lot of considerations we've had to make in the sites that we select
and how we design them, how we make these clusters fit.
And it's not just liquid cooling.
That's one of the new things.
It's also keeping air cooling going for other generation,
seeing how we can push the limits of that and design that sensibly as well.
But yeah, that's the most obvious biggest change.
change as power and calling density.
Now, there's a lot of attention of the GPs themselves and there has been for a number of years,
but I mean, if you follow this talk work lately, you can see that everyone's waking up that
there's a lot more that goes in to building AI data centers.
And your writing often focuses on broader system, networking interconnects, and really the
role of data movement within AI workflows.
Why is it important to think about AI infrastructure as this full system rather than just
individual components that are delivery
white compute.
You're right.
GPUs do get a lot of attention,
mainly because when you look at the financials
or the analysts who look at the markets,
things are often measured in either number of GPUs
or megawatts of compute,
sometimes measured in dollars as well.
But it's always never measured in the amount of networking
you need or the amount of calling that you need.
That's something that comes a little bit later
when you have to consider the practical side of deployments, right?
All that supporting infrastructure
to get the GPUs to a point where you can actually use them,
run something meaningful, full on them, and keep them stable.
All of that is part of my role.
First of all, it's always power and cooling.
So that's the first thing you've got to check is once you have that number in mind,
you want to deploy 10,000 GPUs, let's say,
can you get the power into the site and distribute it and keep all your components cool.
Then it's networking and storage, the management plane for the GPUs.
Can I orchestrate them?
Can I run things and monitor how they're going?
and store all my data, connect everything with a high-speed fabric,
so all the GPs that talk to each other.
Once you have all that planned out,
finally then you can start worrying about the 10,000 GPs
that you heard in the headline of the article.
And we often even abstract away GPUs quite a lot of the time,
so we get told that we're deploying these GPUs in a cluster,
and we say, okay, this is what the GPU rack will look like.
These days it might be NVIDIA's NVL-72s that are quite popular.
So we'll just get that given to us in like a,
template, we'll put all those into places that we kind of need, start working things around
that and then do this kind of cyclical thing where this sort of fits, now we're going
to change a little bit of this and refits it all again. But in that whole process, GPUs are
quite often abstracted away into something that we always either come back to later or figure
out right at the start. So yeah, my job as a GPU cluster architect focuses, a lot of things
that aren't GPUs. Okay, because I was going to be about my next question, Hittash. From your
perspective, where do you see storage kind of fitting into the broader system design, especially as
data volumes grow beyond our wildest dreams, what workloads are becoming more distributed in your
opinion? And what teams are kind of underestimating the value of storage in the overall infrastructure?
Storage has always been really important. More so now. I mean, we're seeing that from the market
side as well, right? I believe people are signing longer and longer term contracts to get hold.
of storage now as well. Storage bandwidths are increasing, storage volumes are increasing as well.
For AI workloads, for training, it's always been obvious, is to store those data sets and then
get the parts of the data you need to your GPUs fast enough so that you don't store the training
on your data feeding. For inference now, actually, it's getting really interesting to it, because
models are getting so big and context sizes are increasing so long, so the amount of conversation
you can have with whatever model you're talking to, for example, because they're getting,
getting so large, we need to start moving the whole context of that conversation, how that's
stored inside memory, needs to start getting moved off the actual GPU and sometimes even off
the host itself. So you'll see this from Jensen Huang's keynote where he was talking about
ICMS or STX or whatever NVIDIA will decide to call it in a couple of weeks. That is essentially
be using storage as a buffer or an intermediate tier between what your GPUs will have very close
to them and what they'll need to read soon or within a certain amount of time. So, yeah,
storage will become more and more important. And the demand for storage is just increasing like
crazy. And I don't really see any change in that. I think storage will, yeah, will continue being
just as important. How do you think it advances in interconnected technologies and cluster architectures
are changing what's possible in AI training in inference today.
I know a lot of this work is built upon high-performance compute clusters and stuff like that,
but I know that we're really pushing the limits of scale.
How do you see that changing both from a standpoint of what's going on in the industry
and what operators like Nebius are bringing to the table?
So interconnects are getting more and more dense too.
We're seeing in the short term, we're moving from what we now use is
plugable optics, which are single transceivers, and those will connect one or two cables,
which will then use optical fiber to connect between servers.
We're moving from smaller plugables to much larger plugables.
And then the next one or two years, I mean, as we've seen recently at, I think,
OFC it was the conference where Arista announced an XPO design,
so that's, I think, extreme plugable optics or whatever it stands for.
They're getting way more dense, and now you'll have up to, I think, 12, 16 or more ports
per a single plugable optic, and even those would be liquid-cooled.
So now the thing that you put into your server to connect your optical fibers will be liquid-cooled too.
That's getting way more and more dense.
And then in the next, say, five or seven or eight years, I think, we're looking at
CPO and NPO.
So that's what NVIDIA has been talking about a lot recently, where you'll have the optical
inches themselves within the server so you can directly attach your fibres to the ports on your
server.
and that's going to be coming in soon where you'll get even more dense.
So our networking density is increasing a lot,
and that brings with all sorts of challenges around cooling those,
cabling, those rags that are now extremely dense,
and then a new supply chain for all those parts that we have to work towards.
So a lot of people in the industry are really having to balance
between sticking with what the mature supply chain is able to do,
which is providing these plug-able optics,
but then also preparing or sort of experimenting
with the newer stuff as it comes along so they don't miss out on that wave.
I would imagine NEPA is evaluating the same.
I wouldn't know for sure, but we definitely need a lot of optics too.
So it is surely on our minds.
Yeah.
And a lot of organizations are really trained to build and access to GPU infrastructure,
as you've talked about.
But what are some of the common kind of misconceptions or oversimplifications you see
when people think about just adding more GPUs?
So yeah, you could just purchase and add more.
GPUs, but then everything else around gets a lot more complicated very quickly. One of the most
surprising things that I think people often find out the hard way is that the failure rates of the
components in your cluster, so as a cluster is running, doing its job or whatever, occasionally different
parts will fail just because they randomly do. The rate at which different parts fail,
now because you have so many more GPUs, so many more optics and all the other parts,
the rate at which they fail grows linearly or sub-linearly, which is manageable.
You can add more staff for your maintenance, have a bigger budget for that.
But the amount that affects your workload now just explodes way faster.
I think meta have probably the best publicly available data set and study on this,
where they published the findings from like this 16,000 GPU cluster they worked on,
which I think was A100s or H-100s, like Nvidia GPUs.
In short, they experienced a,
failure every three hours on average. And just thinking about that, if you're not metta and you
don't have the software and the monitoring systems and the staff to be able to manage all of that,
how do you then run anything meaningful on a cluster that large if you're failing every three
hours, right? At that point, it's the most difficult to keep pausing and restarting.
So your software has to get really fault tolerant, which, you know, we're kind of getting there.
Even some open source stuff is doing that. Obviously, your staff and your operations and everything
and the amount of spares you hold,
the logistics that you do,
all of that has to scale.
Now to match that increased error rate.
And this mostly affects
like the really large workload,
so like training,
where you want thousands of GPUs
working together to improve a model progressively.
For things like inference,
you can break your cluster into small parts
and hide it all behind a CDN,
a distribution network, right?
But for training or really large workloads,
it gets pretty tough to manage that.
That's one thing that really surprises people.
I would say another thing,
thing that I found out quite recently actually is the scale of the power that these clusters now
draw. So when you do add, let's say, 10,000 more GPUs, after a while you get to a point where
the power draw of that site is now so significant, you might actually start affecting the grid
if you're connected to it, right? So imagine your training run just finishes a checkpoint where you now
want to pause and take a recording of the models to save it in case something fails later.
At that point where you do that, a lot of different components, the data transfer, the GPS compute themselves, networking, all of that will just like pause.
And so you're utilising your power usage from your cluster, even say your rack goes from a 110 kilowatts down to about 30, 40 or 50 or something.
That is a huge drop, actually, for a lot of these power systems.
So as you do increase to adding more and more GPUs, you also need to consider components of capacitors or batteries or anything intermediate between.
new in your power sources, ways to mitigate these power drops, whether you want to do that
through software even.
This is something that also catches a lot of people out, which not many even talk about.
I'll stop there.
Those are the main two things I'd say.
You'd go a lot.
And those are like things that I would think that we could actually do follow-up interviews
on those as just important interviews, because there's such pressing challenges.
I do want to say, I think for anybody who reads Tech Arena, I think they know that
we're Nebis fan girls.
We write about you guys.
you're positioning yourselves as an AI-native cloud provider.
How does building infrastructure specifically for AI workbloods differ in your mind
from just building out traditional hyperscaler cloud design?
So for traditional cloud design, let's say, web hosting databases,
or if you want to do scientific computers, so supercomputers for things like weather forecasting
or so on.
Traditionally, what we think of as really big deployments,
you do want the density of your compute, so your CPUs, you want to get as many cores as you can,
or you want as much memory as you can, certain footprint, whether that's a powerful footprint or a space.
Networking and storage haven't really been that important in that era.
Storage has been so, but mostly from a point of, I would say, archive and rather than having warm data access like you would for an AI workload.
And networking, of course, has just exploded in bandwidth in the last, let's say, five or so years.
because of the needs of AI to be shuffling so much data around.
Whereas traditionally, you wouldn't really need to shuffle that much data around
between the actual servers themselves.
You'd want, obviously, for scientific computing, you'd want lower latency and stuff.
You'd have some requirements, but not nearly as much as we do now.
So all of that has exploded when it comes to AI work close,
just because of the amount of data we need to move around between servers.
I think that really is what's triggered that change.
I mean, that's a side effect, of course, of model size
and the amount of context these language.
models have. So yeah, networking and storage have they been really huge. Of course, like before I said
power and cooling, we already went over that. There's a lot of adjustments you need to make for those two.
But yeah, I would say those are the two other biggest points that we've seen. Nice. Amazing.
Yeah, thanks, Detesh. That's actually really insightful. I agree with Allison. We should come back to
some of these topics that you're hitting on. For teams that are beginning to think about scaling
their AI infrastructure and workloads, how should they approach decisions?
around the whole infrastructure, whether it's to build, buy, or even partner.
Yeah, very good point.
So if you are scaling from a sort of a starting position, if you're going from zero to one
rather than one to 100, at that point, you really don't want to be thinking about buying,
I would say, your own infrastructure, because the startup costs and the overheads on that
are quite high.
So you can have a few smaller GPUs, let's say, on your own site.
Maybe you'll have some racks for them, that one, and a proper power supply.
But when you start to scale beyond that, when you want to get these enterprise class nodes in your own site, you then need to start thinking about proper power supply cooling, of course, as we said.
But you need staff to be able to run and maintain those.
Then you need the software for that.
You need licenses and insurance.
And that's the whole minefield that we don't want to step into.
But the costs really scale.
The overheads are quite high and the startup costs are quite a lot.
So at the start, it's probably as to sick with cloud resources.
once you do scale beyond a point where your own financial models, say it makes sense to have your own servers,
then you want to look at some sort of co-location agreement where you can have a few racks and a well-built and maintain data center.
And then when you are really pushing the limits of that, then it makes sense to get your own site and then you have all the infrastructure to support having site.
But it's a progressive thing.
And until you really, really think that you have enough guaranteed demand to scale beyond a certain point,
as best to stick to the cloud and then slowly bring things on site.
When you're thinking about the next few years,
we can imagine models getting more and more powerful enterprises,
really driving broad proliferation of adoption,
both from a standpoint of, you know,
fight tuning models for their utilization,
but also just a huge wave of inference.
How do you see AI cluster design evolving to meet this growing demand?
And what do you think we're going to be talking about
in a couple of years about challenges within the underlying infrastructure.
The first thing I would say is efficiency will definitely improve,
and that's the dollars per flop of performance,
or dollars per performance, you could say.
That would massively improve, and that would improve through the whole cluster.
So every part of your stack will get more and more efficient
and you'll get much more bad for your buck when it comes to running workloads on these
clusters.
The one I find most interesting, though, is how I think networking designs for these will change.
So I think when we say workflows become more distributed, we should clarify, I think that instead the clusters will become more sparsely connected.
And workload will still be quite tightly run on those.
But then you'll have lots of small islands, let's say, within your data center, within your whole multi-site fabric, whatever it might be.
You'll have these lots of small islands of very tightly connected to GPUs.
And then between those islands, you'll have more sparse connections.
We're already seeing this now with how
Nvidia and the industry is moving towards these tightly connected rag-scale designs
where you have the NVL 7-2, which is like the premier product of Nvidia,
the NVN that stands for NVL, right?
NVL, rather.
So it's all connected with a very high bandwidth fabric within that rack.
All the GPUs talk to each other really fast, and that's your sort of island.
And between those racks, you have a scale-out fabric,
which is, you know, let's say an order of magnitude slower.
that then means you have these tightly connected islands with sparsely connected fabric between them
and then as you scale beyond that now we're seeing things like scale across being mentioned
that then means that you have multiple data calls within a site or even multiple sites that could be
as much as kilometers apart these are then weakly connected too so you can see how we have this
sort of fractal pattern where you got tightly connected sparse and then you zoom out more and that's
also tightly connected and you zoom out more that's also sparse we're seeing the same thing with
Google's new TPUs. Huawei's ascend accelerators too. They're following the same sort of pattern
of being quite a high bandwidth in a small region and then more sparse later. I think that will be
a predominant pattern that will be seeing again and again, which are already seeing some of.
I think that would be the most interesting thing for me at least. Yeah. I've got a lot of
interesting things from you in this conversation, Hadesh. And I know our audience and our listeners
would love to learn more. Where should they go?
So I would say the best general purpose source might still be semi-analysis, although hopefully one day there'll be more competition around that.
After that, probably Google and Microsoft have very good posts and blogs that you can learn from.
Some places have pretty good technical documentation too.
I would say just throwing some examples, they're Corning, Naddot, both have really good blogs and technical documentation.
InVidia's own posts and documentation are pretty decent too.
If you want to hear what I have to say, you can check out my LinkedIn, so just my name, Hittesh Kumar,
or my substack at six rack units.
I try to post topics like these there, too.
I'm going to be a reader, Natasha, and thank you so much for being part of the tech arena.
It's been a real pleasure.
I would love to have you back on the show to probe deeper into some of the topics that you brought up.
And Janice, thank you so much for another great Data Insights episode.
We keep bringing the heat with our guests, so thanks so much for the collaboration.
You're welcome. Thank you. Thank you, Alison. Thank you, Hatesh. Thank you, for joining Tech Arena.
Subscribe and engage at our website, Techorina.aI. All content is copyright by Techarena.
