Podcast Archive - StorageReview.com - Podcast #152: Data Center Thermal Dynamics and Cooling Architectures With Tim Shedd
Episode Date: August 27, 2026Brian recently caught up with Tim Shedd, an expert in thermal management and a pioneer in liquid cooling technologies. The podcast covers some of the historical aspects of data center liquid cooling, ...from Tim’s time working with Cray and spraying liquid across tubes to his more recent engagement with Dell’s PowerEdge XE9680. Tim Shedd The post Podcast #152: Data Center Thermal Dynamics and Cooling Architectures With Tim Shedd appeared first on StorageReview.com.
Transcript
Discussion (0)
Tim Shed, thanks for doing this.
You're one of the most prolific writers on LinkedIn about liquid cooling,
and you've been in this industry for a long time.
Nobody can take a simple concept of move heat from this hot thing to outside this area
and put a 96-page slide deck behind it with scientific formulas
and reasoning that the execs you were presenting it to surely didn't understand,
but thanks for doing the podcast.
Of course, happy to be here.
So, Tim, you and I met when you were with Dell, and you were, like I said, one of the early pioneers in trying to understand thermals in the data center because we pretty quickly went from, everything was fine with evaporative coolers and the traditional air-cooled servers.
But the switch from that to HPC clusters to soon-to-be the high-density GPU servers happened very quickly.
Just take us to the early days of liquid cooling and why you were so early on it and why actually Dell was so early on it.
And why was it important to you?
Yeah.
So real briefly, I'll say that I was introduced to liquid cooling while I was a professor at the University of Wisconsin-Madison.
The University of Wisconsin is about two hours away from Cray up in Chippewa Falls.
And so I was brought to it through Cray.
and literally spraying liquid onto chips.
So that was really cool, but of course, no one else in the world needed to do that,
wanted to do that, very, very expensive, kind of crazy.
But the industry was born about that time, about 2001, with Cool IT and Asa Tech both being formed
and starting to do hobbyist machines.
But the pushing to the data center was super slow.
there were very, very few real applications where it made a ton of sense.
But going through the mid-20pens, there was kind of a leveling off of density of racks and heat.
But then after about 2015, we started seeing that rack density climb.
And Dell actually got into this about 2016 to 18
in doing all the hard work to lay out the foundations
for their launching of a very dense rack,
liquid-cooled rack system in 2019 at TAC in Texas.
And that was using a cool IT system into end,
and that was Dell's approach.
But as a result, they got familiar with trying to deploy
thousands of servers with liquid cooling.
Of course, Kray and HPE were doing this for specialized HPC systems as well,
but the volumes were really, in retrospect, quite low, right?
We're talking maybe a few thousand servers a year type of thing.
And now we're shipping a few thousand servers a week.
So that's really the story.
And when I joined Dell in 2022, coming from Motivocal,
where we're, you know, a place where we're creating the CDUs for a lot of these systems,
both Dell and HPE and others.
The CDUs are Coalent Distribution Unit.
What I saw and what I wanted to see within Dell was, okay, how do you go from thousands of something to millions of something?
Because Dell was really good at that.
And, but it just wasn't really necessary as I was surveyed.
the portfolio and what customers wanted, even for the AI chassis that we were putting out
the XE9680, which turned out to be the fastest growing server in double history.
We have one upstairs in the lab.
Yeah.
It's a great, robust system, and it was air cool, right?
Because it was just at the right balance where you could get good performance with air cooling,
Initially, there was the shock at the sound of the fans.
I'm sure you've experienced that.
Believe me.
I'm well aware of when we turn any of our air-cooled GPU servers on.
They are not quiet.
They're not quiet.
And without special care, they draw a lot of power for the fans.
But one of the things that was really key in the early build-out of the AI systems
where it's looking at that, we got this really robust, known technology.
in the server with air cooling.
And we know how to source all the heat sinks.
We know how to source all the fans.
We can build a lot of these fast, and they're very reliable.
And if we couple that with a rear-door heat exchanger,
we essentially have a liquid-cooled rack.
That actually, on the whole, it might have been a little noisy
compared with a coal plate system,
but it actually used less energy total,
both for powering the compute and for the cooling.
And so I give that example to say that a lot of what we were working on in this transition wasn't just,
and we do this because in certain cases it's more efficient.
But what makes sense for the customer or what makes sense at scale?
Yeah, that's a good point, though, because in the early days, when you were evaluating these systems,
I've been in your labs in Texas.
And there was a time where I think, I don't know, you tell me, there was at least half a dozen, maybe a dozen different solutions in that lab, in that time period where you're evaluating, you know, Cool IT was the lead partner.
They were doing the Cdues.
They were doing the cold plates and manifolds and putting all the hoses together.
So that was great.
But on top of that, there was two phase.
There was negative pressure in there.
I remember seeing that.
There was all sorts of other stuff.
you talk about customer readiness from an ISG perspective,
like even your own readiness had to be a bit of a,
yeah, I'm sure it was fun,
but also a bit of a nightmare too to settle these systems up
and try to sort out what's real and what's not
because that's the benefit that any large vendor can deliver their customers
is customer don't worry about it
because we've already logged thousands of hours or whatever against these systems
and we know what works and what doesn't.
And you can trust us to advise you on that.
whether it's Dell or HPE or Cisco or whoever,
it doesn't make any difference.
But in those labs, in those early days,
when you're looking at all those different things,
take us to that moment.
What did that feel like when you were just getting,
I was going to say feet wet,
but that'd be a horrible plumberlandier,
but maybe you literally did that too.
Yeah, so you don't work in liquid cooling
and not have your share of green fluid
and other things all over you.
The epiplasm everywhere, right?
Yeah.
So I've been somewhat noisy on this topic.
In fact, we published a paper at Smythurm with some colleagues from Dell that goes through five technologies and comparison for efficiency.
And I made some folks kind of upset about that.
But the main point was, I think, going from, say, 2010 to 2022,
the actual heat density at the chips was growing,
and every generation seemed like a pretty intense challenge,
but in the end it was pretty straightforward to solve it,
either with air cooling or with certain forms of liquid cooling.
Dell was actually, to my knowledge,
one of the first to commercially deploy
immersion cooling, I think as early as
2010, if my
memory serves correctly.
So Dell has always
been out there to try to solve
the problem the customer needs solved.
The challenge came that
we, as an
industry in particular,
just weren't getting,
weren't having to push hard enough on the
physics of each cooling
system to really see where it would break.
And so part of, and what would actually scale beyond, as we, if we're going to move beyond
500 watt CPUs, what does that look like?
You know, honestly, 1,000 watt GPUs were kind of out, you know, mentioned in the, in the ether,
but it wasn't something we're focused on right away, you know, in 2022.
But this was something that I was focused on, was saying, as part of it was a lot of, you know,
of the CTO's organization, what do we need for 2028?
And if these trends continue, what's going to break?
And I was noticing, hey, you know, immersion is not really great for really intense heat sources.
It's great if I've got a lot of heat spread out over a lot of area, like for crypto or certain other applications.
It's actually pretty interesting.
Right.
And so that was the point.
That was, you know, right about the time you're referring to is this.
transition where we were saying, okay, let's push on this, let's see where it breaks and what it
came down to is, contrary to what people think, I didn't have a preconception here.
We were ready to go whichever direction.
And PG-25 or, you know, water mixed with propylene glycol, going through coplates was just
the technology that clearly could get us to the future.
And we could scale it.
We knew how to scale it.
we had already the makings of this ecosystem that could build everything we needed.
And all of that together said that's the direction we need to have.
Yeah.
Did you, so what were you concerned about in those days then?
Was it leaks?
Was it cost of the components, maybe even interoperability between a CDU and a manifold
and multiple server vendors?
Maybe that's less common, but that's something that I've heard pop up a little bit.
what were your core concerns in those early days about going to cold plates and PG-25 and a CDU?
What were you worried?
So if you come search, you know, I'm on record, especially in that early transition 22, 23, you know, really calling for standards and the ability to source from multiple vendors and mix them.
match gear
because one approach to this
was to do as you said to be able to
buy coal plates from multiple vendors,
quick connects from multiple vendors
and CDUs from multiple vendors
and be able to put them together.
That's a lot harder
than it seems and
in those standards
it didn't happen.
That's a complex
yeah.
Well no, I mean you had you guys were going
with the cool IT integration
Lenovo was doing their thing with Neptune, which was separate.
Super Micro was doing a little both.
I mean, it was all over the road, right?
Yeah.
And this happened for various reasons, but a big part of the reason,
there's a couple really big issues that arose pretty early.
And one of it is just lawyers, right?
If it's liability, if there is a problem, a leak or contamination of fluid
or anything like that, you need a clear place to point the finger.
And a clear, somebody's got to take care of that.
So the customer is not sitting without getting serviced and taken care of.
And that actually is really hard.
And to be honest, today is still an area that I consider not settled.
We don't really deal with that, which is why another reason why today we still have a hard time mixing equipment.
it's not necessarily technological.
It's a little bit of that,
but it's mostly just liability.
If something goes wrong,
who gets blamed?
That's a really important question.
I know it's annoying,
but it's a really important question
because if you don't know who gets blamed,
you don't know who's going to actually service it.
And if you don't know who's going to service it,
you might as well not build it, right?
Well, we've heard, there's plenty of stories of this.
I remember, so you're talking about the early Dell systems,
I mean, even leak detection,
and I don't know how much this has evolved, really,
I haven't looked at it lately,
but there'd be little ropes inside
and it would flag you if something happened.
And ideally, you'd have some trap set up
to shut the system down and then go service it and whatever.
But as that scales, there's so many more components in the systems that can fail.
And we're still talking about hoses that have been around for decades,
really, from automotive and industrial and other use cases,
that have failure rates.
Like every little connector has a failure rate of some sort,
and it's not like a hard driver at SSD
where you can just swap it out and carry on
because you expect it.
There will be problems.
So how does the industry solve that then
if you don't feel like it's there yet?
So let me just say,
we're talking about this like it's a mature product.
This is something that this is an industry
that's from 2023 till 2026 has grown from shipping maybe thousands of servers a year to millions of servers a year.
It's just incredible.
I want to shout out to the hosing manufacturers have done phenomenal.
We were talking about those things before this took off.
Okay, looking at all the different ways you treat EPDM back in the 2019, 2020 timeframe.
and you know what's best and we settled on weed that they settled on peroxide treating
epidium and in the just they the hoses are i just wanted to point the hoses are actually super
reliable now you're probably not going to leak from a hose um you're probably not going to
leak from a joint on the hose where it's joining the quick connects or the whole or the
co plates because they focus so much on that that they've really engineered the crap out of it and
it's really solid.
But, as you said, even if I'm five-nines, even six-nines reliable,
and I'm shipping millions of servers a year,
I'm going to have failures.
And honestly, it's just the reality.
I have to say that something will happen
that somewhere there will be a leak, sometimes quite a dramatic one.
it is generally not an engineering problem.
It is generally a quality,
not even a quality of the part so much as cleaning of the system.
These are complex piping networks that go through
from the origin of the stainless steel to showing up the data center.
There's a whole lot of hands involved,
and there's a whole lot of processes that have to work right.
And we're still seeing pictures.
It's crazy.
At Ashery, they'll throw up pictures of what comes out of these systems when they first flush them.
You'll see gloves, phones, hammers, everything, right?
Seriously?
That's awful.
And it doesn't take, you know, it takes something about, you know, 100 to 200 microns in diameter,
a little fiber from a wire brush getting caught in a QD, and you've got a dead rack, right?
So it's that sort of thing that's, again, in Astray, TC9.9, we're going to publish another tech bulletin here in the next week or two, actually, focusing on this aspect and everything that needs to happen there.
So all that to say is, I would say the industry has matured, the makers of quick disconnects have been really phenomenal, actually, in their innovations and the hoses and the co-places.
just unbelievable what we do with Colpice today.
I don't know how to describe.
It's hard to describe.
This is why I write thousands of words on this.
Because it's, on the one hand, really hard to imagine that we're moving 1,500 watts
with only a 30-degree C penalty or something like that, right?
It's crazy for those of us in the field anyway.
To think that we're accomplishing that today, whereas just a few years ago,
that was really hard to think about.
Right.
Now, honestly, a lot of that has happened by just a whole bunch of people
throw in CFD at it and iterating, right?
There's not quite as much science as maybe I'd like there to be on that,
but it's working.
We're getting there.
Well, so one of the early guys we worked with,
Childine, was going with a negative pressure approach.
And I know they've been since rolled up,
and we should get into the consolidation in a little bit.
But they had a really innovative approach of saying, cut any line you want because the negative pressure, it's never going to leave.
And they were the first guys that we met that actually had a scientist on the team who was only focused on water.
And it was really interesting.
I had not thought about this before.
So you talk about if you get particulate or anything in the lines that even those cold plates, they've got little turbulators and all sorts of little channels and stuff in there.
And if you start gunking those up, then the efficiency of that cold plate goes down sometimes dramatically to, you know, theoretically zero if it's a serious clog.
But what is it about negative pressure that seems intuitively to be a viable solution here but hasn't really caught on at a mass scale?
So there's several things in that question.
And on that last point, it just has to do with pressure availability.
So if I'm trying to drive a system with the pressure difference between atmosphere and a total vacuum,
by definition, I've got 14.69 PSI or about 100 KPA of pressure to push the fluid through.
And that is not even all usable because the water itself will start boiling at some point.
at room temperature.
So I really only have
somewhere around
7 to 9 PSI to push my
fluid through. And if you look at
how big these systems are, that's just not
enough pressure to push the fluid from
the CDU through the heat exchangers, through the
piping, through the quick connects, through the
hoses, through the coal plate and back again.
And I need
to use
a
fluid that is
widely available and somewhat standardized, right?
Fluid isn't a whole other topic.
We could probably spend literally hours on.
But the, and this ties back to,
I do even want to touch on that fact that, you know,
Steve and his team with at Chiline,
you know, focusing on the chemistry was the right thing to do
probably a little early and probably again,
And they just didn't have the market penetration to be able to really drive that piece into the market.
But it's absolutely critical today.
And I think people are learning that actually, you know, folks, you know, neoclouds are probably having to invest in their own chemists on staff because they are deploying enough that they actually, not just, invidia, not just AMD, not just Dell, but like the actual.
customers need to have that talent because of all the things that could possibly go wrong with
the chemistry and the interaction with the materials.
Yeah, the water thing was something that really opened my eyes to this too, or the fluids that
you're using.
Depending on how you're set up, environmental conditions can be different in different, you know,
in Cincinnati where I am or in, you know, Texas or California, altitude can make a difference.
the just the thermals of your region, humidity.
I mean, there's so many things at play.
Has the industry gotten better at that, do you think?
They're getting better really fast,
but it's still a learning experience.
So one of the things that we see is there's an easy thought that,
well, if my climate is actually not that hot,
I'm good, right?
So I know folks who've deployed, for instance,
in northeast United States,
and climate's not too bad,
but it turns out that every morning in March
is going to be 100% humidity pretty much
if you're near Lake Erie or something, right?
It's just, it just is,
even though the temperature is not that bad.
That's a problem.
And that's a problem that has to be managed
separately from everything
else. Not just the water coal and everything else,
but I can't be condensing it's on my service.
So there's just so much more
that is involved
in deploying these systems
than deploying air-cooled systems.
And
I have to think about
all of these things. Now,
I don't like to over-complicate it because I think
liquid cooling is brilliant, and I think we should
do more of it. But
it also,
we need to make sure we have
people who understand the big picture and who are, you know, trained to deploy these systems
and deploy them well, or you end up risking much larger losses because when they fail,
as you've pointed out already, they can fail catastrophically.
The leak detection, I circle back real quickly.
The leak detection has progressed amazingly in three years, you know, instead of just
leak ropes everywhere.
Leak ropes are actually a huge problem because they ended up being bottlenecked by just a couple
people in the world and make those
the fundamental materials and the
ropes and that's expanded
but they're also not great detectors
are very slow and they take a lot of liquid
to kick them off and so
there's a lot of work I know
various
developments you know I can speak to the Dell
developments but I know others are doing similar
of very very sensitive
basically
traces printed
on polymer so that they're
flexible and you can put them wherever you need them
and they can be extremely sensitive,
very small amounts of liquid.
And there's been work on optical leak detection,
so sensing the dyes in the coolant,
even essentially sniffers that smell the components of the coolant
and help to find small leaks before they can get big
and try to treat them.
And then there's innovations, for instance, again, you know,
if I sense a small leak,
well, maybe I can adjust the pressure of my CDUs,
so I minimize the chance of leaking.
I may leak, but at much lower values.
There's a lot that's going on.
Dell deployed, what they call, an integrated rec controller
to start managing some of this stuff actively.
I know other folks are doing this as well.
And I think all of this is important.
it's adding to complexity of the system, which is not great, right?
And that has to cost.
But because the risk is pretty high, it's justified.
And it's just one of the, I think one of the prices we have to pay to deploy liquid cooling at scale.
One rack where you can babysit it like we were doing in 2020 or, you know, a relatively few racks.
That's okay.
You know, talk to the national labs.
Lots of experience babysitting racks.
They're good at it, but they had people who got good at it because they had to do it.
We just can't do that when we're deploying thousands and thousands of wrecks a year.
No, and it's a big ask for the customer, but they're going to have to figure it out.
All the neoclouds, as you point out.
But, I mean, this is one of the things, too, where the enterprise is, you know, being forced with this decision as well.
and trying to figure out how much more they can get away with air versus having to overhaul existing data centers, if they can, to get the power and the floor space and the plumbing required to run these liquid loops.
I do want to talk more about air, but before we get there, since we were talking about negative pressure, you talked a little bit about full immersion, which we've looked at too, and there's some interesting possibilities there at a smaller,
scale. But what do you, I know you're, you're close to two phase also. And that's one that's
gone through cycles of being really exciting, then being scary with the forever chemicals, the
P-FAS, but it seems like that hasn't gone away. And there might still be some, some life
in two-phase. What's, what's the latest there that you know about? So there's a, yes, there's,
there's a lot that is being discussed around two-phase.
It hasn't gone away.
There's been intrepid work by Excelsius and Zudor in particular,
keeping it in the limelight.
But also, there's a lot of organizations that have at least bench-stop experiments going on in two-phase.
And the reason is it's one of the biggest parts of the thermal resistance at the chip and the cold plate is this.
caloric resistance or the fact that when I heat a single phase fluid, like air, like liquid,
like water, when I add heat to it, its temperature goes up. That's a pretty big penalty.
And generally, it's at least a 5 degree C penalty. So that's not insignificant these days when we're
trying to keep our water temperatures really high. In theory, two phase doesn't happen. It boils at the same
temperature, whether I'm boiling 10 watts or 1,000 watts.
So that's one key thing.
And it just looks cool and you imagine that anything is bubbling like that.
It's got to be great.
It turns out that there may be some advantages to two phase at the chip.
On the condensation side, it's actually a lot harder to manage.
So there's still some challenges in engineering this.
The key thing ties together some of the points of the discussion we've had so far.
And the key thing is if we think about this issue with water, water quality, leaking, all this stuff, they don't go away, but they are dramatically minimized, right?
So I actually believe, think, you know, believes a strong word, but I think there's a lot of potential in the enterprise space.
If they can get some water, they may have water on the floor for, that's a bad word, they may have piping that exists for, say, rear door heat exchangers.
But they don't want to take the risk with water in their rack.
Well, two-phase might be a great answer to them.
Even if, you know, it's not necessarily performing better than a water-based DLC system,
it takes that risk out, right?
And the same thing could be true for the neoclodes.
It's a risk mitigator if we can design systems that can actually deploy a scale.
And that's been the challenge, right?
Right now, to deploy a scale,
you know, we're going to be, if we're not already,
producing nearly a million coal plates a month for AI-based servers.
That's a huge scale.
And so right now, a lot of these two-face systems are end-to-end design
such that the coal plate is integral to the manifold
is integral with the CDU,
and they all have to work together because you're balancing
all the instabilities and the two-face flow and all this stuff.
That can't work at scale.
You've got to be able to take a co-plate from some vendor in Taiwan
and plug it into a manifold that was made in Canada,
and that's got to work with the CDU made in Germany, right?
We've got to really figure that out, or it's just not going to scale.
And so that's where the work's happening, but it is happening.
And so I don't see it ever displacing air or water completely,
but I do see that there's probably an opportunity because of the challenge,
of dealing with water and water chemistry and the risk of water, I do see that there's a niche
and potentially the chance for some companies to do pretty well in that space.
Are we still concerned as an industry about the chemicals used in two phase?
So I won't say no, but there have been developments in the refrigeration space of
compromise between global warming potential and minimizing the impacts of PFAS.
to the point where regulation-wise anyway,
it's believed that they can get the concentration down
to where they're not damaging.
I am not a specialist in that area,
and so I haven't formed a personal judgment on that.
But what we're seeing is the regulatory blocks are being lifted
with the new refrigerants, like R515B, R1,25B,
R-2-3-3-ZD, things like this that are a balance of good global warming potential, low-epass content, and zero flymobility, which is another key thing.
So it's lots of great to refrigerants, but they burn, and that's not good in the data center.
No, not great.
That's why they took fuel stops out of F-1 because it can be problematic.
Talk about the CDUs a little bit then, because that's another area that I think is confusing and continues to change as the historical data center guys have now invested heavily via acquisitions to be able to supply their customers with these end-to-end solutions.
I think I still kind of get wrapped up on where are we going with CDUs?
How do we get to an HA-style scenario where if we have a CDU fail,
we're not losing racks and racks and racks of gear?
And again, we keep hitting on it, but the compatibility issue,
is it going to be possible to have multiple different CDUs on your floor?
And how are you thinking about that piece of it?
Yeah, this is super important.
And it seemed like forever ago.
It wasn't that long ago.
2023, I think a small group of us in Ashera said, you know,
it's kind of a nightmare if you're at Dell or someplace else
in trying to buy a CDU because the specifications are all over the place.
And whatever the vendor says is probably not what we're going to see in the field.
So one of the key things that I will push is Ashtray Standard 127.
We've published a method of test.
that standardizes the way that you measure the efficiency of a CDU and how it operates.
And that's super important because that just didn't exist before.
There was no apples-to-apples comparison of CDUs.
Now there is.
So that's the first step and an important step in allowing customers to be able to understand,
can I take this vertive CDU and put it out on the floor with this Motivir CDU and they'll work together?
Generally speaking, that is becoming much more reality.
So in one sense, I think there's been a lot of progress there.
And credit to the manufacturers as well,
have been very much part of this process and want to see this happen.
On the other hand, there's
so there's been a lot of progress as well in just engineering CDUs.
And CDUs are thought of as a pump and a heat exchanger.
They've got to do a lot more than that.
one of the things that people don't realize is every small variation in flow is instantaneously
uh sorry instantaneously appears as a temperature fluctuation at the cold plate and so your cdU has to manage that
really well um and to be honest people just didn't pay a lot of attention to it because the the risk was
low now the risk is high um if i have a 10% change in my um flow
that can actually cause throttling events, right?
So the CDU is harder to design than people think,
and some folks have gotten really good at it.
Heat exchanger technology has come a long ways in what we're deploying,
as well as control algorithms and control vowels even
are much better than they were even just three or four years ago.
So all of this is improving.
I think CDU technology is improving.
one of the things that I
personally want the caution about
this is an opinion that may not be popular
No, that's why you're here
We're here for your popular opinion.
But we've been talking about
the issues with water and piping
and all this stuff
and there's been this move to go to bigger and bigger
CDUs, you know, having 10 megawatts
CDUs or whatever. And I understand that
from the facility point of view, from the
Neocloud point of view.
but it means I've got this big blast radius
and in theory I can put a bunch of CDUs in parallel
and I've got good uptime
because they're all supporting each other
but there are going to be times when
one of those CDUs or one server
even contaminates that whole loop
and what do I do? If I have a leak in one of the main feeds
everything comes down right
my volume of liquid
itself goes up non-linearly
when I, you know, change the size of my pipe.
So I've got these big pipes feeding all these racks.
I've got a ton of coolant in there.
And I can't mix coolants.
Very few coolants from different makers can be mixed because there's proprietary ingredients,
just a small percentage, but we have found that even that small percentage can potentially cause reactions that may be bad.
We don't know for sure.
There's still more work to be done, but right now it generally is recommended you don't mix
coolant. So what that means is, okay,
you're locked into a coolant,
and if your supplier can't
get you a truckload in time, then you're down
too, right? So if I have these big systems,
I now have
this risk of
going down and taking a lot
of compute with it. Whereas if we can
keep the system smaller
at a row level or even a rack level,
then
I now
control and contain
my, you know, blast
rate is that contain what's impacted. And it also has a huge benefit in time to deploy.
I'm not waiting for all these big piping systems to get built and cleaned. I can, you know,
in the case of an in-racks you, I'm rolling a rack onto the floor and that same day running it.
So I'm glad you brought that up because that's something that I know you're not at Dell now,
but at Dell Tech World in May, they were showing the evolution of their,
in rack CDU, basically a little mini-CDU just for that rack.
We saw it at Supercompute and then the evolution in May.
And that's sort of counter to where the industry has been going with these larger and larger,
you know, sub-zero-sized refrigerator, you know, kind of massive units that would power
or cool multiple racks.
Do you see a fragmentation in where we're going here?
And maybe there's not a right answer.
and it just depends on the customer.
But the in-RAC CDUs seem like they make logical sense.
So we'll see where that goes.
That was motivated by customer demand.
We had customers for this, in particular for the serviceability, deployability, and time to deploy.
They wanted in RAC CDUs.
And they wanted in RAC CDUs that could cool Vera Ruman with no compromise.
Right?
Those CDUs didn't exist.
So we made one.
And that can do 220 kilowatts with 4 degrees C approach.
In other words, if you give me 41 degree C facility water,
I can generate 45 degrees C for Vera Rubin at 1.5 liters per kilowatt, the full flow rate.
And so great work by the engineering teams.
That was actually developed internally to Dell.
It's a pretty unusual project.
But just amazing engineering work with them and our partners that made that happen.
And that is exactly the sort of thing that I think, personally, that accelerates deployment.
It allows folks to deploy very quickly.
It's got a lot of redundancy.
It is not easily serviceable.
I mean, yeah, I mean, you have to pull it out if you're going to change a pump, for instance, or whatever.
But it's got two pumps that one pump can drive all the cooling.
So the chances of you going down at a critical time are very small.
And if you do, you're just...
on one rack, not your entire.
That's right.
Yeah.
And so I, it is my belief that there will be a significant portion of the market that
will see that advantage and go that way.
There will be some like hyperscalers and certain neoclods that will still prefer the really
large CDUs and that infrastructure investment and they have the staff to make it work.
And they're building on long timeframes and they don't care, you know, but those folks who
want to order a Vera Rubin and get it.
it up and running, like within a month or two of the announcement, I think that they're going
to want that in RAC CDU. Now, right now, DEL is the only one that provides that class of CDU.
So we'll see if anyone else can get there.
But it is my belief that that is a wise way to go with the CDUs just from, again, a service, time
to deployability standpoint.
So I'm a little biased there.
And it's just minimizing, well, just minimizing all that effort of keeping all those pipes clean and so on.
If you can read that Astro-T-Bolton is coming out.
And you need to do everything that's in there.
And it's a lot of work to put that big piping network together.
And if I can eliminate the big piping network and just have it all contain the rack,
it just seems to me to be a big plus.
So that's a good segue then.
I said I wanted to talk about air still.
And let's get back there because for all of the buzzy stuff around agentic AI and enterprise inference and tokenomics and all these other things,
you know, we're seeing these air-cooled servers with four or eight RTX Pro 6,000s doing a lot of work in a lot of businesses, education institutions, and other places where it's not a training task.
It's something else. And these systems are great for that.
they're almost all air-cooled.
And the way vendors get there is a little bit different.
Some are a little more dense.
Some are taller, eight or ten-U enclosures for the servers.
But do we have runway?
Do you think on air-cooled for, and if we do,
obviously we have some, but how long do you think we can continue to get away
with air-cooled servers with the GPUs inside?
Really good question.
So again, I'll refer back to another aspect publication.
Just because, I mean, that's our role as DC-Denpling is to get that information out there and try to educate.
And one of the things that I really want to make clear is somewhere around 90% of all server shipments today are air-cooled.
And a significant portion, actually, of AI server shipments are air-cooled.
Not the majority, but a significant number.
And the reason is that you just look in the news.
how many new 500 megawatt data center sites are going to be deployed this year?
I mean, a lot, more than you would imagine, but not the capacity that's needed.
Also, inferencing occurs, just to borrow from a former employer,
inferencing occurs where the data is generated.
You don't want to send the data to some central site.
So we want to put inferencing compute out in the field at 15.
year old data centers that have air cooling, that have some capacity.
And so I personally anticipate that that is only going to expand.
Now, everything in this market is going to expand.
So, you know, maybe as a share of the market, it won't seem like a big clip.
But if we look at, I think, capacity that's going to be deployed in, like you said,
university data centers and bank data centers and even bank branch
data centers or rooms, right?
We're going to see these servers appear
that have to be air-cooled.
Now, one of the challenges is a lot
of those data centers are
kind of power locked, right?
So they are not just landlocked,
but they've
designed for, I don't know,
500 kilowatts.
And that's all they've got.
And so one of the things that I really
have as a mission
is to try to help
people see, okay,
if you've got an older 10-year-old data center that's 500 kilowatts,
you're probably spending a significant portion.
Somewhere around 200 kilowatts on cooling on average over the year,
that's power you really want to send to your inference box.
So what can we do?
So that might be a relatively low lift of,
let's go to Reardoner Heat Exchange for some of our computer.
that can increase my efficiency and save me, you know, tens of kilowatts if not hundreds of kilowatts.
Things like this that I think are going to be really important to shift cooling kilowatts to compute kilowatts
so that we can deploy these 10, 15 kilowatt inference boxes in these remote locations.
So I do think there's going to be some modification needed in order to make power space, you know,
to make space in the power budget for these systems.
Or just, you know, refreshing the kind of Excel servers so that where I needed, you know,
20 in a rack before now I only need five new servers to serve up Excel and I can use that power,
even though each box is more powerful, the total racks, lower power maybe, and use that for interesting.
So I think we're going to see a lot of very careful planning, probably some upgrade.
you know, upgrading a data center to use a cooling tower,
which I realize some people, there's a whole other discussion
and the water you stare, but it's a great cooling device,
very low energy, very inexpensive.
You can upgrade those much cheaper than people think
and be talking about how do we shift those cooling watts
to compute watts and enable the air-cooled,
inferenceing and compute to continue to expand
because we're absolutely we're going to see that happen.
Well, you started by talking about the gaming rigs
and using the all-in-one cooler internal loops.
We've seen that deployed several times on servers before.
Do you think there's a hybrid shot there that makes sense for these systems?
So I'm not a huge fan of those.
They had a lot of complexity.
They take space.
And they tend to only give a marginal improvement over a high-performance air-cooled heat sink.
Now, it's not zero improvement.
I mean, there's been great engineering by folks, you know, UNOVOHP and others have engineered great systems that do provide an advantage over even the best vapor chamber-based air-cold heating.
it's just that the complexity that you introduce for that advantage
never seemed to quite add up to me unless there were some customers
that absolutely needed to push right to that limit and you know I see that
I just don't personally feel like that's a mass market that there's demand
for large volumes of that type of technology when again the best
benefits relatively small.
What you could, you know, if you're going to do that,
I think the right way to do it is just invest in a liquid cool rack
and then have a liquid air CDU, a sidecar next to it.
Much more efficient, quieter.
And it's just a much better way to get that sort of thermal advantage
without all the complexity inside the box and the penalties on service
and everything else.
Let me ask one last really naive question.
For air-cooled servers, is there anything more to get out of the fans?
And I know it's a double-edged sword because the fans are taking up a ginathear, enormous amount
of the power budget in air-cooled servers.
There's many studies on this about the increasing share of the fans.
But can they physically get better, more efficient at moving air?
at moving air?
That's a great question.
And I want to put out there,
our colleagues in the fan engineering space,
and again, shout out to predecessors at Dell,
who really, especially in the mid-2010s,
really pushed fan efficiency.
Back in 2015, 2010, a fan might be 15 to 20% efficient,
so it's turning about 15% to 20% efficient.
So it's turning about 15 to 20% efficiency.
20% of the electricity coming in into air motion.
Today, it is normal for them to be 50% to 60% efficient.
That is actually ridiculous in efficiency.
That is a ton of engineering that's going on across many organizations.
And that's done at scale.
These are mass-produced devices.
That is just mind-bogglingly efficient.
So to answer your question, every generation,
I can recall seeing improved fan profiles.
They are getting better.
I just feel like there's not a whole lot of juice left in that lemon,
but they're still going to squeeze somewhere out.
That said, fans, they're noisy as I'll get out,
but their actual power consumption for what they're doing has gotten a lot better.
And so I would say that, you know, fans are not the evil.
that they sound like.
They're actually pretty energy efficient.
What needs to happen, coupled with the fans,
is really good engineering inside the chassis.
And this is something that I think still can be improved.
Like, I would push, in my role at Dell,
I would push on, like,
you want to use every bit of air to pick up, you know,
that pick your number
but a good number is
100 CFM per kilowatt is a pretty
good design number. It's not a
requirement by any stretch. But
if I'm doing that, that means I need every
cubic inch of air going through
to heat up by 18 degrees. And if it's
not heating up by 18 degrees, or
if some of that air is heating up by
18 degrees Celsius,
if some other chunk of air
is heating up by 25 degrees Celsius, that means
I've done something wrong. I've not used that
efficiently. If I can use all the air
going through that box efficiently,
just like I want to use every drop of water
falling into a cold plate efficiently,
I can do a lot with those fans.
So I think there's just a lot to be done
still on the engineering of the
chassis as the
fans get incrementally better.
But I'm optimistic.
Air cooling is actually
a great, you transfer fluid. People don't think about it that way.
It's been really maligned.
But actually from a, there are certain
characteristics of air, like it's
it's thermal diffusivity, things like this that are actually better than water.
And so I can actually be very effective at cooling with air, but I need to be really smart about it.
You can't just throw fans at it like we used to and say we're done.
Well, it's a good point, and we're already seeing a fragmentation from the large server vendors
for these inference servers in terms of their design.
So this will be a fun one to watch to see how they continue to develop the airflow and
optimizations and those boxes don't all look the same, whereas obviously on the big racks,
NVL, that's pretty prescriptive. So there's a little more room to play on the engineering side
with the more traditional servers. Tim, I mean, this has been amazing. Like I said at the intro,
you're one of the great thought leaders in this space. And for anyone that wants to keep up with
everything around liquid cooling and efficient data center, you need to be following Tim on LinkedIn.
a link to his profile in the description.
But Tim, thanks for sitting down and doing this.
It's great to see you again.
Really appreciate your time.
Thank you very much.
I've really enjoyed it.
