Podcast Archive - StorageReview.com - Podcast #152: Data Center Thermal Dynamics and Cooling Architectures With Tim Shedd

Episode Date: August 27, 2026

Brian recently caught up with Tim Shedd, an expert in thermal management and a pioneer in liquid cooling technologies. The podcast covers some of the historical aspects of data center liquid cooling, ...from Tim’s time working with Cray and spraying liquid across tubes to his more recent engagement with Dell’s PowerEdge XE9680.  Tim Shedd The post Podcast #152: Data Center Thermal Dynamics and Cooling Architectures With Tim Shedd appeared first on StorageReview.com.

Transcript
Discussion (0)
Starting point is 00:00:03 Tim Shed, thanks for doing this. You're one of the most prolific writers on LinkedIn about liquid cooling, and you've been in this industry for a long time. Nobody can take a simple concept of move heat from this hot thing to outside this area and put a 96-page slide deck behind it with scientific formulas and reasoning that the execs you were presenting it to surely didn't understand, but thanks for doing the podcast. Of course, happy to be here.
Starting point is 00:00:32 So, Tim, you and I met when you were with Dell, and you were, like I said, one of the early pioneers in trying to understand thermals in the data center because we pretty quickly went from, everything was fine with evaporative coolers and the traditional air-cooled servers. But the switch from that to HPC clusters to soon-to-be the high-density GPU servers happened very quickly. Just take us to the early days of liquid cooling and why you were so early on it and why actually Dell was so early on it. And why was it important to you? Yeah. So real briefly, I'll say that I was introduced to liquid cooling while I was a professor at the University of Wisconsin-Madison. The University of Wisconsin is about two hours away from Cray up in Chippewa Falls. And so I was brought to it through Cray.
Starting point is 00:01:32 and literally spraying liquid onto chips. So that was really cool, but of course, no one else in the world needed to do that, wanted to do that, very, very expensive, kind of crazy. But the industry was born about that time, about 2001, with Cool IT and Asa Tech both being formed and starting to do hobbyist machines. But the pushing to the data center was super slow. there were very, very few real applications where it made a ton of sense. But going through the mid-20pens, there was kind of a leveling off of density of racks and heat.
Starting point is 00:02:17 But then after about 2015, we started seeing that rack density climb. And Dell actually got into this about 2016 to 18 in doing all the hard work to lay out the foundations for their launching of a very dense rack, liquid-cooled rack system in 2019 at TAC in Texas. And that was using a cool IT system into end, and that was Dell's approach. But as a result, they got familiar with trying to deploy
Starting point is 00:02:51 thousands of servers with liquid cooling. Of course, Kray and HPE were doing this for specialized HPC systems as well, but the volumes were really, in retrospect, quite low, right? We're talking maybe a few thousand servers a year type of thing. And now we're shipping a few thousand servers a week. So that's really the story. And when I joined Dell in 2022, coming from Motivocal, where we're, you know, a place where we're creating the CDUs for a lot of these systems,
Starting point is 00:03:26 both Dell and HPE and others. The CDUs are Coalent Distribution Unit. What I saw and what I wanted to see within Dell was, okay, how do you go from thousands of something to millions of something? Because Dell was really good at that. And, but it just wasn't really necessary as I was surveyed. the portfolio and what customers wanted, even for the AI chassis that we were putting out the XE9680, which turned out to be the fastest growing server in double history. We have one upstairs in the lab.
Starting point is 00:04:04 Yeah. It's a great, robust system, and it was air cool, right? Because it was just at the right balance where you could get good performance with air cooling, Initially, there was the shock at the sound of the fans. I'm sure you've experienced that. Believe me. I'm well aware of when we turn any of our air-cooled GPU servers on. They are not quiet.
Starting point is 00:04:33 They're not quiet. And without special care, they draw a lot of power for the fans. But one of the things that was really key in the early build-out of the AI systems where it's looking at that, we got this really robust, known technology. in the server with air cooling. And we know how to source all the heat sinks. We know how to source all the fans. We can build a lot of these fast, and they're very reliable.
Starting point is 00:05:02 And if we couple that with a rear-door heat exchanger, we essentially have a liquid-cooled rack. That actually, on the whole, it might have been a little noisy compared with a coal plate system, but it actually used less energy total, both for powering the compute and for the cooling. And so I give that example to say that a lot of what we were working on in this transition wasn't just, and we do this because in certain cases it's more efficient.
Starting point is 00:05:34 But what makes sense for the customer or what makes sense at scale? Yeah, that's a good point, though, because in the early days, when you were evaluating these systems, I've been in your labs in Texas. And there was a time where I think, I don't know, you tell me, there was at least half a dozen, maybe a dozen different solutions in that lab, in that time period where you're evaluating, you know, Cool IT was the lead partner. They were doing the Cdues. They were doing the cold plates and manifolds and putting all the hoses together. So that was great. But on top of that, there was two phase.
Starting point is 00:06:11 There was negative pressure in there. I remember seeing that. There was all sorts of other stuff. you talk about customer readiness from an ISG perspective, like even your own readiness had to be a bit of a, yeah, I'm sure it was fun, but also a bit of a nightmare too to settle these systems up and try to sort out what's real and what's not
Starting point is 00:06:31 because that's the benefit that any large vendor can deliver their customers is customer don't worry about it because we've already logged thousands of hours or whatever against these systems and we know what works and what doesn't. And you can trust us to advise you on that. whether it's Dell or HPE or Cisco or whoever, it doesn't make any difference. But in those labs, in those early days,
Starting point is 00:06:53 when you're looking at all those different things, take us to that moment. What did that feel like when you were just getting, I was going to say feet wet, but that'd be a horrible plumberlandier, but maybe you literally did that too. Yeah, so you don't work in liquid cooling and not have your share of green fluid
Starting point is 00:07:15 and other things all over you. The epiplasm everywhere, right? Yeah. So I've been somewhat noisy on this topic. In fact, we published a paper at Smythurm with some colleagues from Dell that goes through five technologies and comparison for efficiency. And I made some folks kind of upset about that. But the main point was, I think, going from, say, 2010 to 2022, the actual heat density at the chips was growing,
Starting point is 00:08:00 and every generation seemed like a pretty intense challenge, but in the end it was pretty straightforward to solve it, either with air cooling or with certain forms of liquid cooling. Dell was actually, to my knowledge, one of the first to commercially deploy immersion cooling, I think as early as 2010, if my memory serves correctly.
Starting point is 00:08:19 So Dell has always been out there to try to solve the problem the customer needs solved. The challenge came that we, as an industry in particular, just weren't getting, weren't having to push hard enough on the
Starting point is 00:08:37 physics of each cooling system to really see where it would break. And so part of, and what would actually scale beyond, as we, if we're going to move beyond 500 watt CPUs, what does that look like? You know, honestly, 1,000 watt GPUs were kind of out, you know, mentioned in the, in the ether, but it wasn't something we're focused on right away, you know, in 2022. But this was something that I was focused on, was saying, as part of it was a lot of, you know, of the CTO's organization, what do we need for 2028?
Starting point is 00:09:14 And if these trends continue, what's going to break? And I was noticing, hey, you know, immersion is not really great for really intense heat sources. It's great if I've got a lot of heat spread out over a lot of area, like for crypto or certain other applications. It's actually pretty interesting. Right. And so that was the point. That was, you know, right about the time you're referring to is this. transition where we were saying, okay, let's push on this, let's see where it breaks and what it
Starting point is 00:09:44 came down to is, contrary to what people think, I didn't have a preconception here. We were ready to go whichever direction. And PG-25 or, you know, water mixed with propylene glycol, going through coplates was just the technology that clearly could get us to the future. And we could scale it. We knew how to scale it. we had already the makings of this ecosystem that could build everything we needed. And all of that together said that's the direction we need to have.
Starting point is 00:10:14 Yeah. Did you, so what were you concerned about in those days then? Was it leaks? Was it cost of the components, maybe even interoperability between a CDU and a manifold and multiple server vendors? Maybe that's less common, but that's something that I've heard pop up a little bit. what were your core concerns in those early days about going to cold plates and PG-25 and a CDU? What were you worried?
Starting point is 00:10:43 So if you come search, you know, I'm on record, especially in that early transition 22, 23, you know, really calling for standards and the ability to source from multiple vendors and mix them. match gear because one approach to this was to do as you said to be able to buy coal plates from multiple vendors, quick connects from multiple vendors and CDUs from multiple vendors and be able to put them together.
Starting point is 00:11:19 That's a lot harder than it seems and in those standards it didn't happen. That's a complex yeah. Well no, I mean you had you guys were going with the cool IT integration
Starting point is 00:11:33 Lenovo was doing their thing with Neptune, which was separate. Super Micro was doing a little both. I mean, it was all over the road, right? Yeah. And this happened for various reasons, but a big part of the reason, there's a couple really big issues that arose pretty early. And one of it is just lawyers, right? If it's liability, if there is a problem, a leak or contamination of fluid
Starting point is 00:12:02 or anything like that, you need a clear place to point the finger. And a clear, somebody's got to take care of that. So the customer is not sitting without getting serviced and taken care of. And that actually is really hard. And to be honest, today is still an area that I consider not settled. We don't really deal with that, which is why another reason why today we still have a hard time mixing equipment. it's not necessarily technological. It's a little bit of that,
Starting point is 00:12:37 but it's mostly just liability. If something goes wrong, who gets blamed? That's a really important question. I know it's annoying, but it's a really important question because if you don't know who gets blamed, you don't know who's going to actually service it.
Starting point is 00:12:51 And if you don't know who's going to service it, you might as well not build it, right? Well, we've heard, there's plenty of stories of this. I remember, so you're talking about the early Dell systems, I mean, even leak detection, and I don't know how much this has evolved, really, I haven't looked at it lately, but there'd be little ropes inside
Starting point is 00:13:07 and it would flag you if something happened. And ideally, you'd have some trap set up to shut the system down and then go service it and whatever. But as that scales, there's so many more components in the systems that can fail. And we're still talking about hoses that have been around for decades, really, from automotive and industrial and other use cases, that have failure rates. Like every little connector has a failure rate of some sort,
Starting point is 00:13:37 and it's not like a hard driver at SSD where you can just swap it out and carry on because you expect it. There will be problems. So how does the industry solve that then if you don't feel like it's there yet? So let me just say, we're talking about this like it's a mature product.
Starting point is 00:13:59 This is something that this is an industry that's from 2023 till 2026 has grown from shipping maybe thousands of servers a year to millions of servers a year. It's just incredible. I want to shout out to the hosing manufacturers have done phenomenal. We were talking about those things before this took off. Okay, looking at all the different ways you treat EPDM back in the 2019, 2020 timeframe. and you know what's best and we settled on weed that they settled on peroxide treating epidium and in the just they the hoses are i just wanted to point the hoses are actually super
Starting point is 00:14:43 reliable now you're probably not going to leak from a hose um you're probably not going to leak from a joint on the hose where it's joining the quick connects or the whole or the co plates because they focus so much on that that they've really engineered the crap out of it and it's really solid. But, as you said, even if I'm five-nines, even six-nines reliable, and I'm shipping millions of servers a year, I'm going to have failures. And honestly, it's just the reality.
Starting point is 00:15:16 I have to say that something will happen that somewhere there will be a leak, sometimes quite a dramatic one. it is generally not an engineering problem. It is generally a quality, not even a quality of the part so much as cleaning of the system. These are complex piping networks that go through from the origin of the stainless steel to showing up the data center. There's a whole lot of hands involved,
Starting point is 00:15:51 and there's a whole lot of processes that have to work right. And we're still seeing pictures. It's crazy. At Ashery, they'll throw up pictures of what comes out of these systems when they first flush them. You'll see gloves, phones, hammers, everything, right? Seriously? That's awful. And it doesn't take, you know, it takes something about, you know, 100 to 200 microns in diameter,
Starting point is 00:16:13 a little fiber from a wire brush getting caught in a QD, and you've got a dead rack, right? So it's that sort of thing that's, again, in Astray, TC9.9, we're going to publish another tech bulletin here in the next week or two, actually, focusing on this aspect and everything that needs to happen there. So all that to say is, I would say the industry has matured, the makers of quick disconnects have been really phenomenal, actually, in their innovations and the hoses and the co-places. just unbelievable what we do with Colpice today. I don't know how to describe. It's hard to describe. This is why I write thousands of words on this. Because it's, on the one hand, really hard to imagine that we're moving 1,500 watts
Starting point is 00:17:04 with only a 30-degree C penalty or something like that, right? It's crazy for those of us in the field anyway. To think that we're accomplishing that today, whereas just a few years ago, that was really hard to think about. Right. Now, honestly, a lot of that has happened by just a whole bunch of people throw in CFD at it and iterating, right? There's not quite as much science as maybe I'd like there to be on that,
Starting point is 00:17:28 but it's working. We're getting there. Well, so one of the early guys we worked with, Childine, was going with a negative pressure approach. And I know they've been since rolled up, and we should get into the consolidation in a little bit. But they had a really innovative approach of saying, cut any line you want because the negative pressure, it's never going to leave. And they were the first guys that we met that actually had a scientist on the team who was only focused on water.
Starting point is 00:18:00 And it was really interesting. I had not thought about this before. So you talk about if you get particulate or anything in the lines that even those cold plates, they've got little turbulators and all sorts of little channels and stuff in there. And if you start gunking those up, then the efficiency of that cold plate goes down sometimes dramatically to, you know, theoretically zero if it's a serious clog. But what is it about negative pressure that seems intuitively to be a viable solution here but hasn't really caught on at a mass scale? So there's several things in that question. And on that last point, it just has to do with pressure availability. So if I'm trying to drive a system with the pressure difference between atmosphere and a total vacuum,
Starting point is 00:18:58 by definition, I've got 14.69 PSI or about 100 KPA of pressure to push the fluid through. And that is not even all usable because the water itself will start boiling at some point. at room temperature. So I really only have somewhere around 7 to 9 PSI to push my fluid through. And if you look at how big these systems are, that's just not
Starting point is 00:19:20 enough pressure to push the fluid from the CDU through the heat exchangers, through the piping, through the quick connects, through the hoses, through the coal plate and back again. And I need to use a fluid that is
Starting point is 00:19:37 widely available and somewhat standardized, right? Fluid isn't a whole other topic. We could probably spend literally hours on. But the, and this ties back to, I do even want to touch on that fact that, you know, Steve and his team with at Chiline, you know, focusing on the chemistry was the right thing to do probably a little early and probably again,
Starting point is 00:20:06 And they just didn't have the market penetration to be able to really drive that piece into the market. But it's absolutely critical today. And I think people are learning that actually, you know, folks, you know, neoclouds are probably having to invest in their own chemists on staff because they are deploying enough that they actually, not just, invidia, not just AMD, not just Dell, but like the actual. customers need to have that talent because of all the things that could possibly go wrong with the chemistry and the interaction with the materials. Yeah, the water thing was something that really opened my eyes to this too, or the fluids that you're using. Depending on how you're set up, environmental conditions can be different in different, you know,
Starting point is 00:20:57 in Cincinnati where I am or in, you know, Texas or California, altitude can make a difference. the just the thermals of your region, humidity. I mean, there's so many things at play. Has the industry gotten better at that, do you think? They're getting better really fast, but it's still a learning experience. So one of the things that we see is there's an easy thought that, well, if my climate is actually not that hot,
Starting point is 00:21:33 I'm good, right? So I know folks who've deployed, for instance, in northeast United States, and climate's not too bad, but it turns out that every morning in March is going to be 100% humidity pretty much if you're near Lake Erie or something, right? It's just, it just is,
Starting point is 00:21:54 even though the temperature is not that bad. That's a problem. And that's a problem that has to be managed separately from everything else. Not just the water coal and everything else, but I can't be condensing it's on my service. So there's just so much more that is involved
Starting point is 00:22:11 in deploying these systems than deploying air-cooled systems. And I have to think about all of these things. Now, I don't like to over-complicate it because I think liquid cooling is brilliant, and I think we should do more of it. But
Starting point is 00:22:29 it also, we need to make sure we have people who understand the big picture and who are, you know, trained to deploy these systems and deploy them well, or you end up risking much larger losses because when they fail, as you've pointed out already, they can fail catastrophically. The leak detection, I circle back real quickly. The leak detection has progressed amazingly in three years, you know, instead of just leak ropes everywhere.
Starting point is 00:22:59 Leak ropes are actually a huge problem because they ended up being bottlenecked by just a couple people in the world and make those the fundamental materials and the ropes and that's expanded but they're also not great detectors are very slow and they take a lot of liquid to kick them off and so there's a lot of work I know
Starting point is 00:23:17 various developments you know I can speak to the Dell developments but I know others are doing similar of very very sensitive basically traces printed on polymer so that they're flexible and you can put them wherever you need them
Starting point is 00:23:33 and they can be extremely sensitive, very small amounts of liquid. And there's been work on optical leak detection, so sensing the dyes in the coolant, even essentially sniffers that smell the components of the coolant and help to find small leaks before they can get big and try to treat them. And then there's innovations, for instance, again, you know,
Starting point is 00:24:02 if I sense a small leak, well, maybe I can adjust the pressure of my CDUs, so I minimize the chance of leaking. I may leak, but at much lower values. There's a lot that's going on. Dell deployed, what they call, an integrated rec controller to start managing some of this stuff actively. I know other folks are doing this as well.
Starting point is 00:24:26 And I think all of this is important. it's adding to complexity of the system, which is not great, right? And that has to cost. But because the risk is pretty high, it's justified. And it's just one of the, I think one of the prices we have to pay to deploy liquid cooling at scale. One rack where you can babysit it like we were doing in 2020 or, you know, a relatively few racks. That's okay. You know, talk to the national labs.
Starting point is 00:25:00 Lots of experience babysitting racks. They're good at it, but they had people who got good at it because they had to do it. We just can't do that when we're deploying thousands and thousands of wrecks a year. No, and it's a big ask for the customer, but they're going to have to figure it out. All the neoclouds, as you point out. But, I mean, this is one of the things, too, where the enterprise is, you know, being forced with this decision as well. and trying to figure out how much more they can get away with air versus having to overhaul existing data centers, if they can, to get the power and the floor space and the plumbing required to run these liquid loops. I do want to talk more about air, but before we get there, since we were talking about negative pressure, you talked a little bit about full immersion, which we've looked at too, and there's some interesting possibilities there at a smaller,
Starting point is 00:25:57 scale. But what do you, I know you're, you're close to two phase also. And that's one that's gone through cycles of being really exciting, then being scary with the forever chemicals, the P-FAS, but it seems like that hasn't gone away. And there might still be some, some life in two-phase. What's, what's the latest there that you know about? So there's a, yes, there's, there's a lot that is being discussed around two-phase. It hasn't gone away. There's been intrepid work by Excelsius and Zudor in particular, keeping it in the limelight.
Starting point is 00:26:35 But also, there's a lot of organizations that have at least bench-stop experiments going on in two-phase. And the reason is it's one of the biggest parts of the thermal resistance at the chip and the cold plate is this. caloric resistance or the fact that when I heat a single phase fluid, like air, like liquid, like water, when I add heat to it, its temperature goes up. That's a pretty big penalty. And generally, it's at least a 5 degree C penalty. So that's not insignificant these days when we're trying to keep our water temperatures really high. In theory, two phase doesn't happen. It boils at the same temperature, whether I'm boiling 10 watts or 1,000 watts. So that's one key thing.
Starting point is 00:27:29 And it just looks cool and you imagine that anything is bubbling like that. It's got to be great. It turns out that there may be some advantages to two phase at the chip. On the condensation side, it's actually a lot harder to manage. So there's still some challenges in engineering this. The key thing ties together some of the points of the discussion we've had so far. And the key thing is if we think about this issue with water, water quality, leaking, all this stuff, they don't go away, but they are dramatically minimized, right? So I actually believe, think, you know, believes a strong word, but I think there's a lot of potential in the enterprise space.
Starting point is 00:28:10 If they can get some water, they may have water on the floor for, that's a bad word, they may have piping that exists for, say, rear door heat exchangers. But they don't want to take the risk with water in their rack. Well, two-phase might be a great answer to them. Even if, you know, it's not necessarily performing better than a water-based DLC system, it takes that risk out, right? And the same thing could be true for the neoclodes. It's a risk mitigator if we can design systems that can actually deploy a scale. And that's been the challenge, right?
Starting point is 00:28:46 Right now, to deploy a scale, you know, we're going to be, if we're not already, producing nearly a million coal plates a month for AI-based servers. That's a huge scale. And so right now, a lot of these two-face systems are end-to-end design such that the coal plate is integral to the manifold is integral with the CDU, and they all have to work together because you're balancing
Starting point is 00:29:14 all the instabilities and the two-face flow and all this stuff. That can't work at scale. You've got to be able to take a co-plate from some vendor in Taiwan and plug it into a manifold that was made in Canada, and that's got to work with the CDU made in Germany, right? We've got to really figure that out, or it's just not going to scale. And so that's where the work's happening, but it is happening. And so I don't see it ever displacing air or water completely,
Starting point is 00:29:43 but I do see that there's probably an opportunity because of the challenge, of dealing with water and water chemistry and the risk of water, I do see that there's a niche and potentially the chance for some companies to do pretty well in that space. Are we still concerned as an industry about the chemicals used in two phase? So I won't say no, but there have been developments in the refrigeration space of compromise between global warming potential and minimizing the impacts of PFAS. to the point where regulation-wise anyway, it's believed that they can get the concentration down
Starting point is 00:30:26 to where they're not damaging. I am not a specialist in that area, and so I haven't formed a personal judgment on that. But what we're seeing is the regulatory blocks are being lifted with the new refrigerants, like R515B, R1,25B, R-2-3-3-ZD, things like this that are a balance of good global warming potential, low-epass content, and zero flymobility, which is another key thing. So it's lots of great to refrigerants, but they burn, and that's not good in the data center. No, not great.
Starting point is 00:31:06 That's why they took fuel stops out of F-1 because it can be problematic. Talk about the CDUs a little bit then, because that's another area that I think is confusing and continues to change as the historical data center guys have now invested heavily via acquisitions to be able to supply their customers with these end-to-end solutions. I think I still kind of get wrapped up on where are we going with CDUs? How do we get to an HA-style scenario where if we have a CDU fail, we're not losing racks and racks and racks of gear? And again, we keep hitting on it, but the compatibility issue, is it going to be possible to have multiple different CDUs on your floor? And how are you thinking about that piece of it?
Starting point is 00:32:02 Yeah, this is super important. And it seemed like forever ago. It wasn't that long ago. 2023, I think a small group of us in Ashera said, you know, it's kind of a nightmare if you're at Dell or someplace else in trying to buy a CDU because the specifications are all over the place. And whatever the vendor says is probably not what we're going to see in the field. So one of the key things that I will push is Ashtray Standard 127.
Starting point is 00:32:35 We've published a method of test. that standardizes the way that you measure the efficiency of a CDU and how it operates. And that's super important because that just didn't exist before. There was no apples-to-apples comparison of CDUs. Now there is. So that's the first step and an important step in allowing customers to be able to understand, can I take this vertive CDU and put it out on the floor with this Motivir CDU and they'll work together? Generally speaking, that is becoming much more reality.
Starting point is 00:33:04 So in one sense, I think there's been a lot of progress there. And credit to the manufacturers as well, have been very much part of this process and want to see this happen. On the other hand, there's so there's been a lot of progress as well in just engineering CDUs. And CDUs are thought of as a pump and a heat exchanger. They've got to do a lot more than that. one of the things that people don't realize is every small variation in flow is instantaneously
Starting point is 00:33:39 uh sorry instantaneously appears as a temperature fluctuation at the cold plate and so your cdU has to manage that really well um and to be honest people just didn't pay a lot of attention to it because the the risk was low now the risk is high um if i have a 10% change in my um flow that can actually cause throttling events, right? So the CDU is harder to design than people think, and some folks have gotten really good at it. Heat exchanger technology has come a long ways in what we're deploying, as well as control algorithms and control vowels even
Starting point is 00:34:21 are much better than they were even just three or four years ago. So all of this is improving. I think CDU technology is improving. one of the things that I personally want the caution about this is an opinion that may not be popular No, that's why you're here We're here for your popular opinion.
Starting point is 00:34:39 But we've been talking about the issues with water and piping and all this stuff and there's been this move to go to bigger and bigger CDUs, you know, having 10 megawatts CDUs or whatever. And I understand that from the facility point of view, from the Neocloud point of view.
Starting point is 00:34:57 but it means I've got this big blast radius and in theory I can put a bunch of CDUs in parallel and I've got good uptime because they're all supporting each other but there are going to be times when one of those CDUs or one server even contaminates that whole loop and what do I do? If I have a leak in one of the main feeds
Starting point is 00:35:18 everything comes down right my volume of liquid itself goes up non-linearly when I, you know, change the size of my pipe. So I've got these big pipes feeding all these racks. I've got a ton of coolant in there. And I can't mix coolants. Very few coolants from different makers can be mixed because there's proprietary ingredients,
Starting point is 00:35:42 just a small percentage, but we have found that even that small percentage can potentially cause reactions that may be bad. We don't know for sure. There's still more work to be done, but right now it generally is recommended you don't mix coolant. So what that means is, okay, you're locked into a coolant, and if your supplier can't get you a truckload in time, then you're down too, right? So if I have these big systems,
Starting point is 00:36:07 I now have this risk of going down and taking a lot of compute with it. Whereas if we can keep the system smaller at a row level or even a rack level, then I now
Starting point is 00:36:23 control and contain my, you know, blast rate is that contain what's impacted. And it also has a huge benefit in time to deploy. I'm not waiting for all these big piping systems to get built and cleaned. I can, you know, in the case of an in-racks you, I'm rolling a rack onto the floor and that same day running it. So I'm glad you brought that up because that's something that I know you're not at Dell now, but at Dell Tech World in May, they were showing the evolution of their, in rack CDU, basically a little mini-CDU just for that rack.
Starting point is 00:37:01 We saw it at Supercompute and then the evolution in May. And that's sort of counter to where the industry has been going with these larger and larger, you know, sub-zero-sized refrigerator, you know, kind of massive units that would power or cool multiple racks. Do you see a fragmentation in where we're going here? And maybe there's not a right answer. and it just depends on the customer. But the in-RAC CDUs seem like they make logical sense.
Starting point is 00:37:34 So we'll see where that goes. That was motivated by customer demand. We had customers for this, in particular for the serviceability, deployability, and time to deploy. They wanted in RAC CDUs. And they wanted in RAC CDUs that could cool Vera Ruman with no compromise. Right? Those CDUs didn't exist. So we made one.
Starting point is 00:37:56 And that can do 220 kilowatts with 4 degrees C approach. In other words, if you give me 41 degree C facility water, I can generate 45 degrees C for Vera Rubin at 1.5 liters per kilowatt, the full flow rate. And so great work by the engineering teams. That was actually developed internally to Dell. It's a pretty unusual project. But just amazing engineering work with them and our partners that made that happen. And that is exactly the sort of thing that I think, personally, that accelerates deployment.
Starting point is 00:38:34 It allows folks to deploy very quickly. It's got a lot of redundancy. It is not easily serviceable. I mean, yeah, I mean, you have to pull it out if you're going to change a pump, for instance, or whatever. But it's got two pumps that one pump can drive all the cooling. So the chances of you going down at a critical time are very small. And if you do, you're just... on one rack, not your entire.
Starting point is 00:38:58 That's right. Yeah. And so I, it is my belief that there will be a significant portion of the market that will see that advantage and go that way. There will be some like hyperscalers and certain neoclods that will still prefer the really large CDUs and that infrastructure investment and they have the staff to make it work. And they're building on long timeframes and they don't care, you know, but those folks who want to order a Vera Rubin and get it.
Starting point is 00:39:26 it up and running, like within a month or two of the announcement, I think that they're going to want that in RAC CDU. Now, right now, DEL is the only one that provides that class of CDU. So we'll see if anyone else can get there. But it is my belief that that is a wise way to go with the CDUs just from, again, a service, time to deployability standpoint. So I'm a little biased there. And it's just minimizing, well, just minimizing all that effort of keeping all those pipes clean and so on. If you can read that Astro-T-Bolton is coming out.
Starting point is 00:40:07 And you need to do everything that's in there. And it's a lot of work to put that big piping network together. And if I can eliminate the big piping network and just have it all contain the rack, it just seems to me to be a big plus. So that's a good segue then. I said I wanted to talk about air still. And let's get back there because for all of the buzzy stuff around agentic AI and enterprise inference and tokenomics and all these other things, you know, we're seeing these air-cooled servers with four or eight RTX Pro 6,000s doing a lot of work in a lot of businesses, education institutions, and other places where it's not a training task.
Starting point is 00:40:50 It's something else. And these systems are great for that. they're almost all air-cooled. And the way vendors get there is a little bit different. Some are a little more dense. Some are taller, eight or ten-U enclosures for the servers. But do we have runway? Do you think on air-cooled for, and if we do, obviously we have some, but how long do you think we can continue to get away
Starting point is 00:41:14 with air-cooled servers with the GPUs inside? Really good question. So again, I'll refer back to another aspect publication. Just because, I mean, that's our role as DC-Denpling is to get that information out there and try to educate. And one of the things that I really want to make clear is somewhere around 90% of all server shipments today are air-cooled. And a significant portion, actually, of AI server shipments are air-cooled. Not the majority, but a significant number. And the reason is that you just look in the news.
Starting point is 00:41:50 how many new 500 megawatt data center sites are going to be deployed this year? I mean, a lot, more than you would imagine, but not the capacity that's needed. Also, inferencing occurs, just to borrow from a former employer, inferencing occurs where the data is generated. You don't want to send the data to some central site. So we want to put inferencing compute out in the field at 15. year old data centers that have air cooling, that have some capacity. And so I personally anticipate that that is only going to expand.
Starting point is 00:42:34 Now, everything in this market is going to expand. So, you know, maybe as a share of the market, it won't seem like a big clip. But if we look at, I think, capacity that's going to be deployed in, like you said, university data centers and bank data centers and even bank branch data centers or rooms, right? We're going to see these servers appear that have to be air-cooled. Now, one of the challenges is a lot
Starting point is 00:42:58 of those data centers are kind of power locked, right? So they are not just landlocked, but they've designed for, I don't know, 500 kilowatts. And that's all they've got. And so one of the things that I really
Starting point is 00:43:14 have as a mission is to try to help people see, okay, if you've got an older 10-year-old data center that's 500 kilowatts, you're probably spending a significant portion. Somewhere around 200 kilowatts on cooling on average over the year, that's power you really want to send to your inference box. So what can we do?
Starting point is 00:43:40 So that might be a relatively low lift of, let's go to Reardoner Heat Exchange for some of our computer. that can increase my efficiency and save me, you know, tens of kilowatts if not hundreds of kilowatts. Things like this that I think are going to be really important to shift cooling kilowatts to compute kilowatts so that we can deploy these 10, 15 kilowatt inference boxes in these remote locations. So I do think there's going to be some modification needed in order to make power space, you know, to make space in the power budget for these systems. Or just, you know, refreshing the kind of Excel servers so that where I needed, you know,
Starting point is 00:44:27 20 in a rack before now I only need five new servers to serve up Excel and I can use that power, even though each box is more powerful, the total racks, lower power maybe, and use that for interesting. So I think we're going to see a lot of very careful planning, probably some upgrade. you know, upgrading a data center to use a cooling tower, which I realize some people, there's a whole other discussion and the water you stare, but it's a great cooling device, very low energy, very inexpensive. You can upgrade those much cheaper than people think
Starting point is 00:44:59 and be talking about how do we shift those cooling watts to compute watts and enable the air-cooled, inferenceing and compute to continue to expand because we're absolutely we're going to see that happen. Well, you started by talking about the gaming rigs and using the all-in-one cooler internal loops. We've seen that deployed several times on servers before. Do you think there's a hybrid shot there that makes sense for these systems?
Starting point is 00:45:33 So I'm not a huge fan of those. They had a lot of complexity. They take space. And they tend to only give a marginal improvement over a high-performance air-cooled heat sink. Now, it's not zero improvement. I mean, there's been great engineering by folks, you know, UNOVOHP and others have engineered great systems that do provide an advantage over even the best vapor chamber-based air-cold heating. it's just that the complexity that you introduce for that advantage never seemed to quite add up to me unless there were some customers
Starting point is 00:46:21 that absolutely needed to push right to that limit and you know I see that I just don't personally feel like that's a mass market that there's demand for large volumes of that type of technology when again the best benefits relatively small. What you could, you know, if you're going to do that, I think the right way to do it is just invest in a liquid cool rack and then have a liquid air CDU, a sidecar next to it. Much more efficient, quieter.
Starting point is 00:46:57 And it's just a much better way to get that sort of thermal advantage without all the complexity inside the box and the penalties on service and everything else. Let me ask one last really naive question. For air-cooled servers, is there anything more to get out of the fans? And I know it's a double-edged sword because the fans are taking up a ginathear, enormous amount of the power budget in air-cooled servers. There's many studies on this about the increasing share of the fans.
Starting point is 00:47:33 But can they physically get better, more efficient at moving air? at moving air? That's a great question. And I want to put out there, our colleagues in the fan engineering space, and again, shout out to predecessors at Dell, who really, especially in the mid-2010s, really pushed fan efficiency.
Starting point is 00:47:58 Back in 2015, 2010, a fan might be 15 to 20% efficient, so it's turning about 15% to 20% efficient. So it's turning about 15 to 20% efficiency. 20% of the electricity coming in into air motion. Today, it is normal for them to be 50% to 60% efficient. That is actually ridiculous in efficiency. That is a ton of engineering that's going on across many organizations. And that's done at scale.
Starting point is 00:48:26 These are mass-produced devices. That is just mind-bogglingly efficient. So to answer your question, every generation, I can recall seeing improved fan profiles. They are getting better. I just feel like there's not a whole lot of juice left in that lemon, but they're still going to squeeze somewhere out. That said, fans, they're noisy as I'll get out,
Starting point is 00:48:56 but their actual power consumption for what they're doing has gotten a lot better. And so I would say that, you know, fans are not the evil. that they sound like. They're actually pretty energy efficient. What needs to happen, coupled with the fans, is really good engineering inside the chassis. And this is something that I think still can be improved. Like, I would push, in my role at Dell,
Starting point is 00:49:28 I would push on, like, you want to use every bit of air to pick up, you know, that pick your number but a good number is 100 CFM per kilowatt is a pretty good design number. It's not a requirement by any stretch. But if I'm doing that, that means I need every
Starting point is 00:49:48 cubic inch of air going through to heat up by 18 degrees. And if it's not heating up by 18 degrees, or if some of that air is heating up by 18 degrees Celsius, if some other chunk of air is heating up by 25 degrees Celsius, that means I've done something wrong. I've not used that
Starting point is 00:50:06 efficiently. If I can use all the air going through that box efficiently, just like I want to use every drop of water falling into a cold plate efficiently, I can do a lot with those fans. So I think there's just a lot to be done still on the engineering of the chassis as the
Starting point is 00:50:22 fans get incrementally better. But I'm optimistic. Air cooling is actually a great, you transfer fluid. People don't think about it that way. It's been really maligned. But actually from a, there are certain characteristics of air, like it's it's thermal diffusivity, things like this that are actually better than water.
Starting point is 00:50:41 And so I can actually be very effective at cooling with air, but I need to be really smart about it. You can't just throw fans at it like we used to and say we're done. Well, it's a good point, and we're already seeing a fragmentation from the large server vendors for these inference servers in terms of their design. So this will be a fun one to watch to see how they continue to develop the airflow and optimizations and those boxes don't all look the same, whereas obviously on the big racks, NVL, that's pretty prescriptive. So there's a little more room to play on the engineering side with the more traditional servers. Tim, I mean, this has been amazing. Like I said at the intro,
Starting point is 00:51:24 you're one of the great thought leaders in this space. And for anyone that wants to keep up with everything around liquid cooling and efficient data center, you need to be following Tim on LinkedIn. a link to his profile in the description. But Tim, thanks for sitting down and doing this. It's great to see you again. Really appreciate your time. Thank you very much. I've really enjoyed it.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.