In The Arena by TechArena - HPE's Trish Damkroger on the Convergence of HPC and AI

Episode Date: August 11, 2026

TechArena's Allyson Klein sits down with Trish Damkroger, SVP and GM of HPC and AI Business at HPE, for a look at the machines behind modern science and AI. They cover the convergence of HPC and AI, t...he engineering behind the Cray GX 5000 platform, 100% liquid-cooled racks reaching 400 kilowatts, the rise of sovereign AI, HPE's lead on the Green 500, the new NVIDIA collaborations, and what it takes to deliver a supercomputer customers commit to years before the silicon ships.

Transcript
Discussion (0)
Starting point is 00:00:00 Welcome to Tech Arena, featuring authentic discussions between tech's leading innovators and our host, Alison Klein. Now, let's step into the arena. Welcome in the arena. My name is Allison Klein. Tech Arena is turning to high-performance computing and everything that's going on in the scientific realm to see what's happening next in tech. To do that, I have an exciting guest with me. I'm so excited to have her on the show. Trish Dam Proger, Senior Vice President and General Manager for HPC and AI business at HPE is with us. Trish, it's such a delight to have you on the show. Thank you.
Starting point is 00:00:43 Super excited to be here. So Trish, you and I have known each other for a long time, but you are new to Tech Arena. So we're just going to start with an introduction. You lead the team behind the world's three fastest X-Scale computers. That's amazing. Frontier Aurora and El Capiton. Can you walk us through your role at HBC and how you and your students, team work together to deliver AI and HPC solutions to the marketplace?
Starting point is 00:01:08 Sure. I have an amazing team. Seriously, it's over 2,000 folks plus 140 countries. So we do have supercomputers all around the world, truly global. We go from product concept all the way through end of life of services. So I have a wonderful organization that, you know, product development, engineering, deploying these big systems, which we have over 50 years of experience, and then servicing. So truly into end and just working hand in hand with our customers, which I think is what makes us very unique, because we're truly living on these customer sites. And we do both the traditional high performance computing and these large AI infrastructure buildouts. Now, super computing and international super computing are two of my most favorite times of the year
Starting point is 00:02:03 because we get to tell amazing stories about how technology is being used for such great reasons. And ISC really sits at the intersection of science, sovereign research, and a wave of AI investment that I think is, to be honest with you, has surprised all of us about the speed and scale that it's moving. Where do you see HPC heading over the next 18 months? And what's shifting fastest from where you sit? Yeah, so I agree with you as far as the most exciting time of the year is these conferences. And the other thing that I think is so fascinating to me is how just at SC in November, how they couldn't fit all the vendors in the vendor space anymore.
Starting point is 00:02:48 I mean, it was like it was shrinking a couple years ago and now it's just boomed because of AI. And I think that is the theme. Well, we go back to what is the definition of high performance computing? It is modeling and simulation, AI, and data analytics. Ten years ago is modeling and simulation, and now we're definitely seeing true convergence of AI along with modeling and simulation. Talking about the vendor expansion, it's really about the need for supercomputers to power AI. It's the same infrastructure.
Starting point is 00:03:25 It's that dense liquid cooling that's required across the board in order to power these, the TDP, or how much power the CPUs these days take along with the GPUs. It's so funny because I think about myself as somebody that's lived in Data Center for over two decades is, you know, the geek in the room that's finally having their moment. And for you being in HPC so long, the small community that has driven. in HPC, really having their moment in terms of AI, and it's really all about AI infrastructure and what we can leverage from HPC. When you look at what you're engineering inside of HPE, and you've got the Kray systems now, the GX-5,000 platform, how do you see that engaging
Starting point is 00:04:15 to do things like traditional classical HPC, which you've been talking about, and then also deliver the trillion parameter AI training with the same system? What kind of engineering is it taking to deliver that type of capability? Yeah, I think it's what we've been doing again for 50 years at Cray. It's continuing to look at how do you get the density and the power delivered to these systems so that they can do these world-class time-to-solution development. So we still do have the densest solutions out there. The current portfolio is 100% liquid quill, which is unique in the market.
Starting point is 00:05:00 And that is how we can make them denser so that you can fit more on your compute floor, which is important. And we're also make it vendor agnostic. So you can choose which hardware you want, software, open source all the way through, more proprietary, and even interconnect. We would go to the GX. We're really open it up so that. you can choose what you need for your workload.
Starting point is 00:05:28 It's so impressive. What you've been able to engineer is just really, it's astounding, but keep going. Yeah, and you talk about being that geek. I mean, I do. I kind of geek out about this. I'm going to like that to go to your background-wise. So I get super excited when I can see how we can truly power those big machines. And we just have that experience to deploy and scale class systems,
Starting point is 00:05:52 which I think is also a very big machine. unique advantage. Now, I will tell you that the industry is definitely with Nvidia and AMD pushing both Rackscale design and hopefully Intel was also talking about it and getting into that market, that it really helps all of us because it increases the supply change and our ability to have even more options. Yeah, no, for sure. And that's such an interesting design point change, and hopefully we'll get back to that. I do want to talk to you about, I think the word that is on my buzzword bingo card more than anything else in 2026, which is sovereign AI. And I know that this is going to be a theme at ISC. It was certainly a theme at OCP in Barcelona
Starting point is 00:06:39 earlier. And we see it from a number of players. There's HLRS and Stuttgart, Kiste and Korea, Argonne in the U.S., who are all making moves in the space. What does this require for? from an infrastructure perspective, and why are national labs and governments turning to HPE to build it? Well, I, you know, I laugh because we've been doing sovereign forever, right? It's just where the national labs have sat for their history. And just thinking about the different needs,
Starting point is 00:07:11 if you look at the EU regulations and the stuff about data security, governance and control, and the regulatory requirements that are coming down, This is just adding to the need for sovereign. And it's not just for nations any longer. It's truly any regulated industry, telco, pharmaceutical, all of that, is requiring a sovereign approach. And you do this through a number of mechanisms. You can do somewhat through networking, et cetera.
Starting point is 00:07:42 But it's really about developing that software management and monitoring that allows you to know what's going on in your system. and allow that deeper visibility into your data and your data management and control. Now, we all know that HPE tops the performance vector of the top 500, but something that we don't talk about as much and we should is how you suite the green 500 too. You've got 10 of the top 20 most energy efficient systems on that list, and you have over 300 liquid cooling patents. So you must know something about driving heat mitigation from an energy efficient standpoint as well. When you think about power climbing to hundreds of kilowatts per rack, how is your team rethinking the physics of rack scale and how you can get so much compute and compute efficiency within those racks?
Starting point is 00:08:41 Yeah. So our racks are up to 400 kilowatts, which again is the highest out there in the industry. And so that allows us, again, to empower a lot, but we've got to do it efficiently because power is very expensive and also hard to get in some places. So liquid cooling is core there. And you do it through a lot of that liquid cooling expertise that we have been developing, Craig I was liquid cooled. So way back then, we were doing liquid cooling and we've gone through so many iterations of liquid cooling from emerging cooling. to direct liquid cooling. And I think that's super important as we look forward. We're always looking at what is that next phase.
Starting point is 00:09:29 We don't see immersion yet and going for two different phases as the answer yet. But these are the things that are going to start happening. You're seeing people actually putting liquid cooling into the chips. They're experimenting with that. I think we're going to get down to that level where we're doing. doing maybe even at that point where it's a fully integrated package that you get with liquid cooling when you actually buy these high performance chips coming. I mean, it's been all day and what it means from a standpoint of the cooling market and the
Starting point is 00:10:05 semiconductor market. That's so interesting. But I wanted to get you in terms of performance. I know that you've got a lot of industry collaborations and you're doing a great job of offering a choice of systems to the market. But I did want to talk to you. a little bit about Nvidia in particular. You recently made some announcements with Nvidia about bringing that Vera CPU blade and the Quantum X-800 Infiniban to the Kray line.
Starting point is 00:10:29 But you also announced the Vera-Rubin MVL-72 system for the neocloud market. How do you ensure that you're delivering a full portfolio that makes sense to the market where you're addressing so many needs with such myriad solutions? How do you make sense of it? It's a lot.
Starting point is 00:10:47 I will tell you, my roadmap used to be one page and now it's many pages. And you just have to obviously look at what your customer wants. And that's the end. We look at different customer segments and what we're selling to these different customer segments and prioritize by what our customers are asking for. So our more traditional, the GX, X, X, X line that's currently on the GX coming is very much more that traditional market where they do still want FP64.
Starting point is 00:11:18 That's still important to them. They are going to do both modeling simulation and AI, so a convergence, is needed. And they're going to look for different software packages that are important. They're also going to look at a network that needs to scale kind of to those large exascale type class systems. Where the NBL 72 or those RAC scale are really going into, we're working with customers that are standing up these, structures quickly. They want them stood up in hours. They want them running in hours too, not days even. They have just that different requirement because it's a lot of our neoclouds, our CSP customers, that just a sensational demand to get the latest technology and get it up and
Starting point is 00:12:08 running. So it's a different beast in some ways as far as the customer expectations. But again, And it all comes back to our fundamental understanding of how to build large, dense infrastructure, liquid cooled, and how to deploy it quickly with some of the tricks that we've learned over the years. And then making sure it stays up. So it's only good if it's reliable. So really putting in the monitoring capabilities and understanding what's going on in your system. So you have those flats if anything needs attention. I have a follow-up to that.
Starting point is 00:12:48 For those who are following HPC, and they may understand the convergence of HPC and AI, but don't really understand the nature of the various workloads that you're talking about. Can you just give a little bit more on the difference
Starting point is 00:13:02 between floating point and vector-based computing and why this is such an inherent existential focus of these customers? So traditional modeling and simulation has always been floating point 64. So very standard. All of their codes are written to take advantage of that. The investment you make in your code base,
Starting point is 00:13:28 and if you look at some of these dash labs, they have been working on this for years and years. And so to pivot is a huge investment, but it's also the fidelity of the calculation. So you get 64 bits of fidelity in that calculation. When you go to your AI systems, you may have eight fits. And so you don't have the fidelity of the calculation. So the traditional Wadalina simulation community or that traditional HBC customer is definitely using AI to steer their workloads.
Starting point is 00:14:03 They're replacing some of their physics-based models that take a long time to converge with some data models that they've developed out of AI. So they're substituting in things to be faster, but right now they still need that fidelity of 64 bits. Interesting. And when you talk about that years and years and years, that really leads into my next question because customers don't just buy a supercomputer from you. They're actually making a multi-year or multi-decade commitment to you. How do you work alongside teams? And you know, you were on the other side of this earlier in your career.
Starting point is 00:14:38 Like when you look at teams like Argonne, HLRS, and Okra, how do you work with them from architecture to code modernization to actually operational deployment to make sure that these solutions functioned in the ways that they have envisioned? Yes. And I mean, it's such a different even buying behavior for these customers because a lot of times these customers are buying a system two to three years before the silicon is productized. And believe it or not, things change in two to three years. So you have to be working closely with the customer as all of these changes are happening. And again, maybe moving the customer from one codebase to another, whether it's a tender change or even moderizing as you're talking about.
Starting point is 00:15:27 We have Centers of Excellence that we have as part of this buying behavior where from day one, as soon as that contract is signed, we are meeting with them. We have people on site working with their algorithm developers, their application developers, to make sure that they are able to pivot as soon as the system is ready, working closely with the whatever silicon provider they have chosen. We also, in fact, have these corral reviews. I know I'm going to be in Houston in a meeting with a core customer on Monday doing exactly that. What has changed? What do we need to talk about?
Starting point is 00:16:07 where are we headed to make sure that we're aligned because when you are two years out, it does take a village to land these systems up and running, especially when we're using brand new silicate. That's incredible. Yeah, no, I can't even imagine the complexity of what you're trying to land from a standpoint of landing the plane of one of these supercomputers. And it's so impressive. You and your team have the privilege of having supercomputers that you're bringing to the market
Starting point is 00:16:36 that have been actually recognized by the U.S. government in a coin. How many people have that? That's so cool. I've written about that on social. It's so awesome to see the recognition, not just of Cray and the foundational role that Cray has played, but of the importance of supercomputers to our society. I'm sure that our listeners have learned a ton. Every time I talk to you, and I've known you for years.
Starting point is 00:16:59 Every time I talk to you, I learned something, Trish. And today was no different. I know that the others who are listening or watching online are the same. where can they go to find out more about what HPE is doing in the high performance computing space and how to engage with you and your team? Obviously, our web pages are the best to keep up to date on where we're going, what's new, what's the latest. We're constantly trying to make sure that we are getting you the latest. I think as far as the Craig Coin, we are definitely going to give that to our core customers. We had a great instant, and I won't tell you which competitor, but even at one of the recent competitors,
Starting point is 00:17:43 somebody stood up and started talking about their Kray system. And I was like, oh, I love this so much because Kray is just so symbolic of HPC and the industry. So I am sending that gentleman a Kray coin. So yes, mention Kray and let's talk about it. And you too can win a Kray coin. That's so awesome. Well, thank you, Trish, for your time today. It's such a pleasure.
Starting point is 00:18:07 Thank you. Thank you so much. Thanks for joining Tech Arena. Subscribe and engage at our website, Techorina. All content is copyright by Tech Arena.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.