In The Arena by TechArena - Google's Bill Magro on AI, HPC, & the Future of Scientific Computing

Episode Date: August 6, 2026

In this episode of In the Arena, Allyson Klein sits down with Bill Magro, Global Director and Chief Technologist of High Performance Computing at Google, for a wide-ranging look at where scientific an...d technical computing is headed.Bill explains how Google Cloud approached HPC by meeting customers where they are, letting researchers bring their existing binaries and workflows without rewriting for the cloud.

Transcript
Discussion (0)
Starting point is 00:00:00 Welcome to Tech Arena, featuring authentic discussions between tech's leading innovators and our host, Alison Klein. Now, let's step into the arena. Welcome in the arena. My name is Allison Klein. Today, I am delighted to be having a episode about high performance computing and how technology is used to advance scientific discovery. It's one of my favorite topics. And to do so, I've got one of my good friends, Bill Magro with me. Bill is the global director of high performance computing at Google, and Google is really transforming the way people think about HPC. Welcome to the program, Bill. Thanks, Alison. I'm glad to be here and thank for the invitation. Now, I know that Google really
Starting point is 00:00:50 needs no introduction. I think everyone knows who Google is, but they may not be as familiar with Google's HPC business. And you've never been on the show before. Could you just walk us through how Google looks at HPC, what you're delivering to the market, and what is your role in all of that? As you said, Google's a household name. It's a part of a larger company called Alphabet. And within Google, we have a number of different divisions. And Google Cloud is where I sit. Our approach to high performance computing is maybe a little bit different than some others in the market. Google came to Cloud and started its cloud offering a little bit later than some of the other players in the market.
Starting point is 00:01:29 But one of the things that our customers really appreciate about us is a very thoughtful approach that Google has taken to the architecture and just the whole user experience. So I would say that Google Cloud, something that's not very necessarily very familiar to a high-performance computing user, but maybe more broadly in the IT community, is something known as the magic quadrant, right? And Google's been doing steadily up and to the right. And up into the right means capability and strategy, right? your ability to execute and then also your long-term strategy.
Starting point is 00:01:58 So Google was making great progress, or Google Cloud rather, in establishing itself as a leader in the cloud industry, but in terms of high-performance computing, really didn't have a dedicated product group, at least when I joined the company. So in terms of my role, it really was to come in and bring some of the high-performance competing experience that I had and look at Google Cloud's capabilities and offerings
Starting point is 00:02:21 through the lens of an HPC user and someone would spend the time in the industry building and delivering HPC systems to customers and ensure that we had the right products. And product doesn't just mean things like virtual machines, storage, networking. It also means the tools, the engineering teams, the internal practices that really give you the confidence. Say, we can do high performance computing. We can, you know, do things that are comparable with what folks can do on premise and then putting that narrative together so that we go talk to customers.
Starting point is 00:02:52 So my role initially was to come in under a chief technologist title, but really look at things holistically and start to establish the right people, products, practices in order to approach high-performance computing users and addressing the things that really they know and are important to them to open the door for Google to bring in the things that was already, I'll say, great at, broadly in South. You know, I think that a lot of people who are in the industry think about high-performance computing is kind of monolithic clusters that exist in national labs, the top 500. People also think of hyperscalers like Google as running massive data centers with incredible
Starting point is 00:03:34 compute capability, but high performance workloads are unique. How do you tackle high performance workloads within a cloud environment? And how is that different? Well, I think the key difference, and this is one of the things that kind of differentiates us maybe from some of the other folks out there is that Google has really been innovative in the way it approaches building workloads. But because Google built it, services up with a lot of internal systems and a lot of inventions, there's a lot of innovation at Google. And Google's fairly open, right? Very big participant in the open source community. So you can see
Starting point is 00:04:10 what some of those innovations are, from networking to storage architectures to analytics and so on. Those, however, do require that you build the application in a way to take advantage of those innovations. And so fighting that impedance matching of how do people who are doing high performance computing in the traditional ways with compiled programming languages like 4Trend C, C++, Kuda, and running them with message passing systems such as MPI and nickel, how does that compose with something that's truly a cloud native architecture? Sure. And so one of the biggest challenges, it wasn't a huge technical challenge. It's just a matter of focus, really, in bringing in, I'd say, folks with the right experience, was bringing those worlds together so that people can approach Google Cloud and not have to rewrite their workloads, not have to make really any changes. So our customers are able to bring their binaries in, their actual ex-keptials and also off-the-shelf binaries and achieve good compatibility and performance. That's what I was talking about earlier in terms of building up the right team. and the right tests and the right tools. And then once they're able to do that, they're not breaking their workflows from on-premises. That then opens the door for them to explore things like
Starting point is 00:05:21 re-architecting maybe a subset of their workflow to take you there as a some of Google's advantages. So a lot of the work here really was around just meeting customers where they are versus saying, hey, come do it the Google way or come do it the cloud way. Yeah. Now, one of the things that I wanted to talk to you about in particular,
Starting point is 00:05:38 in particular because you've been in HPC for a really long time. And I was at super computing last year and noticed that the entire energy of the conference shifted. HPC seems really hot right now in terms of a focus for the industry. What is going on in HPC right now that is driving so much industry attention and where is it going next? There's really two parts to it and it's also two letters. It's AI, right?
Starting point is 00:06:04 So the first thing is these large language models, There's been a lot of innovation over the last decade now. It's amazing that it's been a decade almost since some of the seminal papers came out around Transformer and attention architecture coming out on Google's deep mine and other researchers around the world. It's just been an explosion in AI research. But fundamentally, artificial intelligence training of these large models is an HPC workload, meaning they require computers.
Starting point is 00:06:30 And that's why we've seen this explosion in the industry of demand for accelerated systems that are able to run these models. So we see it, I'd say most prominently, with things like GPUs coming from Nvidia at NEMB. We also see specialized processors that have been purpose built for AI, such as Google's own, cancer processing, or as TPU.
Starting point is 00:06:49 We just announced our eighth generation, right? The other thing, though, that's what really happening in high-performance computing community, and when I say HPC community, I really mean scientific and technical computing community, people who are solving problems that are grounded in threats, and the laws of physics, the laws of chemistry,
Starting point is 00:07:04 all the way up the static biology, the heart sciences, or other technical computing areas like rendering special effects in films or program trading in financial markets, weather prediction, climate and model drug discovery, the applications are endless. And what's happening is people are looking at AI as a tool to accelerate their discovery. Folks initially thought of it as artificial intelligence was going to be kind of an assistant. It would help you take through papers faster. It would help you analyze results more quickly. It would help you. prepare inputs to your applications more quickly. But what's really been exciting over the last of a year or so is the development of agenic systems and multi-agentic systems where we can do
Starting point is 00:07:45 hypothesis generation. We can take ideas that are grounded in a researcher's own field and she may have graduate students who need a novel idea on being able to take an existing idea and vet it or take an area and say, what are some new things I could explore that are consistent with by groups' capabilities of research. So that's really been where the conversation we've been lately is around these multi-agentic systems. And then finally, just as large language models have become incredibly powerful by training on massive data sets of text or images. What is the analog of that in science? Can you train them up on the body of scientific literature, experiments, experimental data, simulation data, and now get predictive capabilities. So that's another huge focus
Starting point is 00:08:29 in the industry right now. And it's causing seismic shifts. in the way people are thinking about how they do science. Now, you talk about AI integration in math. Can you just take one step back and say, you mentioned Google DeepMind and obviously very famous efforts from Google DeepMine in this arena, including winning a Nobel Prize for some of the advancements that they've driven into the scientific community? Can you talk a little bit about the difference in math and AI
Starting point is 00:08:54 and traditional HBC and why this is such an attractive culmination? And then also from an agentic stage, standpoint, we're thinking about creating hypotheses. How does that implement in terms of how you deliver the services that you're delivering in Google Cloud? I think the main difference is, and again, I think of it all as high performance computing. In the early days, scientists were doing things on pencil and paper, right? Or programming single machines to do the math.
Starting point is 00:09:23 And at some point, I think most people know that Fortransians for formula translator, right? It was built to help scientists put their math formulas into computers. in the 1950s. And obviously, the field of computational science has developed and really got sophisticated with solvers for finite element, analysis, and different methods to solve partial differential equations, working directly with the laws of physics or phenomenological data and methods that have developed.
Starting point is 00:09:51 So I would say the workhorse of HPC for the last decades really has been modeling and simulation, a direct numerical simulation. It's been augmented, though, by other things, such as visualization, analytics, right? So I think of high performance computing as a practice. It's an activity that someone is doing to try and gain insight where very powerful computers are involved. So AI is really just a new tool on that landscape.
Starting point is 00:10:18 It hasn't only changed the types of problems that people are solving or their desire to solve them. It's offering an ability to become more productive. And one of the key ways is through its predictive power. So I think a good example would be some. of the research that came out of Google DeepMind recently, when I say recent, I mean in the last couple of years, around weather forecasting. This is another place where we have a vast store of historical data, right? We have both the optional data and then we also have
Starting point is 00:10:44 some of the predictions. So being able to provide something like that, those data sets that come from organizations like the National Weather Service in the U.S. and ECMWS in Europe and training the AI models on them so that they can actually do multi-day forecasts, it has been very successful. But the thing that makes it most exciting is that a traditional numerical weather simulation might take hours on hundreds of processors to generate, say, a seven-day forecast, and they need to do this every four hours or so. These AI models that have been trained, you can make those predictions in a matter of minutes with much, much smaller amount of computer resources.
Starting point is 00:11:20 So AI holds the promise to accelerate the exploration of, say, the design space, automobiles, aerodynamics, or quickly, give you an estimate of the weather. Ultimately, if you're doing something where safety is involved, or you're going to build a product, or you're going to stamp out thousands of thousands of things, you still want to go back to get as close to the ground truth as possible by running numerical simulations for validation. But to play a role that helps you really widow down the space to the most promising candidates saying desires very quickly or get a quick look at what the weather forecast is going to be. Now, I know that Google got into cloud computing relatively late compared to the other hypers, but you've really had a
Starting point is 00:12:02 meteoric rise in terms of your capabilities for HBC. You've really stood out in the marketplace. How have you done that in terms of what your key moves were in terms of establishing and growing services or any particular inflection points that allowed you to engage the HBC community in unique ways? The approach that we took, and this was my focus when I arrived, was recognizing, recognizing that HPC users are not single application users. Most HPC centers are shared resources. Unlike IT, where maybe you have a collection of servers and it could be a small,
Starting point is 00:12:38 number, a very large number, tend to be devoted to individual applications or individual services, tend to run at relatively low utilization on average in order to absorb spikes that happen from time to time. Cloud was really built around that notion of how do we consolidate those workloads onto fewer servers taking advantage of techniques like virtualization. And that's kind of like the opposite of what high performance computing is, right? High performance computing is about aggregating resources to solve a single problem more quickly. But because they're such expensive and capable resources, they have to be time shared. They have to be space shared, right? So anyone
Starting point is 00:13:14 from there with HPC will know that you go to a computing center and you wait your turn in the queue so that you can get access to a slice of the machine for a certain amount of time. So those machines have to be typically more heterogeneous in terms of their design. They have to be a very general purpose that can serve many applications. So a lot of making cloud ready for HPC really came down to ensuring that we're building compatible systems and doing it in an automated way so that people could bring their workloads and not have to become cloud experts. And so that's where a lot of the focus was. In order to do that, we had to build benchmarking teams. We had to establish the right relationships with software vendors in the industry, we had to go look at the difference between
Starting point is 00:13:57 object storage and traditional parallel file systems, bringing in high-performance networking in a way that's compatible with cloud security, building MPI that is fast and runs RDMA, but doesn't necessarily mean overhauling Google's network with an Infiniman card, instead adopting with principles in Infidivant, making them compatible with our network. And then finally, I think one of the biggest open challenges is still storage. But generally just looking at the way people solve the problems in HPC, high utilization, getting the economics right is really important. HPC users have some very different expectations about what compute and high performance computing should cost in, say, a traditional IT shop. And it comes down to those utilization rates I was talking about earlier.
Starting point is 00:14:40 One thing that we do differently is that we really approach our customers acknowledging the value of their on-crime systems, acknowledging pride that they have in systems they built, and then come in with the question of how can it, help as opposed to necessarily replace those systems. That's different approach has also been a great way to engage our customers a little bit differently than some of the others in the market. No, that's an awesome explanation. And one of the things that I think it'll bring it to life is if you're working with a large supercomputing center that's looking at easing your services or their workloads, how do you partner with them to get this actually stood up and on real-time work?
Starting point is 00:15:20 We approach this with multiple phases. there's the introduction, and I think the most important thing there is ensuring you show up with people that have experience in high performance computing and can speak the language. So the language of high performance competing, just the vocabulary that's used. The way we talk about is very different than the world of IT. And so building credibility by coming in and talking to customers and having them very quickly understand that we, this is the first time we've seen their type of environment, their type of problem. That's the first thing that really needs to be put in place. and making sure that we have that conversation where we acknowledge the value of what they've done
Starting point is 00:15:55 and their role in doing that. We're not here to replace you or threaten your job. We're here to actually make you more capable in getting your richer set of tools to offer your users. That's really important. So after establishing that credibility, then we really need to bridge the gap between their deep understanding
Starting point is 00:16:12 of their domain, high performance computing, our understanding high performance computing, and then a specialist that we have of deep understanding of cloud computing, and then bridge that. So a lot of times it will involve doing the very early things, what we call establishing a foundation or sometimes a land zone. Once that's in place, we built tools that allow customers to express their high-performance
Starting point is 00:16:34 competing environments in very familiar terms and do it in a very modular way. And this is another thing that we found a little bit differently than others in the market, simply because we came along later and were able to look at what customers were liking and not liking. So we have a product called Google Cloud Plusker Toolkit. This actually allows you in very natural language, what's done as a YAML file with many, many different examples to just very modularly put together systems that you understand. So you can say, this is my NetApp or this is my home directory with NFS, this is my network, and build that virtual data center. And then the automation of the tools goes and writes all the low-level code to interface to the cloud APIs.
Starting point is 00:17:10 So as a result, many of our customers never write a single line of cloud code or make a single cloud call. It studies these tools to create the environments, including the schedulers. And then finally, the way a cloud business works is that we realize revenue when customers use our services. It's not like the traditional super computing purchase where you deliver, accept, and get the revenue. And so post-agreement, I would call it, we have teams that are dedicated to our customer success who will help them unblock, whether it's a technical issue or guiding them to
Starting point is 00:17:41 the right regions to achieve the right performance, cost, availability, resources. So it's a full life cycle engagement with the customer. Bill, it's been great to talk to you. I think I've got one more question for you. It's interesting to listen to you because I know you've worked in so many aspects in high performance computing from developing the foundational compute for high performance computing. You now deliver a cloud services that are uniquely designed for high performance computing. Can you just give us a preview of what comes next?
Starting point is 00:18:11 You know, we talked about AI, Agentic. What are you looking forward to from hearing from the community, and how will that inform what you do next with your team at Google Cloud? We really tend at these trade shows like ISC and supercomputing. They're academic conferences, and they're also trade shows. We do a mix of activities. We engage with our customers, oftentimes directly at our offices. We do it through our receptions where we engage directly with the community, meeting leaders and bringing them up to date. A lot of times the most prominent leaders are also the ones who are busy at.
Starting point is 00:18:43 themselves and getting briefs. And then, of course, we have our breaching rooms where we meet with customers and share information that's maybe more sensitive that we're not ready to discuss publicly, but we can tell them as coming. I think some of the most exciting work that we're doing right now is working towards making productization of Google's Deep Minds co-scientist under the Google Cloud Product Umbrella and also Alpha Evolve, which is a technique that existing solutions to problems. And you use AI-driven evolutionary algorithms to optimize.
Starting point is 00:19:13 We've got some really exciting case studies in that front. So I think what we're going to be looking at is a couple things. One is really hearing to what extent people are adopting these things and having success. What are their outstanding needs? The AI revolution has caused us to really do some innovative things, for example, in file systems where we can have storage that thousands of clients can attach to and read the same software library or dataset. Very important for machine learning.
Starting point is 00:19:39 What are the applications of that in areas such as drug design, electronic design, automation, and anything where you have large shared libraries and large shared datasets. The other thing I think that's really going to be a big part of the conversation is AI is now driving the hardware roadmap. And we have a panel that will be leading. And the question is, how do scientists survive in a world where the machines are getting faster, but they're getting faster on low precision math?
Starting point is 00:20:04 But in order to really get the right answers in scientific computing, we need high precision math. So how is the application and library and methods landscape changing? And does a scientific community have a handle on that? Or do we need to be changing our trajectories in industry to ensure that they have access to the high precision compute that they've had in the past? That's amazing.
Starting point is 00:20:24 I can't wait to learn more about what you're doing in terms of delivering value to your customers. And I'm sure that the folks who are following online want to engage with your team. Where can listeners go to learn more about Google Cloud, the services you're operating for HBC, and engage in conversations? about adopting your services?
Starting point is 00:20:45 I think one of the things that folks will find most interesting is, you know, I've just touched on a handful of these topics in the last 20 minutes with you, but we have weekly meetings of what we call the Google Cloud Advanced Computing Community, and those are hosted by Jay Blassov, who is a prominent member of the HBC community as well in our team. And every single session involves either leaders from Google, leaders from the community, or some of our own customers sharing their stories. And every episode is recorded and posted to YouTube.
Starting point is 00:21:12 So not only can you sign up for the community and start to hear more in depth about the topics I've been talking about from what's happening, what's coming, but you can also go to the library historically. And the other opportunity that I mentioned was if you sign up for the Google Cloud Advanced Computing community, you will also have the opportunity to come attend our briefings and our receptions at ISC this year. That'll be a great place to get engaged and learn more. Awesome. Well, Bill, it's always a pleasure to talk to you. And today was no different. learned a lot. So excited to see what you're doing at Google. Keep up the fantastic work. And thank you so much for your time. Thank you for the opportunity. Alice. That's great to see you again. Thanks for joining Tech Arena. Subscribe and engage at our website, Techorina.ai. All content is copyright by Tech Arena.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.