In The Arena by TechArena - Google's Bill Magro on AI, HPC, & the Future of Scientific Computing
Episode Date: August 6, 2026In this episode of In the Arena, Allyson Klein sits down with Bill Magro, Global Director and Chief Technologist of High Performance Computing at Google, for a wide-ranging look at where scientific an...d technical computing is headed.Bill explains how Google Cloud approached HPC by meeting customers where they are, letting researchers bring their existing binaries and workflows without rewriting for the cloud.
Transcript
Discussion (0)
Welcome to Tech Arena, featuring authentic discussions between tech's leading innovators and our host, Alison Klein.
Now, let's step into the arena.
Welcome in the arena. My name is Allison Klein.
Today, I am delighted to be having a episode about high performance computing and how technology is used to advance scientific discovery.
It's one of my favorite topics. And to do so, I've got one of my good friends,
Bill Magro with me. Bill is the global director of high performance computing at Google,
and Google is really transforming the way people think about HPC. Welcome to the program, Bill.
Thanks, Alison. I'm glad to be here and thank for the invitation. Now, I know that Google really
needs no introduction. I think everyone knows who Google is, but they may not be as familiar
with Google's HPC business. And you've never been on the show before. Could you just walk us
through how Google looks at HPC, what you're delivering to the market, and what is your
role in all of that? As you said, Google's a household name. It's a part of a larger company
called Alphabet. And within Google, we have a number of different divisions. And Google Cloud is
where I sit. Our approach to high performance computing is maybe a little bit different than
some others in the market. Google came to Cloud and started its cloud offering a little bit
later than some of the other players in the market.
But one of the things that our customers really appreciate about us is a very thoughtful
approach that Google has taken to the architecture and just the whole user experience.
So I would say that Google Cloud, something that's not very necessarily very familiar
to a high-performance computing user, but maybe more broadly in the IT community,
is something known as the magic quadrant, right?
And Google's been doing steadily up and to the right.
And up into the right means capability and strategy, right?
your ability to execute and then also your long-term strategy.
So Google was making great progress, or Google Cloud rather,
in establishing itself as a leader in the cloud industry,
but in terms of high-performance computing,
really didn't have a dedicated product group,
at least when I joined the company.
So in terms of my role,
it really was to come in and bring some of the high-performance competing experience
that I had and look at Google Cloud's capabilities and offerings
through the lens of an HPC user
and someone would spend the time in the industry building and delivering HPC systems to customers
and ensure that we had the right products.
And product doesn't just mean things like virtual machines, storage, networking.
It also means the tools, the engineering teams, the internal practices that really give you the confidence.
Say, we can do high performance computing.
We can, you know, do things that are comparable with what folks can do on premise
and then putting that narrative together so that we go talk to customers.
So my role initially was to come in under a chief technologist title, but really look at things
holistically and start to establish the right people, products, practices in order to approach
high-performance computing users and addressing the things that really they know and are important
to them to open the door for Google to bring in the things that was already, I'll say, great at,
broadly in South.
You know, I think that a lot of people who are in the industry think about high-performance
computing is kind of monolithic clusters that exist in national labs, the top 500.
People also think of hyperscalers like Google as running massive data centers with incredible
compute capability, but high performance workloads are unique.
How do you tackle high performance workloads within a cloud environment?
And how is that different?
Well, I think the key difference, and this is one of the things that kind of differentiates
us maybe from some of the other folks out there is that Google has really been innovative
in the way it approaches building workloads. But because Google built it, services up
with a lot of internal systems and a lot of inventions, there's a lot of innovation at Google.
And Google's fairly open, right? Very big participant in the open source community. So you can see
what some of those innovations are, from networking to storage architectures to analytics and so on.
Those, however, do require that you build the application in a way to take advantage of those innovations.
And so fighting that impedance matching of how do people who are doing high performance computing in the traditional ways with compiled programming languages like 4Trend C, C++, Kuda, and running them with message passing systems such as MPI and nickel, how does that compose with something that's truly a cloud native architecture?
Sure. And so one of the biggest challenges, it wasn't a huge technical challenge. It's just a matter of focus, really, in bringing in, I'd say, folks with the right experience, was bringing those worlds together so that people can approach Google Cloud and not have to rewrite their workloads, not have to make really any changes. So our customers are able to bring their binaries in, their actual ex-keptials and also off-the-shelf binaries and achieve good compatibility and performance. That's what I was talking about earlier in terms of building up the right team.
and the right tests and the right tools.
And then once they're able to do that,
they're not breaking their workflows from on-premises.
That then opens the door for them to explore things like
re-architecting maybe a subset of their workflow
to take you there as a some of Google's advantages.
So a lot of the work here really was around
just meeting customers where they are
versus saying, hey, come do it the Google way
or come do it the cloud way.
Yeah.
Now, one of the things that I wanted to talk to you about in particular,
in particular because you've been in HPC for a really long time.
And I was at super computing last year
and noticed that the entire energy of the conference shifted.
HPC seems really hot right now in terms of a focus for the industry.
What is going on in HPC right now that is driving so much industry attention
and where is it going next?
There's really two parts to it and it's also two letters.
It's AI, right?
So the first thing is these large language models,
There's been a lot of innovation over the last decade now.
It's amazing that it's been a decade almost since some of the seminal papers came out
around Transformer and attention architecture coming out on Google's deep mine
and other researchers around the world.
It's just been an explosion in AI research.
But fundamentally, artificial intelligence training of these large models is an HPC workload,
meaning they require computers.
And that's why we've seen this explosion in the industry of demand for accelerated systems
that are able to run these models.
So we see it, I'd say most prominently,
with things like GPUs coming from Nvidia at NEMB.
We also see specialized processors
that have been purpose built for AI,
such as Google's own,
cancer processing, or as TPU.
We just announced our eighth generation, right?
The other thing, though,
that's what really happening in high-performance computing community,
and when I say HPC community,
I really mean scientific and technical computing community,
people who are solving problems that are grounded in threats,
and the laws of physics,
the laws of chemistry,
all the way up the static biology, the heart sciences, or other technical computing areas like
rendering special effects in films or program trading in financial markets, weather prediction,
climate and model drug discovery, the applications are endless. And what's happening is people are
looking at AI as a tool to accelerate their discovery. Folks initially thought of it as artificial
intelligence was going to be kind of an assistant. It would help you take through papers faster.
It would help you analyze results more quickly. It would help you.
prepare inputs to your applications more quickly. But what's really been exciting over the last
of a year or so is the development of agenic systems and multi-agentic systems where we can do
hypothesis generation. We can take ideas that are grounded in a researcher's own field and she may
have graduate students who need a novel idea on being able to take an existing idea and vet it
or take an area and say, what are some new things I could explore that are consistent with
by groups' capabilities of research. So that's really been where the conversation we've been
lately is around these multi-agentic systems. And then finally, just as large language models have
become incredibly powerful by training on massive data sets of text or images. What is the analog of
that in science? Can you train them up on the body of scientific literature, experiments,
experimental data, simulation data, and now get predictive capabilities. So that's another huge focus
in the industry right now. And it's causing seismic shifts.
in the way people are thinking about how they do science.
Now, you talk about AI integration in math.
Can you just take one step back and say, you mentioned Google DeepMind
and obviously very famous efforts from Google DeepMine in this arena,
including winning a Nobel Prize for some of the advancements
that they've driven into the scientific community?
Can you talk a little bit about the difference in math and AI
and traditional HBC and why this is such an attractive culmination?
And then also from an agentic stage,
standpoint, we're thinking about creating hypotheses.
How does that implement in terms of how you deliver the services that you're delivering
in Google Cloud?
I think the main difference is, and again, I think of it all as high performance computing.
In the early days, scientists were doing things on pencil and paper, right?
Or programming single machines to do the math.
And at some point, I think most people know that Fortransians for formula translator, right?
It was built to help scientists put their math formulas into computers.
in the 1950s.
And obviously, the field of computational science has developed
and really got sophisticated with solvers for finite element,
analysis, and different methods to solve partial differential equations,
working directly with the laws of physics or phenomenological data
and methods that have developed.
So I would say the workhorse of HPC for the last decades
really has been modeling and simulation,
a direct numerical simulation.
It's been augmented, though, by other things, such as visualization, analytics, right?
So I think of high performance computing as a practice.
It's an activity that someone is doing to try and gain insight where very powerful computers
are involved.
So AI is really just a new tool on that landscape.
It hasn't only changed the types of problems that people are solving or their desire
to solve them.
It's offering an ability to become more productive.
And one of the key ways is through its predictive power.
So I think a good example would be some.
of the research that came out of Google DeepMind recently, when I say recent, I mean
in the last couple of years, around weather forecasting. This is another place where we have
a vast store of historical data, right? We have both the optional data and then we also have
some of the predictions. So being able to provide something like that, those data sets that come
from organizations like the National Weather Service in the U.S. and ECMWS in Europe and training
the AI models on them so that they can actually do multi-day forecasts, it has been very successful.
But the thing that makes it most exciting is that a traditional numerical weather simulation
might take hours on hundreds of processors to generate, say, a seven-day forecast,
and they need to do this every four hours or so.
These AI models that have been trained, you can make those predictions in a matter of minutes
with much, much smaller amount of computer resources.
So AI holds the promise to accelerate the exploration of, say, the design space,
automobiles, aerodynamics, or quickly, give you an estimate of the weather. Ultimately, if you're doing
something where safety is involved, or you're going to build a product, or you're going to stamp out
thousands of thousands of things, you still want to go back to get as close to the ground truth as
possible by running numerical simulations for validation. But to play a role that helps you really
widow down the space to the most promising candidates saying desires very quickly or get a quick look
at what the weather forecast is going to be. Now, I know that Google
got into cloud computing relatively late compared to the other hypers, but you've really had a
meteoric rise in terms of your capabilities for HBC. You've really stood out in the marketplace.
How have you done that in terms of what your key moves were in terms of establishing and growing
services or any particular inflection points that allowed you to engage the HBC community in
unique ways?
The approach that we took, and this was my focus when I arrived, was recognizing,
recognizing that HPC users are not single application users.
Most HPC centers are shared resources.
Unlike IT, where maybe you have a collection of servers and it could be a small,
number, a very large number, tend to be devoted to individual applications or individual
services, tend to run at relatively low utilization on average in order to absorb spikes
that happen from time to time.
Cloud was really built around that notion of how do we consolidate those workloads onto fewer
servers taking advantage of techniques like virtualization. And that's kind of like the opposite
of what high performance computing is, right? High performance computing is about aggregating
resources to solve a single problem more quickly. But because they're such expensive and
capable resources, they have to be time shared. They have to be space shared, right? So anyone
from there with HPC will know that you go to a computing center and you wait your turn in the queue
so that you can get access to a slice of the machine for a certain amount of time. So those machines
have to be typically more heterogeneous in terms of their design. They have to be a very general
purpose that can serve many applications. So a lot of making cloud ready for HPC really came down
to ensuring that we're building compatible systems and doing it in an automated way so that people
could bring their workloads and not have to become cloud experts. And so that's where a lot of the
focus was. In order to do that, we had to build benchmarking teams. We had to establish the right
relationships with software vendors in the industry, we had to go look at the difference between
object storage and traditional parallel file systems, bringing in high-performance networking in a way
that's compatible with cloud security, building MPI that is fast and runs RDMA, but doesn't
necessarily mean overhauling Google's network with an Infiniman card, instead adopting with principles
in Infidivant, making them compatible with our network. And then finally, I think one of the
biggest open challenges is still storage.
But generally just looking at the way people solve the problems in HPC, high utilization, getting the economics right is really important.
HPC users have some very different expectations about what compute and high performance computing should cost in, say, a traditional IT shop.
And it comes down to those utilization rates I was talking about earlier.
One thing that we do differently is that we really approach our customers acknowledging the value of their on-crime systems, acknowledging pride that they have in systems they built, and then come in with the question of how can it,
help as opposed to necessarily replace those systems.
That's different approach has also been a great way to engage our customers a little bit differently
than some of the others in the market.
No, that's an awesome explanation.
And one of the things that I think it'll bring it to life is if you're working with a large
supercomputing center that's looking at easing your services or their workloads,
how do you partner with them to get this actually stood up and on real-time work?
We approach this with multiple phases.
there's the introduction, and I think the most important thing there is ensuring you show up with people that have experience in high performance computing and can speak the language.
So the language of high performance competing, just the vocabulary that's used.
The way we talk about is very different than the world of IT.
And so building credibility by coming in and talking to customers and having them very quickly understand that we, this is the first time we've seen their type of environment, their type of problem.
That's the first thing that really needs to be put in place.
and making sure that we have that conversation
where we acknowledge the value of what they've done
and their role in doing that.
We're not here to replace you or threaten your job.
We're here to actually make you more capable
in getting your richer set of tools to offer your users.
That's really important.
So after establishing that credibility,
then we really need to bridge the gap
between their deep understanding
of their domain, high performance computing,
our understanding high performance computing,
and then a specialist that we have
of deep understanding of cloud computing,
and then bridge that.
So a lot of times it will involve doing the very early things,
what we call establishing a foundation or sometimes a land zone.
Once that's in place, we built tools that allow customers to express their high-performance
competing environments in very familiar terms and do it in a very modular way.
And this is another thing that we found a little bit differently than others in the market,
simply because we came along later and were able to look at what customers were liking
and not liking.
So we have a product called Google Cloud Plusker Toolkit.
This actually allows you in very natural language, what's done as a YAML file with many, many different examples to just very modularly put together systems that you understand.
So you can say, this is my NetApp or this is my home directory with NFS, this is my network, and build that virtual data center.
And then the automation of the tools goes and writes all the low-level code to interface to the cloud APIs.
So as a result, many of our customers never write a single line of cloud code or make a single cloud call.
It studies these tools to create the environments, including the schedulers.
And then finally, the way a cloud business works is that we realize revenue when customers
use our services.
It's not like the traditional super computing purchase where you deliver, accept, and get the
revenue.
And so post-agreement, I would call it, we have teams that are dedicated to our customer
success who will help them unblock, whether it's a technical issue or guiding them to
the right regions to achieve the right performance, cost, availability, resources.
So it's a full life cycle engagement with the customer.
Bill, it's been great to talk to you.
I think I've got one more question for you.
It's interesting to listen to you because I know you've worked in so many aspects in high
performance computing from developing the foundational compute for high performance computing.
You now deliver a cloud services that are uniquely designed for high performance computing.
Can you just give us a preview of what comes next?
You know, we talked about AI, Agentic.
What are you looking forward to from hearing from the community, and how will that inform what you do next with your team at Google Cloud?
We really tend at these trade shows like ISC and supercomputing.
They're academic conferences, and they're also trade shows.
We do a mix of activities.
We engage with our customers, oftentimes directly at our offices.
We do it through our receptions where we engage directly with the community, meeting leaders and bringing them up to date.
A lot of times the most prominent leaders are also the ones who are busy at.
themselves and getting briefs.
And then, of course, we have our breaching rooms where we meet with customers and share
information that's maybe more sensitive that we're not ready to discuss publicly, but
we can tell them as coming.
I think some of the most exciting work that we're doing right now is working towards
making productization of Google's Deep Minds co-scientist under the Google Cloud Product
Umbrella and also Alpha Evolve, which is a technique that existing solutions to problems.
And you use AI-driven evolutionary algorithms to optimize.
We've got some really exciting case studies in that front.
So I think what we're going to be looking at is a couple things.
One is really hearing to what extent people are adopting these things and having success.
What are their outstanding needs?
The AI revolution has caused us to really do some innovative things, for example, in file systems
where we can have storage that thousands of clients can attach to and read the same software
library or dataset.
Very important for machine learning.
What are the applications of that in areas such as drug design, electronic design,
automation, and anything where you have large shared libraries and large shared
datasets.
The other thing I think that's really going to be a big part of the conversation is AI is now
driving the hardware roadmap.
And we have a panel that will be leading.
And the question is, how do scientists survive in a world where the machines are getting
faster, but they're getting faster on low precision math?
But in order to really get the right answers in scientific computing, we need high
precision math.
So how is the application and library and methods landscape changing?
And does a scientific community have a handle on that?
Or do we need to be changing our trajectories in industry
to ensure that they have access to the high precision compute
that they've had in the past?
That's amazing.
I can't wait to learn more about what you're doing
in terms of delivering value to your customers.
And I'm sure that the folks who are following online
want to engage with your team.
Where can listeners go to learn more about Google Cloud,
the services you're operating for HBC,
and engage in conversations?
about adopting your services?
I think one of the things that folks will find most interesting is,
you know, I've just touched on a handful of these topics in the last 20 minutes with you,
but we have weekly meetings of what we call the Google Cloud Advanced Computing Community,
and those are hosted by Jay Blassov,
who is a prominent member of the HBC community as well in our team.
And every single session involves either leaders from Google,
leaders from the community, or some of our own customers sharing their stories.
And every episode is recorded and posted to YouTube.
So not only can you sign up for the community and start to hear more in depth about the topics I've been talking about from what's happening, what's coming, but you can also go to the library historically.
And the other opportunity that I mentioned was if you sign up for the Google Cloud Advanced Computing community, you will also have the opportunity to come attend our briefings and our receptions at ISC this year.
That'll be a great place to get engaged and learn more.
Awesome. Well, Bill, it's always a pleasure to talk to you. And today was no different.
learned a lot. So excited to see what you're doing at Google. Keep up the fantastic work. And thank you so
much for your time. Thank you for the opportunity. Alice. That's great to see you again.
Thanks for joining Tech Arena. Subscribe and engage at our website, Techorina.ai. All content is
copyright by Tech Arena.
