In The Arena by TechArena - OCP's Zane Ball on the Open AI Era and Chip-to-Grid Vision
Episode Date: August 17, 2026In this episode of The Control Plane sponsored by AMI, recorded at the OCP EMEA Summit in Barcelona, host Allyson Klein sits down with Zane Ball, Chief Technical Officer of the Open Compute Project Fo...undation, and Colin Brix, Vice President of Marketing at AMI, for a wide-ranging conversation on the state of open hardware, the rising importance of firmware across AI infrastructure, and the emerging chip-to-grid era reshaping how the industry designs data centers.
Transcript
Discussion (0)
I'm Alison Klein, and we're coming to you from OCP Summit, Emia, in Barcelona.
Today, I'm joined by two fantastic guests.
First, Zane Ball, CTO of OCP.
Welcome.
Thank you.
And also, Colin Brick's, VP of Marketing at AMI.
Thank you.
So let's, we'll just start with introductions.
Zane, you've been on the program before.
Why don't you just give a quick introduction about your background and what your role is at OCP
and how it relates to the conversation of the day.
Sounds good.
I'm Zane Ball. I worked at Intel for almost 30 years doing all kinds of things. But most recently, I led the data center and AI engineering team. And after retiring from Intel, joined OCP as chief technical officer. And we're pursuing this big vision of an open data center for AI beyond our traditional focus on racks and IT. We're boldly moving into facilities to grid, to chip and everything else. And it's a real exciting time.
Colin back on the program, why don't you go ahead and introduce your background and what you're
driving at AMI. Yeah, thanks, Alison. I'm Colin Bricks, and I'm two months into my tenure at AMI.
Previous to that, I was also at Intel helping to manage their partner marketing program.
And before that, I was working for some hardware manufacturers. I was in Taiwan for about
10 years working in the ODM-O-DM space. I'm super excited to be at AMI. I think I've had a long history
with BIOS and working with that firmware solutions. And I think now is a really interesting
inflection point in the industry where firmware can take a more prominent role. We've always been
this background player. And I think now we have to be much more upfront, really expressing the
end stream, the customers, what is the problems we're trying to solve? And how do we get upstream
those right answers, right? So excited to be here. Awesome. Thank you so much, guys. This week is all about
open innovation. And one of the things that I wanted to talk about is just to start us off and we'll
start with using open source tends to evoke software, but the organization and the community around
it is very much focused on open hardware. Tell me why open hardware has become such a central
focus of the industry. And how does AI and the AI inflection point relate to that?
It's a good question to have now, because we just celebrate
our 15th birthday. And if folks recall, the birth of OCP 15 years ago was when
Matt on Facebook back then recognized that big hyperskillers like Facebook and the others
had started designing their own computers. And they were going direct to the ODM ecosystem in
Taiwan for manufacturing. Because they found that for their really large scale applications,
there were a lot of return on investment and benefits to customizing things and building something
specific to their needs, no more, no less. And they also realized that there are other people in the
industry doing the same thing. And they'd had so much success building on open source software.
They're like, I don't see why we can't open source the hardware. Why don't we bring that to a community?
They founded OCP with some partners like Intel, like Rackspace and some others. And no one had
ever done anything like that before. And it started getting a bit of traction. And a conference went
from a couple hundred people to our OCP in San Jose, this past year was like 14.
thousand. We have well over 500 members from all kinds of ducts and crannies of the industry,
all working together in an open source environment to collaborate on solutions. And one thing
that really makes OCP unique is that there's a diversity of collaboration models. Some of the
projects look more like a rigid standard, but some of them are more like just rough coloring
inside the lines alignment. And that makes a very flexible tool. And kind of companies can come
together and collaborate within the OCP community in a lot of ways. A lot of ways to solve a lot of
different kinds of problems. And it's just turned out to be a recipe that's incredibly flexible and has
been a resource for the whole industry. That gamut of engagement opportunities is really interesting
to me. One of the things that I think about with Open is that a lot of companies throw out the term
we're doing open innovation, but it's more of a marketing term. How do you know when innovation is truly open
and there's a community that is growing versus maybe selectively open.
It's a good point.
You do see quite a flesion open source software.
You do see quite a few projects that are, yeah, there isn't necessarily like a big contributor base.
But I think it's all about community and contribution.
If there's only a couple of companies working on and making contributions, yeah, that's not super open.
But where the magic really happens is when lots of people are contributing and bringing their insights and experience to the table.
That's the number one sign that open is really open and really working in the intended way.
You described the great growing of collaboration with an OCP and the scale of folks attending the summit.
What is the risk, the downside risk of the industry at this point in the market becomes too dependent on proprietary stacks?
And why do you think the industry is drawn to the magnet of OCP?
If you want to go fast, go alone.
If you want to go far, go together.
a little quick. And there's some truths of that. And definitely an AI with its tightly coupled nature,
where how you design the chip, how you design the model, how you design the software, how you design
the data center, how you interact with even the power grid, are all tied together. And that's what
makes AI computing unique compared to cloud computing. And so in some ways, it's easier to solve
that problem than a proprietary vertically integrated way. But then you realize that not every
data centers should be built for just one model. Not every building can be built the same way
in different geographies and different climates and different power grids have different opportunities
in terms of energy and optimization points. And you realize, wow, if in order to take advantage
of energy opportunities in West Texas, I got to rethink how I architect the chip, it's fragile.
And so its fragility is the reason proprietary solutions can be challenged, right? You get the benefits
of deep integration, but now you're kind of locked in to kind of one way of doing things,
and figuring out where the right interface is where the magic of openness is.
And engineers, they have to partner and work together to find the right interfaces when
they're doing something new. It isn't obvious necessarily where the cut line needs to be,
where we can follow the spec, where things maybe need to be a porous boundary, more proprietary.
You have to figure those things out as you go along. But when you find the right interface
points, you really unleash a lot of competition and innovation and a lot of unique solutions
for a lot of different cases in the industry can really grow. So that's the process we're going
through with AI because we just haven't built machines like this before. So we're working our
way through finding what can be open and what's not open. And honestly, I think you see
some of the very largest players who you would think would be the ones to benefit from a
proprietary approach. They're often the ones leading the charge.
on trying to unleash the power of a more open ecosystem.
Now, you talked about cross-geography deployments,
and this is one of the most interesting things
about this era of compute is that when you see
the strategic opportunity of AI,
we're transitioning from a technical to a geopolitical pursuit.
And how much has hardware trust become a geopolitical issue
in the AI era in your mind?
Well, to be glad I'm not a public policy expert,
So I'm not in the room where governments are debating these things, but it's pretty obvious from what a lot of the discussion that you see in a lot of the moves that you see different countries making that AI is obviously really important.
And having some level of sovereignty around AI is important to a lot of countries.
And if this resource becomes increasingly important for the functioning of economies, the ability to constrain that resources or cut off that research is something people are going to be worried about.
That could be access to software or a model.
That could be access to just the horsepower to run the stuff,
which means data centers and chips and those kind of things.
So I think it makes sense that those kind of concerns are top of mind
for a lot of decision makers around the world.
Now, I know that you're former silicon guy.
So this is going to be a question that's near and dear to your heart.
When we talk about AI infrastructure,
a lot of the conversation comes down to command and control and we get into firmware.
This is something that most people have not historically thought a lot about,
but it's becoming more and more essential to the delivery and management of AI factories.
How do you think we should be thinking about firmware at this moment
and is it getting the attention that it needs at this point?
We're entering kind of a golden age of firmware.
I think we need every firmware engineer we can get
and we need every firmware engineer powered by the absolute best AI support they can get
because we have big challenges ahead.
And so what is firmware?
firmware is a kind of very low-level software
that lets us work close to the heart.
And sometimes the firmware maybe aren't really changing all that much.
Like BIOS and a server, it's like, yeah, you still need BIOS,
you always need BIOS.
And maybe AI working with an EDA vendor, we can automate some of that,
make that a little bit easier.
Maybe we can make that cheaper.
So that's great.
But that isn't, I think, the thing that's so different now.
The thing that's so different now is a tightly coupled AI data center involves
all these different machines coming from all these different industries,
and they all have to work together.
So you may have the best battery energy storage system
or a low-voltage DC power generation system
or a coolant unit that all have to work together,
all have controllers inside, all of which run firmware.
They need to talk to the other firmware.
So an IT rack needs to talk to the CDU.
The facility needs to talk to the CDU.
Normally we don't have IT rack.
talking to facilities because we worry about security and things like that and just the complexity
at all and figuring out the standards interfaces for all these different industries to talk together
becomes kind of a hairball. But the benefits of being able to at minimum at least read the same
telemetry but actually act on that and work in concert at a data center level is going to be
pretty enticing. One of my favorite examples, and this would be a big one, is that
if the data center can be responsive to the power grid, we can deliver enormous amount of value.
So one data point, often citing, based on a Duke Nicholas Institute study.
The power grid is mostly not very well utilized, most of the time.
You have in North America, I think it's similar in Europe, but we have a lot of good data on North America.
If you could just curtail 0.25% of time, that's less than 24 hours in a year.
there would be 76 gigawatts available today.
Wow.
If you could curtail 0.5%,
there'd be almost 100 gigawatts available today.
A huge part of the grid is built for a very small set of events.
But they're very important, right?
Because you never ever want the grid to go down.
Well, if my data center could be integrated with the grid.
And imagine I could just flexibly,
when that nightmare scenario occurs,
those few hours of the year,
I could just move my workload to another data center.
Or maybe I turn on more on-site generation.
Or maybe I have some extra battery capability.
Or maybe there's some special sauce in our software and our firmware.
But if we could actually have the grid say,
hey, there's a weather event and 3.30 tomorrow afternoon,
you need to moderate for 23 minutes.
As a manageability firmware system management person,
you're like, well, that sounds like a nightmare.
But if you can do that, the benefits are enormous.
And think about like people worry about things like AI driving up electric bills.
If you're utilizing the infrastructure you already have better, then you're actually bringing revenue into an existing infrastructure, which is going to bring electric bills down, not up.
Because data centers have been become the world's best electric utility customer because they turn off in the peak and they burn hard during the drop.
So they're your best friend.
So now it's like you want more data centers, right?
And how would you ever make that happen?
You're going to make that happen with systems manageability and orchestration and all these things.
Yeah, so that's one benefit.
I'm sorry, I'm just kind of rambling on and on, but another big benefit is reliability.
AI systems are not super reliable.
Cloud systems, because you can isolate every piece of the machine from every other piece,
99.999% reliable.
AI systems aren't like that because one GPU goes down over here, it causes everything.
If you have better manageability, better firmware, better collaboration across all these systems, guess what?
You can build more reliable systems.
So if a billion-dollar supercomputer has better uptime, it turns out that's pretty valuable, right?
So Colin, Zane just said that firmware is pretty essential to the future of AI and data center deployments.
How does AMI plan to lead that charge?
First of all, we think the same, obviously.
But Zane pointed out some interesting things, right?
I think one is AMI has been in this industry for, you know, 40 years.
And that disaggregation is something we've had to deal with for a very long time, right?
Making sure that when a platform launches on the silicon side,
that we're managing all of these disparate parts of hardware to work together, right?
And I think, obviously, as that data center has changed and grown to what it is to
day and who knows what it's going to be tomorrow, right? Those things are changing very quickly.
It's like AMI has this understanding of how do we take that control that is necessary at the
endpoint and spec and build that in at the very beginning. Things like reliability and security.
How do you get that top of mind when people are doing design ins? And sometimes it's not an easy
conversation to have, right? Making sure that you're the voice of the endpoint rather than
at the starting point. And AMI plays a really critical role. We're kind of like this glue and this
ecosystem, bringing these parts together and saying, hey, did you realize this is going to be
super critical in that back end stage? Making sure that we're working in partnership, right? We're
sort of like the Switzerland of all of these competing companies. We bring things to work together,
right? And I think that's super critical in the open source community. And we were talking about earlier of
What makes a good open source collaborator, right?
And I do think it's realizing that working together as a community solves a lot of challenges downstream,
which creates a lot more opportunities for innovation.
Companies who can be successful are moving towards open because they're realizing that.
There's a base level of standardization that some companies shy away from because they're like,
what's my secret sauce in this?
and how do I differentiate my product?
But what they're not realizing is further downstream,
there's all these other innovations that can be enabled
if we're all working together as a community, right?
So I think EMI is in the past,
we've been this quiet player that just makes things work.
And I think in the future, we're really looking at
how do we start leading some of these conversations, right?
Because when it gets to that endpoint,
if it wasn't designed in the right way, it's too late.
Now, Colin, like a follow-on.
question for you there.
firmware has been a quiet space in the industry,
and it's becoming a much louder space in the industry.
How do you ensure that the broader industry landscape,
the entire value chain, understands the importance of this technology,
and how does AMI help connect the dots there?
Yeah, I think one is really how do we pull out the success stories, right?
The things that we can say and speak to that when we're working,
in a collaborative environment, it leads to some sort of a solution downstream that shows value,
right? And I think that's one of my opportunities here at AMI. It's really like, how are we
collaborating with the partners, making sure that at the end stage, we're showing values so that it
spurs that innovation again, right? And we get other people to realize the value of being able to
collaborate and work together. I think a lot of times everybody is doing.
their own thing in a silo. And a lot of times, even the manufacturers upstream don't see the
end product and they don't see how it's actually being deployed. If we can bring those stories
front and center and really start leading with that, leading with the problems and the challenges
that we're solving downstream, I think we just have a much healthier ecosystem in the end.
One more question on that is that interoperability and openness are so critical, yet a lot of people see firmware as their competitive moat in terms of ensuring that their solutions are the ones that are deployed again and again.
How do you work with value chain players to really embrace open innovation and interoperability?
Yeah, that's a good question. I think even internally at AMI, there had been a lot of discussions of how and what it's the right way,
for us to be open, right?
And I think in a very baseline level,
it's making sure that there's some base level of standards
that we're all operating off of,
making sure that one, even from the silicon stage,
that there's so much fragmentation now.
And that development cycle used to be two years to one year to six months.
Keeping up with that cycle is really near impossible
for a lot of companies, right?
So I think open allows us to be able to, one, enable solutions faster together.
But then also there's a value add that AMI can bring in terms of, hey, we have all these other
services on top of that, right?
If you're worried about security, we can help provide that extra layer that you need to get you
there, right?
If you're wanting to do it alone, great.
We have open source.
We're making sure that we're open source first.
We're making the contributions to the community.
But then there's all these other things that go into enabling a solution that we can help guide the customers to.
So I think there's a lot of opportunity.
And when you start to bring people together in that collaborative mode, it's a much better kind of conversation than it's like, hey, this is our walled garden.
Only certain people can play in it.
And that's it.
If you don't want to pay, you can't play, right?
No, Zane, I think that one of the questions that I have is do you think operators are
going to start in their RFPs demanding open source firmware as a baseline requirement,
especially when you're considering some of the AI infrastructure that you discussed before.
I think there's a lot of cases where that's already the case. There's a lot of firmware.
There are a lot of devices and there's a lot of different context. So it's a pretty vast space
and so you don't get the same answer everywhere. But like famously most cloud companies use
open BMCs and their servers. They don't use proprietary.
regulatory manageability for
and that's because that gets them
a platform to innovate on top.
It also helps them.
They generally don't want to have code
in their data center that they haven't seen
from a security standpoint
because you're responsible for the outcome.
And I think regardless of
a lot of things in the end,
a given company is responsible for an outcome,
whether it's security or performance or quality.
I think there's some limits to where
open source and firmware makes sense.
I think it's getting the interfaces right,
which is what matters the most.
if I'm thinking of like a really complex piece of silicon,
you might need a BIOS to configure a bunch of registers
and things that are super close.
And you wouldn't necessarily want to open source.
You wouldn't necessarily want to create the largest possible security surface
and for certain kinds of like really super close to the metal kind of things.
But effectively what the BIOS is doing for you is it's simplifying the hardware.
It configures it and it gives a uniform experience for the drivers and the operating system.
to experience the hardware in a completely stable way.
And so some of that might be super unique
to a particular bit of IP.
And that's probably fine.
They're probably people that disagree with me about that.
But far more important is an extra-off abstraction above that,
where we're talking about security features,
manageability features,
where the firmware is going beyond
configuring a piece of silicon to providing value-add services
that could be super important
out of the band of the actual operation of the machine to provide important functions for making
the entire data center function better.
And something you said, Colin, earlier really resonated with me that the depth of the supply chain
is pretty amazing and it's a lot to expect that somebody, I don't know, designing a board
that goes in a switch is going to understand the importance of how the data center is
interacting with the grit.
There's a totally unrealistic expectation.
But if through collaboration we've established the right interfaces and the right standard,
they only know they need to meet the interface.
They don't need to know anything about all the rest of this.
And that's what we want, right?
We create sandboxes where innovation can occur.
And as long as they respect the boundaries, an old colleague of mine, Jim Keller used to say,
good fences make good neighbors.
And that's what standards in open search are about.
It's like, let's have clear boundaries or have clear responsibilities between layers of the stack.
And when we do that, people can innovate.
and they can do so without impacting downstream players.
And AI for that to work, though, because it's such an integrated system,
we have to do some work to create those interfaces in front.
And that's a lot of the work of the next few years, I think.
Now, I want to look ahead for a second.
I'm going to give you a bonus question, which is,
if you think about open silicon, open firmware,
composable infrastructure, or maybe even something else
that we haven't even started talking about yet, saying,
what are you excited for in the next five years
in terms of the innovation that is going to deliver what operators are looking for
to unleash all of this AI goodness.
I am super interested to see how chip-to-grid really works.
I don't think we've ever seen the industry the kind of innovation happening across
such a broad set of disciplines.
And if I look at the chip side, one of the things I'm most excited about is seeing how
AI workloads are getting broken down into components.
and unique silicon solutions for components.
So I always think of this in Microsoft paper called Splitwise,
and they might be the only ones that understood this insight,
that lots of companies are acting on it,
that in inference you can separate generation from pre-fill,
and you can get significant benefits,
but it means you can have a unique chip architecture for one.
You're dividing inference, so we get training inference within inference.
We're breaking it down.
I think you're going to see more and more of that,
and you're going to see more and more innovation,
and then that's going to have implications at the system level.
That's kind of implications at the data center level.
And so seeing how that plays out, and maybe that'll be chiplets,
maybe there'll be more of a chipplet ecosystem where people can buy and sell individual
triplets that have some of this key functionality all the way now to the grid side.
And you heard my spiel about the flexible data center and being a grid asset and integrating
data centers with the grid.
The breadth of that is just overwhelming to think about.
I don't think any of our last 30 years of technology history,
we've ever seen innovation happening on such a broad reach.
We had big things happen like the birth of mobile,
the birth of wireless internet.
But all those were kind of like particular steps.
You know, I don't think we've ever seen so much innovation
across so many layers simultaneously.
So that's kind of what I'm most excited about.
This was a fantastic conversation.
Thank you guys for being here.
I have one final question for you.
where can folks find out more about the programs and opportunities that will help drive this industry forward,
both from the standpoint of AMI and from the standpoint of OCP?
Colin, do you want to kick us off?
Yeah, sure.
So I think we're super glad to be able to support OCP.
We're going to be at the future OCP events around the world.
You want to learn more.
Go to AMI.com.
We have information about what we're showcasing here, along with some open source solutions.
that we're launching, our MegaRack One Tree Community Edition,
as well as our Rack Manager community edition.
So go to AMI.com to learn more.
And saying?
Overcomput Project website, great resource.
Just kind of a reminder, open compute is not a show.
It's not a marketing event.
It's a community of thousands of engineers working together, and it's open.
So we have over 150 technical projects that
have sub-projects, all those are open. You can go find a workshop or a regular call, and you can go
download the spec, download the white paper, jump on the monthly call, learn about what other
companies are doing, reach out to the project leader, get yourself on the agenda and share your
ideas, your company's idea. You don't even have to be a member of LCP to participate. And then I
would say, and just use your chatbot. Since it's all open source, you can ask your favorite chatbot,
anything you need to know about what the OCP community is doing on any given problem
set that you're involved in. And since it's all open source, it's all out there. So it's for
everyone's use. And people should just take advantage of it. Some of our calls will just get
hundreds of people dialing in to hear the latest on low voltage power distribution or a new
telemetry white paper or something like that. So it's really easy to get the information.
That's awesome. Well, that is a wrap. Thank you so much for being on. It was a fantastic discussion.
