In The Arena by TechArena - OCP on How Open Compute Is Shaping the Future of AI Clusters
Episode Date: July 16, 2026In this episode of The AI Hedge, host Marc Austin (Founder and CEO of Hedgehog) interviews James Kelly, VP of Market Intelligence and Innovation at the Open Compute Project (OCP) to explore how OCP ev...olved from Meta's open-source server initiative and discuss AI infrastructure trends, including disaggregated inference, specialized AI accelerators, heterogeneous AI clusters, open networking, and OCP's Open Cluster Design initiative.
Transcript
Discussion (0)
Hello, my name is Mark Austin. I'm founder and CEO of Hedgehog. This is the Hedge where I talk with
industry leaders about AI and AI infrastructure. AI promises a lot of benefits. It also has a lot of
risks along the way. Part of this conversation is about how do you hedge the risks and reap the benefits.
So on the show today with me today is James Kelly. He's VP of Market Intelligence and Innovation
at the Open Compute Project, working with technology leaders in the community
and augmenting the foundation's directional strategy, innovation efforts, and marketing.
His contributions to the OCP proceed from over two decades of experience in technology
engineering, product management, and marketing.
Prior to working at OCP, James was a product management senior director at Juniper Networks,
leading teams responsible for its AI cluster solution, cloud native products, and automation tooling in the data center and cloud business.
In the past, James is also a tech researcher, software developer, interesting one I find hedge fund founder and executive technology consultant.
So welcome to the show, James.
Hey, thanks for having me, Mark.
It's a really impressive resume.
And you're now at the Open Compute Project.
A lot of people have never heard about OCP.
Can you give us an idea what it is
and why somebody would leave a perfectly good job
at Juniper to go to OCP?
Yeah, absolutely.
So I was trying to expand a little bit outside of networking.
Obviously, there's so much going on with AI,
and that's revolutionizing everything in the data center,
infraspace and the data center of physical space as well, frankly.
The network is super important.
I'm sure we'll get into that for AI clusters,
but it was an opportunity for me to have a different
perch across a wider set of infrastructure and the entire solution with the data center,
rather than just staying focused on the network.
But I obviously bring my network expertise with me to OCP.
There's a lot of new projects, new work streams inside of the OCP community.
And so I'm helping to, you know, foster those, move them along, if you want to say,
obviously there's a plethora of neoclods or some people call them neoscalers that have kind of
started up.
We're trying to welcome as many of those.
So I wear a lot of different hats in the small 20-odd-person staff of the foundation,
but I'm one of the technical folks there.
So how was the OCP formed?
What was its origin story?
Yeah, good question.
I mean, like for those people that don't know the OCP, kind of like going back to 2011,
15 years ago, it was formed by met at that point, Facebook, and Intel, a handful of companies
that came together to do open source hardware for the data center.
But the problem was really.
one that was initiated by Meta or Facebook at that time.
I'll just say meta now.
So they were buying hundreds of thousands of servers.
And as anybody that's on the customer side kind of knows,
you don't necessarily want to be beholden to your vendors for forklift upgrades.
And they figured out that there's a better way that they can design their servers.
They said that they wanted to design their own servers and effectively did just that.
And then said, anybody that wants to build these servers,
for us, we're going to buy hundreds of thousands, millions of these things.
And because they were effectively open sourcing the hardware design for the server,
they needed a place for that.
So the OCP Foundation was born as a nonprofit to steward the community.
In that 15 years, the scope grew from the server to open source network switches,
to storage specifications.
The community now really is responsible for specifications across the entire spectrum
of all different kinds of data center technologies.
We say from grid to chip or even concrete to chip,
like literally the very, very foundations of a physical data center building.
People are collaborating on open specifications for those all the way from that kind
of technology into the data center, all of the mechanicals, the electrical, the water,
the cooling, everything inside of the rack and down even to chipplet
intercommunications and chiplet design and test and validation right down to the semiconductor level.
It's grown to thousands and thousands of engineers.
And one thing that a lot of people often think about the foundation, because we do have members,
you don't have to be a member to participate in the RCP community.
Part of its success is just free and open for anybody who come and collaborate.
So is it fair to just kind of think about it as a supply chain specification marketplace for hyperscalers?
Yeah, that's one thing I didn't say.
So the benefit of these open source designs is that multiple vendors implement them.
And this really fosters this multi-vender ecosystem of effectively the hyperscale innovation.
The actual community tagline is community-driven hyperscale innovation for all.
So obviously those innovative specs get adopted by the hyperscalers,
but they're open for everybody else to benefit from as well.
So yeah, we see certainly the neoclouds and data.
down into the enterprise and smaller cloud and SaaS companies, for example.
Yeah.
Our tagline at Hedgehog is network like a hyperscaler.
Yeah.
And our charter, at least from inception, was to do exactly what you just said,
which is to make it possible for everybody to network like a hyperscaler.
And you mentioned neoclouds and enterprises.
So is OCP relevant for a neocloud or an enterprise or is hyperscale just too hard?
And they just have to accept it.
No, I don't have the talent to do this.
And so I should really just buy vertically integrated solutions
from their proprietary from a particular vendor,
like your former.
I mean, after all, even if you're buying vertically integrated solutions,
there's a very good chance that there's tons of OCP specified technology in it.
It's almost impossible to build IT systems for the data center
without implementing specs they're in the OCP.
But it's a good question, right?
The hypers come across the hardest problems first.
And that's why the tag.
line is what it is. But of course, everybody can come together. And a lot of people use the
community by simply being an adopter. Maybe those enterprises or the eoclods don't have the resources,
don't have the staff to do a lot of volunteer work or to open source things. But for them to
simply adopt the technology or effectively passively benefit from it, it's absolutely fine.
So whether you're just trying to keep your ear to the ground, maybe to see what's kind of
you're on the corner in terms of technology.
You might participate as an engineer in the projects,
or if you're just like a buyer,
you might just prefer to buy stuff that's OCP specified,
like OCP accepted and recognized equipment
within our OCP marketplace catalog,
because you then have the evolvability of your infrastructure, right?
Because it's based on open source,
you know that there's plenty of vendors that are behind that.
Yep.
And so you mentioned OCP acceptance.
What does OCP accepted mean?
Yeah, so I just mentioned it in passing.
So to kind of back up and actually talk about what it is,
so the way that the command of you worked is the engineers kind of work,
I say upstream in projects and work streams,
various areas on the data center technologies,
and they work together to collaborate on white papers,
reference architectures, and specifications, right?
And then those specifications are implemented by vendors, of course.
and when they are, those vendors are allowed to list their products in our OCP marketplace.
It's a catalog, but it's not a place of business.
It just redirects into whoever's vendors' website or their product pages.
But it's a nice way that people can find OCP recognized equipment.
So anybody that has something that adheres to an OCP spec,
have we even loosened it at this point, any products that effectively align with an OCP white paper,
perhaps those things can get the OCP accepted badge and sit in our marketplace.
And our marketplace also includes more integrated solutions as well.
We might talk about that later as we probably talk about the hedgehog.
Yeah, yeah, we'll talk about it.
Okay, so if I've got something I want to contribute to the community,
I want to have that solution in the OCP marketplace,
what's the process for getting accepted for a reference architecture in particular?
Because we're going to talk about reference architecture.
Yeah, exactly.
So the reference architecture contribution or any contribution to the OCP goes through a standard process.
We've got a few staff that help people through the process and a lot of it's automated.
You sign the contributor license agreement, which is Open Web Foundation, I think, 0.9 agreement.
OCP doesn't acquire any intellectual property as part of that.
It stays with the contributors.
But they do sign that saying that it's effectively going to be opened and cataloged.
in the OCP contribution database.
Then there's, obviously, if the work has already been done,
it gets presented within the relevant project at OCP,
and then it's presented to the steering committee.
The steering committee is basically one representative on that committee per project.
So it's a group of experts across the data center and obviously technical experts.
And they, in turn, with our principal engineer at OCP, Russ,
they basically make sure that the specification looks good.
And once they rubber stamp it,
then it's free to go into our contribution database.
And at that point,
then anybody that builds a product that uses that specification,
they're eligible to have that product in the OCP marketplace.
So the contribution database and the marketplace upstream, downstream of the same overall process.
Cool.
Awesome.
Okay.
Thanks for that background.
So I want to share.
shift to the topic of, it's a hot topic right now, it's disaggregated inference. So what is
disaggregated inference and why is it important to the future of AI? Yeah. So I mean,
fundamentally, the servers aren't necessarily ubiquitous anymore. We have known about the difference
between training and pre-training and inference for a long time in the AI conversations. But over the
past a year with these specialized chips, whether it's like GROC or Serbress or others, more and more
we have specialization within the inference workload. If you are, for example, uploading a document
to chat GPU to your cloud, you're providing a ton of input tokens. That's part of the pre-fill
phase. And then actually where you're getting a lot of output tokens doing the work, doing coding,
and whatever have you as part of your request, your prompt.
To fulfill it, that's basically evoking the decode phase.
There's different performance characteristics that be optimized for there.
So pulling apart the hardware architecture to match those two different phases,
what the disaggregation is effectively doing with different memory technologies
and different accelerators or GPUs, XPUs.
Got it.
So the general idea is that I use two different AI accelerator.
to get faster, more efficient inference.
Exactly.
Yeah.
One through the pre-fill, one through the decode.
Cool performance.
So then that means really that we should expect AI cloud builders going forward
to have more and more diversity in their data centers
of different AI accelerator types, right?
Yeah, absolutely.
There's a whole bunch of reasons driving diversity.
So part of it is what we just talked about,
disaggregated inference.
Obviously, people that are doing training in inference,
may do that on different platforms.
When you hear about the frontier models getting released,
you'll often hear, oh, such and such,
Chaggbti 5.5 was trained on Hopper,
but then it was tuned on GV200.
So there's things like that,
and obviously they'll have different infrastructure
for inference to run it and optimize for pre-fill and decode.
Besides that, I think GPU infrastructure is not going away
as fast as people thought.
Because there's still a lot of A100 stuff,
around that people are not decommissioning because at this point, there's so much more demand
than there is supply. Right. And then obviously, AMD, with their chips, are coming along, getting
deployed massively. There's tons of announcements. Obviously, we heard VDio acquired Rock
and we're going to be deploying that. Jets's last GTC was showing five different types of wrecks
that basically constitute an AI factory. So even if you're allie deploying in video,
Accelerators are still to play lots of different generations and lots of different variants.
Yeah, and I think we saw Cerebrus IPO last week.
Yeah, IPO price was, or valuation was just under 100 billion.
I think only META and Alibaba had exceeded that in IPO prices.
That's an example as well of an XPU that might be in an AI data center alongside
invidia GPUs, yeah?
Yeah.
Cool.
Okay.
So then that leads me to the next question, which is the OCP OpenClU.
There's a specification for open clusters.
There's another specification for open pods.
What is a cluster?
What's a pod?
Why do they need to be open?
Yeah.
So pod is basically a building block of a cluster, right?
A cluster for most people, it's the distributed,
but acting as one computer for potentially a single problem.
If you're training on a cluster,
often using the entire cluster.
Otherwise, the cluster is basically your infrastructure,
that maybe used for different workloads.
Obviously, if you're somebody like a hyperscaler or a neocloud
that's providing that as a service,
then you've got multi-tenant concerns on that shared infrastructure
and you're dividing it up and selling it.
So I think what a cluster is is well understood.
POT is a building block of that.
People think of even a sub-element of the pod
as sort of the scale-up domain, I suppose.
But to answer your question about the open bits of it
that have started within the OCP,
OCP traditionally people built specifications for specific point technology needs,
like a server, a disk drive, or power supply, all these different things.
And one of the hard problems that people were facing when they're trying to build AI clusters
is that they're reinventing the wheels, solving the same problems again and again.
And there's no point of designing a snowflake for every AI clusters.
So a bunch of collaborators created a project,
specifically we called it a strategic initiative within the OCP,
and it's got a long name.
It's a mouthful open cluster design for AI.
And they created a specification for a pod,
and the pod is a building block for the open cluster.
So they've effectively open sourced an architecture
or a design of how to build in a pod multiple racks.
They actually started with air-cooled and 400-gid connect.
RACs. Now they're working on liquid-cooled 800-gig scale-out connected racks. And there's a
specification for exactly how that pot is designed so that you know everything from all of the
cables that you would need, the number of racks, the exact number of compute and accelerators
within that. And they tried to do it in a way that multiple vendors can kind of slot in
and fill the different roles for like servers and switches and racks. Obviously, wherever they
can. They base those decisions on other OCP specifications. But yeah, now somebody can take that
and build an AI cluster from it. So it's basically everything that's on the IT side. It does
include like the OT reference architecture, how to build your floors or how to build your
overhead designs and space like that. But in terms of the row layout of multiple racks and everything
within them and how to connect them, how to configure them, that's basically what it is.
So if I want to build an AI cloud that say has
NvidiaGPUs, AMDGPUs, Google TPUs,
maybe some cerebrous accelerators,
each one of those accelerator types or XPU types
would have its own pod and they'd all work together
in the same cluster.
Is that the general idea?
Yeah, I suspect that most people would probably deploy them
in different pods.
When it comes to the disaggregated inference bits,
I suppose you could have the pre-fill of the decode kind of working together.
They could be deployed in the single pod.
At this point, the architecture doesn't have that in its specification.
But is that pre-fill decode desegregation is a fairly new idea, right?
I think it's less than a year old.
Yep.
So let's get into pre-fill and decode for a minute.
Let's assume I'm using multiple accelerators or different accelerators for the pre-fill and the decode.
They need to share memory.
so they need remote direct memory access on a network.
Is that going to work if each XPU vendor shows up with their own proprietary,
vertically integrated network stack?
And I guess the second part of that question is,
is there a reason to have an OCP reference architecture for open AI networking
that enables those diversified clusters to interoperate?
Yeah.
Well, I mean, to go back to one other words I said before,
evolvability is really important, right? The answer is not really. You may start with a network that
might work today, but as soon as you start adding different accelerators, if those vendors have
different opinions on what networks you should be deploying, that's kind of a non-starter. You
want your infrastructure to be able to talk to each other. So obviously the pre-fell of the decode
infrastructure has to be able to talk to each other to work together for a common influence workload.
So that's absolutely true. But there's other times when you'll have infrastructure that
You just simply want to all work together.
And the network is what stitches everything together, right?
There's the old saying the network is the computer,
and that's especially true for AI clusters.
So it's a non-starter for you to have different silos, so to speak, in your network.
You want your network to ubiquitously connect up whatever you decide to deploy next.
We're just a few weeks past the OCP, the Mia Summit,
where Hedgehog announced our contribution for this OCP reference,
architecture for AI networks. There's actually two RAs, one for training, one for inference,
and you and I were on a panel talking about it. So if somebody listened to this show as an AI
cloud builder and they want to find those OCP reference architectures, how do they find them?
They're in the OCP contribution database. An AI query or Google search would definitely turn
them up, I'm sure, as well, but a prominent part of the OCP website, opencomputt.org, is the
contribution database. You can go to opencomputt.org slash contributions.
Because they're fairly new, they're probably sitting right at the top.
But if you typed in reference architecture, you'd find them very quickly.
Also, there's ways to filter things by what project an initiative they fall into.
So that open cluster design initiatives that we talked about is kind of what these specific reference architectures you guys contributed are tagged under.
And I think at the same time at that show, you also had a contribution from OpenAI on MRC for, like,
really big training clusters.
Can you talk about that a little bit?
Multipath, Reliable connection.
Yeah.
So sort of another method of building an AI network for really large training clusters.
It's a host-driven protocol to optimize and to allow for the network to be a little bit
dark, so to speak, or to allow the high performance needs and the optimizations that you
need in an AI back-end network without implementing all of those optimizations in the fabric
itself. Again, because you want your fabric to effectively be ubiquitous,
to be able to connect whatever you want. So yeah, OpenAI, Microsoft, AMD as well,
for that matter, and Nvidia and Intel work together on the MRC spec.
Yeah, so that is an example of a spec that's a PDF in our contribution database.
The reference architectures you guys contributed to did in GitHub, right? And I know that in the panel
we saw people can create one of those reference architectures in a virtual way, too,
if they don't even have their cluster today,
there's ways they can play around
and actually see what it would look like.
So, James, you're, like, in a really interesting position
where you just have visibility on what the world's biggest hyperscalers are doing,
what the world's largest infrastructure providers are doing
for those hypers, for neoclodds, for enterprises,
across all the different layers of the AI infrastructure stack.
What's coming in the future?
What are we going to be talking about,
next year at OCP Global Summit and what should people be thinking about?
So some of these things will absolutely see continued adoption of them.
So we absolutely love having adoption.
We have a full adoption track at all our OCP summits.
So hearing about the success stories of these things getting deployed,
open rack wide was announced at last Global Summit and the specifications available,
but I know we're going to start seeing open rack wide flavors from,
All of the usual suspects, right?
So the Helios from AMD, from HPEE,
some from Super Micro, I'm sure.
Those are the double wide racks that people are talking about.
And those scale of domains have a scale out
that needs to be compassionate with everything else that you're deploying, right?
You're not going to deploy ubiquitous open rack wide.
You're going to deploy that with GV200, 300 racks,
VR racks in the future from Nvidia,
and then, yeah, probably some surrogous, et cetera.
So I think you're going to continue to see an expansion of the flavors of AI infrastructure
that's deployed.
Hopefully, we're going to see some hardening of the decisions on the facility level
so that the facility can be a little bit more fungible for these sort of like, as the
open letter at OCP says, late-binding decision-making on all of the evolving IT
specifications.
A court that houses the IT, right?
Is that what you're talking about?
Yeah, exactly.
Yeah, a big part of the OCP that is not my wheelhouse.
So as you introduced me, I'm a network walk, but learning about everything that's
happening in liquid cooling, whether it's called plate or immersion, whether single
phase, two phase, all of the stuff that's happening for low voltage direct client at
800 or they're talking about how do they go to 1500 in the future.
We just had a workshop on that.
So seeing the new busways and bus bars for some of this higher voltage,
but still LVDC stuff is fascinating.
So, yeah, there's things happening all around.
If you're on the networking side at OCP,
there's a brand new optic circuit switching project for OCS.
There's tons of stuff happening to try to drive more stability and reliability of the optics.
Got a new work stream there.
That one bites us all the time.
It's a lot of people all the time.
But yeah, we're deploying millions of these things.
It turns out that, yeah, if you have even a small nugner of them, their fault,
you know, yeah, reliability in these huge clusters because there's thousands of moving parts.
So I woke up this morning, checked my email.
I had a survey request in my inbox from OCP asking if you should move the location of OCP Global Summit.
So I went last year to San Jose Convention Center.
Of course, we go to Nvidia GTC every year.
I think OCP Global Summit last year was the only show I've been to.
We do a lot of shows, like probably two a month, that felt as big.
It's not quite as big, but maybe number two to Nvidia GTC.
Is the community growing?
And is it too big now for San Jose?
It's growing massively.
Some data points on that.
So the recent contribution for LVDC was the most collaborative contribution we've ever received.
It had 70 different companies and hundreds of people that worked on this.
Amazing.
Obviously, I said the OCP has grown in scope to all things data center.
And with that, all of the OT vendors and solution providers are joining OCP.
That's another way that OCP is growing massively.
But yeah, although a lot of the work happens virtually in projects on calls and mailing,
lists. One of the ways that everybody comes to an OCP is by coming to one of the summits,
the conferences. And the biggest one is the global summit happens every October. And it's
happened in San Jose, every October at San Jose Convention Center. But it's getting to the point
where we're busting at the seams of the convention center. Yeah. You know, over 11,000. We're probably
going to be 13, 14,000 this year. So we're like approaching GTC levels. I don't think we're going to
have the shark tank for our keynotes. It's pretty tough. I think Taylor Swift's probably the only
person on the planet who could top Jensen in a keynote. Yeah. James, thank you so much for talking with
me today. Thank you so much for being a part of the OCP community and welcoming Hedgehog into it.
Honestly, when I first started thinking about this problem space that we're solving at Hedgehog,
I was still working at Cisco. And I was talking about it with my boss. He's like, you know what you
should do, just go to the OCP Global Summit and check that out and then come back to me and report
on Open Networking. So that was really the genesis of Hedgehog was attending that event and
forgetting now how many years ago. It's a long time ago now, but we've been to every single
global summit since and they're fantastic events. I encourage everybody listening to this to
go to an OCP event, be a global summit or one of the regional summits, because I guarantee
you we'll learn a ton about AI infrastructure, which is a growing.
community for a very rapidly growing market. So thanks, James. Yeah, thanks for having me.
It's a pleasure.
