In The Arena by TechArena - OCP on How Open Compute Is Shaping the Future of AI Clusters

Episode Date: July 16, 2026

In this episode of The AI Hedge, host Marc Austin (Founder and CEO of Hedgehog) interviews James Kelly, VP of Market Intelligence and Innovation at the Open Compute Project (OCP) to explore how OCP ev...olved from Meta's open-source server initiative and discuss AI infrastructure trends, including disaggregated inference, specialized AI accelerators, heterogeneous AI clusters, open networking, and OCP's Open Cluster Design initiative.

Transcript
Discussion (0)
Starting point is 00:00:00 Hello, my name is Mark Austin. I'm founder and CEO of Hedgehog. This is the Hedge where I talk with industry leaders about AI and AI infrastructure. AI promises a lot of benefits. It also has a lot of risks along the way. Part of this conversation is about how do you hedge the risks and reap the benefits. So on the show today with me today is James Kelly. He's VP of Market Intelligence and Innovation at the Open Compute Project, working with technology leaders in the community and augmenting the foundation's directional strategy, innovation efforts, and marketing. His contributions to the OCP proceed from over two decades of experience in technology engineering, product management, and marketing.
Starting point is 00:00:47 Prior to working at OCP, James was a product management senior director at Juniper Networks, leading teams responsible for its AI cluster solution, cloud native products, and automation tooling in the data center and cloud business. In the past, James is also a tech researcher, software developer, interesting one I find hedge fund founder and executive technology consultant. So welcome to the show, James. Hey, thanks for having me, Mark. It's a really impressive resume. And you're now at the Open Compute Project. A lot of people have never heard about OCP.
Starting point is 00:01:23 Can you give us an idea what it is and why somebody would leave a perfectly good job at Juniper to go to OCP? Yeah, absolutely. So I was trying to expand a little bit outside of networking. Obviously, there's so much going on with AI, and that's revolutionizing everything in the data center, infraspace and the data center of physical space as well, frankly.
Starting point is 00:01:41 The network is super important. I'm sure we'll get into that for AI clusters, but it was an opportunity for me to have a different perch across a wider set of infrastructure and the entire solution with the data center, rather than just staying focused on the network. But I obviously bring my network expertise with me to OCP. There's a lot of new projects, new work streams inside of the OCP community. And so I'm helping to, you know, foster those, move them along, if you want to say,
Starting point is 00:02:11 obviously there's a plethora of neoclods or some people call them neoscalers that have kind of started up. We're trying to welcome as many of those. So I wear a lot of different hats in the small 20-odd-person staff of the foundation, but I'm one of the technical folks there. So how was the OCP formed? What was its origin story? Yeah, good question.
Starting point is 00:02:30 I mean, like for those people that don't know the OCP, kind of like going back to 2011, 15 years ago, it was formed by met at that point, Facebook, and Intel, a handful of companies that came together to do open source hardware for the data center. But the problem was really. one that was initiated by Meta or Facebook at that time. I'll just say meta now. So they were buying hundreds of thousands of servers. And as anybody that's on the customer side kind of knows,
Starting point is 00:03:01 you don't necessarily want to be beholden to your vendors for forklift upgrades. And they figured out that there's a better way that they can design their servers. They said that they wanted to design their own servers and effectively did just that. And then said, anybody that wants to build these servers, for us, we're going to buy hundreds of thousands, millions of these things. And because they were effectively open sourcing the hardware design for the server, they needed a place for that. So the OCP Foundation was born as a nonprofit to steward the community.
Starting point is 00:03:35 In that 15 years, the scope grew from the server to open source network switches, to storage specifications. The community now really is responsible for specifications across the entire spectrum of all different kinds of data center technologies. We say from grid to chip or even concrete to chip, like literally the very, very foundations of a physical data center building. People are collaborating on open specifications for those all the way from that kind of technology into the data center, all of the mechanicals, the electrical, the water,
Starting point is 00:04:08 the cooling, everything inside of the rack and down even to chipplet intercommunications and chiplet design and test and validation right down to the semiconductor level. It's grown to thousands and thousands of engineers. And one thing that a lot of people often think about the foundation, because we do have members, you don't have to be a member to participate in the RCP community. Part of its success is just free and open for anybody who come and collaborate. So is it fair to just kind of think about it as a supply chain specification marketplace for hyperscalers? Yeah, that's one thing I didn't say.
Starting point is 00:04:47 So the benefit of these open source designs is that multiple vendors implement them. And this really fosters this multi-vender ecosystem of effectively the hyperscale innovation. The actual community tagline is community-driven hyperscale innovation for all. So obviously those innovative specs get adopted by the hyperscalers, but they're open for everybody else to benefit from as well. So yeah, we see certainly the neoclouds and data. down into the enterprise and smaller cloud and SaaS companies, for example. Yeah.
Starting point is 00:05:22 Our tagline at Hedgehog is network like a hyperscaler. Yeah. And our charter, at least from inception, was to do exactly what you just said, which is to make it possible for everybody to network like a hyperscaler. And you mentioned neoclouds and enterprises. So is OCP relevant for a neocloud or an enterprise or is hyperscale just too hard? And they just have to accept it. No, I don't have the talent to do this.
Starting point is 00:05:48 And so I should really just buy vertically integrated solutions from their proprietary from a particular vendor, like your former. I mean, after all, even if you're buying vertically integrated solutions, there's a very good chance that there's tons of OCP specified technology in it. It's almost impossible to build IT systems for the data center without implementing specs they're in the OCP. But it's a good question, right?
Starting point is 00:06:12 The hypers come across the hardest problems first. And that's why the tag. line is what it is. But of course, everybody can come together. And a lot of people use the community by simply being an adopter. Maybe those enterprises or the eoclods don't have the resources, don't have the staff to do a lot of volunteer work or to open source things. But for them to simply adopt the technology or effectively passively benefit from it, it's absolutely fine. So whether you're just trying to keep your ear to the ground, maybe to see what's kind of you're on the corner in terms of technology.
Starting point is 00:06:47 You might participate as an engineer in the projects, or if you're just like a buyer, you might just prefer to buy stuff that's OCP specified, like OCP accepted and recognized equipment within our OCP marketplace catalog, because you then have the evolvability of your infrastructure, right? Because it's based on open source, you know that there's plenty of vendors that are behind that.
Starting point is 00:07:12 Yep. And so you mentioned OCP acceptance. What does OCP accepted mean? Yeah, so I just mentioned it in passing. So to kind of back up and actually talk about what it is, so the way that the command of you worked is the engineers kind of work, I say upstream in projects and work streams, various areas on the data center technologies,
Starting point is 00:07:32 and they work together to collaborate on white papers, reference architectures, and specifications, right? And then those specifications are implemented by vendors, of course. and when they are, those vendors are allowed to list their products in our OCP marketplace. It's a catalog, but it's not a place of business. It just redirects into whoever's vendors' website or their product pages. But it's a nice way that people can find OCP recognized equipment. So anybody that has something that adheres to an OCP spec,
Starting point is 00:08:07 have we even loosened it at this point, any products that effectively align with an OCP white paper, perhaps those things can get the OCP accepted badge and sit in our marketplace. And our marketplace also includes more integrated solutions as well. We might talk about that later as we probably talk about the hedgehog. Yeah, yeah, we'll talk about it. Okay, so if I've got something I want to contribute to the community, I want to have that solution in the OCP marketplace, what's the process for getting accepted for a reference architecture in particular?
Starting point is 00:08:40 Because we're going to talk about reference architecture. Yeah, exactly. So the reference architecture contribution or any contribution to the OCP goes through a standard process. We've got a few staff that help people through the process and a lot of it's automated. You sign the contributor license agreement, which is Open Web Foundation, I think, 0.9 agreement. OCP doesn't acquire any intellectual property as part of that. It stays with the contributors. But they do sign that saying that it's effectively going to be opened and cataloged.
Starting point is 00:09:12 in the OCP contribution database. Then there's, obviously, if the work has already been done, it gets presented within the relevant project at OCP, and then it's presented to the steering committee. The steering committee is basically one representative on that committee per project. So it's a group of experts across the data center and obviously technical experts. And they, in turn, with our principal engineer at OCP, Russ, they basically make sure that the specification looks good.
Starting point is 00:09:44 And once they rubber stamp it, then it's free to go into our contribution database. And at that point, then anybody that builds a product that uses that specification, they're eligible to have that product in the OCP marketplace. So the contribution database and the marketplace upstream, downstream of the same overall process. Cool. Awesome.
Starting point is 00:10:07 Okay. Thanks for that background. So I want to share. shift to the topic of, it's a hot topic right now, it's disaggregated inference. So what is disaggregated inference and why is it important to the future of AI? Yeah. So I mean, fundamentally, the servers aren't necessarily ubiquitous anymore. We have known about the difference between training and pre-training and inference for a long time in the AI conversations. But over the past a year with these specialized chips, whether it's like GROC or Serbress or others, more and more
Starting point is 00:10:44 we have specialization within the inference workload. If you are, for example, uploading a document to chat GPU to your cloud, you're providing a ton of input tokens. That's part of the pre-fill phase. And then actually where you're getting a lot of output tokens doing the work, doing coding, and whatever have you as part of your request, your prompt. To fulfill it, that's basically evoking the decode phase. There's different performance characteristics that be optimized for there. So pulling apart the hardware architecture to match those two different phases, what the disaggregation is effectively doing with different memory technologies
Starting point is 00:11:25 and different accelerators or GPUs, XPUs. Got it. So the general idea is that I use two different AI accelerator. to get faster, more efficient inference. Exactly. Yeah. One through the pre-fill, one through the decode. Cool performance.
Starting point is 00:11:43 So then that means really that we should expect AI cloud builders going forward to have more and more diversity in their data centers of different AI accelerator types, right? Yeah, absolutely. There's a whole bunch of reasons driving diversity. So part of it is what we just talked about, disaggregated inference. Obviously, people that are doing training in inference,
Starting point is 00:12:04 may do that on different platforms. When you hear about the frontier models getting released, you'll often hear, oh, such and such, Chaggbti 5.5 was trained on Hopper, but then it was tuned on GV200. So there's things like that, and obviously they'll have different infrastructure for inference to run it and optimize for pre-fill and decode.
Starting point is 00:12:25 Besides that, I think GPU infrastructure is not going away as fast as people thought. Because there's still a lot of A100 stuff, around that people are not decommissioning because at this point, there's so much more demand than there is supply. Right. And then obviously, AMD, with their chips, are coming along, getting deployed massively. There's tons of announcements. Obviously, we heard VDio acquired Rock and we're going to be deploying that. Jets's last GTC was showing five different types of wrecks that basically constitute an AI factory. So even if you're allie deploying in video,
Starting point is 00:13:04 Accelerators are still to play lots of different generations and lots of different variants. Yeah, and I think we saw Cerebrus IPO last week. Yeah, IPO price was, or valuation was just under 100 billion. I think only META and Alibaba had exceeded that in IPO prices. That's an example as well of an XPU that might be in an AI data center alongside invidia GPUs, yeah? Yeah. Cool.
Starting point is 00:13:27 Okay. So then that leads me to the next question, which is the OCP OpenClU. There's a specification for open clusters. There's another specification for open pods. What is a cluster? What's a pod? Why do they need to be open? Yeah.
Starting point is 00:13:44 So pod is basically a building block of a cluster, right? A cluster for most people, it's the distributed, but acting as one computer for potentially a single problem. If you're training on a cluster, often using the entire cluster. Otherwise, the cluster is basically your infrastructure, that maybe used for different workloads. Obviously, if you're somebody like a hyperscaler or a neocloud
Starting point is 00:14:09 that's providing that as a service, then you've got multi-tenant concerns on that shared infrastructure and you're dividing it up and selling it. So I think what a cluster is is well understood. POT is a building block of that. People think of even a sub-element of the pod as sort of the scale-up domain, I suppose. But to answer your question about the open bits of it
Starting point is 00:14:31 that have started within the OCP, OCP traditionally people built specifications for specific point technology needs, like a server, a disk drive, or power supply, all these different things. And one of the hard problems that people were facing when they're trying to build AI clusters is that they're reinventing the wheels, solving the same problems again and again. And there's no point of designing a snowflake for every AI clusters. So a bunch of collaborators created a project, specifically we called it a strategic initiative within the OCP,
Starting point is 00:15:08 and it's got a long name. It's a mouthful open cluster design for AI. And they created a specification for a pod, and the pod is a building block for the open cluster. So they've effectively open sourced an architecture or a design of how to build in a pod multiple racks. They actually started with air-cooled and 400-gid connect. RACs. Now they're working on liquid-cooled 800-gig scale-out connected racks. And there's a
Starting point is 00:15:37 specification for exactly how that pot is designed so that you know everything from all of the cables that you would need, the number of racks, the exact number of compute and accelerators within that. And they tried to do it in a way that multiple vendors can kind of slot in and fill the different roles for like servers and switches and racks. Obviously, wherever they can. They base those decisions on other OCP specifications. But yeah, now somebody can take that and build an AI cluster from it. So it's basically everything that's on the IT side. It does include like the OT reference architecture, how to build your floors or how to build your overhead designs and space like that. But in terms of the row layout of multiple racks and everything
Starting point is 00:16:24 within them and how to connect them, how to configure them, that's basically what it is. So if I want to build an AI cloud that say has NvidiaGPUs, AMDGPUs, Google TPUs, maybe some cerebrous accelerators, each one of those accelerator types or XPU types would have its own pod and they'd all work together in the same cluster. Is that the general idea?
Starting point is 00:16:49 Yeah, I suspect that most people would probably deploy them in different pods. When it comes to the disaggregated inference bits, I suppose you could have the pre-fill of the decode kind of working together. They could be deployed in the single pod. At this point, the architecture doesn't have that in its specification. But is that pre-fill decode desegregation is a fairly new idea, right? I think it's less than a year old.
Starting point is 00:17:15 Yep. So let's get into pre-fill and decode for a minute. Let's assume I'm using multiple accelerators or different accelerators for the pre-fill and the decode. They need to share memory. so they need remote direct memory access on a network. Is that going to work if each XPU vendor shows up with their own proprietary, vertically integrated network stack? And I guess the second part of that question is,
Starting point is 00:17:42 is there a reason to have an OCP reference architecture for open AI networking that enables those diversified clusters to interoperate? Yeah. Well, I mean, to go back to one other words I said before, evolvability is really important, right? The answer is not really. You may start with a network that might work today, but as soon as you start adding different accelerators, if those vendors have different opinions on what networks you should be deploying, that's kind of a non-starter. You want your infrastructure to be able to talk to each other. So obviously the pre-fell of the decode
Starting point is 00:18:16 infrastructure has to be able to talk to each other to work together for a common influence workload. So that's absolutely true. But there's other times when you'll have infrastructure that You just simply want to all work together. And the network is what stitches everything together, right? There's the old saying the network is the computer, and that's especially true for AI clusters. So it's a non-starter for you to have different silos, so to speak, in your network. You want your network to ubiquitously connect up whatever you decide to deploy next.
Starting point is 00:18:47 We're just a few weeks past the OCP, the Mia Summit, where Hedgehog announced our contribution for this OCP reference, architecture for AI networks. There's actually two RAs, one for training, one for inference, and you and I were on a panel talking about it. So if somebody listened to this show as an AI cloud builder and they want to find those OCP reference architectures, how do they find them? They're in the OCP contribution database. An AI query or Google search would definitely turn them up, I'm sure, as well, but a prominent part of the OCP website, opencomputt.org, is the contribution database. You can go to opencomputt.org slash contributions.
Starting point is 00:19:25 Because they're fairly new, they're probably sitting right at the top. But if you typed in reference architecture, you'd find them very quickly. Also, there's ways to filter things by what project an initiative they fall into. So that open cluster design initiatives that we talked about is kind of what these specific reference architectures you guys contributed are tagged under. And I think at the same time at that show, you also had a contribution from OpenAI on MRC for, like, really big training clusters. Can you talk about that a little bit? Multipath, Reliable connection.
Starting point is 00:20:00 Yeah. So sort of another method of building an AI network for really large training clusters. It's a host-driven protocol to optimize and to allow for the network to be a little bit dark, so to speak, or to allow the high performance needs and the optimizations that you need in an AI back-end network without implementing all of those optimizations in the fabric itself. Again, because you want your fabric to effectively be ubiquitous, to be able to connect whatever you want. So yeah, OpenAI, Microsoft, AMD as well, for that matter, and Nvidia and Intel work together on the MRC spec.
Starting point is 00:20:39 Yeah, so that is an example of a spec that's a PDF in our contribution database. The reference architectures you guys contributed to did in GitHub, right? And I know that in the panel we saw people can create one of those reference architectures in a virtual way, too, if they don't even have their cluster today, there's ways they can play around and actually see what it would look like. So, James, you're, like, in a really interesting position where you just have visibility on what the world's biggest hyperscalers are doing,
Starting point is 00:21:09 what the world's largest infrastructure providers are doing for those hypers, for neoclodds, for enterprises, across all the different layers of the AI infrastructure stack. What's coming in the future? What are we going to be talking about, next year at OCP Global Summit and what should people be thinking about? So some of these things will absolutely see continued adoption of them. So we absolutely love having adoption.
Starting point is 00:21:37 We have a full adoption track at all our OCP summits. So hearing about the success stories of these things getting deployed, open rack wide was announced at last Global Summit and the specifications available, but I know we're going to start seeing open rack wide flavors from, All of the usual suspects, right? So the Helios from AMD, from HPEE, some from Super Micro, I'm sure. Those are the double wide racks that people are talking about.
Starting point is 00:22:06 And those scale of domains have a scale out that needs to be compassionate with everything else that you're deploying, right? You're not going to deploy ubiquitous open rack wide. You're going to deploy that with GV200, 300 racks, VR racks in the future from Nvidia, and then, yeah, probably some surrogous, et cetera. So I think you're going to continue to see an expansion of the flavors of AI infrastructure that's deployed.
Starting point is 00:22:34 Hopefully, we're going to see some hardening of the decisions on the facility level so that the facility can be a little bit more fungible for these sort of like, as the open letter at OCP says, late-binding decision-making on all of the evolving IT specifications. A court that houses the IT, right? Is that what you're talking about? Yeah, exactly. Yeah, a big part of the OCP that is not my wheelhouse.
Starting point is 00:23:04 So as you introduced me, I'm a network walk, but learning about everything that's happening in liquid cooling, whether it's called plate or immersion, whether single phase, two phase, all of the stuff that's happening for low voltage direct client at 800 or they're talking about how do they go to 1500 in the future. We just had a workshop on that. So seeing the new busways and bus bars for some of this higher voltage, but still LVDC stuff is fascinating. So, yeah, there's things happening all around.
Starting point is 00:23:34 If you're on the networking side at OCP, there's a brand new optic circuit switching project for OCS. There's tons of stuff happening to try to drive more stability and reliability of the optics. Got a new work stream there. That one bites us all the time. It's a lot of people all the time. But yeah, we're deploying millions of these things. It turns out that, yeah, if you have even a small nugner of them, their fault,
Starting point is 00:24:01 you know, yeah, reliability in these huge clusters because there's thousands of moving parts. So I woke up this morning, checked my email. I had a survey request in my inbox from OCP asking if you should move the location of OCP Global Summit. So I went last year to San Jose Convention Center. Of course, we go to Nvidia GTC every year. I think OCP Global Summit last year was the only show I've been to. We do a lot of shows, like probably two a month, that felt as big. It's not quite as big, but maybe number two to Nvidia GTC.
Starting point is 00:24:38 Is the community growing? And is it too big now for San Jose? It's growing massively. Some data points on that. So the recent contribution for LVDC was the most collaborative contribution we've ever received. It had 70 different companies and hundreds of people that worked on this. Amazing. Obviously, I said the OCP has grown in scope to all things data center.
Starting point is 00:25:02 And with that, all of the OT vendors and solution providers are joining OCP. That's another way that OCP is growing massively. But yeah, although a lot of the work happens virtually in projects on calls and mailing, lists. One of the ways that everybody comes to an OCP is by coming to one of the summits, the conferences. And the biggest one is the global summit happens every October. And it's happened in San Jose, every October at San Jose Convention Center. But it's getting to the point where we're busting at the seams of the convention center. Yeah. You know, over 11,000. We're probably going to be 13, 14,000 this year. So we're like approaching GTC levels. I don't think we're going to
Starting point is 00:25:42 have the shark tank for our keynotes. It's pretty tough. I think Taylor Swift's probably the only person on the planet who could top Jensen in a keynote. Yeah. James, thank you so much for talking with me today. Thank you so much for being a part of the OCP community and welcoming Hedgehog into it. Honestly, when I first started thinking about this problem space that we're solving at Hedgehog, I was still working at Cisco. And I was talking about it with my boss. He's like, you know what you should do, just go to the OCP Global Summit and check that out and then come back to me and report on Open Networking. So that was really the genesis of Hedgehog was attending that event and forgetting now how many years ago. It's a long time ago now, but we've been to every single
Starting point is 00:26:25 global summit since and they're fantastic events. I encourage everybody listening to this to go to an OCP event, be a global summit or one of the regional summits, because I guarantee you we'll learn a ton about AI infrastructure, which is a growing. community for a very rapidly growing market. So thanks, James. Yeah, thanks for having me. It's a pleasure.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.