Latent Space: The AI Engineer Podcast - DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever

Episode Date: October 7, 2025

At OpenAI DevDay, we sit down with Sherwin Wu and Christina Huang from the OpenAI Platform Team to discuss the launch of AgentKit - a comprehensive suite of tools for building, deploying, and optimizi...ng AI agents. Christina walks us through the live demo she performed on stage, building a customer support agent in just 8 minutes using the visual Agent Builder, while Sherwin shares insights on how OpenAI is inverting the traditional website-chatbot paradigm by embedding apps directly within ChatGPT through the new Apps SDK.The conversation explores how OpenAI is tackling the challenges developers face when taking agents to production - from writing and optimizing prompts to building evaluation pipelines. They discuss the decision to adopt Anthropic’s MCP protocol for tool connectivity, the importance of visual workflows for complex agent systems, and how features like human-in-the-loop approvals and automated prompt optimization are making agent development more accessible to a broader range of developers.Sherwin and Christina also reveal how OpenAI is dogfooding these tools internally, with their own customer support at openai.com already powered by AgentKit, and share candid insights about the evolution from plugins to GPTs to this new agent platform. They discuss the surprising persistence of prompting as a critical skill (contrary to predictions from two years ago), the challenges of serving custom fine-tuned models at scale, and why they believe visual agent builders are essential as workflows grow to span dozens of nodes.Guests:* Sherwin Wu: Head of Engineering, OpenAI Platform https://www.linkedin.com/in/sherwinwu1/ https://x.com/sherwinwu?lang=en* Christina Huang: Platform Experience, OpenAI https://x.com/christinaahuang https://www.linkedin.com/in/christinaahuang/Thanks very much to Lindsay and Shaokyi for helping us set up this great deepdive into the new DevDay launches!Key Topics:• AgentKit launch: Agent SDK, Builder, Evals, and deployment tools• Apps SDK and the inversion of the app-chatbot paradigm• Adopting MCP protocol for universal tool connectivity• Visual agent building vs code-first approaches• Human-in-the-loop workflows and approval systems• Automated prompt optimization and “zero-gradient fine-tuning”• Service Health Dashboard and achieving five nines reliability• ChatKit as an embeddable, evergreen chat interface• The evolution from plugins to GPTs to agent platforms• Internal dogfooding with Codex and agent-powered supportFull Video EpisodeTimestamps00:00 Welcome to the OpenAI Dev Day Studio01:11 Dev Day Evolution and Community Growth03:08 Apps SDK and ChatGPT Distribution Strategy05:27 MCP Protocol Integration Decision09:26 Agent Kit Launch and Platform Vision11:33 Agent Builder Canvas and Visual Workflows17:22 Evaluations and Agent Testing Evolution19:20 Automated Prompt Optimization and Research26:35 Connector Registry and MCP Servers34:10 Chat Kit as Consumer-Grade Infrastructure39:13 Codex Power User Tips and AI-Native Development42:27 Service Health Dashboard and Reliability Journey This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

Transcript
Discussion (0)
Starting point is 00:00:04 Hey, everyone. Welcome to the Late in Space podcast. This is Alessio from the R kernel Labs, and I'm joined by Swix, editor of Layden Space. Hello, hello, and we are here in the Open AI Dev Day studio with Sherwin and Christina from the Open Eye Platform team. Welcome. Thank you for having us. Yeah. It's always... It's such a nice thing.
Starting point is 00:00:24 We've been, we've covered like three of these Dev Days now. And this is like the first time it's been like so well organized that we have our own little studio podcast studio in the Dev Day venue. And it's really nice to actually get a chance to sit down with you guys. So thanks for taking the time. Yeah, I feel like we, Dev Day is always a process. And like, we've only had three of them and we try to improve it every time.
Starting point is 00:00:44 And I actually, I know for a fact that I think we have this podcast studio this time because the podcast interviews and the interviews and the interviews with folks like yourselves last time went really well. And so I want to lean into a little bit more. I'm glad that we were able to have this studio for you all. We were kneeling on the ground interviewing like Michelle last year. I fell in the living here.
Starting point is 00:01:01 I just saw it post production. I thought it was. We had to have people like. cordoned off the area so they wouldn't walk in front of the cameras. People just come up, hey, good to, I'm like, we're like recording. I guess if you guys have been to three, like what, what stood out from today or what, what's your favorite part? I feel like the vibes are just a lot more confident.
Starting point is 00:01:20 Like, you are obviously doing very well. You have the numbers to show it. You know, I just, every year in death day, you report the number of developers. This year is four million. I think last year was like three. And I have more questions about. that kind of stuff. But also like just like very interesting, very high confidence launches.
Starting point is 00:01:41 And and then also like I think that this is the community is clearly much more developed. Like I think there's just a lot more things to dive into across the API surface area of OpenEI than I think last year in my mind. I don't know about you. Yeah. And we were at the OG Dev Day, which was the Dali Hacknight at OpenAI in 2022. And I think Sam spoke to like 30 people. So I think it's just crazy to see the...
Starting point is 00:02:07 Yeah, honestly, I think it's like, it's kind of similar to this podcast studio, which is I think we've had a number of dev days now. We honestly were like slowly figuring things out as a company over time as well, and both from a product perspective and also from a like how we want to present ourselves with Dev Day. And at this third only, at this point, we've had a lot of feedback from people. I actually think a lot of the attendees you'll get like an email with like a chance for feedback as well. And we actually like do read those and we act on those. And like one of the things that we did this year that I really liked were all of those,
Starting point is 00:02:34 There was like some art installations and like the little arcade games that we did, which was, you know, came up with, via like engaging with the feedback from the game. Yeah, the arcade games were so fun. I loved like the theme of all the ASCII art throughout. This is my first SF dev day. But I've been to the Singapore one. That was actually my first week. Oh, yeah, that's the one I spoke. Yeah, I saw you there.
Starting point is 00:02:54 That was my first week of Open AI. So really in the defense. Put around a plan to Singapore. Yeah. Yeah, that's awesome. Well, so, you know, that's congrats on everything. And like, kudos to the organizing team. We should talk about some developer API stuff.
Starting point is 00:03:07 Yeah. So we're going to cover a few of the things. You're not exactly working on apps SDK, but I guess what should people just generically take away? What should developers take away from the apps SDK launch? Like, how do you internally view it? So the way that I think about it is I actually view Open AI since the very beginning as the company that is really valued, kind of like opening up our technology and like
Starting point is 00:03:31 bringing it out to the rest of the world. One thing we talk about a lot internally is, you know, our mission at Open AI is to, one, build AGI, which we're trying to do. But two, you know, potentially, you know, just as important is to bring the benefits of that to the entire world. And one thing that we realize very early on is that we as a company, it's very difficult for us to just bring it to every, truly every corner of the world. And we really need to rely on developers, other third parties to be able to do this, which is, you know, Greg talked about the start of the API and like kind of how, you know, that was formulated. But that was part of, you know, that mentality, which is we needed to rely on developers and we need to open up our technology to the rest of the world so that they can partake for us to really fulfill our mission. So the API obviously is a very natural, you know, a way of doing that where we just literally expose API endpoints or expose tools for people to build things. But now that we have, you know, chat to BT with its, I don't know, like 800 million weekly active users.
Starting point is 00:04:25 I forgot the stat that we share. I think it's like now the fifth or like sixth largest website in the world. And the number one and number two, most downloaded on the Apple App Store. Oh, yeah, with Sora. Yeah, but that one, like, it moves around all the time, so it's kind of hard to celebrate. You just screenshot it when it's good. Yeah, yeah, we definitely screenshot it when it was good. But kind of going back to my main point is, like, we've always kind of engaged with developers
Starting point is 00:04:50 as a way for us to bring the benefits of AGI to the rest of the world. And so I view this is actually a natural extension of this. Candidly, we've actually been trying to do this, you know, a couple of times with the last dev day with GPTs, two dev days ago with, I'm sorry, two devs ago with GPs and plugins, which was, I think, not tied to a dev day. So I view this as like, again, we love to deploy things so iteratively. And I view it as like just a continuation of that process and also engaging deeply with developers and helping them benefit from some of the stuff that we have, which in this case is chat GPT distribution. And when, so apps has the case built on the MCP protocol.
Starting point is 00:05:27 when did OpenEAAid become MCP-pilled? I'm sure internally you must have had, you know, designed discussions before about doing your own protocol. When did you buy into it? And how long ago was that? I think it was in March, I want to say. It's hard for me to remember kind of like the exact. March was the takeoff of MCP.
Starting point is 00:05:44 Okay, yeah, yeah. So we built the agents SDK and we launched that alongside the responses API in early March. And I think as MCP was growing, that felt like a really, and, you know, we're building kind of a new agentic API that can call tools and just be much more powerful. MCP was kind of like the natural protocol that developers were already using to bring all the tools into their system. And I think like in March is when we added an MCP to agents SDK first and then soon after
Starting point is 00:06:10 with kind of our other products. Yeah, I think there was like a tweet or something we did. There was like opening I, you know, is. Yeah, there was definitely a moment. I think there was a specific moment in a specific tweet. But what I will say though is like, and this is honestly that credit to the team at Anthropic that kind of created MCP is I really do think they treat it as an open protocol. Like, we work very closely with, I think, like, David and the folks on the, like, you know, consortium.
Starting point is 00:06:33 And they are not, you know, really viewing it as this, like, thing that is specific to Anthropic. They really view it as this open protocol. There is, like, it is an open protocol. The way in which you make changes feels very open. We actually have a member of our team, Nick Cooper, who is sitting on kind of like that steering committee for MCP as well. And so I think they are really treating it as something that is easy for us and other companies, you know, everyone else to embrace, which I think they should because they do want it to be something that is very embraced by all. And so because of that, I think it makes it a little bit easier for us to embrace it.
Starting point is 00:07:05 And honestly, it's a great protocol. It's a very general. It's already solved. Why would you make it? Yeah, yeah, it's very general. There's obviously still more to do with it. But it was very easy for us to, you know, integrate because of how streamlined and how simple it was. Yeah.
Starting point is 00:07:18 My final comment on apps SDK stuff and then we'll move to Agent Kit is, you know, like, I always see like in abstractly when you sort of watch. wireframe a website or an AI app. It used to be that the initial AI integration on the website would be you have the normal website and then you have a little chatbot app. And now it's kind of like inverted where there's chat GBT at the top layer and then it's like to know the website embedded inside of it. And it's kind of like that inversion that I honestly have been looking for for a little bit.
Starting point is 00:07:48 And I think it's really well done. Like actually all like the integrations and the custom UI components that come up, you had like Canva on the keynote there, and it looks like Canva, but like you can chat with it in all your, the context of your chat GBT. That is an experience I've never seen. Yeah. And I think that's kind of back to the iterative like learning that we've had. That I think was because we've learned a lot from plugins. So like when we launched plugins, I remember one of the feedback that we got. I don't know if, you know, if people here really remember plugins, it was like March 23. Yeah. But like one of the points of feedback was like, oh, you can integrate, we tell, we told like,
Starting point is 00:08:23 you know, all these companies that you can integrate these plugins in a chat GPT, but they really didn't have that much control over how exactly it was used. It was really just like a tool that the model could call. And you were just like really bound by a chat CBT. And so I think like you can kind of see the evolution of our product with this. And like this time we realized how important it was for companies for third-wide developers to really own and like steer the experience to make it feel like themselves, to help them, you know, like really preserve their own brand.
Starting point is 00:08:47 And so, and, you know, I actually don't think we would have gotten that learning had we not, you know, had all these other steps. beforehand. Awesome. Christina, you were to start today on stage with the Agent Kit demo. You had eight minutes to build an agent. You had a minute to spare and then you have some issues. Yeah, I wasn't sure.
Starting point is 00:09:06 Honestly, I was like, let's do a little bit less testing and maybe we, I don't know how much time I killed on the, on the widget. Yeah, I was stressed out. I was stressed out. If a UI bug is what like takes the demo down and be so sad. I think it was a full screen, yeah, like focus. I heard the window wasn't in focus or something. Yeah. Maybe you want to introduce Agent Kit to the audience.
Starting point is 00:09:26 Yeah, so we launched Agent Kit today. Full set of solutions to build, deploy, and optimize agents. I think a lot of this comes from working with API customers and realizing how hard it actually is to build agents and then actually take them into production, hard to get kind of that confidence and the iterative loop and writing prompts, optimizing them, writing evals, all takes a lot of expertise. and so kind of taking those learnings and packaging them into a set of tools that makes it a lot easier and kind of intuitive to know what you need to do. And so there's a few different building blocks that can be used independently, but they're kind of stronger together because you then get the whole end-to-end system and releasing that today for people to try out and see what they build. Yeah, so I find it hard to hold all the building blocks in my head. But actually chronologically, it's really interesting that you guys started out with the agent SDK first.
Starting point is 00:10:23 And then you have agent builder. You have a connector registry. You've chat kit. And then you have the Eval stuff. Am I missing any major components? Those are the main moving parts, right? Yeah, I think that's it. And then, I mean, we also still have like the RFT, like fine-tuning API.
Starting point is 00:10:38 But we technically group it outside of the agent kit umbrella. Got it, got it, got it. Yeah. So, like, it's weird how it develops. and it's now become the full agent platform, right? And I think one thing that I wasn't clear about when I was looking at the demo was, it's very funny because what you did on stage was build like a live chat app for Dev Days website. Yeah, did you get a chance to try it out?
Starting point is 00:11:03 Yeah, it was awesome. And actually I kind of wanted to ask like how to deploy. Where's merch? Yeah, exactly. I was like, where did you click the merch? Anyway, and this is very close to home because I've done it for my conferences. and like it's it's a very similar process but like um i think what it was not obvious is like how much is going to be done inside of agent builder i see there's some actually very interesting
Starting point is 00:11:24 nodes that you didn't get to talk about on stage like user approval that's like a whole thing and uh you know like transform and set state like there's there's like a kind of like a touring complete machine in here yeah yeah so i mean i think again like this is the first time that we're showing agent builder and so it's definitely the beginning of what we're building and um Human approval is one of those use cases that we want to go pretty deep on, I think. The node today that I showed is pretty simple, like binary approval. It's similar to kind of what you'd see for MCP tools, of approving that an action can take place. But I think what we've seen with much more complex workflows from our users is that it's actually quite advanced, like, human-in-the-loop interaction.
Starting point is 00:12:05 Sometimes these could be over the course of weeks, right? It's not just kind of simple approval of the tool. There's actual decision-making involved in it. And I think as we work with those customers, we definitely want to continue to go deeper onto those use cases too. Yeah. What's the entry point? So are developers also supposed to come here and then do the two code export, like just segment like the use cases? Yeah.
Starting point is 00:12:31 So I think the two reasons that you would come to Agent Builder are one kind of more as a playground, right, to kind of model and iterate on your systems and write your prompts and optimize them and test them out. and then you can export it and run it in your own systems, using agents SDK, using kind of, you know, other models as well. The second would be kind of to get all of the benefits of us deploying that for you, too. So you can kind of use maybe like natural language to describe what type of agent you want to build, model it out, bring in subject matter experts so that you really have this canvas for iterating on it and getting feedback, you know, building datasets and kind of getting feedback from those subject matter experts as well. And then being able to deploy it all without needing to handle.
Starting point is 00:13:13 that on your own. And that's a lot of the philosophy around how we're building it with chat kit as well, right? You can kind of take pieces of it. You can have a more advanced integration where it's much more customized. But you also get a really natural path of going live without like with really kind of easy defaults as well. Yeah. Do you see it as a two-way thing? So I build here, I go to code, then maybe I make changes in code and then I bring those changes back to the agent builder. Eventually, like that's definitely what we want to do. do. So maybe you could start off in code. You could bring it in. We'll also probably have like ability to, you know, run code and in the agent builder as well. And so I think just a lot of
Starting point is 00:13:53 flexibility around. The one thing I'd say, too, is a lot of the demos that we showed today, I think we're like, you know, aired on the side of simplicity just so that the audience could kind of see it. But like if you talked to a lot of these customers, like they're building like pretty complex. Like you got to like zoom out on that canvas quite a bit to kind of like see the full flow. And that and then for us, you know, we were kind of like working with a lot of customers who were doing this. And then, you know, if you turn that into like an actual agent's SDK like file, it's like pretty, it's pretty long. And so we saw a lot of like benefit from having the visual setup here, especially as the as the setup grows grows longer and longer. It would have been a
Starting point is 00:14:25 little difficult to kind of showcase this. But even on like some of the, right, yeah, it can do it in eight minutes. But like even with some of the presets that we have on yeah. So one of the things. Yeah. One of the things that, um, we launched today as well alongside just like the canvas is a set of templates that we've actually gathered from our engineers who are working in the field with customers directly of like the kind of common patterns that they have in our own basically like playbooks when we're working with customers on customer support, document discovery. And so kind of publishing those as well.
Starting point is 00:14:54 Data enrichment, planning helper, customer service, structured data Q&A, document comparison. That's nice. Internal knowledge assistant. Yeah. Yeah. And I think like we just plan to add more to those as we can kind of build those out. I always wonder if there should be. So we're not the only agent builders.
Starting point is 00:15:10 But obviously by default of being an open AI, you are a very significant one. Any interest in like a protocol or like interrupt between different open source implementations of this kind of pattern of agent builder? I think we've thought about it, especially around, I'd say, agents SDK. I would actually say maybe even like zooming out a bit more from just this is like, yeah, we were like, we're also sitting here and kind of like observing like things being made over and over again. Even like besides like agent workflows, we're kind of want. what the industry is trying to do with responses, like what we've done with responses API, like stateful APIs. And so, you know, obviously we were the first one to launch responses API, but like a couple of other people have kind of adopted. I think I think GROC has it in their
Starting point is 00:15:53 API. I think I saw LMSS just at something you're seeing walls, but not, you know, not everyone. And so unfortunately, I don't have a great answer today of like yes or no, but we are kind of like assessing everything and trying to see like, hey, you know, there has been a lot of value with MCP, with, hopefully with our, with our commerce protocol as well. ACP, yeah, it's, I definitely did not forget the name. And so, like, even thinking about, like, what we want to do with agents, with the agent workflow, the portability story around that, as well as the portability, I'd say even of, like, responses API would be great if, you know, that could be a standard or something, and developers
Starting point is 00:16:32 don't need to, you know, like build three different stateful API integrations. if they want to use different models. Yeah, and I think that's one of the, so it's not exactly a protocol, but one of the things that we launched today with Eval's too is ability to use like third-party models as well and kind of bringing that into one place. And so I think definitely kind of see where the ecosystem is at, which is, you know, using multi-models and kind of having...
Starting point is 00:16:55 Third-party models is in non-open-open-air models? Yeah, yeah. It'll work with E-VALs starting today. Okay, got it. We have a really cool setup with open router where we're working with them, and then you can bring your open-router. setup. And then with that, you can actually, you know, you write your evils using our data sets tool or user dataset tool to create a bunch of evals. And you'd actually be able to hit a bunch
Starting point is 00:17:16 of different model providers, you know, take your pick from wherever, even like open source ones on together and see the results in our product. Yeah, that's awesome. Speaking more about eVals, right, like I think I saw somewhere in the release docs that you basically had to expand the evils products a little bit to allow for agent evils. Maybe you can talk about what you had to do there. Yeah. Yeah, I was going to say, so I actually think Asian evils is still a work in progress. So I think we've made maybe 10% of the progress that we need here.
Starting point is 00:17:52 For example, I think we could still do a lot more around multimodal evils. But the main progress that we made this time was kind of allowing you to take traces. So the agents SDK has like really nice traces feature where if you run, if you define things, you can have like a really long trace, allowing you to use that in the Eval's product and be able to grade it in some way, shape, or firm over the entire entirety of what it's supposed to be doing. I think this is a step one. Like, I think it's good to be able to do this. But I think our roadmap from here on out is to, you know, really allow you to break down the different parts of the trace and allow you to eval and like kind of like measure each of those and optimize each of those as well. A lot of the times this will involve human in the loop as well, which is why we have the human in the loop component here too. But if you kind of look at our Eval's product over the last year, it's been very simple.
Starting point is 00:18:41 It's been much more geared towards this like simple prompt completion setup. But obviously as we see people doing these longer gentic traces, like, you know, how do you even evaluate a 20-minute task correctly? And it's like it's a really hard problem. We're trying to set up our Evalds product and move in that way to help you not only evaluate the overall trajectory, but also individual parts of it. Yeah. I mean, the magic keyword is Rubrics, right? Everyone wants LMS judge Rubrics. Yeah, yeah, yeah. Obviously, where this will go. Okay, great. The other thing I think online, I see the developer community, very excited about is sort of automated prompts optimization, which is kind of e-vails
Starting point is 00:19:17 in the loop with prompts. What's the thinking there? Where's things going? Yeah, so we have automated prompt optimization, but again, like, I think this is an area that we definitely want to invest more in. We, I think did a pretty big. a launch of this when we launched GPD5 actually because we saw that it was pretty difficult as new models come out to kind of learn all the quirks about a new model. Yeah, the prompts. Right. There's like, we have a big prompting guide, right, for every model that we launch. And I think building out a system to make that a lot easier. Um, we definitely want to tie that in like completely with evals. We should be able to kind of improve your prompts over time, improve your agents over time as well. They're kind of made in the agent builder based on the evils that you've set up. And so I think we see this as like a pretty core part of, of the platform of basically. suggested improvements to the things that you're building. I actually think it's a really cool time right now in prompt optimization. I'm sure you guys are seeing this too.
Starting point is 00:20:09 It's like not only there are a lot of products kind of like gearing around this, so like kind of what we're thinking about, but I also think like there's a lot of interesting research around this, like the data breaks folks are actually doing really cool stuff around. That's, we're obviously not doing any of the cool GEPA optimization right now in our product, but would love to do that soon. And also it's just an active research area. So like, you know, whatever Matei and the data,
Starting point is 00:20:31 Databricks folks might think about next, what we might, you know, think about internally as well. Whatever new prompt optimization techniques come out, I think we'd love to be able to have that in our product as well. And yeah, and it's interesting because it's coming at a time when people are realizing that prompt, you know, like, I feel like two years ago, people were like, oh, at some point prop, like prompting's going to be dead. No. Like, you know, and it's like, you know. It's gone up.
Starting point is 00:20:53 Yeah, yeah. Yeah. And if anything, it is like become more and more entrenched. And I think that, you know, there's this interesting trend where like it's becoming more and important and then there's also interesting cool working done to like further entrenched like prompt optimization. And so that's why I just think it's like a very fascinating, you know, area to follow right now. And also it was an area where I think a lot of us were wrong two years ago, because if anything, it's only gotten more important. Yeah, I would say like what, shouldn't you
Starting point is 00:21:20 used to work at opening? I know it was an MSL. We call this kind of like zero gradient fine tuning or zero gradient updating because you're just tweaking the prompts. But like, it is so much prompt that is actually, like, you end up with a different model at the end of it. There's a lot of, like, things that make it more practical, too, just like, even from our perspective, like, we, we have a fine-tuning API. And, like, it is extremely difficult for us to run, you know, and serve, like, all of these different snapshots. Like, you know, Laura's great, MSL just, you know, or sorry, Thinking Labs just, just published, John Schoomler just had a cool blog post about this. But, like, man, it is, like, pretty difficult
Starting point is 00:21:54 for us to, like, manage all of these different snapshots. And so if there is a way to, like, hill climb and, yeah, do this, like, zero, uh, gradient. like optimization via prompts. Like, yeah, I'm all for it. And I think developers should be all for it because you get all these gains without having to do any of the fancy, fancy fine-tuning work. Since you are part of the API, you lead the API team, and since you mentioned thinky, I got to throw a cheeky one in there.
Starting point is 00:22:17 What do you think about the Tinker API? Yeah, it's a good one. So it's actually funny. When it launched, I actually DM John Schulman. I was like, wow, we finally launch it. Because you used to work with him. Yeah. Yeah, yeah. So we, is that, it's actually funny. So at, yeah, so right when I joined Open AI, like, this has actually been, I think, a passion project of Johns. Like, he's been talking about doing something in this, like, in this shape for a while, which is like a truly, like, low-level research, like, fine-tuning library. And so we actually talked about it quite a bit when he was at Open AI as well. It's actually funny. I talked to one of my friends who said that when he was at Anthropic, he also.
Starting point is 00:23:00 you know, worked on this idea for a bit. He's a man on a mission. Yeah, I mean, John's, like, so great in this regard. He's, like, so purely just, like, interested in the impact of this because it's, one, it's like a really cool problem. And then, two, it also empowers builders and researchers. But you saw all the researchers who, like, express all this love for Tinker because it is a great, great product.
Starting point is 00:23:18 And so I'm just really happy to see that they shipped it. And I think he was really happy to kind of get it out there in the world as, as well. Yeah, this is probably, this is very much a digression. But, like, it's weird. as somewhat passionate about API design, that it took this long to find a good fine-tuning API abstraction, which is effectively all he wanted.
Starting point is 00:23:36 He was like, guys, like, I don't want to worry about all the infra. Like, I'm a researcher. I just want these four functions. And it's kind of interesting. Yeah. Yeah. Cool. Before the opening icons team barges in the room. I know. So what feedback
Starting point is 00:23:51 do you want from people like the agent builder? For example, the thing I was surprised by was the if-else blocks not being natural language and using the common expression language, I'm sure that's something already on your roadmap. What are other things where you're kind of like at a fork that you will love more input on? I think like one of the things that we spent a lot of time discussing
Starting point is 00:24:11 was like whether we want kind of more of like the deterministic workflows or more LLM driven workflows. And so I think like getting feedback on that, honestly having people model existing workflow. A lot of what we did was kind of work with our team on, especially with engineers who are working with customers, like modeling the workflows that already exist in the agent builder and like what gaps exist, like what types of nodes are really common and how can we like add those in? I think that would be
Starting point is 00:24:38 like the most helpful feedback to get back. And then as we expand kind of from just like chat-based, like right now the initial deployment for agent builders through chat kit, we plan on kind of releasing more standalone like workflow runs as well and kind of the types of like tasks that people would like to use in that type of API. So like more modalities, for example. Yeah, I mean, I think, like, for sure, like, more modalities. Like, you know, I think kind of voice would be, is already something that a lot of people have talked to us about, even today at Dev Day.
Starting point is 00:25:13 So I think modalities, for sure, but also more like the logical nodes of what can't be expressed today. Yeah. Well, you know, you're building a language, right? You have common expression language, which I never heard of prior to this. I thought it was this Python, this JavaScript, and then there was a whole link in there. Was that a big decision for you guys? I think that was more just kind of like a way that we thought we could kind of represent a mix of like the variables and I don't know, like conditional statements.
Starting point is 00:25:42 The other thing I'll also mention is that you let once you, so there's a trope in developer tooling where like anything that can be, that can store state will eventually be used as a database, including DNS. So to be prepared for your state store to become a database, I don't know if there's like any limits on that, because people will be using it. It's actually funny. I'd heard this quote before, and there's definitely some truth to it. I don't know if our stateful APIs have become a database, just quite yet, but like, who knows? Like, you know, I mean, conversations. Well, you charge for it. You charge for assistance. Storage, yeah. The storage. Right. So there's some limit on that, but like. Yeah, but it's very cheap. It's like, I remember we pressed it. I think if you wanted to kind of like dump all your data somewhere, I don't know.
Starting point is 00:26:24 This is like the most like transforming it all into this shape. It's useful. It's easy. It's the best place or whatever. But also please don't do this because I think it'll put quite a bit of strain on on Ventot and our info team and what we try and do. So, yeah. How do you think about the MCP side? So you have open AI first party connectors.
Starting point is 00:26:41 You have third party preferred, I guess, servers you will call them. And then you have open-ended ones. Do you see that part of registry like functionality? expanding or do you see most of it being user-driven? OTH is like the biggest thing. Like if you add Gmail and calendar and drive, you have to like ought each of them separately. There's not like a canonical odd. What's the thinking there?
Starting point is 00:27:04 Yeah, I mean, I think definitely for the registry, that's why we want to make it a lot easier for like companies to kind of manage what their like developers have access to, managing kind of the configurations around it. And I think in terms of like first party versus third party, like we want to support both of those. We have some direct integrations, and then anyone can kind of create MCP servers. I think we want to make that a lot easier to, like, establish kind of private links for companies to use those internally. So I think, like, just really excited about that ecosystem growing. Yeah. I think one of the coolest things observed, too, is just I actually think we,
Starting point is 00:27:38 we as an industry are still trying to figure out the ideal shape of connectors. So, I mean, part of why I think the 1P connectors exist, too, like we end up storing quite a bit of state. It's like a lot of work for us. But, like, by having a lot of social. state on our side. We call them sync connectors. We can actually end up doing a lot more creative stuff on our side when you're chatting with chatDB and using these connectors to kind of boost the quality of how you're using it. If you have all the data there, you can do all this re-ranking. You can like do we can put it in a vector store if you want to put it anywhere else. Whereas and so there's some inherent tradeoffs here where like you put in a lot of work
Starting point is 00:28:09 to get these like 1P connectors working, but because you have the data, you can do a lot more and get higher quality. But then but then the question is like, oh my God, there's like such a long tail of other things, which is where the MCP and, like, the third-party connectors come in. But then you have the trade-off of, like, you're beholden to, like, the API shape of the MCP creator. It might actually work well. It might not work well with the models. And then what happens if it doesn't work well, then you kind of have to, like, you know,
Starting point is 00:28:32 you're kind of like at the mercy of this. And MCP, by the way, is like really great because it already does some layer of standardization, but my senses are still going to be more evolving here. And I think, you know, we want to support both of them because we see value in both right now, especially working with developers you want to have kind of like all options kind of on the table here, but it will be interesting to see how this evolves over time. Yeah, when I saw about three, four months ago when you launched a forum for like signing with chat, GPT interest, I think to me that's kind of like the vision where I log in and I have the MCPs tied in and then I sign in which IDPD somewhere and I can run these workflows in that app where I'm logging in.
Starting point is 00:29:10 So yeah, I think Sam, you know, said in an interview that he's, chat GPT as your personal assistant. So I think this is like a great step in that direction. Yeah, I think there's a lot more to go in that direction. But so far, no plan on like chat GPT or opening I as IDP, right, which is a different role in the off ecosystem. Yeah, it's interesting because so direct answer is like no plans right now, of course. But I actually think we currently have some version of this, which is our partnership with Apple.
Starting point is 00:29:42 Because with Apple, you can actually sign in to your chat. IPD account and some of that identity does carry with you into your iOS experience with the area, right? Like if you, if you, I don't know if you've actually used this, the Sierra integration. I actually use it quite a bit, but if you sign into your chaty bt account, the Siri integration will actually use your subscription status to decide what type of model to use when it, when it passes things over to chat chbt. And so if you're, you know, just a free user, you get, you know, the free model. But if you're a plus or pro subscriber, you get routed to GPD 5, which is, I think, what they...
Starting point is 00:30:17 I think we also recently announced the partnership with cacao. Oh, yeah, cacao's another one. Yeah, where I think you... It's a similar thing where you can sign in with chat GBT. Cacao is one of the largest, like, messenger, yeah, absent in Korea, and kind of interact with Cacao directly there. Yeah, I mean, Sam's been talking about it for a while. It's a very compelling vision.
Starting point is 00:30:34 We obviously want to be very thoughtful with him and how we do it. You know, now you have a social network, you have a developer platform. Like, you know, my... At the beginning is a social network. Very, very valuable. Yeah. Yeah, exactly. Okay, so and then on the other side of off is something I was really interested to look at, and I couldn't get a straight answer.
Starting point is 00:30:50 Is there some form of bring your own key for Agent Kit? Like when I expose it to the wider worlds, obviously, like, I mean, by default, I'm paying for all the inference. But it'd be nice for that to have a limit. And then if you want more, you can bring your own key. Yeah. I mean, we don't have something like that yet. But I think, yeah, it's definitely an interesting area, too. Yeah, it doesn't do it out of the box today, but, you know, developers have been asking about it for forever.
Starting point is 00:31:19 Like, it's a really cool concept because then as a developer, you, especially indie developer, you don't need to bear the burden of inference. Yeah, I think, like, when you get into the business of, like, agent builders that are publicly exposed where you have, like, and allow list of domains. Like, this is, this is the, it rhymes with this exact pattern of, like, someone has to bear the cost it. Like, sometimes you want to mess around with, like, the different levels of responsibility. Yeah. I will say in general, like, if you kind of look at our roadmap, we engage a lot with developers. We kind of hear what are the pain points, and we try and build things that address it. And, you know, ideally we're prioritizing in a way that's helpful.
Starting point is 00:31:55 But, yeah, we've definitely heard from a good number of developers that, like, the cost is, or like all of the, like, copy paste your key, like solutions right now, which are, like, huge security issues, like hazards, because developers don't want to bear the burden of inference. You know, hopefully we make the cost cheaper. So it's a model's keep getting cheaper. Yeah, yeah. So hopefully, you know, that helps. But what we realize is as we make it cheaper, you know, the demand for that goes up even more and you end up, you know, still spending quite a bit.
Starting point is 00:32:19 But yeah, so we definitely heard this from a lot of developers. And it's definitely something top of mind. Do you see this as mostly like an internal tools platform, though? Like to me, like you've been doing a big push on like the more forward deployed engineering things. It's almost like, hey, we needed to build this for ourselves as we sell into these enterprises. Might as well open it up to everybody. What drives building these tools? Like you think of people building tools to then expose or mostly on the internal side?
Starting point is 00:32:45 Yeah. I mean, so like I think our, again, our first deployment is chat kit, which is kind of one of, it's intended to be for external users. But I think one of the things that we also did see a lot as we were working with customers is that a lot of companies have actually built some version of an agent builder internally to kind of manage prompts internally, to manage templates that they're sharing across, you know, the different developers that they have, maybe the different product areas. And we were seeing that kind of like over and over again as well and really wanted to like build a platform so that this is not, you know, an area that every company needs to invest in and like rebuild from scratch, but that they can kind of have a place where they can manage that these templates, manage these prompts and really focus on the parts of agent building that is more unique to their like business. It is interesting too. Like from a deployment perspective, it is like it has spanned both internal and external use cases, right? Like kind of like these internal platforms, people use it for like data problems. processing or something, which is an internal use case. But if you saw some of the demos today,
Starting point is 00:33:43 like there have been a huge number of companies that are trying to do this for external-facing use cases as well. Customer service is one temporary service. The like ramp use case. We use this internally and externally. Like our customer support, help to open outer.com, already powered on agent kit and then various other like internal use cases as well. And one of the things that I actually think the team has done a really great job of, so like Tyler, David and G1 on the team, they built the, especially the chat kit, components, they built it to be like very consumer grade and like very polished. Like you kind of look at that, there's like a whole grid of like the different widgets
Starting point is 00:34:16 and things that you could create there. Like ideally people see it and they see it as like these very polished like consumer grade ready, external facing things versus like, you know, you think of internal tools and like the UI is always, like the last thing that people care about. But like you really, you know, push the team and I think they did a really great job of making the chat kit experience like really, really consumer grade. And it should feel almost like chat GPT or, um, and with like really buttery. smooth animations and like really responsive designs and all of that. Yeah, I think your point on widgets is like definitely like really resonates, right? Because chat kit, it handles the chat ux, but we're also just building like really visual
Starting point is 00:34:53 ways for you to represent like every action that you want to take and that is definitely like very high polished. Yeah. And when working with customers, like those have been the most helpful customers for us to work with because, you know, when Ramp is thinking about, you know, how what what they want to publicly present to people, like they have a pretty high bar. as they should, as well as, you know, all the other customers that have been iterating on it. And so that kind of feedback from our customers has really helped us up level the general product
Starting point is 00:35:17 quality of the launch that we had today as well. Yeah. Would you ever, would you open source check it? Talked about it. We've talked about it. There are a bunch of tradeoffs. I think so check it itself is like an embeddable eye frame. And so I think the actual.
Starting point is 00:35:35 It was an eye frame. Yeah. And so that helps us keep it like evergreen, right? So if you are using ChatKit and we come up with new, I don't know, a new model that reasons in a different way, right, or a kind of new modalities that you don't actually need to rebuild and like pull a new components to use it in the front end. I think there's parts of, you know, widgets, for example, that is much more like a language and can definitely, um, is something that is easier to explore that for as well as kind of the design system that we've built, um, for ChatKit. Um, but I think like as part of, yeah, the actual eye frame itself, I think there's a lot of value in that being, more evergreen, more evergreen experience. That is a pretty opinionated.
Starting point is 00:36:13 Like there would be no point being open source. Right. You want to. Then you don't get the benefits of it. You know, being Stripe Alumni's like Stripe Checkout. Like it's also optimized for you to like. So I'm not a Stripe alum, but Christ. And the team actually is the team that built.
Starting point is 00:36:29 Stripe Checkup? Yeah. So it's very similar philosophically. Right. So Stripe, you know, can build elements and checkout and not every, business needs to rebuild, right, the pieces that are really common. And I think we see the same with chat. We see chat being built over and over again, especially as we kind of come up with new modalities, like reasoning, everything. It's not really something that is easy to keep up to date.
Starting point is 00:36:55 And so we should just do that. And we've kind of the hard parts of building agents again to the developers. Does it feel, I mean, I know WordPress is like a bad connotation in a lot of circles, but to me it almost feels like the WordPress equivalent of like chat. It's like, hey, this is like drop in thing. And then you have all these different widgets. Do you see the widget becoming a big kind of like developer ecosystem where people share a widget? Is that kind of like a first party thing?
Starting point is 00:37:23 And then what's like the MCP versus widget forest? No, exactly. I mean, it's kind of like, it seems great for people that are like in between being technical and like not really being technical enough. Yeah. Yeah. Yeah. I mean, I think that's a big part of building widgets, right?
Starting point is 00:37:37 Like it's already kind of in the language that is very consumer-friendly. You can use, in our widget builder already, you can kind of use AI to create those widgets, and they look pretty good. I don't know if you guys have gotten a chance to try that out yet, but definitely see kind of, I don't know, a forest. If you haven't tried out the widget studio and the demo, like apps as well, yeah. You got a custom domain like widget. Dot studio. Actually, don't know how we got that.
Starting point is 00:38:04 Yeah, everything's in Chackett. studio and then we have like the playground there so you can try out what chaka would look like with all the customizations we have check it dot world which is a fun site we built i was like spinning the globe for a while this morning it was um i think kasha also like uploaded some of her solar system stuff and yeah yeah all the demos as well yeah and then that's where like the widget um builder yeah so it's like it's really come together like it's taken like almost more than a year to like come together and like build all this stuff but it's coming together it's like really yeah yeah that we like like you definitely planned all of this up front oh yeah yeah yeah we have the master plan
Starting point is 00:38:40 from you know three years ago um no but like i think uh especially on this stuff i think there was like an arc of a general like you know platform that we did want to kind of build around and um it takes a while to build these things obviously codex helps speed it up quite a bit now but um it yeah i will say it does seem great to kind of like start start to have all the pieces start fitting together yeah i mean you saw we launched e-vals and we got the fine-tuning API for a while and um and we laid all the groundwork for some of the stuff over last year. And we're hoping that we can eventually, you know, make it into this full feature platform that's helpful for people.
Starting point is 00:39:12 I think you have. Since you did the Codex mentioned, maybe a quick tip from each of you on Codex Power User Tools or Tips. So there's actually a funny one that one of the new grads, I think, like, taught our team in general. And I think this is like a point for like just how like new grads, and younger generation people are actually more AI native. So one of them is to really lean in to like push yourself to like trust the model to do more and more.
Starting point is 00:39:45 So like I feel like the way that I was using Codex. And so for me, it's usually for my personal projects. They don't let me touch the code anymore. But you give it like small tasks. So you're like not really trusting it. Like I view it as like this like intern that I like really don't trust. But what a lot of the like, so we had an intern class this year, but a lot of the interns would do is just like full. YOLO mode, like, trust it to, like, write the whole feature.
Starting point is 00:40:08 And it, like, it doesn't work for worse. It, like, doesn't work sometimes. But, like, I don't know, like, 30, 40 percent of the time, it's just, like, one-shots it. I actually haven't tried this with, like, codec, GPD5 codex. I bet it probably, like, one shots it even more. But one tip that I'm, like, starting to, like, I feel, like, undo this, like, relearn things here is to, like, really lean into, like, the AGI component of it and
Starting point is 00:40:28 just, like, really let the model rip and, like, kind of trust it. Yeah. Because a lot of times, they can actually do stuff that surprises me. And then I have to, like, readjust my priors. whereas before I feel like I was in this safe space of like I'm just treating this I'm giving this thing like a tiny bit of rope. Yeah. And because of that, I was kind of limiting myself with how effective I could be. Like, sure, but okay.
Starting point is 00:40:45 But also, is there an etiquette around submitting effectively, you know, vibe coded PRs that someone else now has to review, right? And it's like, it can be offensive. We have codex to reviews now. Okay. It actually reviews itself. Does Codex approve its own PRs a lot more than humans? It doesn't get to prove them. I was going to say, I think like the Codex, PR,
Starting point is 00:41:04 reviews are actually one of like the things that my team like very much relies on. I think they're very, very high quality reviews. On the Codex PR side, like for the visual agents builder, we only started that probably less than two months ago. And that wouldn't be possible without Codex. So I think there's definitely a lot of use of Codex internally and it keeps getting better and better. And so yeah, I think people are just finding they can rely on it more and more. And it's not, you know, totally vibe-coded. It's still, you know, checked and edited, but definitely has a kicking off point. And I think I've heard of people on my team, it's like on their way to work, they're like kicking off like five codex tasks because the bus takes 30 minutes, right? And you
Starting point is 00:41:47 get to the office and it kind of helps you orient yourself for the day. You're like, okay, now I know the files. I have the rough sense. Like maybe I don't even take that PR and I actually just like still code it, but it helps you just context switch so much faster to and be able to like orient yourself in a code base. There are so many meetings nowadays where I have like one-on-ones with engineers and I walk into the room. They're like, wait, wait, wait, give me a second. I got to kick off my like Codex thing. I'm like, oh, sorry. We're about to enter async zone. It's like almost like your notes, right? You're like, let me. And they're like, okay, now we can start our 101 because now it's great. Yeah. Cool. We're almost out of time. I wanted to leave a little
Starting point is 00:42:19 bit of time for you to shout out the Service Health dashboard because I know you're passionate about it. Well, tell people what it is and why it matters. Yeah. So this is a launch that we actually didn't, you know, It didn't get any stage time today, but it's actually something I'm really excited about. So we launched this thing called the Service Health Dashboard. You can now go into your usage or like your settings account and kind of see the health of your integration with our Open AI API. And so this is scope to your own org. So basically, if you have an integration that's running with us doing a bunch of tokens per minute or a bunch of queries, it's now tracking each of those responses, looking at your token velocity, TPM, that you're getting the throughput, as well as.
Starting point is 00:43:00 the responses, the response codes. And so you can see kind of like a real-time personal SLO for your integration. The reason why I care a lot about this is, obviously over the last year, we've spent a lot of time thinking about reliability. We had that really bad outage last December, you know, longest like three, four hours of my life and then had to, you know, talk to a bunch of customers. We haven't had one that bad since, you know, knock on wood. We've done a bunch of work. We have an infer team led by Venkat. And They've been working with Jana on our team, and they've just been doing so much good work to get reliability better. And so we actually, again, knock on wood, we think we've got reliability in a spot where we're, like, comfortable kind of putting this out there and kind of like letting people actually see their SLO.
Starting point is 00:43:47 And hopefully, you know, it's, you know, three, four, soon to be five nines. But the reason why I cared a lot about is because we spent so much time on it. And we feel confident enough to kind of have it behind a product now. Five nines is like two minutes of outage or something. Yeah, yeah. We're working to get to five nines. What is an extra nine take? It's exponentially more work.
Starting point is 00:44:10 So, you know, and then, but like we always, you know, in the last couple of days, you were talking about like hitting three nines and hitting three and a half nines and then hitting four nines. But yeah, it's exponentially more work. I could go for a while on the different topics. We'll have to do that in a follow up. I mean, that's all, that's the engineering side, right? Yes, yes, yes. Like you're surveying six billion tokens per minute.
Starting point is 00:44:32 We actually zoom past that. Yeah, that's the, that's the, that's outdated. Yeah, but yeah, it's been crazy, the growth that we've seen. Awesome. I know we're out of time. It's been a long day for both of you, so we'll let you go, but thank you both for joining us. Yeah. Yeah.
Starting point is 00:44:45 Thanks for having us. Thanks. Thank you. That's it. How was that? That's great. Okay. We have the mic's offer.
Starting point is 00:44:55 I didn't want to say on the podcast was on the tinker thing so we actually go

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.