How I AI - Build your own company brain: the enterprise AI playbook from Stripe’s engineering team | Sharadh Krishnamurthy

Episode Date: September 7, 2026

Sharadh Krishnamurthy is an engineering manager at Stripe, where he helped build Kai, the company’s internal AI agent used by more than 10,000 employees every week. He’s worked across several of S...tripe’s core infrastructure teams, including data and developer experience, which gives him a grounded, systems-level perspective on what it actually takes to make AI work at enterprise scale. He’s currently focused on the governance, skills, and infrastructure layers that let every Stripe employee use AI safely and effectively, regardless of their technical background.What you’ll learn:Why Stripe built Kai from scratch instead of buying, and what tipped the decisionWhat Kai knows about you by default and what you actually controlWhy “projects” at Stripe are a governance mechanism, not just a folderHow Stripe structured its data layer so agents can query safely at scaleWhy the infrastructure Stripe built for human developers turned out to be exactly what agents neededHow Kai’s skills platform lets any employee package a workflow, and what happens when you have 2,000 of themWhat Sharadh learned the hard way when agents nearly took down production systems—Brought to you by:DX—Engineering intelligence for the AI eraHyperagent—Deploy fleets of agents that handle real work—In this episode, we cover:(00:00) Introducing Sharadh(02:46) Why Stripe built an AI agent (Kai) instead of buying tools(05:18) What Kai knows about you (and what you can turn off)(06:51) Projects as a governance layer(10:04) Live demo: Kai builds a dashboard(12:18) Tools, skills, and the secure sandbox(17:22) Why Stripe has benefited so much from AI(19:20) Agentic identity, load shedding, and rogue agents(20:41) Iterating on the dashboard(25:01) How they rolled out Kai across the team(29:07) How projects work(34:18) Bespoke agents for bespoke use cases(35:58) The skill builder workflow(40:40) Skill quality, evals, and telemetry(43:01) Recap(45:13) Lightning round—Tools referenced:• Trino: https://trino.io/• Anthropic: https://www.anthropic.com/• Gemini: https://gemini.google.com/• Cursor: https://www.cursor.com/—Where to find Sharadh Krishnamurthy:LinkedIn: https://www.linkedin.com/in/sharadhk—Where to find Claire Vo:ChatPRD: https://www.chatprd.ai/Website: https://clairevo.com/LinkedIn: https://www.linkedin.com/in/clairevo/X: https://x.com/clairevo—Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Transcript
Discussion (0)
Starting point is 00:00:00 Agents are very creative at bringing your infra down. It turns out that agents just like dial up all your failure mode. It just multiplies the amplitude of problems you can get. There are agents that went rogue. There are agents that may have almost taken down core systems. One of the cool things about projects that can be very concrete for people is the idea of tool policies. Let's say you're a person on the HR team who's dealing with a bunch of sensitive information. You really don't want the agent to son of a go rogue and put that sensitive data into some public
Starting point is 00:00:30 Google document that all stripes can access. But you also don't want to tell then, oh, you can't use any tools because your work clothes are too sensitive. I love this idea of this like three layer triage that a data agent can go through. And that's really smart. Find existing reports then use the analytics layer to find the right query. Then it's like you really have to fall down to the data catalog and write your own query.
Starting point is 00:00:50 Your data warehouse has to be very resilient to high volume queries because when in doubt an agent will just brute force it. We could test this out and it's going to pick off what is a human-in-the-loop workflow. Create a calendar invite for me and Wong tomorrow at 11 a.m. Pacific. This isn't the fun part. The fun part is what I showed before. Agents are really good. They're very creative.
Starting point is 00:01:15 So we got to put some restrictions on them so they don't go rogue. Welcome back to How IAI. I'm Claire Vow, a product leader and AI obsessive here on a mission to help you build better with these new tools. Today we have Sherrod, an engine. manager at Stripe and part of the team who built Kai, their internal company brain and company agent. He's going to show us why you might want to build your own custom agent for your company. What are the governance and control mechanisms of Kai that make it super special and how to build not just a skill building skill building skill, but a skill building platform for your team to share
Starting point is 00:01:52 their automations and workflows with the rest of the company. Let's get to it. This episode is brought to you by DX. In a recent study across more than 500 engineering organizations, DX found that spend on AI tools has grown 28X over the last year. The share of AI authored code is climbing, but overall innovation has remained flat. As teams generate code faster, new friction in code review and validation is offsetting those early velocity gains. DX tracks speed, quality, and cost together across the software development lifecycle, giving engineering leaders clear visibility into how AI impacts the delivery.
Starting point is 00:02:31 and whether those investments are translating into real value. Download the full report at getdx.com slash how IAI. That's GETDX.com slash how IAI. Shared, it's so nice to have you here and I had to reach out to the Stripe team because I wanted to learn not just about how Kai, the company brain, the company agent, works at this big company, but I really wanted to understand why in the world you built it yourself. And so I wonder if you just start us there. Why build Kai? What was the problem you all were trying to solve? We wondered about this a lot before we built it because we have all this like
Starting point is 00:03:18 this avalanche of AI tools. This was like early 2026, huge number of tools coming out. Clot code had just like taken the world by storm. Like a lot of like really cool things were happening. And the problem that me and my and my colleague, Anupam, we're dealing with is how do they get AI to everyone, right? And we quickly realized it's not just an engineering or technical problem. The harder problems are in trying to replicate the way a company works at scale. And Stripe is an incredibly complex business. Like all around the world, like multitude of products, so many, so many processes that keep us in a shape so that we can help our users. And so we quickly realize that it's not about providing AI. It's about providing the correct governance structures so that everyone can just go
Starting point is 00:04:07 use AI and know and do the right thing for them. So some of the things that we really thought on and why we built Kai is governance is a big thing for us like I just said. The second thing is like really interesting. It's context aware. It knows who you are and what you do with the company and it knows what are the things that your colleagues are mostly interested in and it has access to the org chart and your projects and all the things that are going on, which means that it has a lot more mechanisms to do the right thing for you and understand what you're trying to do. And the other thing that I really like and our security team really likes is that Kai is all hosted on the cloud.
Starting point is 00:04:47 It's always on. It's behind our standard security boundaries. And it's very stripy in that everything that we use to build Kai actually ends up being standard infrastructure that helps us build great agents for our users. And ultimately, while it's great for me that we're really helping Stripe be more effective and enabling everyone, I really am happy that what we're doing here is helping build the rails that make our Stripe users get better products out of this using great agents that are coming out. So just to kind of repeat back what I heard and some of the unique things that you built into Kai is one, this sort of
Starting point is 00:05:28 bounded context engine. So, you know, I'm Claire. I work at Stripe. Kai knows about me. It knows about my place in the org chart. It knows, does it know kind of like strategic projects that I'm working on? Does it know like conversations about having? Like, how do you, how do you ingest that knowledge?
Starting point is 00:05:45 What does it know about an individual? We allow people to select how much they, they give Kai access to. But out of the box, they, Kai knows like who you are and where do you sit in the art chart. And it also knows some helpful other things like, like, what day is it and what time is it and so on. But relevant to you, personalized context is like who you are in by the art chart. From there on, the standard tools that we have connected to Kai let you talk to our project management system that figures out like all the OKRs and all the recent ship emails and all the projects you're part of. And it lets you connect if you choose to let it to your Google Drive, to your Slack, right, to private methods. messages, things that are pretty sensitive, and we really like keeping a tight boundary around,
Starting point is 00:06:35 you can choose to let Kaino as much as that as you want to. Some people choose not to, and I'm actually one of those. I turn on and turn off my access every session or every day. But many people are like, let the AI figure it out and more context is better. Great. So you have this context engine, but controlled by the end user to some extent. So you can decide as an employee how much you want to automatically ingest into the system. You have these, you're also building it on some fundamental agent building blocks. And, you know, this is one of the, you know, there are lots of reasons you might build one of these things at a company. One is just to teach your team how to build great agents.
Starting point is 00:07:18 It's like the perfect dog fooding agent experience, which is if you're going to build a great agent for your customers, you should learn to build a great agent for yourself. and then it sounds like you're reusing some of that infrastructure, which is nice because then you can pressure test it against Stripe employees or external customers. What else is unique about Kai? What are the places where you're like this? We went a little extra on building this. So as folks that are listening,
Starting point is 00:07:44 thinking about building this internally, they can decide where they want to differentiate their own kind of internal agents. The two things that we were very intentional about is the idea of projects. And projects are primarily a governance mechanism. them, but they also let you, you have this context engine as you said, right? But projects are almost like intentionally the user is telling you what they are trying to do, and that's a very strong signal of intent, and that lets the AI perform a lot better.
Starting point is 00:08:12 But projects, you can, if you see my screen in a second, there's the Hawaii AI demo project that I have right here. Projects are some things that strikes and create. It's all open, right? And what they can do is, we have projects that have 500 people in them. We have projects that have five people in them. The idea is that you create something where someone decides what's the appropriate set of things that you should need for the AI to function well. And what are the appropriate safety controls?
Starting point is 00:08:42 Like, as token spend is top of mind for a lot of companies, a project can kind of say, like, hey, here's the default model we want people to use. We don't even want to let them use these, like, super expensive models, right? because the job that you're going to do here doesn't need, you know, one of these supermodels to look at them. So projects as a governance mechanism, I think we did a lot there because, again, Stripe is a very complex business. Enterprise-scale AI requires these sort of mechanisms to make sense of things.
Starting point is 00:09:10 And I like the fact that we can have a few people who are very, you know, knowledgeable about the AI and the trade-offs between cost, performance, and latency. And they can sort of set the stage for everyone to just go used, right? We should try to minimize the number of people who have to actively make these choices every day. And just it should do the right thing for them. And projects are one way that we can do that. The other one is kind of related. We'll see that as well, our skills.
Starting point is 00:09:39 The way that it builds skills and our skill routing, there's a lot there and how we think about skill quality, skill governance. A lot of good stuff there that we can get into, that we've been with a lot in. I think the team is shipped. It's going to be interesting. It's going to be a live demo. But I think they shipped a feature where automatically people, everyone who authors a skill get suggestions on how to make their skill better,
Starting point is 00:10:02 how to help climb built into the platform. I like this because I've seen, you know, we've seen projects for holding context or just simply as an organization structure to AI work, right? Like, here are the files you need. Just put all my chats in this group. What I haven't seen anybody talk about, which I actually think is really interesting,
Starting point is 00:10:20 is using projects as a configuration layer and a governance layer on how your team actually uses AI to get a specific job done. And so I like that idea of project-based model routing, project-based, I'm sure, connectors or approvals or all those sorts of things, and I'm sure there's more and more and more you could do in the future. Which really interesting.
Starting point is 00:10:43 So let's, I mean, show us what's Kai good at? And why do you just walk us through a common kind task and how this customization benefits the end user, in particular, maybe like someone who's a little less technical. Yeah, absolutely, happy to do it. So I have a few trustee prompts here. What we're going to be doing today is we're actually going to be creating a dashboard because Stripes, we love our data,
Starting point is 00:11:09 and something that literally everyone at Stripe does with Kai is create a bunch of dashboards, right? So they're going to do this life and they're going to ask Kai to create a dashboard for us. And, you know, in the interest of vanity metrics or things that I can speak to easily, I'm going to ask it to go create something related to Kite itself, right? So Hubble is our internal sort of like data querying Leo. And I'm going to tell it, hey, go find these queries that I usually use to trap Kai adoption. And I want you to like create a dashboard for me. And this is really interesting. I really like this because some things that I've found, most people you really use AI for it.
Starting point is 00:11:52 And I talk about most people, I'm like, so I'm an engineer, I've been an engineer all my working life, but I'm super passionate about how we scale AI to like everyone. And not everyone is necessarily an engineer, even if they're super technical at Streda. So something that people really have been using the AI for is creating dashboards that communicate a point.
Starting point is 00:12:12 And that's what I wanna show how we will be doing that using Kai today. And there's a lot of stuff you could get into along the way. And I see it pulling for tools and skills. So is this part of the harness, which is it's kind of like tuned to go find what it can use to solve this problem? Absolutely. So there are two things. We have, of course, we all have, we all know what tools are and we all know what skills are.
Starting point is 00:12:41 Tools basically being the things you can do. And skills are like essentially a way to package up relevant tools so that we can find that easier. So the thing that you see Kai doing immediately is that it's going to find the skill that says ask data. And that's a skill that's going to go figure out how to answer any strike data and SQL-related question using our standard tools. Using that skill, which it's loaded up, it's got access to a bunch of internal tools. And it's going to use those tools and go ahead and do more things, right? So the tools are a combination of a bunch of different things. there are tools that we set up in the harness.
Starting point is 00:13:20 Example, discovering a skill is itself a kind of tool that we give the harness, right? But you also give it things like there's a sandbox that's secure for your session, where you can do grip and stuff. Because, again, unlike our typical agent, this isn't running on your laptop, is running in the cloud. And so how do we make sure that your session, Claire, and my session don't like eat each other? So we set up the secure sandbox and give the harness tools to interact with the sandbox, tools that can securely send data to it,
Starting point is 00:13:48 tools that the sandbox can use, the agent can use to search things and do things in the sandbox, and tools that can get data out of the sandbox, right? So that's a bunch of tools, bunch of tools all over the place, but this one is pretty cool because, again, we have a sandbox, you don't even need to know that the sandbox exists,
Starting point is 00:14:07 but the agent is going to be writing some sort of script on your behalf. I don't even know what it does half the time. I know it's secure, but it's going to go in, It's from the data. Looks right. Yeah, it looks right. I told it not to make any mistakes. So that's what these tools and skills are doing.
Starting point is 00:14:26 Yeah, I have a kind of separate question while this is running, specifically about making great data agents because I talked to a lot of companies and almost universally, the first internal agent they build is specifically for this use case, it is like data querying dashboard visualization agent. And so I'm curious in Hubble, which I think you said is like kind of your data store query engine, were there anything that high level you had to do to make Hubble agent ready? Because I see like query metadata and, you know, ask data and these skills. I'm just curious, like if you were building an agent and you needed to ready the data warehouse and the query layer, what are a couple of key things that you think are super important
Starting point is 00:15:15 for folks to think about? 100%. That is such a great question. So like I said, luckily, at Stry, we care about our data so much that we've invested a lot into both the data querying layer. We use Trino as our data sort of querying layer in our warehouse in that perspective. They've invested a lot into making that super resilient, right? And those investments have helped agents like slam it like crazy and not bring it down, right?
Starting point is 00:15:45 We've invested a lot our data, the data platform side of things. We've invested a lot in a catalog of data and tiering of data. So we have access to schema that can quickly tell us, oh, these are the relevant data sets that you might want to find and use and how would you use it, right? But even there are higher level investments as well. There is a blessed analytics layer where like the really key. metrics go in, right? And like there's a tiering system where you, there's an analytics layer. If you fail that, you go look at all the standard data dashboards that we have and you use the
Starting point is 00:16:20 queries from that. And if you fail that, then you use the data catalog and search through for the high quality datasets and figure out what to use it. Agents are incredibly good at figuring this out. However, the key part, and you asked about the ask data skill, the key part is we have some really smart data scientists as well who sort of said, hey, this is a, this is probably the right way that most data query should be handled. And what that skill does, if we dig into the cast data skill itself, what it's going to be saying here is route to direct artifacts first, use the analytics layer first,
Starting point is 00:16:53 and if that fails and fall back and fall back and fall back, and until you actually hit the data catalog directly, right? So these investments were made for humans, but have held up really well for agents because terms out the reasoning through it, agents have the same problem. They can answer the question, but they have no idea if it was the right query
Starting point is 00:17:16 or the right table. And these investments have paid off in helping guard that. I want people that are listening to hear a couple of things in, you know, I'm going to make the Stripe Team Blush. I say this specifically about Stripe a lot, which is, I think one of the reasons why Stripe has been able to benefit so much from AI
Starting point is 00:17:33 is prior to AI. There's been a commitment to, developer experience, developer platform, data platform, analytics layers, like all these things that made humans really efficient at the company pre-AI are foundational investments that now give you extreme leverage when you throw agents at it. And so, you know, when people ask me like, Claire, what can I do to ship more product with AI? They think I'm going to say something about product development. And I say double the size of your devX team. Double the size of your devex team. double the size of your data team.
Starting point is 00:18:09 Like work on platform investments, good for humans, good for agents, and that's what we'll let you run. The other thing you said, and I don't want people to miss, because I love this idea of this like three layer triage that a data agent can go through, and that's really smart.
Starting point is 00:18:25 Like, find existing reports, please. Then use the analytics layer to find the right query. Then if like you really have to fall down to the data catalog and write your own query, the thing that I also heard you say is your data warehouse has to be, very resilient to high volume queries because when in doubt an agent will just brute force it. And so again, this is like infrastructure, hardening investment, performance investment,
Starting point is 00:18:52 not sexy, not what people are thinking about when you're building these data agents, but actually allow agents to do a really effective job because you don't worry about like, you know, turning over your data warehouse because a agent is hammering it. 100% and like everything that you said makes so it's it resonates so much with with all of that my my personal history at strike has actually been on each of the kind of teams that you've referenced so I'm like yes someone gets it so this is great the the thing about resilience agents are very creative at bringing your intra down what can I say they're like it's almost like all these scripts that they were trained on just teach them to be script kiddo agrees or
Starting point is 00:19:35 thing, right? The thing that we really did well is thinking about agentic identity, like, we haven't solved this yet, right? But thinking about how do we say that, you know, this is the, this is an agent and this is what it's trying to do. Like, what is a use case it's trying to use as it goes around doing its thing in our infrastructure. And using that as a way to think about priorities and load shedding and all of that good stuff, again, not super sexy, very deep infra stuff, but the same principles apply. It turns out that agents just like dial up all your failure modes. It's just, it's just, it just multiplies the amplitude of problems you can get.
Starting point is 00:20:20 Right. And the investments, I wouldn't claim that we did not have any issues. We definitely had a bunch of issues where when we started doing this, like there were agents that went rogue. There were agents that, you know, may have almost taken down. four systems, but we caught it in time, and now we've hardened those systems as well. I love it. Okay, so we've yapped while Kai ran. Let's show what Kai actually generated using these skills and tools in Sandbox. Yeah, of course. So here's what you see. You see that, you know, Kai adoption is looking good. And this is something I'm personally super happy about,
Starting point is 00:21:00 like pretty much everyone at strike uses Kai. Like, 86 plus percent of the company. And, you see that, now, so really AI for everyone, which is how we started out this process, and this is a ramp that's gone from a fairly low number. I think if we had done this a couple of weeks ago, which I've been in the hundreds, up to a very high number. So happy to talk more if you're interested, if viewers are interested into how we manage that. But, okay, we have a dashboard. Dashboard looks good. It also looks like vaguely stripy. So I need to go back and see how the agent figured out that it needs to make things global. So I've got to go for it. figure it out. But it has a bunch of things here. It's an interactive dashboard and has links to a
Starting point is 00:21:40 bunch of things, right? That's fine. This is great. We can already see how this can be useful for like, I now generate a dashboard every meeting I go to because it's so easy and it helps me drive, drive the meeting a lot better. But the dual power here starts to come in and you talk about multi-turn conversations, right? So great, we have a dashboard. Awesome. But let's do something more. Let's sort of like get Karai to iterate on this for us, right? So, hey, I love this dashboard, but let's do some more here and use this buddy, get a breakdown, yada, yada, yada. And it's going to do some interesting things here. So I'm going to take this off, but I'm going to talk through what I'm doing, right?
Starting point is 00:22:26 A, the dashboard isn't like, is the artifact isn't like created and like it's not fired and forget, right? We give a chance for people to do deep work by iterating on their artifacts, and that's really powerful. It's better for token efficiency. You don't want to be throwing away a HTML dashboard every turn, but it's also really moving into this idea where the agent and you are collaborating on a task, right? And we have turns that are like super deep, like hundreds of turns over multiple weeks. So the idea here is you have somewhat like a somewhat like a
Starting point is 00:23:02 pretty smart collaborator who has some artifacts and you can iterate with them on it. I'm going to add this query and I'm going to do some really interesting things and this is something that I think it's worth getting into. I'm telling it, okay, it's not just pulling the data. It's not about pulling the data and displaying it. I want you to do things with the data. I want you to like munge the data in some way or form so I get what I want. And the reason why I'm touching from this is a lot of the data sort of things, that people want to do end up being last mile data. You think about people's workflows.
Starting point is 00:23:37 It's so different. It's so hard to build a dashboard for everyone to do every part of their job. And then you have like a gazillion dashboards and how do you manage them? You can't keep the right dashboards of the right level of quality. Using an AI like Kyle to do this means that you can create like light apps, almost like the whole like the lovable style thing where people are creating apps to just hyper-optimized for their workflow. And the fact that they have a sandbox
Starting point is 00:24:07 that anybody, regardless of whether they're an engineer or not, can get the agent to write code for them and do whatever the heck they want with the data, it's really powerful. And I'm pretty sure that, again, as expected, it's gone in, it's all like, said, okay, here's the actual data, and I want you to go do some summation somewhere
Starting point is 00:24:26 to do the other tab. And out again, so super interesting. And if I open up the updated dashboard, it's the exact same dashboard, and you should now see this really cool little segment below that. So I could keep yapping about this dashboard. I love the fact that our marketing theme is like 100% all in. They need it, right? I don't know a single marketing person that doesn't either want some sort of app built
Starting point is 00:24:57 or some sort of dashboard. So I think you have product market fit. That leads me to my next question, which is, how do you roll out? I'm just curious kind of, you know, inside the doors of Stripe. How do you roll something like this out? Is it really organic adoption? How did it get built? How did it get shared to the team?
Starting point is 00:25:16 Was this like 20 engineers, like, how did this come to be? Definitely not 20 engineers. Stripe, we run fairly lean and very nimble and very fast. We built us super quick. it's a very interesting case because so we had like say me and I call it, we had the idea I was moonlighting as an
Starting point is 00:25:40 as an engineering individual contributor again trying to get this out of the door and so it took us like one and a half people over two weeks to get vis-vis out of the door and something we've realized was a lot of the questions became the answers became apparent once we could show people something it was very hard to tell people
Starting point is 00:25:58 why something like this is required in a world where you had the coding agents around and they could be super powerful. But the moment we got that V0 out, super inexpensive, one and a half engineers for like two weeks, right, V0 out. And then we moved into a pilot stage where we started seeing a lot of interest
Starting point is 00:26:19 from primarily, we have a great collaborator, Alia on the GTM team, who builds like AI for Gtm, right? And they were like super interested in this because like marketers are go-to-market function at strike. Like, they are extremely, like, they're looking for whatever they can do to reach more people, to reach them more in the right manner and so on.
Starting point is 00:26:41 So these are a lot of our options from them. And that's how it kicked off. It went into this pilot stage. We still had about 200 to 300 users. At this point, we had like two and a half, three people working on it. And this was the next, like, month or so. things really ramped once we did a company-wide demo saying that, hey, we built this thing, we invited to use it, and it just clicked for everyone.
Starting point is 00:27:06 And people started, that's when you see the really steep ramp up somewhere, I feel, and everyone started using it. Even then, the team itself, I wouldn't say, is humongous. Like, yeah, 10,000 plus people use it every week, but the core team that manages the experience is still like less than 10 people. And we have a lot of other things we have going on with those 10 people as well. And the things that let us build it, of course, coding agents and the productivity that they've given us and our DEFRA team are incredible. We've spoken about minions on the show before.
Starting point is 00:27:41 We have incredible tools at Stry to get more from who we have. And we have all this infrastructure that you referenced. And all that has helped us. So I would say it's less than 10 people, but I also want to give credit where it's due. There are a lot of people helping those 10 people do what they can do. I love it. This episode is brought to you by HyperAgent, the platform for deploying always-on agents that actually run your business.
Starting point is 00:28:07 With HyperAgent, you build agents in the cloud and deploy them where your work already happens, like Slack, Telegram, or Email. An agent will scan your inbox and draft replies to vendor follow-ups, another monitors, competitors, and spins up rich ad kits and landing pages. A third notices a deal going close. cold in Salesforce and writes the save email with full account context. These aren't chatbots waiting for a perfect prompt. They're proactive, learning your preferences, retaining your playbooks, and getting better with every run. One user built four agents to run an outbound sales pipeline,
Starting point is 00:28:42 prospecting, outreach, follow-ups, CRM updates, all in a single afternoon. No local setup, no VPS bills, no fragile permissions on your laptop, just powerful agents with full control over skills, tools, and guardrails. How IAI listeners get $100 in free inference to start building. Claim yours at hyperagent.com slash how IAI. What else, maybe one or two other things that you think are worth pointing out in Kai that you think make it pretty unique or at least, you know, fun and easy to work with? I'm going to do two things. I'm going to show skills and I'm going to show projects, right? So when I come to say skills, we have this. thing, it's great, everyone loves the dashboard, but I don't want to be creating this dashboard
Starting point is 00:29:28 and paying a bunch of tokens and time every time. So the thing that I think I did really well, and one reason for its product market fit was I can create a skill that basically takes what I've done in this session and packages it up so that it can become a load-baring, repeatable workflow. And that's when the AI goes from here's something I'm just like iterating with on the site, like a chat interface to share something I can trust to run my, to sort of like run my business or run my workflows or help me, help me do that. And a little bit after,
Starting point is 00:30:04 you're going to see this kick off this like skill creator skill. It's a skill that the harness has that's going to go in and create, take all the things it's learned from this session, from its interaction with me, and package that up nicely into something that I can just like pull up at any time. And we'll talk a little bit more about how, we do the skill retrieval. I think that's a really cool part of the system as well.
Starting point is 00:30:27 But while that's cooking, right, let's actually look at this other thing I love about how we built Kai, which is this notion of projects, right? So the thing about projects is there's so much stuff happening at stuff. We have like 2,000 skills, right? Projects are this really nice packaging mechanism where we can draw boundary around those skills and say these are what most people who are doing this workflow
Starting point is 00:30:50 must be using. We have projects that are created for projects like short-lived things. They have projects created for teams. Like the People team has a super secure version of Kai in a different project. That's backed by a totally secure backend and stuff. The thing about this is, again, it lets one person or a few people who are DRIs of the space to figure out how to get the agent to perform well for everyone. The really interesting thing I have on projects, there's this thing called settings that, like I said,
Starting point is 00:31:20 You can use a custom agent to power your project. It doesn't have to be archive, which is pretty good and general purpose. But let's say you have something really bespoke. You can use all the same features we have, but just backed by a different API in the backend and a different harness, right? So really, again, when we talk about AI at the enterprise, there's going to be heterogeneity. It's going to be a lot of different cases.
Starting point is 00:31:43 And building this in layers so we can give maximum leverage and customizability. one of the cool things about projects that can be very concrete for people is the idea of tool policies. Now, I mention the people team. Let's say you're a person on the HR team who's dealing with a bunch of sensitive information, right? You really don't want the agent to send of a go-rogue and put that sensitive data into some public Google document that all stripes can access. That seems like an accident waiting to happen. We don't like that. But you also don't want to tell them, oh, you can't use any tools because you're,
Starting point is 00:32:18 workloads are too sensitive, right? So what projects let us do is to say, for this workflow, I'm going to set up a tool policy that says, in this case, I set the run Hubble tool so that I don't inadvertently put some confidential information into the demo, right? But we could test this out with some other tool, and it's going to kick off what is a human-in-the-loop workflow. So create a calendar invite for me and Wong Morrow at 11 a. in the Pacific, right?
Starting point is 00:32:53 And I set this up ahead of time just to show what a human of the loop flow would look like. But you can see it extends to any other kind of tool. What this is going to do is going to tell me, hey, should be familiar to most people who've used like cursor or
Starting point is 00:33:09 the other large, you know, big products out there. But I want to do this, right? The interesting part and why this isn't the fun part. The fun part is what I showed before, which is that someone who is the DRI of a space can decide that certain tools are kind of sensitive for the workloads that these people are going to be using.
Starting point is 00:33:29 So we need a human interduke to confirm if that action can be taken by the agent. Agents are really good. They're very creative. So we've got to put some restrictions on them so they don't go rogue, right? And projects help us decide. I also don't want this to be happening for every person at Stripe.
Starting point is 00:33:43 That would be kind of frictionful. So projects are, again, are drawing the boundary around it. Yeah, I love this because I think a lot of the existing tools let you maybe configure some of this at the individual level, but then it applies to every session. It's not contexted to what you're working on and you can't share that permission set across different users. And so what I think is interesting about Kai is the like permission and context boundaries are very purpose built for how your company works. on things. And I think, you know, when people are asking themselves either should I build something myself and does that make a lot of sense or do I need to pluck something off the shelf for my
Starting point is 00:34:27 enterprise use case, again, you need to ask yourself how much appropriate or inappropriate friction will this put in everybody's day-to-day work? Because at the end of the day, what you want to do is make everybody's life easier without causing chaos or trouble. And you want the management requirements to go down really low, right? You don't want everybody to think every task, like, do I need to turn on this connector, off this connector, connect to this data. And so I do think one of the benefits right now of teams building their own thing is they can really think about bespoke agents for bespoke use cases, but kind of hide all that complexity from the end employee, the end teammate, and just let them get to work. That's super insightful, because as you're
Starting point is 00:35:15 speaking about bespoke agents, and I showed a little bit of this earlier, Kai looks like a single product. It really isn't. It's like the icing on top of a multi-layer cake, and each of those layers can be, like, customized to work at the enterprise, right? So 100% agree that the notion of both customization, but lowering the cost of management, the cost of ownership, and just a friction. If you put too much friction in front of people, they're just going to do unsafe things,
Starting point is 00:35:43 because that's how humans are, right? We don't, if I showed you this, every single session for every single tool, eventually you're going to press the wrong button, right? So really thinking through that is a big part of what we're trying to do here. So that's projects and why I love projects as a unit of governance. Let's hop back really quickly to the skill builder flow that I spoke about. Again, we made that dashboard.
Starting point is 00:36:08 We want to make this something that I can reuse, right? And it's not just me. It could be my entire team. Why do we have to keep things close to ourselves, right? And so I can now go, Kaya's created a skill for me. It's like a standard open spec skill that you can use on any of your harnesses, right?
Starting point is 00:36:28 But for the purpose of this, I'm going to go click this button. And Kaya is going to say, okay, what do you want me to do? I'm going to go ahead. I could choose to push it to an area and we'll talk a little bit about area skills, but for now I want to keep it private to myself. right and of course this is like nothing special about me here any user of kai can do this and that's why
Starting point is 00:36:47 we have a lot of skills now that are making people more effective so it's filled in the description for me it's filled in like when kai could use it so this is important we'll come back to this in like a second right and i'm just going to go ahead and export draft oh man i already created one again this is what happens if you prepare too well for devils. This is when you do it live. I mean, we believe you. I think what you're showing here, great. You have kind of a skill creator,
Starting point is 00:37:22 skill or tool. You have a specific spec that you're using to ensure that it's both written well generally for agents, but also written well very specifically for the Kai harness. And then I love this idea of a draft kind of skill
Starting point is 00:37:38 editor that you can test and edit and optimize and manage. It's quite nice. I know people just love fussing around in markdown in, you know, Python files, but just a little quality of life UI here can go a long way. 100%. We really invested a lot in making this feel like a little bit like an IDE so that everyone can get access to that quality of life improvement. Going back here for just a second, now that I did this, this is the magical part.
Starting point is 00:38:13 This is the part that I'm really excited about. And the team has really kicked us here. So give me the latest I adoption dashboard. Right? So I'm going to say. And this is the part where the magic of Kyle really shines.
Starting point is 00:38:31 Now, when you're in a coding agent, you can see what it's picked up. It's picked up the skill that we just, like literally just created. Right. And the reason why this is interesting is, when I sit all the way at the top of this, that you can just go in and start using Kai
Starting point is 00:38:43 and it knows what to do. This is how it knows. Looks like my tool policies are too secure, right? So now, the thing that we've done is, when you're a coding agent, right, you're in a repository, you're in a folder, you have this natural structure to what you're trying to do, right?
Starting point is 00:39:04 And so you can pick up the skills in the hierarchy of where you're working, and you get the right set of skills required to do your job. When you're at an enterprise and you're starting to work, you don't know, like, there is no hierarchy or what you're not doing. You're frequently trying to fit to like five different systems. And so a large part of the investments have done, and what we've managed to give to strike is the ability to package skills
Starting point is 00:39:29 and retrieve them. And we do this like really rigorous flow of knowing when the right skills are being invoked so that can perform at a high level, right? And the number of skills that we have, we've got to do some pretty interesting things to ensure that we keep those skills in tip-top shape. We've also built out those interesting thing where it's not enough to enable people to build a bunch of skills. How do you make sure that they actually know what it's doing and how you keep them in top shape? And that's the other thing that we're really investing in. This automatic platform-driven suggestions for how to improve your skills so that, especially since we let anybody publish a,
Starting point is 00:40:09 shared skill, right, that anybody else can pick up, it's become really important to ensure that we can give people the tools to keep those in tip-top shape. Again, something that you probably don't worry too much about if you're just using AI for yourself. But the moment you introduce that sharing and the enterprise, quality, governance, policies all become something really important. And they're usually something pretty specific to the company you're working at and and Stripe is no different. You know, the only other thing that I've seen here that I'm curious, maybe you have, but you haven't shown, is we see a lot of folks that are building these internal harnesses do skill and tool telemetry and observability and see where, like, tool calls are failing a lot
Starting point is 00:40:55 so they can auto-eval that. And then they also have a deprecation policy for skills. So if skills have not been invoked for like 30 days, you get a little notice and it's like, hey, you haven't used a skill, if it's dead, maybe we archive it. And if they don't get a response, it goes into like deprecation status and then they delete it two weeks later just to like prune all this stuff that's happening. And so I think this enterprise level maintenance of the skill library is really important, not just from a quality perspective, which is what we see here with the evals, but honestly from a quantity perspective, like just is any of this useful? anymore.
Starting point is 00:41:37 100%. And it's, I almost think you can't separate quality and quantity when it comes to these systems because, you know, context is everything. The more underrated context you throw into the AI, the less good your results to come. So quantity is almost a facet of quality.
Starting point is 00:41:59 We've started off with like a bunch of different skills. And we do have telemetry on, we have, let's say, 50 skills that are used, like, hammered every day across the company. We have this long tail of 100 to 150 other skills that are used by sub-categories, you know, parts of the arc chain. And we have a bunch of tools that are used by two or three people, right? We need to respect that all of these exist.
Starting point is 00:42:23 There are teams that are three member squads doing some really bespoke thing, and they want to share between themselves. But then, again, what is the telemetry we have? And part of the process I'm showing here is this is the user-facing. part of it, but part of the process is telling us as harness owners, these are the kind of skills you probably want to promote up into like a general workflow. And these are the skills that you want to like delegate out and move out of the general workflow because it's just taking up context. Right. So definitely, unfortunately, don't have something cool I can show around that,
Starting point is 00:42:57 but it's, it is something that there's an ETL pipeline happening somewhere that's doing this. Amazing. Well, I just want to recap for folks because this has been awesome. Just high-level things about Kai, personalized context for people, org chart, awareness, you know, tune tools, a sandbox that you can put data in, a sandbox, you can pull data out. Shared artifacts, projects, which I feel like if you miss that part, rewind, go back to it, because projects are not just how you organize chats across a team, but how you give a specific space, whether that's a team and or an initiative, access to tools, access to data. data permissions and controls, including what requires a human in the loop, a skill builder's skill, but not just a skill builder skill, a skills platform for the company that allows you to build skills, edit skills, eval skills, share skills, and then, you know, AI that just works for the things that matter, including data analysis, which benefits not just from this
Starting point is 00:44:03 tuned harness and skills around it, but investment from infrastructure all the way up the stack on a great agent-ready data layer. And it took one and a half agents or one and half humans, sorry, probably a million agents. Many more agents. A couple weeks to get V1 going and now is serving 10,000 stripes with less than 10 people plus a bunch of great infrastructure that you've been investing pre-imposed AI. That's it. that's all. That's all. Not much to it.
Starting point is 00:44:38 But yeah, it's, it's, I think that was a fantastic summary. A lot of good stuff. I think it's important also. I would be remiss if I didn't say, it sounds like we figure this all out. We absolutely haven't. Like,
Starting point is 00:44:50 we are very cognizant that we are in the earliest parts of this journey. And we're hoping that, like you said, the strong foundations we have help us iterate and move forward with the, with the times. But yeah, it's been a fantastic journey. far. And maybe we'll be back in a year showing you something completely different because that's how quickly the space moves. I really hope. I hope sooner than a year. Well, before we get out of
Starting point is 00:45:14 here, let's do two lightning round questions. My first one is let's just put Kai aside for a minute. Let's put aside Stripe Blurple. What are personal AI things that are fun that you're doing or that you're excited about as an engineer? You know, when you shut up, You know, as you say, we like shut the work laptop and open the fun laptop on the weekends. What are you excited about? What's cool? I'm super boring, but so I'm going to like say something less boring first, which is it's helped me like not sound super dumb to my to my six-year-old who's right at the stage where he's asking me all these complex questions about like exoplanets and like galaxies far away. And I'm like on the side, you know, Gemini on my phone.
Starting point is 00:46:02 like, hey, can you tell me what's happening? And then I act like I know the answer. So it's helped me, you know, keep up to his model of dad knows everything. So that's good. That's perfect. But on the workfront, honestly, like the stuff that I'm not doing when I'm building Kai, it's like my, I have this whole workflow now, which is around using the AI to make sure that I don't miss things.
Starting point is 00:46:26 It's so boring, but it's so good. We didn't get to show you Kai schedules, but I basically used. who's Kai is like my personal assistant. It just tells me things that I'm supposed to be doing. And so, it's so basic. I'm almost embarrassed to say it out loud. But it's been the biggest life hack. Just not having to keep it all in the brain.
Starting point is 00:46:48 It's been awesome. Yeah. What I tell people is we think a lot about how to put agents to work. I want the agents to put me to work. I want them to say, Claire, please fill out this form. Claire, please do. this thing you said you were going to do. And so I think it's a it's a give and get relationship and I love that. And I also have many children who ask me really existential questions about the universe and about
Starting point is 00:47:15 dinosaurs and about history. And I agree. Intelligence on demand helps us keep our superiority in that parental child relationship. Yes. Yes. For a few more years at least until they figure out what they're all doing. So last question. When kind of, well, maybe not Kai. Maybe you're very sweet to Kai. But when AI is not listening, what do you do? How do you prompt? Are you a yeller? I don't know. I feel like I'm maybe I'm just subconsciously afraid of what it's going to do to me when it figures out, you know, knows where I live or something. But I'm very nice to the AI. I just say, hey, that's not what I wanted. Here, I'm going to say it again. And then maybe like all caps it. But I don't know. I don't like shouting at me. It feels,
Starting point is 00:48:01 it feels wrong. It's almost like I'm shouting about people who built the AI. So maybe. But I just insist. I just, I say yell harder to my team, but I actually end up just like saying a lot of please, if you will, read the thing I said better. But it does mess up and it's very frustrating. Like many of our guests, you gentle parent, the AI, which is, this is true. I know, I know you can do better. I believe in you. I'm not. I'm not. I'm not. Not mad, I'm disappointed. I'm just disappointed, Kai. Like, you should have done better.
Starting point is 00:48:37 I love this. Well, this has been super helpful and interesting for me. It's given me so many ideas about just my own use of AI and how I talk to people in enterprises about their use of AI. Where can we find you? And how can we be helpful to you in the strength team? Well, I'm on LinkedIn. And I'm happy to connect with anyone who's, like, super interested in learning more about
Starting point is 00:49:00 what we've built here. and how we think about, you know, scaling AI for the enterprise. And Stripe is always hiding. We are always in the lookout for people who want to, you know, join this crazy band of people trying to build amazing things for the world. So please look out on the Stripe Carriers page. But otherwise, my email is Sharad at Stripe.com. And I'm happy to engage with anybody who has questions
Starting point is 00:49:28 about anything we covered today. But yeah, keep those questions. awesome. Well, thank you to you and thank you to the Stripe team for being so generous with all the things that you've shared with the audience. We really appreciate it. And thanks for joining how IAI. Thank you for having me. And this was fun. I don't know if I mentioned, you're a minor celebrity on the team. So I now have some reflected glory. And yeah, it has been amazing time. Thank you so much. Thank you. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find
Starting point is 00:50:04 this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at how IAIIPod.com. See you next time.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.