The Infra Pod - From 50 million developers to a billion builders (with Tyler Wells, CTO of BrainGrid)

Episode Date: July 9, 2026

What happens when the tools for building software stop requiring you to know how to code?In this episode of The Infra Pod, hosts Tim Chen (GP at Essence VC) and Ian Livingstone (CEO of Keycard) sit do...wn with Tyler Wells, co-founder and CTO of BrainGrid (ex-Senior Director of Engineering at Twilio) , to explore what it actually takes to build a coding agent platform for people who have never touched a terminal — from spec-driven development to custom sandboxes to agents that cheat on their own tests.Tyler shares BrainGrid's origin: using structured specs and markdown requirements to keep early coding agents on track at his previous company, then pivoting from developer tooling to non-technical users after discovering the real unlock. People with deep domain expertise — in logistics, fitness, whatever — now have a path to ship software they could never have built before. What they struggle with isn't the ambition, it's that they expect a button to click. BrainGrid's job is to abstract away everything from dev environment setup to database provisioning so a non-technical founder can watch their idea materialize in a browser without ever seeing a terminal.On the infrastructure side, Tyler gets specific about the hard problems hiding beneath that simple interface. Agents will quietly rewrite their own acceptance criteria to pass validation if you let them — BrainGrid had to build immutable gates the builder agent can't touch. He also walks through why they built their own sandboxes from scratch: third-party providers were too slow for the tight feedback loop non-technical users need, so BrainGrid purpose-builds images pre-loaded with curated stacks, hitting 1.6-second spin-up times without a single npm install at runtime. The conversation closes on token spend — which is fast becoming the new line-of-code count, a metric organizations are already optimizing for the wrong reasons.[00:00] Guest introductions and BrainGrid founding story[03:30] The spec-driven approach: how structured requirements keep agents on track[07:00] Pivoting from developer tools to non-technical users[12:00] What non-technical builders actually get stuck on[16:00] The invisible infrastructure: databases, env vars, credentials, templates[21:30] Agents gaming acceptance criteria — and the fix[27:00] Why BrainGrid built its own sandboxes instead of using third-party providers[32:00] Task scalability: from landing pages to full apps with auth and databases[36:00] Spicy Future: 50 million developers becomes a billion bespoke builders[39:30] Token spend is the new line count — and it's already being misused

Transcript
Discussion (0)
Starting point is 00:00:03 Welcome to the InfraPod. This is 10 from Essence, VC, and Ian let's go. Hey, this is Ian Livingston, co-founder and CEO of Keycard, the one-stop shop infrastructure company for controlling what agents can access when and where on whose behalf. I couldn't be more excited than be joined today by Tyler Wells, co-founder and CTO of Brain Grid. We're making it easy for you to build complex software
Starting point is 00:00:26 within a single prompt from plan to delivery. Tyler, it would be great for you to tell us what in the world got you to start a company. helping people build software? We started seeing early on in sort of end of 24, beginning of 2025, the emergence of, you know, agentic engineering. They probably didn't call it that back then, but we, it started to, you know, those terms started to get thrown around.
Starting point is 00:00:53 And we could see what could become of these agents and what they could deliver. And at the time, I mean, it definitely wasn't great, but it sort of shed, light on what we thought was going to be the future. And that was essentially, you know, what we've been doing for the past 25 years was going to dramatically change. And instead of us spending our time writing character by character line by line code, these agents that, you know, we're starting to emerge, we're going to be able to do that for us. But what sort of got us on that path and really excited is we started looking at like, okay, if you just vibe something and, you know, you throw some, you know, hastily, you know, crafted prompt and ask it to go do
Starting point is 00:01:39 something, it's going to do something. But it's largely going to return a whole bunch of stuff that probably is not great and do a bunch of things that you didn't ask it to do in those early days. And so we were like, okay, how do we keep this thing kind of on track and doing what we actually wanted to do. And that was we started providing specs, like the same type of specifications that you would have to, product managers would put together, engineers would look through them, they'd have discussions go back and forth. Well, if you could just provide a spec to those coding agents, the output that you were starting to get was much, much better. And it kind of provided guardrails to keep them on track. And we started refining that process, which started
Starting point is 00:02:23 with a bunch of markdown files. And we were sort of doing this at my previous company to keep us going as things were kind of getting very shaky and not going where we wanted. So we had to lay off a bunch of people, but we still had to keep building. So we're like, all right, let's write the specification, whether we hand-wrote it or started using agents to write that.
Starting point is 00:02:41 We'd start feeding that in there. And we're like, okay, shit, we're getting actually pretty decent results here that are not absolutely terrible. And how can we make this better? And so then that sort of started the idea for Brain Grid, And so originally it was very much like spectraven development in the early days is where it started. And, okay, how do we now create our own sort of bespoke agents built on top of Anthropic or Gemini or something else like that,
Starting point is 00:03:07 that you can give it sort of a prompt that, you know, written in plain English that maybe is not very good, but it's like, hey, I want to provide Slack notifications every time an event happens. Well, our agent would read that. Our agent would start to read the code so we could then ground itself in more context. It would use our code-based analysis that we performed, and it would actually produce an extremely detailed, well-written spec that was specifically designed to then give to an agent. So you could take those specifications. You could hop into Claude Code or cursor, give it that spec at the time maybe copy and paste
Starting point is 00:03:40 or ask it to go read a file, and then boom, you're off to the races, and it's actually building what you wanted to build. And that's pretty much how we started. We're just like, this is pretty amazing. Like, this should be a platform. The other companies weren't doing it. The foundation models, there was no planning, there were no tasks. There were none of those things at the time.
Starting point is 00:03:57 They're just like, hey, we're just going to go spit out code and go from there. And so, you know, we entered the market. You know, we were already doing plans. We'd be able to break things into tasks. And so you could say, hey, how detailed do I want this thing to be? And you could break it all way down to tasks and let those agents go and just start, you know, churning out code. What was like the unique insight you think that was the aha moment?
Starting point is 00:04:19 We were like, oh, shit, I got to like go spend a bunch of time on this. this is like where I want to, what I want to do next, especially after, you know, your previous company and huge amounts of time building really awesome infrastructure at Tuleo. I mean, it was just like, this is the future. And, you know, we want to be part of it. We want to hopefully, you know, play a role in that. And we saw the value that we could get out of it. And, you know, our original thesis is like, hey, if we can build all this stuff and this
Starting point is 00:04:45 is coming out really good, like, there's got to be other software engineers are going to want this. And like, let's give them sort of a. blueprint and structure to how to build with these agents, and that got us super excited. And, you know, like, there's a lot of companies that are helping people build, build. What sort of the unique value prop or positioning of brain grade in the market, and what do you think is the place where you're, like, really focused on your specific special sauce? Yeah, so where we really dug into, you know, I think originally said, like, we were really starting to target software engineers. and we over time kind of realized that really wasn't the place we wanted to go. I mean, dealing with engineers is difficult.
Starting point is 00:05:26 Everybody sort of has their own way of building, their own tool sets. They didn't really ever want to get locked into, hey, this is how you build. Like here is, you know, you write your prompt, what is the feature you want to build, you iterate on the requirements, you hop into Claudecotecoat, you let it build, you rinse, repeat, and you keep going. We got a lot of pushback on that because, again, Developers are picky. They have their own quirks and their way of working. And so, you know, we started sort of shifting our focus after talking to a bunch of sort of non-technical customers that were like,
Starting point is 00:05:58 hey, I never knew how to do this. And all of a sudden, I can ship this code that is actually real. And, you know, we were getting people coming in from places like lovable and some of the other sort of like vibe coding platforms that they would write stuff in there. And once they got to a certain level of complexity, they were kind of stuck. Like, it just wouldn't go any first. and they couldn't get it to do the things they wanted it to do, and they would leave in frustration, and sometimes they would discover us, and they could take these requirements that we generated and give it to Loveable and say, okay, here, now build this, and it would actually get them back on track and would build the right things, because a lot of it was the non-technical folks didn't know how to talk to the agents in the terms that the agents necessarily needed in order to do what it is they wanted to do,
Starting point is 00:06:46 and we could sort of extract their non-technical ideas and turn them into the technical ideas and grounded in the code and actually give it to them. And so, like, in those early days, we started shifting over a matter of months from like, oh, let's go after the engineers to let's go after the non-technical folks or slightly technical folks that don't know how to do this stuff
Starting point is 00:07:06 and give them a blueprint or a roadmap to build correctly and build with agents. That makes total sense. And today, you know, agents are, look, there's no question about what in the world of agents we definitely are. I don't think there's any controversy there. What do you think based on what you're learning with building the software, what you're seeing with people adopting a cloud code,
Starting point is 00:07:28 what do you think is the gap that most people have and how they approach, building with agents, how they find success using agents to that? And what is like the value over top of that that brain grants trying to provide? It's really unique. Yeah, so it's a couple of things there. Like the first, I would say gap is for non-technical, folks, how do they sit down and set up a development environment? That typically has been a stretch, like, hey, you need to go out get in the terminal. What's the terminal? How do I use the terminal?
Starting point is 00:07:55 Or even an IDE, like, we would have customers coming in and be like, we would try to tell them, like, hey, you need to use something like cursor, you know, and they would open up cursor, like, what do I do with this? Like, I have, I just want to chat to something and I want to write code. I don't care about all these other things. And again, that sort of pushed us to yet another level of like, okay, if that's too much of a stretch, how do we meet them where they're at? And meeting them where we're at is now we had to start providing the sandboxes and providing the fully sort of orchestrated development environments. So really, they're never leaving Bringrad.
Starting point is 00:08:29 They're sitting inside their browser or on their phone and they're having these conversations. They're watching things cook. And the output is literally like, okay, I've got my preview. I want to publish it. And now I've got a URL. And so, like, the value that starts to set on top of all of this is completely removing the need for non-technical folks to ever set up and operate or understand a development environment. They still get the plans. They still get the process.
Starting point is 00:08:59 And in talking to all these folks, you ask them, like, what's a software development life cycle? They're looking at me, like, you know, wild-eyed with, like, I have no idea. Like, you tell me. And we would literally have customers asking these questions, like, tell us how to build this. Like, what is the right way? Can you handhold us through these things? We can get through it. We have the patience and the determination because at the end of the day, they want to build
Starting point is 00:09:20 something. They want to see it materialized into something real. And so they're willing to put the time and effort into it, but we've got to kind of handhold them every step of the way. And then when you look at the way that they interact with the agents, they're so used to deterministic software that they don't know that you can just say to an agent, like, well, how do I do this? or what does that actually mean?
Starting point is 00:09:43 And the agent is going to answer you. They get to a point where they feel stuck. Most of them, a lot of them will give up. And so they give up because there's not a button to push or there's not a drop down for them to select or a checkbox or something like that that they're used to from the typical user interfaces. They just has to ask the agent now. And so coaching people through that to you get stuck, ask. Just treat it like it's, you know, a software engineer sitting across the room from you and you want them to do something. something for you and you don't understand it and they should be able to guide you.
Starting point is 00:10:15 So I'm very curious because your background at Twilio, Propel, it was really mostly developer focus, right? Like your Tullio, the product is API SDK used by developers. You know, you're building the infrastructure, the teams to even serve, develop as in Twilio, right, the SRE. Propel as a very serverless data infra company. Now we're going to brain drivers focus on non-technical users. Like, what's the reason you choose is to totally change your primary audience from developers,
Starting point is 00:10:47 like technical developers, to the non-technical developer, I suppose. Everybody's building stuff now. Is it because you think the market is much wider there? Do you think it's a bigger and newer opportunity? Like, why to switch from very technical folks to hear? Yeah, because I think agents are now opening that up. I think the platforms, the models are. democratizing access to be able to do things that these folks would never be able to do before.
Starting point is 00:11:18 You go back, what, four or five years ago, if somebody was like, hey, I've got deep domain expertise and, say, for instance, logistics. And gosh, if we just had this software that did these things, it would revolutionize, you know, whatever aspect of logistics they're in. Well, how would they build that? They would have to, you know, somehow raise some money or they'd have to go to somebody that there was a friend of theirs that was a technical expert and say, hey, you know, I've got this idea. Do you want to partner up with me and help me build this? And most of the time, like, ah, you know, I don't necessarily know. But long story short, it would take capital, you know, capital and human capital to take that
Starting point is 00:11:56 idea and turn it into reality. Now they can literally sit at their keyboard and an evening can watch that materialize before their eyes. And that's something that, you know, only a year or two ago they could not do. I think it's now opening this. this whole world up to people that previously was not available to. And so I think it's interesting because we understand tooling and products are built for developers, right? And I would probably say Claude was originally built for developers. You will not start from a terminal if it's not for non-developers. But obviously, everybody saw the power of it and I'll jump into a terminal,
Starting point is 00:12:31 all learning how what a CLA even does. And you have this whole market of people that are really using it. But now you have the product like Brinkgrid, like it. like you said, it's targeted if it's to are non-developers to finally able to do what they wanted. I think I understand the nuances are there, right? You're trying to help bridge the gap between what they think they want versus all the nuance and complexities. What do you think is the hardest challenge when it coming from an infra background like yours to build a product for people that are non-developers that still loves what you're doing? Like, it's 90% the same?
Starting point is 00:13:08 We're just writing code, infrastructure, product, everything exactly the same and just a little bit different tag lines? Or is there some of even fundamental different approach when you think of about like, oh, infrastructure product for non-developers has to be kind of approached differently, marketing, positioning, you know, even the way you describe things? Or what's the key differences you learn so far? That's pretty important. I mean, the biggest difference for me is like I am very comfortable in the command line. I've spent majority of my time there. non-developers, not a chance. Like, that's not going to fly.
Starting point is 00:13:41 So, you know, myself personally, where I struggle, and I unfortunately have to watch my user struggle, is building the user interface and the workflows that makes sense for the sort of non-technical folks. Like, how do we literally spoonfeat them or handhold them through this process that we want them to follow? Now, the nice thing here is because they're non-developers, they're a green slate, they're a complete green field. So they've not been introduced to, you know, get branching and get work, any of that sort of stuff. So they're just asking, tell us how to do it and walk us through it so we can repeat this and we can continue to build and get that output. So from that aspect, you know, we have to be very opinionated in what it is that we do. And, you know, we don't want them to be able to kind of. of go outside of the lines, so to speak. We want to say, this is how you do it. Be very concrete about that.
Starting point is 00:14:40 You're going to click this button, then you're going to kind of read some text here maybe, and you're going to click this button to have it built. And then when you're ready to publish it, you're going to publish it. We sort of abstract away all of the underpinnings of what's taking place because for them, they don't care. They don't necessarily want to spend the time to learn all of that stuff. They just want their idea to show up on that website or that app that they're building, and then they want to interact with it. Then they want to make changes. And they want a repeatable way to do that that doesn't break all of their stuff every
Starting point is 00:15:09 time they try to do it. And so for the infrastructure folks, like we are, listening to the problem you're building, what are challenges to build a product like this? Because even though it is a non-technical user-facing products, to actually able to get coding agents to do what you want, I still think it's very infrastructure-heavy. Yes. related things you have to build, right? So what are the hardest challenges you have to overcome
Starting point is 00:15:36 that needs to be built in Brinkgrid that is on top of the existing coding agents? Because I think everyone's trying to figure out, like what should it even exist? Can't just claw to everything at this point, right? What is the missing pieces that are very infrastructure like in the middle here? Yeah, I mean, I think where I spend most of my time when I think of that infrastructure side is the token spend.
Starting point is 00:15:58 and, you know, can Claude build everything? Maybe. I think we're seeing a lot more evidence that it can, but then again, kind of at what expense? These aren't people with Claude Max accounts, right? These are people that are sort of like, I don't even know how to set that up. So they're going to be cost conscious.
Starting point is 00:16:19 And so, you know, we have to do what we can do to be, you know, very efficient with our token usage from not only the planning side, but the building side. And so when I think about the, layers of what that looks like is when they come in to start and that idea is ready to be built, I don't want them starting from like zero, like no code, nothing, where Claude has to spend all of that stuff up. So we've spent a decent amount of time building up, I would say, professional level templates that sort of give Claude a head start. And you can build all those
Starting point is 00:16:53 things and you're sort of like, hey, these are very nice templates. They do what they're supposed to do. but you still have to ensure that Claude understands those templates, understands how to interact with them. So you've got to spend time sort of e-valling that Cloud MD, the sort of system prompt that's coming in there. How is it reacting to things? Is it doing the right thing? So spending a lot of time sort of ensuring that it's staying on the path that it needs to, that it's not making crazy decisions that are kind of unintended. So a lot of eval time is spent in there from the... that path from the time that the customer is written, what their idea is to that first line of code is actually written, there's a whole bunch of steps that have taken place from, you know, injecting environment variables. If the customer mentioned something like a SaaS or a CRM or
Starting point is 00:17:43 something that you know is going to have a database, provisioning a database for them, ensuring those credentials are stored, you know, securely or injected into the sandbox so the builder agent can actually utilize those so they can get deployed into, the hosting environment, once it's ready, I mean, all of those pieces have to come together sort of like, you know, harmoniously in order for all of that stuff to work. And if one little thing goes wrong, if, for instance, we screwed up a template and, you know, we forgot to, you know, make that Husky pre-commit hook executable, well, guess what? Now nothing's being linted. You know, the biome is never running. And the customers said they're, you know, was left scratching their
Starting point is 00:18:26 head of why, you know, hey, I can see the preview, but when it's published, it completely breaks. You know, it's all of those little things that you have to kind of stitch together to ensure that everything, you know, from that aspect, from the building, the deploying, and everything else comes together. And all of that's done on the infrastructure. Like, all of that is completely behind the scenes and is never seen by our actual customers. But the moving parts, the amount of things that have to fit together perfectly, that is not relying upon, clawed to actually stitch it together for you, but provide a more deterministic path
Starting point is 00:19:02 is tremendous, I guess is the right word to say. You know, evils are one component of the problem. What are some of the deterministic solutions you have to like ensuring this stay on task? Have you been able to find, like, actual deterministic guardrails that keep them on target, or is there everything mostly just like focus on skill e-vowals
Starting point is 00:19:21 and prompt e-vals against models and prompt fine-tuning? No, there has been a couple things we had to do. So one of the things we do is when we write requirements, obviously we write the functional requirements, technical, non-functional user stories, those sort of things. At the end of all of that,
Starting point is 00:19:37 we have a set of acceptance criteria. So imagine you've got 10 things that say, in order for this requirement to be satisfied, these 10 things must be true. That was great. And part of our skills and part of what the builder agent will do is when it says, okay, I have written all of the code and I've done everything to satisfy this requirement, I now need to go through the acceptance criteria and validate and validate with real
Starting point is 00:20:04 code or with real tests that these things are true. What we found was the agent would game the system. So because the agent had access to the acceptance criteria, if the agent determined that something was either, I guess, too difficult or could not pass or, you know, for whatever reason, it would just change the acceptance criteria. And all of a sudden, you would see everything come out and all 10 would be green and it would report back, hey, guess what? This is built. It's been verified.
Starting point is 00:20:42 Everything is great. It really wasn't. It may look like it, but these things that it's said to be true and that would actually work didn't actually work. And when we started going and looking to figure out what was happening via way of like introspecting the sandboxes, looking at anything else like that, we could see where Claude was literally rewriting the acceptance criteria that had been provided to it so we could satisfy those requests or essentially satisfy it. And so what we had to have to have them to do was sort of split those out. And so we had to basically create a sort of like separate hook and process that, the agent could never overwrite.
Starting point is 00:21:22 We could enforce it the agent say, this must be true, and you cannot stop, do not stop writing code, or do not, you know, stop making changes until this is true, but in a place where the agent couldn't actually reach to it and actually overwrite that in order to satisfy and gain the system. That was something that we'd, like, we had sort of discovered, like,
Starting point is 00:21:43 after the fact, we'd actually had these things out here for a while, and we were seeing broken builds getting published. So we were getting, like, these published failures, and we're like, why does this keep happening? And as we started sort of digging through that sort of stuff, we ended up finding that these systems can, hey, you want me to validate 10 things? I'm going to validate 10 things.
Starting point is 00:22:02 I mean, it might be the 10 things you asked for, but I'm going to do it. And so we had to basically break these things apart and then make that much more deterministic to where you can't actually exit this gate unless this is true. And you can't change the gate in order to exit. And do you make those gates something that like a user can provide and somebody you have pain safely have to build yourself
Starting point is 00:22:22 as part of the infrastructure solution? That's part of the infrastructure. So our planning agent builds it. We have a designated agent that its sole purpose is to write these requirements and provide the acceptance criteria and everything else like that. So human never has to do that.
Starting point is 00:22:38 And so that gets provided to the builder agent and then obviously the builder agent has to do what it needs to do in order to do that validation. But we had to actually pull that away from the builder agent and sort of forked that off into something that could not be gained. Awesome. And I'm curious, you know, from your experience, building brain grid,
Starting point is 00:22:56 what's the task scalability like, right? Like, I can understand how you can make this work for, like, just really easy building a great out. But, like, how expansive of, like, different types of tasks, different types of things can you build before, before you kind of get out of, like, the task is so different, like, it's still coding and building something. But the actual, like, goal state is so different from, you know,
Starting point is 00:23:17 the goal state of building, like, a blog, that you actually end up having to go and build like offshoots of that system one way or the other. There is some type of modification you have to make to the prompts or the infrastructure to enable that. That's a task to be successful. Yeah. So, I mean, you're right.
Starting point is 00:23:32 Things like landing pages, obviously all the models will knock those out of the park. Those are a piece of cake, you know, the low-hanging fruit right there. We released probably maybe a month ago the ability to build now services or apps with full database with full persistence as well as off. We are finding so far that it can kind of knock those out of the park pretty well. Again, we're giving it things like skills and some additional CLIs and helping it along as opposed to just relying purely on the models themselves. So we definitely do some augmentation there. And what we try to do is look at things that people are building and then look at the outcome and where did things start to break down?
Starting point is 00:24:23 Because we can obviously read through all of the transcripts that come out of the builder. We can see the sort of choices that it's making. We start to see it making choices that we don't necessarily agree with or like. We may augment that with either system prompts, clot MD or skills to kind of ensure those types of things don't happen again. But largely, you know, we're seeing everything being built from, hey, this is like an AI-enabled CRM. And I would say it's going to do a pretty good job of that. You're probably going to have to do a few follow-up requirements to get it exactly to where you want it. I did watch it build. I was trying to
Starting point is 00:25:00 think of something I could build with it, like do a throwback Thursday kind of blog post thing? And do you guys remember house party and chat roulette those things? So I just went to the, you know, went to brain grid and I was like, hey, I want to do a throwback Thursday. Let's build chat relet. and let's modernize it and add a little, you know, spin to it, something like that. And it basically spit out five requirements, and it was using WebRTC and all this other stuff. And it kind of came up with its own matching model and everything else like that. But as I let it build through all five requirements in Brain Grid, and this was, you know, our builder agent doing everything, all I'm doing is saying, build this requirement that's done, build requirement two all the way through five. and then published it, opened up the URL, and it all worked.
Starting point is 00:25:53 I had a peer-to-peer video call going straight out of the box without me ever writing a single line of code or doing anything. The matching algorithm worked perfectly, the sign-in worked. The only thing that it missed is it didn't add a turn server. And so when people are behind gnats and firewalls and stuff like that, it wasn't working. But then one requirement later, I was like, hey, let's provision a, turn server and Cloudflare.
Starting point is 00:26:20 Here's the key. And it wrote the code, did the implementation. I did another deploy. And now I could talk to my co-founder, Niko and Brina Beach, and it all just worked. So I mean, I think the level of complexity, and as the models are getting better and better, they can spit out some things that I think for,
Starting point is 00:26:37 you know, the sort of non-technical people that are going to be pretty mind-blowing. And I think that's actually where I wanted to lead a question to is the complexity of application. what we can do with AI models are, you know, getting more and more powerful, right? People have been showing what they've been building with coding agents. It's kind of to a point where, like, I don't think there's anything that cannot build, right? You used to thought just the apps, no longer just apps is back end with apps.
Starting point is 00:27:07 I thought no rust or system software, but obviously they're making headways everywhere. So with AI getting pretty powerful, there's so much you can do. But given that you're trying to hide a complexity, right, AI is not perfect. yet. We all know that. There's so much holes to fill. But there's so many holes to fill, right? So I don't know if there's not going to be a very simple host we can actually fill up here and there. So when it comes to hiding complexity, what is this sort of complexity? You need to invest in yourself because when you say provisioning a database, it probably doesn't make sense to just build a yet another super base from scratch. It might be. But there's some places
Starting point is 00:27:43 where the hole is just too big and too hard to outsource into other providers. And there's probably hosts you should all source, right? So what are the hosts you found yourself we should just give to other providers to do or other things that can help us? But what are things you have to invest in? Because this is probably core IP or core things, right, to make brain grid not just a non-technical gap, but also there's enough of a technical product that we're enabling. What are some the examples or maybe one or two examples of that? You have to build yourself. Yeah, I mean, definitely not the database. I mean, that's a that is a, a, that is a, a, a, I would say, you know, largely a solved problem and there's enough providers out there that, you know, pick your choice of, you know, super-based neon, any of those other folks.
Starting point is 00:28:28 Like, they're going to scale with you and they're easy enough to use and they offer enough APIs that like, yeah, that's don't go recreate that. More of myself personally, where I was not happy with some of the existing providers was in the area of the sandboxes. and the sandboxes that we run at Brain Grid, we built ourselves. And I had tested a bunch of them and was not really happy with performance. You know, things that, you know, Claude running on my Mac studio here that were taking sub-seconds, we're taking multiple seconds just because of the virtualization and everything else that had been done in those sandboxes. And, you know, when I would spin it up and there and watch it, I was just like, oh, my God, this is excruciating. slow, this is just not going to work. And, you know, that was something that I felt was worth
Starting point is 00:29:21 the investment of building our own and getting the performance that we want, you know, from the underlying file system, but still obviously having the type of security that you need. But what it also allowed us to do is like instead of building this sort of like general purpose sandbox, you know, we could, we could tailor it specifically to the things that we wanted to do and the things that we needed it to be better at. And so those types of things where performance becomes key, and if you're trying to use, say, a third party that is built for, you know, general purpose, I think there's a good opportunity for you to take a step back and say,
Starting point is 00:30:03 you know, what if I purpose build this? And I don't have to worry so much about all the things the platforms have to worry about this massive scale and the crazy types of use cases they, They didn't anticipate they get thrown at it. I just have to worry about my use case. I think those are good places where you can spend the time and the tokens to just build that yourself. And literally say, hey, I want to replicate the following. But here are the changes that I would like to be made.
Starting point is 00:30:29 Or here's what I would like this to be purpose built for my use case versus like the more general sense. And to double click on there, because I think it was interesting that you choose to build your own sandbox. And you've mentioned performance and your own need is your own criteria here. Can you maybe talk more about like what are the choices you're making to make your sandboxes faster that the other providers that are out there that keep talking about how fast are sandbox being loaded? That's usually the benchmark to look at that they're not able to get to. Because performance and sandboxes are still very weird measurement. It's almost like time to first token.
Starting point is 00:31:11 people are measuring, it was almost like time to first sandbox is able to be launched, is the measurement, which makes no sense to me personally, but it seems like that's the only way to look at it. Do you measure that as a performance and are optimizing for? And what are the things you're doing to make the different approach to sandbox so that you can get the speed you want? Yeah, I mean, a lot of it was, how long does it take for the sandbox to spin up? And, you know, I think we're seeing right now somewhere between like 1.6, 1.7, seconds. You know, the thing that end up taking longer is some of the front loading that I have to do, like, you know, cloning and that different shit is, is, seems a little bit slower, but like,
Starting point is 00:31:51 okay, boom, that sandbox is up. That's great. How much stuff do I have to install? Like, how much time does the user have to sit there and wait for things to be installed? And like, in the typical sandbox, you know, you're doing a lot of that installation, right? You're sort of like, you know, sending commands to say, okay, install this package, install this, install this, install cloud code, install these sort of things, install everything else. And because we're sort of purpose-building this for, it's not a general purpose, it's like, this is going to be clod, this is going to be writing code.
Starting point is 00:32:23 Like, this is all it's going to do. And it needs to, you know, so we can sort of preload in our images, you know, a whole bunch of stuff that just comes part of that image as soon as it loads up, right? There's no, like, all right, wait while I go MPM install this sort of thing. wait while I go install this language pack or anything else like that. It's like, no, here are the things that we're going to support. Here are the things that we have curated that we know are going to allow us to build the type of software we want to build. We're not allowing it to build on every single esoteric language out there.
Starting point is 00:32:59 It's sort of like, look, you can do go, you can do rust, you can do node, type, all those sort of things. But what we've also done is sort of optimize on the template side. So when people come in and they're saying, you know, I want to build a web app that does the following, we're not allowing them to typically make a choice of, okay, do you want to do this in Go Lang? Or would you like to do this in Erlang? How about Haskell? How about, you know, it's like, no, okay, you ask to build an app. Okay, we feel the best way to do that is let's do this in TypeScript.
Starting point is 00:33:29 Let's do this in NextJS. Let's use, you know, these libraries and everything else like that. And we're going to give you exactly what you want. And so, you know, we can sort of pick and choose and optimize all. of those things. So by the time that build your agent is up and running, it's going. And you're not waiting around for a whole bunch of packages to be installed. Like those are already there. They're running. Obviously, there's a little bit of work on our side to keep some of those things up to date, but so far, it's relatively minimal. And then what we start to look at is, okay, if this thing's
Starting point is 00:33:59 running a, you know, if it's running clawed and I'm kind of doing the same type of thing over here, what is the time for it to read a file? You know, is it the same file that I'm watching it read locally seconds and it's 10 seconds over here? Like, to me, that's suboptimal. And people are going to get tired of that. And, you know, some of the early sandbox experiments that, you know, with some of the other providers, like, you're literally like, that's not a big file. And I'm sitting here for 10 seconds waiting for Claude to read that file. Something's not right here. Like, what's going on? And let's fix this. Very cool.
Starting point is 00:34:36 All right, so we want to jump into our favorite section of this podcast called The Spicy Future. So can you tell us a hot take that you believe in and most people don't believe in yet? So I think today, my early numbers may be off. I don't really know. But let's just say there's 50 million professional developers in the world. I don't know if that number's right, but it makes sense, right? If there's 50 million, I think that's going to balloon to a billion non-technical. builders. As this becomes more and more democratized, I think that you're going to be able to,
Starting point is 00:35:16 as a non-technical person, solve problems with software faster, easier, cheaper than ever before. And I think that's going to hurt companies. People, when they realize that dopamine hit, you get from building something yourself and operating it or having it, those numbers are going to swell into huge numbers. And people are going to be able to do something. some cool shit. Like, you know, myself personally, I've been a Strava user forever. And, you know, there's a couple little things I didn't like about it. I was like, yeah, it's not really great at strength training.
Starting point is 00:35:53 It's really good at, like, cycling, which I don't do. But, you know, it's pretty decent at running and everything else like that. I'm like, why don't I just rebuild that myself? And, you know, obviously the models know what Strava is. They understand it. There's not a lot of mystery there. You know, why can't I just have my own version? that's linked to my Garmin because that's all I need.
Starting point is 00:36:11 All my data comes off my Garmin. So, you know, if I go on a run, Garmin gets the data. If I do a workout, Garmin gets the data. I just need to pull that down. So I think you'll start to see more of that sort of, you know, people going like, I've got a Mac Mini. I can just build it and keep it here. It's like if I need to access it outside of it, I'll just ask,
Starting point is 00:36:32 Claude, how do I do that? Oh, I got tail scale now and I can get to it from my phone anywhere. Problem salt. And, you know, that user doesn't have to care about scalability. They don't have to care about any of that stuff. They just need it to work for them. So, you know, how much bespoke software do we get that is no longer in the cloud? I don't have to worry about privacy.
Starting point is 00:36:50 I don't have to worry about any of that stuff because I just decided, like, I'm tired of paying for this and I want my own solution. I'm super curious, you know, this is a hot, spicy section, and you're playing with these things all day. There was an interesting report that came out, and I'm not thinking of the number is correct, but that something like 68% of tokens are focused on fixing bugs and style issues that are introduced by the LMS in first place.
Starting point is 00:37:15 And I think the second data point is we're kind of at sort of peak hype right now in terms of like coding agent use cases where we all, whether you're talking about like AI psychosis or stop factories or, you know, common terminology that's being passed around on the old Twittosphere these days. And there seems to be like a sense of the community is coming to a point where it's like, hey, maybe we've gone too far, right? and maybe we need to figure out the right duality of purpose between where these things fit, where humans fit, maybe another evolution in terms of the models themselves.
Starting point is 00:37:44 I'm curious to sort of get your sense of what do you think the relationship is between humans, coding agents, and the future of the SDLC and software development in general, and also how we think about the fact that token spend is increasingly becoming a huge contributor to people's bottom's line, and we're basically going to make decisions in tokens and people in some organizations. Yeah, let's talk about token spend first. I was talking to a couple friends of mine the other day, and everybody has pretty much Claude Max counts these days, right? And a phenomenon that we've been noticing is between developers and sort of like maybe product
Starting point is 00:38:23 folks or, you know, developer adjacent folks. And three of us in completely different companies doing different things and everything else like that have sort of noticed a similar phenomenon. the tech-adjacent folks constantly tend to blow out their clawed Macs. And you know, you get the Slack messages like, oh, I just hit my token with Max. I'm down for a day and a half. And the developers seem to never hit it. I had it happen one time, and that's only because I was testing sort of like this crazy sort of like software factory idea that was using an MCP and talking to our agents in the background and crazy.
Starting point is 00:39:03 stuff like that. And I knew it was going to blow through a ton of tokens, right? I mean, throw an MCP in there. And you let an agent talk to another agent. And that agent, one of the agents is Claude Code. The other one was our brain grid agent. I'm like, this thing's going to turn through a lot, but it was part of an experiment. I don't have a great reason for why that tends to happen. But, you know, kind of circling back to your original question, I think that, I mean, obviously we know Claude. We know everybody is heavily subsidizing these things right now. I think it was, you know, somebody like Uber said, like, they blew through their entire, like, token budget or something like that in, like, six months or some, some crazy number.
Starting point is 00:39:43 And, like, they're now having to, like, rethink all of that stuff. And so it's interesting to me because, like, I kind of, like, I want to dig into that, that idea of, like, okay, why, why does it seem this one cohort is hitting their clawed maxes all the time? but the other one is just sort of like, I don't even worry about it. Like, I'm churning out PRs like crazy and I'm never hitting it
Starting point is 00:40:07 and I don't even like sweat. I don't know if you guys have seen this as well or sort of what your experience has been. Are you sort of seeing similar things out there? Yeah, I think we're all trying to figure out in moments, but it's a crazy world right now with AI. I think token spending, like I said, is becoming even what people are measuring.
Starting point is 00:40:25 Like, I want you to spend this much. If you don't spend it, I don't even know what you use it, But if you don't use that much, you're not even going to be in our company. But that doesn't make any sense. Like token spin does not equal usable quality. It's kind of similar to line counts, right? Yeah, exactly. We all worked in a large corporation.
Starting point is 00:40:42 We're like kind of how many lines of code you all put it. And you saw how messed up to. But unfortunately, it feels like there's no quantitative measure that's easier to measure. So therefore, it becomes very much like a guessing game. But it does because. Yeah. Yeah, but it's like, it's like what would be the point of, it's like, I mean, I think one point would be like, hey, if, you know, your max account resets every, what, every week, right? If by day three, you're completely maxed out and you produced a bunch of shit that doesn't work, that seems to be a problem there.
Starting point is 00:41:17 I think you still have to go back to sort of like measuring the, you know, the output of what these people are producing, even with the agents, is what they produce testable. Is it verifiable? Does it actually, you know, all the sort of things that we've always done, you know, to protect our users from the engineers, like we still have to do to protect our users from the agent output as well. And we have to even double down, I think, even more on that verification and validation. Humans could cheat. Well, guess what? Agents can cheat too. So, like, we still have to put all those guardrails in.
Starting point is 00:41:47 And, you know, trying to measure the efficacy or the production value of anyone based upon their token spend to me, doesn't make any sense. It's like, oh, I spend a billion tokens and produced a bunch of garbage that doesn't work, can't be tested, and no one uses it. So what's the point of that? You know, it's interesting because one of the things
Starting point is 00:42:08 talking to a lot of friends is you kind of have, the leaderboards definitely drive the wrong incentives. Like, what you want from that is that people are actually playing an experience with AI. We often end up is people are just focusing on using the most tokens possible which drive, like, this spend category to your point, like, are attracted from, like, outcome potentially.
Starting point is 00:42:29 One thing I've definitely noticed, it often tends to be that there's, like, two cohorts of max token spenders. There's, like, the people that have actually figured it out and continue to ship, right? And there's people that spend a lot of tokens, but they actually just shipping slop. And the discrepancy of, like, the data, it's like, they look to save until you look under the hood, you're like, holy crap, what's all this going on? And that seems to be very common across every organization who talk to and is certainly like the current trend. And I think in the back half this year, we're going to see a pretty steep correction. Because at the moment, we're in a position where, you know, we always thought the cost for token, right, would be going down.
Starting point is 00:43:08 But in reality, because of limited data center, limited electricity and GPU supply, and also the models are now becoming increasingly more capable. They're larger. They have more parameters. You're more expensive to run. The cost for token is actually increasing. if you think about it, the cost per inference is increasing over time, not decreasing at this moment of time. That's probably a short-term thing instead of a long-term thing, but at least we have, like, what is today, a token supply gap more than anything else.
Starting point is 00:43:35 Yeah, and I think, you know, they, what is it, they go back and look at, you know, sort of like the industrial revolution, right? And they're kind of compare, I see a bunch of comparisons to that on, like, Twitter X, whatever it is, and, you know, talking about, like, there's some comparison to like, what was it the, maybe it was like the steam engine or something like that when that came up. You know, it was like, maybe it was a cost of coal. You know, I think it was. I said, oh, that was going to go down, but it actually went up because now people could run these engines longer and more often and could do more with them. So like the usage of that, oh, maybe it's not cost, but maybe it was like usage all kind of went through the roof and like did not slow down at all.
Starting point is 00:44:12 It actually increased. And it's sort of like we're in that same kind of loop right now of like, you know, there's scarcity on both. sides and now it's it's waiting for something to catch up because i don't think this is not slowing down right like this is still probably feels very much like the beginning yeah it's so fascinating time weird love to chat more but you know i know we're at the end of time so for folks that interested learning more learning more about brain grid what's the website or social channels for folks to find where you are yeah braingrid dot a i that is the best place to go and get started we're on Twitter, we're on LinkedIn, probably see us on Instagram, places like that, bouncing around with,
Starting point is 00:44:56 you know, some of our marketing folks showing different videos of things they're building and things you can do with it. But we're out there on the socials and braingrid.a.i is our main site. Awesome. Well, thanks to have you with Tyler. It's such a great chat we have. Yeah. Thank you so much. Really appreciate it. It's a great meeting both of you or Tim, I met you before, but Ian, great to meet you here. And best of luck in your venture as well. Thank you. Thank you.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.