The Pragmatic Engineer - Distributed databases with Peter Mattis

Episode Date: September 30, 2026

Brought to You By:• turbopuffer – a vector and full-text search engine built on object storage. It’s fast, cheap, and extremely scalable• Linear – the product development system for teams an...d agents• WorkOS – everything you need to make your app enterprise ready.—How is it that a software veteran who regularly shipped ~100K of database-grade code to production each year, pre-AI, feels like he’s even more productive today, with no drop in quality? Peter Mattis is co-founder and CTO of Cockroach Labs, and an original creator of GIMP. He also worked on Gmail and distributed storage at Google.In this episode, Peter reflects on his journey from open source to Google to founding a database company, and we explore how to keep systems fast, reliable, and correct at scale, from Gmail’s early storage challenges to the tradeoffs in building distributed databases.Peter tells us how AI has brought him back to writing code after his work shifted toward management, and why he believes AI can improve quality and multiply the impact of domain experts. We also consider the future of code review, and Peter has some advice about how to level up our engineering skills.Timestamps00:00 Intro02:42 Peter’s path into tech04:00 Building GIMP09:30 Working on Gmail at Google14:51 Google’s infra: google3, build files, Bazel, and Colossus21:30 Distributed storage bottlenecks23:59 Latency, throughput, and availability30:04 Contributing to libraries41:52 Google Spanner46:10 CockroachDB52:00 Manual vs. automatic sharding55:28 Consistency models and strong consistency1:00:03 Raft consensus1:06:15 How AI brought Peter back to coding1:19:12 Peter’s tools and agentic workflows1:23:08 How AI can improve quality1:26:39 Code reviews: are they done?1:29:17 100x engineers1:35:33 Peter’s advice for leveling up your engineering skills—The Pragmatic Engineer deepdives relevant for this episode:• Inside Google’s Engineering Culture• Resiliency in distributed systems• How to debug large, distributed systems: Antithesis• Pushing software engineering limits with “napkin math”• Designing Data-intensive Applications with Martin Kleppmann• Formal methods with Hillel Wayne—Production and marketing by ⁠⁠⁠⁠⁠⁠⁠⁠https://penname.co/⁠⁠⁠⁠⁠⁠⁠⁠. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com. Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe

Transcript
Discussion (0)
Starting point is 00:00:00 He built a Gimp image editor while at college, designed the original storage system behind Gmail, and built many more large, complex, and widely used systems. This is Peter Mattis, co-founder and CTO of Cockroach Labs. Before we sat down to talk, Peter told me, for the last 30 years, I've always been a prolific coder, but my current output is a bit insane. And this isn't vibe-coded junk, but database worthy, high-quality, high-performance code
Starting point is 00:00:21 thanks to working with strong coding models. Today we cover why bee trees are so important when building databases and why Peter kept breaching for the data structure over and over throughout his career. How he wrote 100,000 lines of production code per year pre-AI, stop coding between 2022 and 2024 and why he is back now. Why he thinks AI agents are lazy about testing and how this is an easier fix than it looks. And many more. If you're interested in distributed databases, distributed source systems,
Starting point is 00:00:49 or knowing how Peter manages to use AI to produce unusually high quality and production ready code, this episode is for you. This episode is presented by Turbo Buffer. A ridiculously scalable, fast and cheap hybrid search engine built on top of object storage by an engineering team I've really gone to like after spending time with them. The turbo buffer engineering team is doing something really, really cool. They're completely redesigning their storage architecture
Starting point is 00:01:10 from first principles to make storage faster, cheaper, and more reliable at scale. If you follow Turbo Buffer, you know that their storage architecture was a massive part of their early success. Redesigning a winning architecture is a big deal. It's one thing for your query plans to pass unit tests. It's another to bring each query plan to performance parity or better while maintaining correctness and reliably in production. Here's a cool part. TurboPuffer is documenting the whole thing.
Starting point is 00:01:36 Their new storage architecture, which they are calling T-Puff V3, demand some hardcore systems engineering, and they're building it in public, sharing design decisions and benchmark results as they ship. They're keeping a work log of their journey and the first post just dropped today. Follow along at turbopuffer.com slash V3. That is turbopuffer.com slash v3. Peter, welcome to the podcast. Oh, I'm happy to be here. This is awesome.
Starting point is 00:02:01 I wanted to get into. How did you get into tech originally? When did you figure out computers are interesting? I figured that out in kind of elementary school, high school. Gaming was a little bit of a gateway drug for me as a software engineer as many other people. I remember, like, early on, my mom did programming at IBM at some point. I'm never even quite sure what she did, but we had computers always around our house, like Apple 2 Plus, Apple 2GS.
Starting point is 00:02:23 I'm from that era on up. And, you know, go to the bookstore. I'd find a book on basic or magazine, type in programs. No idea what it was doing. But, you know, just kind of like I was addicted. Like, you can produce, put stuff into these computers. You'd get interesting stuff out. And then I got to college and I was like, I didn't think there was any money in computers.
Starting point is 00:02:41 I didn't know anything about it. I started a mechanical engineer following the footsteps of my dad. Oh, you started mechanical engineering as your socialization. Yeah, as my major. Yeah. And I got in there and I was like doing, like, I'd done some. computer stuff before. It was a real foolish move to do this, but first semester of doing these homework assignments and mechanical engineering, they were god-awful, like six pages for a single
Starting point is 00:03:00 problem. And I happened to take a CS course at the same time. And it was so easy. And then everyone was failing it. I'm like, I'm in the wrong field. Let me switch. And then you switched. No, I switched, yeah. What was the first software that you built, either at college, it must have been a college, that you were like, all right, this is a piece of software that I'm kind of proud of. That's a complete piece of software. Well, I mean, the big thing that I did, along with my roommate in college, we had this course, was it a compiler's course? I'm quite sure anymore. This is like 30 years ago. And we were kind of bored with it. So we wanted to do something like kind of fun on the side. And I'd done journalism in high school in my senior and senior year. I knew stuff about kind of
Starting point is 00:03:39 computer graphics and wanted to do something like Adobe Photoshop. So started just kicking the tires, and we built up this program that a lot of people know of called the Gimp. And along with the GIMP, I did a lot of the the graphics library, GTK. This has since evolved just massively since then. It's kind of interesting because after college, I kind of stepped away from it. Didn't really stay involved much past my first year out of college, but definitely people still know me and it led to other some interesting events in my career. And it was like starting Gimp, it was literally just you saying, all right, I wanted to do
Starting point is 00:04:09 something like Photoshop, how hard could it be? And then you just, this wasn't through what you learned in college, right? This was like you figuring out of how to build, you know, like a graphical. I guess engine, rendering, drawing, data structures, all of that stuff, right? All that stuff. Figured it all out. I remember trying to look at some papers back then. My roommate was looking at papers.
Starting point is 00:04:29 We're just fingering it all out. And it's one of these things. You almost kind of need to be a little bit naive to start anything like that. Because if you know how hard is going to be, you would never start it. So I think there's this like, you have to have this level of like, if you're really truly wise about the effort, you like, you never get into it. But then when you get going, you just keep on snowballing. It gets far further and further along.
Starting point is 00:04:47 And then at some point, you're like, wow, this is really. awesome. But there's an interesting little tidbit associated with our work on the Gimp, which is kind of a little bit known. I think we've talked about this before, but it's just kind of fascinating where we were getting to the point where we're like, we should release this to the public. And back then there was Usenet news groups where people would post stuff and there's one on graphics. And remember, like, just a couple weeks before we were going to release our first version of the Gimp, someone else came on there like, I've been working on this graphics program. And it did everything the GIMP did, every last thing. And then some. And we're just like,
Starting point is 00:05:18 well that sucks. I guess we'll just keep on working on. It's been fun. And then we release the gimp. Never heard from this other guy again. And I just took a little lesson with that. There's always going to be someone else working on your idea. You can't get dissuaded if they pre-announce it.
Starting point is 00:05:32 Nothing ever comes of it. And a lot of the marketing behind some of that stuff that you might hear is like, I mean, I don't know if we, like, we stole his thunder. I'm not even sure what happened. Never traced down what happened. But there's a real lesson there. So there's a possible future where you read this announcement of someone saying, I'm going to build all of this thing, and you go like, uh, you know,
Starting point is 00:05:52 someone did it and you kind of go back and just do something else and the game never happens. That's right. Wow. Yeah. I guess especially with today with startups, you know, sounds like just do your thing. Put it out there at the very least, right? The advice I give people is like you might have a unique idea, but most likely there's like a dozen people out in the world who've had the same idea. And there might be a couple of them working on it, but a lot of people don't even work on it. There is like whatever they can't get going on their idea. So, you know, I wouldn't be concerned at all. If you hear someone else working on your same idea, that's probably the case.
Starting point is 00:06:22 You know, we're working on some cool stuff now at my current company. I guarantee there's other competitors out there working on the same thing. And just, you know, it's like it's a competition. You got to enjoy that aspect of it, not be afraid of it. And what were some fun things that the Gimp led to? Well, you know, I got kind of tired with graphics. And that's why I kind of moved away from it after college. I got into storage systems and I bounced around, you know,
Starting point is 00:06:43 I worked at one of the early search engines thinking to me. And then I got over to another startup. And I met through that time period in that first startup I was doing, met Lergan Sergei from Google. Because it turns out the learning, I think it was maybe Sergey. I'm not quite sure. The very first version of the Google logo was done in the Gimp. So they knew of us.
Starting point is 00:07:02 Somehow one of them, you know, looked me up, asked me to come interview with Google. I did back in 2001. This is when Google was three years old. Yeah, three years old. And I said no. You did not. I said no. Yeah.
Starting point is 00:07:15 No, here's the calculation in my mind. I was like, Google's down in Mountain View. I was living up in San Francisco, and I didn't want to do the commute. So I went and worked at another startup for a year. And that one, after a year, like, I saw it wasn't going anywhere. And they called me back up and they said, hey, do you want to interview again? I was like, sure. Like, we're not going to be able to offer you the same stock options you got before.
Starting point is 00:07:33 I'm like, okay, I can't remember what it was. And I still can't remember what they offered me the first time. But it probably would have made a lot more money if I'd take that first. But I did all right. I'm not like complaining, but it's one of these things like, yeah. Now I think back about that I'm like, oh, a lot of good stuff came out of it. At some point along the line, Red Hat was going public. They actually offered friends and family stock during their IPO to a lot of people.
Starting point is 00:07:52 We got offered friends. I made a little bit of money off that, not a lot, but it was like just after college. It was actually quite significant at the time, you know. So, I mean, there's a number of good things that kind of came out of that. And I also used the gimp as a, you know, I was looking for like, oh, I can't support Photoshop and here's the gimp. And it did so many things. So like, I'm sure there's so many people like you had a positive impact of being able to use
Starting point is 00:08:13 this thing. And the nice thing that I, what I really liked about it is it was free, but I wasn't like stealing any software from anyone. You see what I mean. This was at a time where free software was not as common as today. Free and open source was not as mainstream as it is today. Yeah. And then the other cool thing I've heard from a lot of people, because people still, like I mentioned this, are like, oh, I've used to it. And then some people, you know, software engineers are like, I learned how to program by looking at your code. And I'm just like, wow, that was the code I wrote 30 years ago. And I wasn't nearly as good as software engineers I am And then you got into Google second time around.
Starting point is 00:08:46 You said yes. What did you start to work on? Yeah, yeah. Well, I got in there. They said like, hey, you know, this was early day, 2002. It was April 1st, auspicious date. Your start date? Yeah, April 1st, 2002.
Starting point is 00:09:00 And I got in there and they're like, well, you know, we're going to actually build email, Google email. And it wasn't called Gmail at the time. There's a code name internally. It's called Caribou. Cariboo. Yeah. And got in there.
Starting point is 00:09:12 I started working on it, and I was kind of tasked with working on the back end threading and message storage and indexing system. And worked on that, you know, pretty hardcore for the first year and a half. Actually, I was on it, I think it was three years on it through the launch, which actually happened to be April 1st, 2004. Do you remember this? There was a, it was like considering April's full joke. Yeah, I remember. So correct me if I'm wrong, but the launch said Gmail with one gigabyte of storage or something or infinite storage. and like star one gigabyte. I'm not sure what it was, but back then, most email providers
Starting point is 00:09:47 would give you about 10 megabyte of storage, the free email providers. And then you could pay for maybe 50 megabytes or 100 megabytes, but that was really expensive. This launch, it looked like an April Fool's joke because who could possibly give you a free email service with 20 to 50 times or 100 times more storage? We were talking about this internally. It was like kind of a shock and awe campaign for the industry. I think it was actually four megabytes for like hotmail or, yeah. our email and then you got this. And not only that we had so much more stores, but it was indexed really fast. So you know, you could do a search, you know, come back almost instantaneously. But can you
Starting point is 00:10:22 tell me internally? Like when the project started, okay, we're going to do email for Google. How did you and the team arrive to the point of like, okay, we will offer all this storage and we'll do it fast? Because this was at a time, if you can take us back. But what I remember is hard drives were still expensive. They were relatively slow. We're talking HDD. We're not talking necessarily. about SSD, if I remember. But if you can take us, can you take us back of what, what it was like,
Starting point is 00:10:47 what the constraints were like, and then how you innovated, like to actually, like, do something that's never been done before. Yeah, yeah. So at the time, Google internally had this large distributed file system called GFS, Google file system.
Starting point is 00:10:59 Yeah. And we were like, basically looking at numbers and being like, yeah, we think we can build on top of this. They had a lot of knowledge about how to do search and retrieval. It ended up being that, like, some of the existing systems
Starting point is 00:11:10 they had for search and retrieval, they started building the prototype on there and then that was completely written. That was actually what I got involved in because I got in there and there's already a prototype and then it was completely rewritten. So I was on that, you know, threading side. We decided to have message threading right from the get-go, which is also kind of innovative. That was in common in email systems. And spreading, do you have the message shred, which I assume needed data structures on the back
Starting point is 00:11:31 that it needed like, you know, storage and figuring out the read intensity of those kind of things. Yeah, you know, there's some bee trees involved in that. I think that might be the second time I implement B-D-D-Sorture. trees and I think I've implemented them like a dozen times now. How do B trees relate to email threads? It's just like you have to have some storage there where you have a thread ID, a message comes in, you have to look up. It was using the search index to match actually the subject or the message ID into the thread. And there are some other things taking place in the B tree as well because we were keeping track of the unread counts of threads and whatnot. So, you know, I can't even remember all the
Starting point is 00:12:05 details now. This is ancient history. This is 2004. What are we in? 2026? 22 years ago. So it's like, left my memory, but there was definitely B trees. There was also this, you know, inverted into next taking place there. All this code has since been completely rear in. What were the economics? Like, in the team, you must have done the economics of being able to offer free email, which of course, I'm sure you did some mats of like how it could be subsidized, but to like not make a terrible, terrible loss. Yeah, yeah, yeah. No, I mean, we were kind of a kicking around ideas. And, you know, one of the ideas that had come up at some point was, hey, maybe we can put ads on this. And it was one of these crazy things. wasn't involved in this. This is Paul Bukite,
Starting point is 00:12:43 went on to do some cool stuff at a friend feed and Facebook and ended up, I think, a partner white combinator at some point. But he just like one night, he's like, no, I think I can just take out of some of our existing ad functionality and incorporated in there. And it just went like gangbusters. And there's
Starting point is 00:12:59 a huge business that, you know, kind of grew up from that, which is kind of incredible. So, I mean, I think there's a lot of things that's just like, you kind of have a sense of what can be alone. Like, oh, this is going to cost, you know, per user a couple of dollars per year. How do we monetize that and we didn't want to charge for it. Eventually they did charge for it because you have the whole Google workspace stuff. But at the time, it was more like, can we do this for free?
Starting point is 00:13:19 Oh, it looks like it can't be economical. And it took a little bit of a leap of faith. And do you remember the launch? The reason I'm asking, because I remember that it was an invite-based system. Like, not everyone could get in. And I assumed that must have been to control the expected demand. Because again, you were, you're offering something for free that was paid before. And it was kind of pretty obvious that there would be massive demand. Yeah. How did you think about it, kind of monitored demand, decide how many people to onboard? Honestly, this is a team effort.
Starting point is 00:13:50 And I wasn't involved in that. I mean, I do remember being worried about like the load that would happen. And then someone came up the idea of like, hey, we should do an invite base system. And this also had like, it played dual role. So it kind of constrained the growth. But it also kind of drummed up this excitement. Oh, can get me a Gmail invite. I remember people asking me at the time was like, yeah, I can get you on this
Starting point is 00:14:08 my Gmail invites as you want. And then after after you built the system, where did you move on to? What was your next project? Was it the built system? Well, there was the build system. I mean, that was kind of muddled in my mind because I was kind of doing that part time when I was doing Gmail as well. At some point, you know, Google actually had this large monor repo.
Starting point is 00:14:26 I think they actually still have the monor repo. They still have the monor repo. And they started out with the Google One repo. That was before my time. Then they moved on to Google 2. That was what was there when I got to Google. And at some point, we saw the strains of Google 2. And Google 2 was just one single monolithic make file.
Starting point is 00:14:40 I think it actually had some sub-make files, but it was this really large unwieldy make file. Someone came to me and was like, I think we could do something. I'd had some interest in the build systems. And I kind of put the foundation in for Google 3. There was a bunch of people involved. I was kind of doing like kind of the initial work on Google 3.
Starting point is 00:14:56 And Google 3, the initial insight was like, hey, make files kind of sucked a right. We introduced this thing called build files. And I decided to do it as just this stripped down kind of Python language. But it was still Python at that point. And what G-Config spit out at the end was a monolithic make file, but one you didn't have to write. Over time, this evolved. It became Blaze internally.
Starting point is 00:15:18 There's some other systems that are associated with now. I don't even know how complicated it's gone. I haven't seen it for a long, long time. Then that became basal externally. It became Buck. There's some other people went to Facebook and they like, that was awesome at Google. What was the reason that make files or make didn't really work? Was it, were we talking about build performance?
Starting point is 00:15:35 Are we talking about maintainability or readability? Yeah. So, like, I think Make itself is like, it's okay in declaring dependencies, but it's a little bit like kind of assembly language. And you didn't want to write, you know, all your dependencies in assembly language. And if you didn't know what you were doing, it was easy to make a mistake and miss dependencies and whatnot. And so the kind of my thought behind the build files is like, hey, you need to express these dependencies, but at a higher level and just like kind of cleaner semantics associated with it.
Starting point is 00:16:03 And from that, you can compile down into the assembly. complete language. And then eventually folks were like, oh, you don't have to compile down to that. We can just kind of implement, you know, the dependency kind of update engine directly. And then that's where the performance improvements were able to come in. So like, because at Google scale, like when you have a large repo, in general, like, my understanding is that the reason Basil and Buck are so popular for large code basis is it can help you improve your build performance. It gives you a lot more levers to play around with from caching, from being smart about cache generation to obviously just raw performance.
Starting point is 00:16:35 Yeah, that's exactly. So you kind of just dabbled and like, okay, I'll make this build file. Did that people take that over? But what was your next main focus? Oh, well, so I mentioned GFS earlier, Google File System. And at some point, we realized that there were some limitations in GFS, scalability bottlenecks. Because I've been working on the storage system for Gmail, I actually dabbled at another storage system, you know, kind of a research thing. They never went anywhere.
Starting point is 00:16:59 But because we're working on that, I got invited to participate in, like, the founding team of Colossus, which was the successor to Giac. And as far as I know, Colossus still exists. It's gone through multiple iterations at Google, but it's like the second generation, you know, distributed file system. So what is Colossus? Yeah. So when I say distributed file system, externally, you might think of something like S3. Kind of blob storage, it had a flat name space, kind of like S3, you give names.
Starting point is 00:17:24 And there's like a minor hierarchy there, but it's like very limited. But it's not like a Puzzix file system. So you don't have the full directory hierarchy. You didn't even have the full like kind of permission system. A lot of that stuff got it added later. But the files are not. stored on your local machine. There is a fleet, you know, kind of a service out there that has all the files. They're writing it down to their hard drives and now SSDs and your client can access that.
Starting point is 00:17:45 And it's all replicated. So if there's any crashes and whatnot, you're not losing your data. I don't know when at some point S3 got erasure coding. We did a ratio coding in classes. That was kind of a big breakthrough. Which coding? We used Reese Solomon. So for the audience who's not familiar with erasure coding, you might think of like, I want to have replicas. And there's this thing in hard disk is called Raid, where it's like, I don't actually have to have full replicas. I can actually, you know, if you take A plus A and B, you can X-Werm together, and then you can have this kind of third kind of version. And there's more and more complicated versions of that. Reed Solomon is kind of, you know, I think there's actually better codes now,
Starting point is 00:18:21 but it's like one of the known ways to do this. One of the known ways, yeah. And we had to, you know, kind of pioneer internally like, oh, how are we going to actually make this work in distributed file system? Where you focused on latency, on being able to store data more efficiently. It's storing it more efficiently. So in GFS, and there's varying costs where it's like, well, you're not storing two copies or one copy of data or two copies is storing three. Triplication. Three times as much storage you're having to use. And with Reed Solomon, you can get that down quite a bit lower. I can't remember offhand exactly what are, what we used for Reed Solomon. I think it was like essentially two X. But you get that with also the redundancy too. So it's like
Starting point is 00:18:56 it's smaller and the redundancy is higher. So because I guess the naive thing, if you're saying, all right, I want my data to be replicated at three places. You take three notes, three machines, physical machines, and you say like copy one, copy one, copy one. I have it three places. Great. If one explodes, I still have two, wonderful. And then you're saying that the algorithm here is you could take not three X
Starting point is 00:19:18 the data, but two X of data, split it smartly across machines, or maybe you could take it and take it lower. And you still have the thing where like, oh, one of them explodes. I still have all my data because it's split enough. That's exactly right. And like kind of the mental model, you know, you just want to understand at a high level, which is essentially you might want to say, like, I want to have eight replicas of this data.
Starting point is 00:19:37 I think, I can't remember offhand. I think S3 might use nine. They've actually talked about this publicly. You have kind of nine chunks of data, but any five of those chunks can be used to reconstruct it. And what this means is you can lose any four copies and you can still reconstruct your data. And oftentimes it's more like, you know, the first five chunks are exact replicas and the other four are kind of parity ones. I've kind of forgot some of the details that's escape my mind.
Starting point is 00:20:00 But it's along this. Yeah, but when you come up with an algorithm, you can then prove that this algorithm will work, right? Like, this is a little bit like, I know in software engineering, like, maths and algorithms a bit out of fashion. But in this case, this is really important because once you, once you can prove it that this algorithm works, it will work. It will work. And, you know, it's like the math behind here is like a Gaw fields over like GF2, something like that. I don't even know that. I never actually understood the full math behind it.
Starting point is 00:20:27 I always regretted not doing more math in college. but you didn't have to you. Like read and solve them and proved how this work. I think it was like back in the 1970s, something associated with like communication network. So you just take that and, you know, kind of use that expertise, but leverage that and have to do all the engineering behind it to make it work in a storage system.
Starting point is 00:20:44 And then when building a distributed storage system like Colossus, what were other things beyond, okay, you want to store data in a resilient but efficient way? What were other things that problems that you needed to solve? I'm thinking things potentially like sharding and resharding, or metadata being important, those kind of things. Particularly for the scale that Google wanted to operate at, GFS kind of had this scalability limit.
Starting point is 00:21:09 I believe it was like, you know, you could have a thousand machines in a GFS cluster which I guess sounds big until like today, it's kind of really small, right? It sounds big for many people. And then Google was like, no, we need to have this scale up to 10,000 machines. And there's some bottlenecks. There's this GFS master, it's a single node.
Starting point is 00:21:27 It was a bit of a bottlenecks. We're like, oh, we need to have a distributed and master to store the metadata for all the objects. And Google at the time happened to have this system called Bigtable. And so Colossist stored its metadata inside Bigtable. And one of the things I'm kind of proud and like, kind of like also, you know, a little bit embarrassed by one of the design choices I went down, but actually worked out is we wanted to use Bigtable for the metadata for Clossus.
Starting point is 00:21:51 And the metadata, just for those of us not as into this reason, what is the metadata and the distributed file system? Yeah, it's like the names of the five. And for each of the files, the files have broken into chunks. What were they, 64 megabyte chunks. And then you have to have the list of chunks for each file. Yeah. And then you have to periodically, you know, the master has to be scanning over this and
Starting point is 00:22:10 doing repair, repair work. And, you know, but there's more metadata in that, but that's in a nutshell. So we wanted Colosses to have this, you know, kind of scalable, you know, service, Big Table to store its metadata. The biggest user of GFS is Big Table. So we want Big Table to work on top of Closses. So you have the circlic defendants. Yeah, well, there's a bootstrapping thing, right?
Starting point is 00:22:31 Boostrapping, yeah. Oh, yes. Yeah. Which one starts up? Like, you need to mock something somewhere, right? Yeah, no. I mean, the way it actually worked at the time, and they've since replaced this, you know, because this is like the way to get started and leverage what you have and eventually you kind of
Starting point is 00:22:45 get rid of it. But there is the, uh, kind of a foundational big table. That big table didn't use classes. There's the classes using the big table. And then there's normal big table sitting on top of classes. And it all worked for years. So I don't even know when they got rid of that. got rid of it at some point, but that was...
Starting point is 00:23:01 But I guess sounds like you can make, like, hacks that go really long, knowing that they're hacks and they get you off the ground, right? They get you off the ground, right? Because if we had to, like, you know, implement that kind of big table layer from the get go, or it just delayed how long it took to get, you know, Colossus built. One of the things that strike me about a system like Colossus is it promises, or this was internal to Google, but even distributed file system that are external, they will promise high things. throughput, high availability, and low latency. And to me, it's always a bit conflicting of like, well, it's pretty easy to, I guess,
Starting point is 00:23:38 build a distributed file system with my limited knowledge. I could probably do something where I have either high throughput but high latency, because whenever I write something, I read it out to all the replicas. You already, like, mentioned one technique of doing it, but how did you kind of reconcile? Like, how did you get, like, low latency while you have high throughput, while you also have a replication going on on the file system? Well, I mean, these distributed file systems, I mean, the latency isn't super low. In particular, when they're running on hard disks, which Colossus was doing at the time, which S3 does, you actually notice the latency.
Starting point is 00:24:08 So, like, S3 is a high performance system. Google GCS, the competitor from Google, which is built on top of Clossus. It's high performance has incredible throughput, but the latency is like 20 to 30 milliseconds for first read. And that is bounded by your hard disk latency. If you put it on SSD, it gets down to closer to SSD latencies. But not actually kind of the state of art. SSD latencies, which is kind of this crazy thing that's been happening in our industry. It's just like how much faster the hardware has been getting.
Starting point is 00:24:34 Reading from a hard drive, maybe five to 10 milliseconds nowadays, and we never touch hard drives. Reading from an SSD over NVME, 30 microseconds, 50 microseconds. So that's a microseconds. There's a thousand microseconds in one millisecond. So we're talking like a huge, huge difference. Well, there are now startups for infrastructure companies that are starting to take advantage of the fact that they can have an NVME layer. and they pull things up either predictively or not.
Starting point is 00:25:00 But as you say, like, when the physical reality changes, you can build systems on top of it, that should take advantage of it. Yeah, absolutely. You know, stuff. The disks are so much faster with SSDs, the networks are so much faster. I mean, just kind of crazy how fast,
Starting point is 00:25:14 like the intra-zone leitzees are to Google or Amazon Center. And it wasn't like that, but when we were building colossus, you know, I can't even remember what the numbers were, but milliseconds to do a network round trip. And now it's down in, you know, 100 microseconds within a zone. I mean, I just look at these things and I'm like, holy crap, you know, the hardware guys have really done a good job.
Starting point is 00:25:32 Yeah, sometimes I feel that our software should feel way more snappy. And there are some snappy software, but sometimes I almost wonder if we're getting too complacent with all the abstractions or not even doing this like napkin maths. Simon Erickson at Turbo Buffer talks about this napkin math word light like you. It was like, all right, here's the theoretical limitation of the hardware. reach from SSD might be, I don't, 30 microseconds. And then, like, how can I build a system that is as close to this as possible as opposed to the other way around saying, okay, like, you know, human will notice like 20 milliseconds, like 100
Starting point is 00:26:05 milliseconds. Let's build around that. Yeah, yeah. And sometimes, like, when you're architecting something, you need to think about the human kind of perceptible latencies. But oftentimes when you're dealing at the storage system layer, you know, I know Simon working on TurboPuffer, they're doing great stuff over there. You have to think about the machine scale and the machine speed, which is a lot.
Starting point is 00:26:22 a lot faster than human perception. Like, a human can tolerate maybe 100 milliseconds of delay, or if you're playing a game, maybe you need to have, like, you know, frame rates of like every four milliseconds. But the machine wants much, much faster in that. Simon says, napkin math, I call this speed of light numbers. And, like, sometimes it's literally the speed of light bottlenecking you. The cross-zone latencies between zones, cross-region latencies, is speed of light and fiber. You want to hear a kind of crazy fact?
Starting point is 00:26:47 I love hearing crazy facts. The fastest way to send active data across the globe is to send in a descent. space. Is it because the speed of light is faster in vacuum? Quite a bit faster. Or, well, it's not full vacuum. No way. So you cover a higher distance. Yeah, well, you actually, I believe the way to do this, you send it straight up. And then with Starlink, you send it straight up, you bounce around between Starlink, you send it down to the other side. So you want to get out of the atmosphere as quickly as possible. But now, if we're effort on my speed of light, this will also go. There's like a digital transformation happening. So you would need to calculate how long it takes for
Starting point is 00:27:20 that system to process and do it. But you're saying that even if you do this like really well, it will be faster than beaming it through an optical cable. And an optical cable slows down the speed of light, right? Yeah, yeah. Yeah. The speed of light is only the speed of light and vacuum in every other medium. It's slower.
Starting point is 00:27:35 I was talking. I did a deep dive on the hedge fund industry. And they didn't tell me exactly. They said that they do use satellites and microwaves and some of these things. They will not tell you because, you know, this is their thing. But I had a suspicion that they might have found a faster way. And I think this is like somewhat well known, but the details they're not going to get into because, again, that. But like, yeah, so they're probably balancing stuff in space.
Starting point is 00:27:58 Yeah, yeah. So, I mean, one of the things that they do in the high frequency train is like between New York and Chicago, it's actually not far enough distance-wise to make it worthwhile to send it in space. So they were doing microwave beams there. Yeah. But, I mean, if you really want to get faster, you need to build like a vacuum to you between them and it's like send a, and maybe they're doing that. Yeah. But, I mean, the thing that I think is like just kind of awesome about performance nowadays is, I mean,
Starting point is 00:28:20 There's so many layers you have to be paid attention to in terms of performance. The rabbit hole just goes so deep there. You're almost certainly running on multi-threaded systems, right? And you're like, oh, I need to have a multi-threaded program. I need to have synchronization in there. Well, you get the best performance if you just kind of avoid the synchronization. And part of this is lock-free programming, but part of it is arranging so you don't need locks at all.
Starting point is 00:28:41 You need to carry about your processor caches. And there's this whole kind of setup of caches above the CPU. Like you can think it registers of the cache, and you have L1, L2, L3. even your memory and onto disk, and there's just you have to pay attention to all those levels. And if you do, your performance gets way better. And if you ignore it and you're like, oh, I'm not worrying about kind of, kind of the cash access is.
Starting point is 00:29:04 I'm just accessing data all over. Your program will just be way, way slower. And some folks pay a ton of attention to this. The high frequency trading, I mean, they do this all day long and in so many other places like no one pays attention to it and you get this kind of gradual degradation of the performance. of the hardware or the software. I did want to talk a bit more about low-level stuff,
Starting point is 00:29:24 but not about the speed of light, but low-level data structures and programming language features. You made some contributions to the standard library, right? Well, I've done a couple. Not quite standard libraries. So, I mean, I just, like, have been always fascinated by data structures. It's kind of awesome. I mean, I think just algorithms in general,
Starting point is 00:29:43 and you're just, like, sorting algorithms, just kind of awesome. You could probably explain, you know, insertion and, you know, to kind of every... Could you explain? quicksort as well. Quick sort's a little bit harder. Right, that's the thing. Yeah.
Starting point is 00:29:55 No, no, but it's a smart one. Yeah, and then you get to these levels of like, you know, it's like, oh, wait, someone really smart came up with this. So one of the things I worked on, just as a little bit of a side at Google at some point, was a colleague came to me, and he was like, you know what? We're using the STL map structure all of the place. And the STL map structure is a balanced binary tree. I can't remember if it was red black trees or one of these other balancing algorithms
Starting point is 00:30:18 that any CS college students implement. And he came to me, he's like, I think we could do better because, you know, there's actually a cash problem here. Every time you're every node you're traversing down, you're going to a different cache line. And he was thinking about using something else, a skip list to do this,
Starting point is 00:30:33 which that's another awesome data structure everybody should kind of look at. But at some point I was like, actually, this feels more like a bee tree. So I implemented bee trees a couple times before and figured out like how to implement a bee tree that implemented almost all the semantics of the STL map.
Starting point is 00:30:50 It couldn't quite do it perfectly. And the reason is when you insert into a B-Tree node, you have to shift stuff around so you don't get pointer stability. This is just kind of fundamental. But if you can, you know, you don't need that for your use case, you actually can pack more data in.
Starting point is 00:31:04 So the thing about a B-tree, like the real easy way to describe this, is like you just have a small list of items, like eight items. The best way to store that, if you want to kind of fast access in sorted order is just to sort the items, right?
Starting point is 00:31:17 Literally not to have a tree at all. Yeah. For like eight items. So I have it in the very simple list. Very, very, just an array, sort the array, and then you can either do a linear scan over it. You can do a binary search. And oftentimes, the linear scan is faster. And then you think about that I'm like, well, if I want to store worn eight items,
Starting point is 00:31:30 I can just have one node that has eight items. And then I have another node. And then you have a parent node that connects them together. And that is essentially like the, you know, you build it, you think about building it bottom up. You start with just one node of eight items. Oh, I need to insert the ninth thing. I'll split into two chunks.
Starting point is 00:31:46 And the two chunks, one will have four, the other will have. five and then you have an apparent node that points to them. And then you just kind of recourse on that. That's the B-tree algorithm in a nutshell. Everybody go implement it. Actually, nobody should implement this anymore because nowadays we have something else that we'll implement this in all the optimizations because there's a crap ton of optimizations that you can do on a B-tree. But I just want to go back to this. Like there was already an existing implementation for maps. And then so your your colleague looked at the code and said, I think we can do better. What I want to figure out is like in my mind, in someone sitting outside,
Starting point is 00:32:18 of, you know, I'm not involved in how some of these libraries are, or data structures are built. I always thought, and again, this might be naive, but really smart people sit down, they kind of look at the state of the art, they implement it, and there's no way it can be faster. In fact, I've had arguments in the past saying, like, oh, let's write a faster sorting thing. Like, it's surely it is the fastest. But if you could bring us a little bit of, like, how, like, you've been inside of how it actually how it happens and how other people like yourself and and your colleague can say like,
Starting point is 00:32:51 oh, what, what if we, what if we try something else? Yeah, yeah. So I mean, my recollection here is he was working on this kind of the big internal system. I think it was called Gaia that actually had the mapping from, you know, you log in, you have your user ID and you have to look this up. And they were storing, you know, all like this, the map from user ID and email to, you know, whatnot to the metadata about the user in STL maps. And you just know, it's like, well, there's a lot of memory usage here and it shows up on profiles. And then we're like, well, what can we do to do better? And that was kind of the genesis of it.
Starting point is 00:33:21 And he happened to be working on it and he happened and they were working with me. Like, we just started kind of noodling on this problem. Like, oh, can we do something better? And it's not one of these things like, I think now with Google, they have a whole team working out of their kind of internal libraries. At the time, it was more of like, you know, everybody working on their own systems and contributing to a shared base. But I guess it still goes back to what you were just saying of like,
Starting point is 00:33:42 just go down the layers, try to understand. And if something just doesn't add up, like, suddenly like, oh, there's this big exposure memory usage, like, you know, just ask the questions, why is this? And if you're able to or you happen to be like, oh, can we do something about it, right? And one of the things that, you know, he observed earlier on. I think part of one of the things was it was like a map from integer ID to something else. And you know, like if you look at red black tree, every node, you have your kind of value that you're storing the map. And then you have two pointers. You might have an energy ID that's like four or eight pipes.
Starting point is 00:34:12 And then two. It's a waste. Yeah. And you look at it and you're like, oh, that seems like a lot of overhead. And you're like, you could just a bit. Well, the bee tree actually has a lot better. It has better spatial locality and that's what made it faster. But it was actually smaller as well at the same time because you had less pointers involved.
Starting point is 00:34:28 You also contributed to go, right? Yeah, yeah. Well, that came later. Yeah, it came a lot later. But can we talk about that? Yeah. I mean, just one of these other things, you know, I pay attention to like, you know, when there's research papers coming out about new data structures and like hash tables.
Starting point is 00:34:41 Hash tables are like the, one of the earliest things you. learn about in college and data structures. Like, how do I map keys to values where the ordering is unimportant? That's when hash tables come in. And there is like, you know, the very earliest ways to do this. I implement hash tables multiple times is like, you take your key and it might be a string and you put through a function and it spits out in integer and then you map that into an array of buckets. And if multiple things map to the same bucket, you have to have a link. You have a linkless. Yeah, this is a naive implementation. naive implementation, used quite frequently.
Starting point is 00:35:11 And over time, people discovered, like, a lot better ways to do hash tables. There's very, like, that's called a chaining of your hash. There's another technique called open addressing, where instead of actually having a linked list, you just kind of hash it again and move on or kind of walk down to subsequent buckets to find out, like, or subsequent slots to find out where you should be. And I remember reading about this new technique and it came out of some folks at Google. I believe it came out of their Swiss office, because it was called Swiss tables. I believe that's where the naming came from.
Starting point is 00:35:42 I'm not 100% sure about that, but I remember reading about it. And then I was working on Go for a long period of time, and Go has this built-in map structure. And it's a hash table. It's a very highly optimized hash table because the Go team is very competent, the Go Run team.
Starting point is 00:35:57 And various folks have taken an attempt at like, you know, putting together a Swiss table implementation for Go. And I tested some of them, and I was like, this is kind of fascinating what Swiss tables do, and I'll explain how it works. in just a second. But I looked at it. It's like, well, it's really hard to beat the performance of the runtime.
Starting point is 00:36:12 The runtime was really good. And I kept on, I noodled on this for a little while. And eventually I ended up having to take this business trip to India, to Bangalore. And so I was on a long flight. No. Yeah. It always starts like this. And I'm just like, I'm just going to try to pull on this.
Starting point is 00:36:29 I pulled on it sufficiently that I can get some of the benchmarks to be faster. Wow. And then I'm like, you know, that like is kind of like catnip for an engineer. Like kind of make it all faster. figure all the rest of it, you know. Got some help from the runtime folks. There's an issue on the Go issue tracker that, you know, where other people have been attempting this.
Starting point is 00:36:45 Because people propose like, hey, let's use Swiss tables. And like, like, oh, first is like, well, you know, we don't quite know all the details. You're going to have to navigate this and that. And there are some ideas there that combined them together and got to the point where it had kind of a complete implementation that was faster on, on most benchmarks, not quite all of them, but most of them. And then the Go folks eventually picked this up and push it over the finish line.
Starting point is 00:37:05 And then so you, you know, like, You came with the idea you got to the point where you were able to show an implementation that showed how some of the benchmarks were faster. And then you started to work with some folks on the goal team too. Well, I didn't, it wasn't quite that. I came up with an implementation that we ended up using at my company. It was good for our use case, but actually putting it into the runtime as a whole other, you know, kind of level. But then you just showed like here is this implementation and then they, they kind of took the inspiration and the ideas. Yeah, they're like, well, this is great. We want to make all. They always are looking for ways.
Starting point is 00:37:37 to make it run time faster and there's like, you know. Oh, wait. And then so, so you did this in this, this is a lot of years after you left Google, right? Yeah, yeah. So this was from the outside. This is from the outside. That's awesome. Yeah, and other people contribute stuff to the outside as well.
Starting point is 00:37:49 You know, we had another colleague, um, he contributed one of the CRC implementations, you know, adapting some stuff. CRC, uh, cyclic redundancy checksum. Mm-hmm. You know, Intel published some papers about here's how to do the CRC and assembly very fast and he contributed one of the implementations. You see a number of those things where, you know, people just like are contributing externally. It's not a lot, honestly. I mean, I actually don't know the full details, but, you know,
Starting point is 00:38:12 people are regularly contributing to these things. Peter just described how the Swiss table work came together using an issue tracker in the GoT tracker with different people contributing to the work and then the Go team pushing all of this over the finish line. This is where I need to mention our season sponsor, Linear, which is a place to coordinate work between humans as well as agents. One thing I've noticed about how most of us work with agents is how it's a pretty single-player thing. You open a terminal UI, go back and forth with an agent, and it usually produces a PR. But the rest of your team has no idea what happened in that chat unless you tell them or copy the whole history. And when everyone in the team works like this, a lot of work happens that's invisible to the rest of the team.
Starting point is 00:38:50 Linear's take is that agent work should be teamwork. Even today, teams already use linear to define the work to be done. Now, Linear can already delegate an issue to a coding agent. This agent could be an AI agent that linear integrates with like Codex or Cursor or Linear's own agent or a custom agent. Either way, the engineer and delegating stays responsible for the outcome. What I really like about how linear works is how the work stays visible. Your teammates can follow the session of the agent, check out the PRF producers, and join their review. We've gone from single-player work to multiplayer engineering work with agents.
Starting point is 00:39:21 Oh, and one more thing I like about linear, a focus on costs. Linear agents' auto-riding chooses a model that is the best suited for the task. Teams can also inspect usage and set limits, track usage, so you can use capable agents without having cost balloon out of control. Hop on board at linear. app slash pragmatic. Peter also previously mentioned Gaia, Google's internal system that map logins to users. It's not surprising that Google custom-built all systems,
Starting point is 00:39:45 including this one, but most of us won't build our internal GAIA. This brings us to our season sponsor, WorkOS. You can think of WorkOS as something like Gaia for the rest of us, identity infrastructure you've otherwise spent quarters building yourself. Workerless includes single sign-on, skim, directory, sync, audit logs, role-based access, control basically everything a big customer security team asks for, delivered as a handful
Starting point is 00:40:07 of clean APIs. It's how companies go from, we have a login, to we can sell to a Fortune 500 without standing up their own internal identity platform. And WorkOS is already building for the next version of the agentic authorization problem. Their newest product is Airlock, the authorization layer for AI agents. Think about what happens when you had an agent attached like clean up the sale opportunities in our pipeline. The last thing you want is for this thing to have standing permissions to delete what
Starting point is 00:40:33 whatever it likes. Airlock sits between your agent and the tools they call. It evaluates every request against the agent's intent and your rules, and then it allows it, denies it, or routes it to human for approval. You write down the policies in plain language, the agent never sees your credentials, and every call and verdict gets logged. It works with coding agents like Cloud Code and Codex and with MCP gateways. So if you're working out how to let agents do real work in production without over-permissioning them, take a look at at WorkOS Airlock at WorkoS.com slash airlock. And with this, let's get back to Peter and why he left Google after building Colossus.
Starting point is 00:41:11 You're at Google, you're building Colossus distributed phall systems. You're at this point probably working on probably the larger system on the planet, honestly. Why did you even consider leaving? Yeah, yeah. Well, after Colossus, I kind of dabbled in this other project called Google Goggles for a little while. Remember the glass holes? Yeah. Keeps coming back, the idea, by the way.
Starting point is 00:41:32 Yeah, yeah, no, it's still here present. Seems like Google was early. Yeah, yeah. And, you know, I think that was a technology before its time. I don't think it was ready to do at that point. It looks like the actually doing the glasses is quite a bit harder. You know, the Android phones we were trying to power it on were, you know, not powerful enough. And then, you know, I just kind of got wanderlust, you know, like, you know,
Starting point is 00:41:51 am I just kind of stagnating here at Google, which is a strange thing to say, but, you know, some other people feel it as well. And decided to go off and try my hand in another startup. That didn't work out. We got Aqua hired by Square. So I want to pause for a second. So this company, what was the company name? The company that we found is called Viewfinder.
Starting point is 00:42:09 Viewfoundure. Yeah. It was in the mobile photo sharing space, which should sound familiar. This is like Instagram. This is like Snapchat. This was in 2012. Yeah. Right as Instagram and all the more we're taking off.
Starting point is 00:42:20 Yeah, yeah. We were right there in the play. And we just didn't have the right go to market kind of like how to track the users, how to get viral growth kind of thing. Because from the outside, like what I read, when I, when I check, you know, the story just like, oh, you know, like you, you co-founded a startup. It got acquired by Square. Hooray, like it sounds like you had bigger ambitions. And this was a decent outcome, but not the dream, right? Yeah, yeah. No, it wasn't dream at all. So the term I used was aquired. So sometimes
Starting point is 00:42:48 a company will get acquired, get bought for, you know, their IP, for their product, for their business. And other times, they get bought just for the talent. The people. The people. And we got bought just for the talent. They acquired the IP, but I don't think wherever did anything with it. It wasn't kind of like where they were working. But we built up a kind of strong technical team and that's what we were hired for. And, you know, like, I can't remember the details. We'd raised a small amount of money. We were able to pay our investors back, make them whole.
Starting point is 00:43:13 Maybe they got a little bit of a haircut, but maybe they actually got a little bit of a, but it was essentially, they got their money back, which is like, you like, you know, as a founder, you, uh, you know, investors are big boys. They're used to losing their money, but you kind of feel bad if you lose a lot of money for them. So, you know, getting them paid back kind of makes you feel a little bit better. Yeah. Yeah. And then you, you spent it a little time at the company that acquired you were just square.
Starting point is 00:43:36 Yeah. And then you started itching that a little bit again. Yeah, yeah. Yeah. Because, you know, we've been working on these, you know, distributed file systems and storage systems, you know, Colossus. One of the kind of sister projects to Clossus is Spanner. And how is Spanner different to Colossus? Well, Spanner is essentially a distributed database.
Starting point is 00:43:55 Colossus is a distributed storage system. And the way I think about the difference between a distributed storage system and distribute database, you might think, oh, they're both storing data. I was about to ask because a distributed database will at some point be a storage system. Right, right. So what's the difference? Yeah. So for Colossus, you know, it was targeting large files, large append only files. You can't update the append only.
Starting point is 00:44:17 Yeah. Yeah. Large append only file. So, you know, 64 megabytes, maybe up to gigabytes in size. But if you're a database, you want to be storing like kind of small, like, you know, kind of You know, if you're using SQL or like the relational data, you might have table with, you know, billions of billions of, you know, rows. Those rows are broken up into columns.
Starting point is 00:44:36 The columns are typed. That just has a very different nature to the engineering challenge for database than it does for a distributed storage system. And usually distributed databases are implemented on top of some kind of distributed storage system. And that was the relationship. So Spanner was implemented on top of Colossus. And some of the design decisions in Spanner kind of directly fell out of the appendage. only nature of the files in classes. You can't update a file in place, so you have to, you know,
Starting point is 00:45:03 make the files immutable in your database. And this is where, like, you know, log-structured merge trees, you know, come into play. And they weren't invented at Google, but Google really popularized them with Level DB, which emerged out of the work on Big Table and Spanner. That got popularized into RocksDB. I subsequently re-implemented one of these things, and this is what we use at Cockroach Labs. It's called Pebble. So I'm very familiar with the internals of that. But it's kind of all based on this idea that the data is kind of stored in the mutable files. So how did you decide to found Cockroch Labs? While we were working at Square, my co-founder and I, there's actually three of us.
Starting point is 00:45:39 We were all at Square. And one of them is Spencer, I mentioned earlier. He was working on the Gimp. He's my college roommate. And he was also at Google. He was also at Google. This other guy, Ben Darnell, who was also at Google. He joined us at Viewfinder and ended up at Square.
Starting point is 00:45:52 We, you know, like, we're kind of just noodling on a project to do. And we'd actually had this design. back in Viewfinder were like, ah, we didn't, we looked around for a data space to be using in order to build viewfinder on top of. We didn't really like the things that were out there. The technologies that side Google looked better. We had the big table. We had Spanner and whatnot. We were looking around and, you know, like H-base existed, but I wasn't quite happy. And there's some other systems like React and others. And, you know, in one point, we're just kind of like, you know, came up with the design for Cockroch to be an initial design. And I was like,
Starting point is 00:46:23 no, no, guys, we're doing a mobile photo sharing site. We shouldn't build a distributed database. So we put it on the back burner, which I think was absolutely the right thing, maybe, or maybe we should just pivoted away from doing the mobile photo sharing site, given the way things worked out. And then we got to square and we saw some of the same problems that they were experiencing with data storage systems. And Spencer is very convincing. My here's convinced some of the management that like, hey, can you just work on this part-time and see if it had life behind the design? And then kind of conscripted Ben and I into it. And eventually it started gaining attention externally. And we're like, hey, can we go and
Starting point is 00:46:54 spin this off into a company? And that's what ended up happening. And so you started a company, but I understand you didn't raise VC funding initially, right? That was the case at Viewfinder. We did it differently. At Viewfinder, we kind of eschewed the VC money. And, you know, in hindsight, I wouldn't recommend that. You wouldn't recommend? Yeah, I would recommend taking the VC money because my experience, the VCs are very, very intelligent.
Starting point is 00:47:21 They can help you navigate a lot of challenges. You know, I think sometimes there's this perception. you know, the VCs will push you into various areas. And maybe there's some bad ones out there that do. The VCs I've had experience with, just like some of the sharpest people, you know, I've ever met. And so you kind of have like an extra like person helping you on the team pretty much. Mentoring you, giving you guidance, telling you what they're seeing. They give you advice seeing what they're seeing in the market where things are going that is very hard for sometimes for a founder.
Starting point is 00:47:47 Especially as a technical founder, right? That you're, your focus on the engineering part. Yeah, yeah. Yeah, no, we actually took money right away for Cockroach Labs. it was almost like as soon as we left we got and did a little road show you know kind of in the Bay Area and got some interest and you know
Starting point is 00:48:00 got an investor right away. I have to ask about the name though. Yeah. How did the name of cockroach. Yeah. So, you know, we named the Gimp. Yes. That was, that was mine.
Starting point is 00:48:12 Pulp fiction had come out in college and like, oh, which we named this thing. Oh, the new image manipulation program. I think we're thinking image manipulation program initially and I'm like, oh, Gimp, it's obvious. It just stuck. And at some point, you know, we're newly on this new database, and you kind of want to give things a name.
Starting point is 00:48:26 You can't just say, oh, we're working on this distributed database. You kind of need something. And the sponsor was like, oh, Cockroach DB. Like, cockroaches are unkillable. I want these things, this database to be unkillable. You know, Cockroach is going to survive the nuclear apocalypse. So that was where the Genesis was and just stuck. Yeah, we're, one of our bunch of our nodes goes down.
Starting point is 00:48:44 This thing will still be up. Yeah, yeah. And, you know, this is where we're at today with Cockroach TV. It's like one of the things that I point out on like, holy crap, this is awesome. You kill a node. We did this whole campaign last year, which was really just to prove out something that already been present. The campaign was performance under adversity. But just like you can run a workload against it, you can kill a node, you can sometimes kill a whole region, and the system keeps on going.
Starting point is 00:49:04 And it's like stories like that, you know, like what we did on the marketing side there, but also we hear this from our customers too. They've had fires and data centers and all the other data systems go down and Cockroach TV keeps on going. I'm like, that's awesome. When you started out, who were companies, startups that who wanted to use Cockroachers DB? be and how has it changed since? Because, you know, like just making the case like, okay, I'm starting a startup. It's a small startup. Like, I will need a database and I'll, I don't know, I'll typically choose a Postgres, right? It's free. Everyone's using it. I'm running it on Node. At what point did you see that typically tech companies are like, okay, like, this is not enough for me that it's
Starting point is 00:49:44 running on a node either because it can go down or because I'm outgrowing. What was it the outgrowing? I'm trying to get a sense of, like, at what point that companies say, like, tell themselves, like, we need something distributed in a database. Yeah. I mean, oftentimes we have companies calling us up after they've had disaster. So, like, a node went up or a hard drive fail, that kind of stuff. You know, it's not quite like we're ambulance chasers, but if you see an outage, like a big outage from some, you know, company, you know, like we'll sometimes be trying to knock them up. But also, they will call us, you know. I mean, like, there's a very big bank who's now a customer.
Starting point is 00:50:22 Don't think I can name them, but you can go read. They had a very serious outage due to a weather event. And after that, there- Which probably knocked down, I'm assuming a region or database or a networking cable or tree fell on something. I think it knocked down a whole region. You know, it was a region-wide power out. You knocked down the region.
Starting point is 00:50:38 And there's a mandate from the CEO. It's like, no, we just have to, you know, be able to survive these things. And that gets pushed down all the way. And you see this in other places where, you know, one of our early customers, they were running on AWS. and they just got to the maximum size you can run in an aurora instance on. And then what would typically happen at that point is then you have to charge your database. This is a very standard practice.
Starting point is 00:50:58 You take your single note database, you create 10 or 20 or 100 charts. And this is what Google has done for some period of time. And that's a heavy burden on the application developer. And the way we always phrase this is like, I mean, the application developer is becoming a database developer at that point, and they're doing it poorly. You know, they're trying to implement distributed transactions or indexes and whatnot. And we felt the burden for that belongs on the database developer. Can we talk about automatic sharding?
Starting point is 00:51:22 I think it's safe to some, most of us will know what sharding is when you're, but actually, let's start from like manual sharding and then how you can implement automatic sharding. And if you can tell us like, you know, tactics that a database like HockwoodDB can do to actually just take that load off of you. Yeah, yeah. So I think the very basic form of sharding is a little bit like the hash table. Let's say, you know, you have a fixed number of shards. Like, let's say it's just a hundred shards. your data model is a user with a lot of data that's said with the user.
Starting point is 00:51:51 You just take the user and you say like, oh, they map them to one of the shards. And, you know, you kind of just rely on the hash function to get like fairly even distribution. The problem with this is at some point, you know, one of your shards will get full and you have to kind of reshard. And that's a very, very onerous process. And reshart, I guess, simple way to do is like if it's just a hard drive, I don't know, per node, where you write the user data, it gets full and you're like, okay, well, I now need to split it somehow. I need to move it. I need to remap it.
Starting point is 00:52:18 I need to reject my metadata, which knows where the data lives, that kind of stuff. Yeah, yeah. And depends on exactly how you're doing that mapping from like, you know, the user ID or whatever your shard key is to the shard.
Starting point is 00:52:27 You might have to remap them all, right? This is very typical. Exactly. So, I mean, this happens in hash tables where, you know, oftentimes in order to grow the hash table, you just have to essentially create a new hash table, double the size and copy all the data over.
Starting point is 00:52:39 Now, that's kind of like the very basic, straightforward way. And there's various levels of complexity you out on it. One of them is called consistent, And there's various techniques to do this. It's kind of fascinating, like how they all work. But in consistent hashing, you can add an additional node, and then it only moves a fraction of the data from each chart over there.
Starting point is 00:52:58 There's various systems to do that. And I believe this is like when it lies Cassandra. The way Cockeridge DB does it is more into big table, more into span, or more into H-base, where instead of actually hashing, we actually take, you know, you can imagine all your keys in a system. And this is always true in a system. You can imagine just in one big, contiguous key space, And then you kind of partition contiguous spans of that.
Starting point is 00:53:19 And then you have to build up an index on top of those contiguous spans. And what I just described there actually sounds a lot like a B tree. So there's this index on top that is like that maps you from, you know, like, I need to have this range, which node is it on? And this is a little bit like a B tree. You know, he kind of squint, you know, it's like, I think you squint and everything's either B tree or it's a hash table. And, you know, that index structure. But it's now I'm starting to make sense because when I remember when I read, it might have been the Wikipedia article on B, trees, it's said, this is a data structure that is frequent to use in databases.
Starting point is 00:53:51 It's now coming back to me because I didn't think too much of it. I'm not a date. I'm not someone who builds databases. But now that we're talking, we just like organically keep touching on trees again and again. And the other place that it comes up in databases. So this is where kind of comes up in distributed databases. And no one ever really calls it a bee tree. I just kind of squint sometimes.
Starting point is 00:54:08 And I see like actually kind of a bee tree. But the other place that comes up in databases is for your indexes. So if you have a table and you have like an index, on your email address. And you want to be able to scan over those email addresses in order. That's a B tree under the hood in a database. And, you know, like any kind of index you have, it usually provides sorted order.
Starting point is 00:54:26 There are hash indexes, but oftentimes it's the B tree index. And they're ubiquitous in databases. There's actually a paper called the ubiquitous B tree. And, you know, basically, like just identified that. I think that paper was written back in the 80s. And they're still ubiquitous today. They are the foundations of single node databases. And pretty much every data system I worked on has had B trees at some point,
Starting point is 00:54:46 it placing them. I want to ask about strong consistency. So CockroarsDB offers strong consistency. Now for people who are a bit more newcomers to distribute a systems, can we talk about the consistency models and then why strong consistency is important and why it's hard to implement it? So, I mean, there's multiple ways to kind of approach this, but I mean, if you've used a database, you've probably heard of transactions. And transactions are a way to perform a whole bunch of mutations atomically. So databases talk about atomicity, consistency, isolation, durability. The durability is really easy. It's like when I write you the database, it has to be durability written. So if anything crashes, it comes back. The animicity is just referring
Starting point is 00:55:26 to the fact I want to do a whole bunch of changes. I want them all committed or all aborted at the same time. I don't want to have like some kind of partial operation. And why is this important? Why is the adamicity important? Well, the adamicity is what gets you to the point where it's like, as an application, I can do a bunch of operations. And if there's an error in, an error occurs, it all kind of gets rolled back, and it's a much simpler development model to work within. And then there's the consistency in isolation, which kind of get, you know, kind of muddled. The isolation is referring to isolation between transactions. I don't just want to run one transaction at a time.
Starting point is 00:55:57 That's easy to do, right? I want to run a lot in parallel. Oh, yeah. And make it so that when they're running in parallel, they are running as concurrently as possible. But you want to have the appearance when they're running as concurrently as possible, that there is kind of a serial order to them. So this is like kind of the whole track, the kind of the gold standard for isolation is called linearizeability. Don't worry about that.
Starting point is 00:56:18 The step down from that is called serializability. And that literally refers to having a serial order of your transactions. But you're having to construct that in a way that you're doing everything as concurrently as possible. And the benefit of this, the serializability, is again, it's a very simple model for the application program. They don't have to worry about weird kind of defects occurring in their program. And some of the ones that can occur, it's like the classic description. is of a bank, right, where I might want to read, you know, like have an operation that reads and says, like, do I have $100 in my bank account to transfer somewhere else? And you could like arrange
Starting point is 00:56:52 for lesser isolation levels that you might be able to subtract that $100 twice. And that's bad, right? You know, we want to keep accurate, you know, track of your bank account or, you know, what's in your shopping cart or, you know, it's kind of like the use cases are endless there. And you do this wrong and you have very egregious bugs. But now going, going, back to weak consistency and strong consistency. Anything less than linearizability or serializability might be considered kind of weak consistency, but there's also like, you know, kind of strong consistency and eventual consistency.
Starting point is 00:57:22 So the eventual consistency is like sometimes I can do an operation and it might not be immediately, I might not be able to see all the updates. The read result will not necessarily give me the current update. But they'll come back, you know, at some point, you know, I've written part of it and the rest of it will show up at some point. And oftentimes when you're doing- It's easiest thing is your credit card balance. right?
Starting point is 00:57:41 Yeah, yeah. Yeah, and it's faster to do it that way, faster in terms of just what the performance you can get out of the system. But again, it's a little bit harder for the application to deal with. One of the places this often comes up in distributed databases or databases
Starting point is 00:57:53 that have any sort of replication is that I could write to the primary replica and then I read from the secondary and it's not there yet. That's eventual consistency. Yeah, that's eventual consistency. And you can often work around this, but you just like puts bigger burden
Starting point is 00:58:06 on the application of Elk because you have to pay attention to that. And then with Carcores DB, you have strong consistency, meaning as soon as you're writing it when you're reading from the database, you already get the written value back. You get the written value back. You know, it's like you read whatever you wrote. It doesn't matter if you're reading from the same, you know, node you wrote it to. If you read it from another node, you still actually get the data you just wrote. Is the trade-off logically not that you would have higher latency? Because clearly,
Starting point is 00:58:33 to implement strong consistency, you would somehow need to, in the naive approach, you would need to write all replicas, right? Yeah. What are you doing? inside Cockworth TV. Well, we are writing to all the replicas. Well, yeah. Yeah, yeah. You're just doing it fast. You're just doing it fast.
Starting point is 00:58:47 You're making that efficient. I think this is one of the things that kind of also fascinates me about the software industry is we keep on finding ways to be more and more sophisticated in order to provide, you know, like do things that make it easier to write the applications, but do it at a very high performance way. And we've gotten, you know, very, very good at this over the years. And this is the area I know about the databases. This is happening everywhere. Like, I'm just fascinated by how fast graphics have gotten where when I should enter the industry,
Starting point is 00:59:11 you were literally writing out each individual pixel, and now you have these GPUs that are doing like billions of triangles per second, whatever the current numbers are, and just like kind of astounded, like, how much sophistication has gotten into every area of computer science, wherever you look at it. One more thing on Cockroach DB, I want to ask about this is RAF consensus. What is the RAF consensus?
Starting point is 00:59:29 Yeah, I mean, consensus protocols, the original consensus protocol is called Paxos, and it was famously hard to implement. Raft was, you might think of it as a variant of Paxos. It was kind of like an alternative to Paxos, But in my mind today. And then because this is algorithm being that you have like a number of noes, like three to a lot more.
Starting point is 00:59:48 And then how do you get them to agree on? What do you typically agree on? Yeah, yeah. So you agree that the right occurred. So like the way to think about consensus. So you might think I want to replicate data and I write it to a primary and write it to a secondary. And you can't actually have consensus when you only have two replicas.
Starting point is 01:00:05 And the reason you can't have consensus is if there's a crash and I come up. Like if I'm on the secondary, how do I know? if something was written to the primary. If I'm on the primary, how do I know it was written on the secondary? Right? And you're always going to be in this kind of confusing place where it's like you either have to roll back a little bit
Starting point is 01:00:19 or like, you know, you lose some data. And consensus requires at least three, but you can have consensus across more than three replicas. And the idea with consensus is I'm going to write to three places. And normally you don't actually, when you're doing a read, you don't read from multiple of them. But if there's a crash, I have to do recovery,
Starting point is 01:00:37 then I'm reading, oh, I can read from any two of the three and I know I can kind of determine what had happened previously. And it's usually just on the recovery time that you're actually doing that consensus read. So reads are typically just happening from one replica. You have to write to all three. And on a crash during that kind of failure is when the consensus read occurs.
Starting point is 01:00:56 And inside Cockroach, DB, how many replicas do you choose for either consensus? It's typically three. For some system tables, it can be five, and then customers also have control of this. At the database level, you can write to five, seven. Five is like, you know, if you're really concerned,
Starting point is 01:01:11 about the data durability, you might use five. But there's a slowdown. You know, the more you're writing to you, it's like the more storage space that. And obviously, like, we're talking like a slowdown with nodes. But of course, if they're like between regions, there's now you have a lot more resilience for, let's say, an earthquake or power hours or whatever. But now you will have additional latency. It's just speed of light, basics, right?
Starting point is 01:01:31 Speed of light latency, right? And it's, you know, tens, you know, or up to hundreds of milliseconds or even higher if you're going across the globe. So, you know, you have to be very careful with that. in terms of how you architect your queries. And one of the things that it's just a general truth of some of distributed databases is you don't want to have a lot of back and forth. You want to kind of do all your reads in one kind of parallel read set, get them back,
Starting point is 01:01:53 then do your rights, right? But if you're doing like kind of serial operations where I read a row, I write a row, I read a row, I read a row, I read a row, I read a row, I mean, the latency's just add up. We talked about founding CockroachDB, but how is the company grown and where are you today? Yeah, yeah. I mean, we're being used in like we're powering mission critical applications. That's our bread and butter. But by the way, can you elaborate a mission critical?
Starting point is 01:02:12 Because it's like, if you're not, you are in the industry where you know what this means, but from the outside it can feel hard to put a thumb on. What is mission critical? Is my SaaS that is like showing as a mission critical? Probably not. Yeah. So mission critical in my mind is like these kind of the other term of art is tier zero applications. The ones that are like just the core crown jewels of what's running a company, you know,
Starting point is 01:02:35 like a trading system. You know, your trading system can't go down. If the trading system goes down, this is. kind of a critical problem or the firm that is running the training system, you know, banking systems. But also, you know, like we work with DoorDash, you know, like some people might think that delivering your burrito is a kind of a mission critical system. It certainly is for DoorDash, right? You know, if that goes down, it's problematic. We power shopping carts, you know, other stuff like this where it's like, well, if the shopping cart goes down, you know, that company is losing, you know,
Starting point is 01:03:04 hundreds of thousands, millions of dollars per hour. So that's kind of the criticality you think about. Yeah, I guess of course that they're losing, but this is like when, yeah, their customers are also like they're used to this just working like running water. And then when it's not the same thing as when your utility breaks, right? Your water or electricity is out, you'll survive, but it's not what you expected. Yeah, yeah, yeah. And everybody's like, what age are we living in? The electricity goes out. We were kind of had this ingrained into our heads at Google.
Starting point is 01:03:31 It's like Gmail cannot get down. People are, you know, relying on it. Search cannot go down, you know. If it goes down too long, people are going to move to other systems. And it's not like, you know, like in some way search isn't his mission critical, except, oh, wait, every single search is ad dollars behind it. And you can actually notice the blip in the revenue. And it's not just the blip in the revenue. It's the blip in reputation as well.
Starting point is 01:03:51 I mean, I think that's the one that really poisons companies is like, if your bank is down for a serious amount of time, the reputational damage there will be horrifically. And, you know, like we often talk about like, oh, and then grandma won't be able to pay her rent and she'll get evicted. You have to take this like responsibility really, really seriously. No, but also like just Gmail being mission critical. just on the way here, we only exchanged numbers later, but we were communicating over email, like, oh, I was telling you that, that you were telling me that you're here, I emailed, and I never for a second thought that it could go down. And I think we were like responding within 30 seconds, right? And it's just, I just know it's there. Like I didn't like bother
Starting point is 01:04:26 setting up a secondary communication line. So. Yeah, yeah. It's like when people just like, when you have that kind of level of trust with your users, you got to maintain it and invest in it. But then it leads to this kind of freedom for the user as well, where you just don't think about it, I don't have to worry about it. It's just going to work. And then in terms of the company, like, how many engineers do you have roughly? We have some hundreds. I don't actually know the precise engineering number 110,
Starting point is 01:04:50 but there might be 150 in R&D overall. There's other folks besides just engineers. I'm clear as engineering managers. Those are engineers as well. And then, you know, like we've been grown steadily. It takes quite a while to build a distributed database. Not for the faint of heart. Yeah.
Starting point is 01:05:05 So it took a couple of years. You've done it a couple times. Yeah. Yeah. Well, I did distribute a storage system. I did a distributed database. It's not for the faint of heart, right? So there's a lot of work getting it to a level of stability, then a level of kind of quality
Starting point is 01:05:17 beyond that level of stability, getting all the bugs out. And then continue to innovate and put more performance into the system, adding functionality to integrate better within enterprises. A revenue has been, you know, kind of steadily growing over the years. And, you know, it's at this place now where we see a path to future successes. as well. And I want to ask about your coding habits. So when you co-founded the company, how much code did you write for the first few years? I wrote a lot. So I've always been a very prolific coder. There were a lot of code early days. And early days, I mean, I was a kind of,
Starting point is 01:05:53 we were all technical co founders, Ben Spencer and I. And we were already a lot of code. And I was no exception. But, you know, I look back at my GitHub output. And, you know, it's like kind of peak years, maybe 100,000. lines of code in a year, which is a lot. Yeah. Yeah. No. So, I mean, rule.
Starting point is 01:06:10 We're talking pre-AI. Pre-AI, right? This is when you're back doing this manually, right? You know, at some point, you know, we started out using a system called RocksDB, which is an L-SM. At some point, I think it's back in 2019, you ran to limitations with it. I decided we wanted to rewrite it and did a big push to, you know, right, that might have been like 40, 50,000 lines of code. And then a bunch of other people have come up and helped and, like, you kind of look at that output, I'm just like, oh my goodness, that was a lot to keep in your head. It's a lot just to type,
Starting point is 01:06:40 you know, 100,000 lines of code. The average kind of like that the industry talks about is 3,000 lines of code from an engineer in a month. And so if you multiply that out, maybe 36,000 in a year, that's good, right? So I was doing a lot. I kind of look at that. It's like, there's kind of a max that you can hold in your head at a time. The tools have gotten a lot better since I first entered the industry. We've gotten better debugging techniques, better testing techniques, but still quite significant. And you were CTO from the beginning, co-founder of CTO, but there was a time,
Starting point is 01:07:11 sometime around like 2002, when you decided to kind of be a bit more hands off, right? Yeah, yeah. Can you tell me about that? I mean, the general rule of thumb for engineering leaders is, well, you got to have your team, you know, to manage your team, and we have a VP of engineering,
Starting point is 01:07:26 but I was kind of getting to the point of like, okay, is my, or my coding days done, you know, like kind of just direct from a higher level. And, you know, I got this advice for a long period of time and I pushed back on it, but, you know, I kind of acquiesced at some point. And I think it was the right advice. I'm not saying that the advice was wrong at the time. But there was time period from about 2022 to 2024, whereas, like, my output declined. I think I actually did the Swiss tables thing in that time period, but I wasn't doing much on the core.
Starting point is 01:07:53 The business. Yeah, the core business. You know, I would get in there and do some work, but like, it's really hard that if you're in meetings all day, to also do coding. I think this is the fundamental. attention. If like, so you kind of took on the kind of the meeting burden, the coordination burn and the stuff that was, if I'm reading correctly, before you spent a lot of your head in
Starting point is 01:08:13 the code and now you're spending a lot of your head like above the code, the business, the engineering, the org, the whatever, customers, that kind of stuff. The customers and just being an executive as well. Yeah. So very hard to wear all those hats simultaneously. And then I got back into it because AI started to emerge. So how, when did you get, when did you start using AI? when you start to find it useful in terms of coding.
Starting point is 01:08:34 Yeah. Well, it's interesting because, you know, those initial versions of like kind of glorified auto complete came out. And we're talking about the GitHub copilot, the cursor, the early version. That was the one I had the first exposure to you. We tabled with cursor at the time. But they're all like kind of glorified auto complete. And it was kind of crazy that you could just start typing something.
Starting point is 01:08:52 You know, like, don't the rest of the function. You look at me. I'm like, wait, you kind of got that right. This is crazy. Right. And, you know, we're trying to encourage our engineers to use this. And at some point, you know, I can't remember if this is my idea or my co-founders or someone basically like, you know, like in order to like really guide people about how to use it,
Starting point is 01:09:11 you have to be a user yourself. You know, I think this is true of like engineering management in general. Like you want to like manage engineers. You have to know how to be an engineer. Like if you if you don't know how to be a good engineer, it's like really hard to manage other engineers. I feel like you'll have a hard time. I'm like just relating to them at the very least.
Starting point is 01:09:27 Exactly. Exactly. So kind of took it on me like no. I mean like, I mean, this is like, it was. clear very early on, like this is probably going to go somewhere, but it wasn't quite clear how far, how fast it would go. And you start dabbling this and I was like, oh, okay, well, it's not quite good enough. It's not quite good enough. But, you know, let me start getting back into the coding very rapidly, you know, you started seeing the signs of life, like, you know, kind of the opus models coming out.
Starting point is 01:09:51 It was sonnet first, then the opus. And you're like, you're looking at these and I'm like, oh, wow, okay. Well, they seem to be able to do quite a lot, but, you know, the code is still not great. But then it was just like just this cadence of continuing improvements. And, you know, I was starting to do a lot with Sonnet and then starting to use Opus. And, you know, I had that same moment everybody else did. And this was last year, last November. November, December, winter break, right?
Starting point is 01:10:19 Yeah, it was Thanksgiving. I distinctly remember it because, you know. You were not chilling. You were coding, weren't you? You were agenting. I was agenting, just like everyone else. I had this thing I'd wanted to do for a long. period of time on CockroachDB, which is like, I wanted to like test like all these
Starting point is 01:10:36 configurations of Cockeridge DB across like, you know, different vertical scaling, like how many CPUs you have on a node, how many stores you, discs you have on a node, how many nodes you have in the system and tested across all this huge matrix of it. This is one of these things they couldn't ever quite get prioritized appropriately because it never seemed like kind of critical, but I've like, I always had this intuition there was something there. And then over this like four day span, the code just like materialized as I was using, I that was Opus 4-7, or is it 4-5, whatever the number was.
Starting point is 01:11:03 Like, the recollection I had is, you know, like, a pretty fast typeer. And then I just remember having this feeling of, like, well, the code is just, like, materializing before my eyes. You know, you would ask for these things. I got out of the habit of actually typing, and you just kind of, like, ask for something materializes. If you've ever seen, like, some of those, you know, a movie where they had this kind of archetype of a software engineer gets in front of the keyboard,
Starting point is 01:11:28 and they start typing, show the screen, it's just like, it's going way beyond. on human speed. It was gone that, right? Like, it was, I think a software engineer is like, we used to laugh at these. Like, I still remember Swartfish, the thing when they're visualizing, the things are going for, or the code appearing, you know, you're seeing that the person's typing and then it's a big line.
Starting point is 01:11:48 And a software engineer is like, we're laughing because it's not how it is. But it's crazy that, like that effect, right? Yeah, yeah. But now it's. It's even better than that effect, right? Because it actually works this time. And it's, that's slow in comparison. It materializes faster than that.
Starting point is 01:12:03 Like, you can literally go. Like, we've been talking about bee trees a whole bunch. I implemented another bee tree in the past month. Of course he did. Yeah. And it took about 30 minutes to implement probably 10,000 lines of highly optimized rest. I mean, it's just like it boggles the mind. I mean, we should look up later at 10,000 lines.
Starting point is 01:12:20 It's like a crazy amount. You physically cannot type that fast. You started getting back to coding. Like, are we talking about kind of vibe coding prototyping, or are we actually talking you started to contribute like proper production ready code that is up the level of what you're doing a whole to be. Well, it started with this tool that was like this kind of benchmarking tool that tested this
Starting point is 01:12:41 matrix. I have the CTO. You have an office of the CTO. The office of the CTO's mandate is to innovate. And I was looking for places like we can have innovation. And one of the things that we want to innovate in was, you know, better auto scaling of a Cockroarstabee cluster. and kind of in the January time frame,
Starting point is 01:13:00 I came out with, like, what I would think is, like, kind of a research breakthrough, you might say. And it came about, because I was dabbling this area and just working with the models trying to understand it. And they're not just good at coding. They're also good at, like, helping you explore design ideas. And I think this is kind of the fascinating thing where it's, you have to get out of the mindset of, like,
Starting point is 01:13:17 I know exactly what I'm going to build, but more like, hey, we have this problem, talk through it, be a partner with me. And it would be a sparring partner. You know, like, there's this advice you might have heard that like, you know, if you're stuck on a problem, you should say, go yellow duck it, go rubber duck it. I think I've heard it as yellow duck, you know, just go talk to something.
Starting point is 01:13:36 You don't even need it to respond. Now you can talk to this system, this intelligence, and it will give you back stuff. And it's not always right. I mean, this is the thing. Even today, these models are fantastic, fables fantastic, Astros is fantastic. And they're not always right, but they kind of like, they know so much. It's just encyclopedia. And you can explore ideas, super, super, super.
Starting point is 01:13:58 fast and then be pointing out, well, that doesn't sound right to me and it'll be like, oh, yeah, you're absolutely right. You know, like, I hate that sick fancy. It's like it kills me. Or you're right to push on back on. You were right to push the back on that. That's a new absolutely right. Yeah. And yet, like, just able to make such fast progress. And I, I find it absolutely incredible. So it quickly moved from just doing kind of side things, building up some tools and whatnot to building, you know, essentially over the last eight months since January. We've been building towards a, you know, a new launch. And, you know, like a lot of that has been produced a gently powered engineers. And like the entire company has gone on board with this now,
Starting point is 01:14:38 where I think it's like probably 100% of engineers are using it. To a greater or lesser degree, and a lot of code is being produced. I think it's very high quality code. These models not test things adequately. They don't look quite intently enough about performance. But if you're like, you can guide them in the right way. And I think there's actually a huge advantage. to anybody who's done management before. It's a little bit like being a manager of people where you're a manager of a large group. You're not looking to every line of code,
Starting point is 01:15:05 but you're definitely kind of helping architect the system. I think there's a very strong analogy there. I sometimes push back on this analogy. The reason being I was an engineering manager. My take is that working with these agents, it's not like management because management has so much of the human stuff. Like as a manager, when I think of all this, a bunch of stuff I dealt with,
Starting point is 01:15:22 which was the people side of things, the conflicts between people, the performance reviews, the meeting, et cetera, and you have none of that. You do have the orchestration. Like, again, this is like almost like, I guess a really like naive way of management where like, you know, they don't push back.
Starting point is 01:15:37 They start to do it. Sometimes they're like unreliable, but you know, I almost like to use like orchestration a bit more because I feel management is so much more involved. Like I think this is saying where like a tech lead who has no management's responsibilities, but they have a group of interns, but they don't need to do with their performance, with their anything.
Starting point is 01:15:52 Like it's a lot closer to that if you know what I mean. Yeah, I do agree. I use the engineering manager shorthand, but it's really about being a tech lead for like a 30 or 40 person organization, you know, or being an architect for one. I think the term architect gives me a little bit of a distaste, but if being an architect, he's also on the ground. A hands-on architect. And you don't have any of the management stuff, which is a blessing and curse. But it's kind of remarkable that you can spin these things up. If they make a mistake, you can keep on correcting them until they get it right. We've always been able to do that on the human side. I can do it faster. I think the ultimate result for me is you just have to be more ambitious about everything you do. You can produce more, higher performance, higher quality, more secure. So our ambitions have to raise up. You also mention that with AI, like you're ever using and you're building stuff,
Starting point is 01:16:42 but also with CockwoodsDB, you're now building something that is also related to AI. Can you talk about that? Yeah, yeah. We're building multiple things. I mean, AI is the future. It's here to say, right? Yeah, it's here to say. Yeah, I mean, like, you know, one of our thesis, which is not crazy, every application in the future is going to be written by AI.
Starting point is 01:17:02 You know, I think there will be some kind of bespoke software, you know, handcrafted software. I think we'll see that continue to exist, just as people write assembly still, you know, but it's going to be diminishing in size. So, you know, asymptotically approaching 100% of software will be written by AI. Generated, right? Yeah. It's hard to say if, like, when it's going to end where the humans are the agent, you know, kind of like supplying the agency and the vision. behind it. I think that might exist for many, many years, but I think the code will fundamentally be written by the AI. And I think we're going to just see this explosion of applications. And we're
Starting point is 01:17:36 seeing that inside Cockrookrookers lives. We've seen it elsewhere. Earlier this year, we kind of rolled out this internal platform where non-engineers could write kind of many applications. I mean, this is like, you're hearing this at other companies. We did the same thing. And over the course of just a couple months, you know, 500 applications, a thousand applications. By non-engineers. By non-engineers, primarily by non-engineers. And I thought it was awesome. Like our HR team is building like these little applications.
Starting point is 01:18:02 Like, this is the thing I've always dreamed about doing. And like, they were never serviced. Like, like your CFOs all over the place. And ours is no exception. Producing dashboards that they can never create it before. And I think it's very empowering. I'm married. I have a wife.
Starting point is 01:18:14 She needs software. She cannot produce that software in her own. And I have never actually helped her produce the software, which is my own failing. But I think there's a world in the future where she gets, like, everyone's game custom software. software built for them. And you're just going to also see greater and greater applications and higher quality systems being produced by every company as well. Today, what is your stack? What do you work with in terms of harness model, how you run
Starting point is 01:18:39 agents? Is it one agent? Is it multiple? What can terminal do you use? Yeah, yeah. It's evolved over time. So, you know, like when I first started dabbling the AI stuff again, it was, you know, GitHub co-pilot. It was an Emacs user for like 20-something years. I got convinced to move to VS code. but that's all gone now. At some point, I moved to using cloud code. That was the thing that, whatever reason I just got to start using it, it was cloud code in the terminal. Of late, I use a mixture.
Starting point is 01:19:04 So I'll just grab my current setup. I use the cloud desktop app, cloud code via the cloud desktop app. It's fantastic. Good job, Anthropic. I also sometimes use the Kodak's desktop app, just so I have an alternative model to turn to. For some very critical things we're working on, I will get, you know, one of those models,
Starting point is 01:19:22 sometimes like the best model, Oftentimes I'm using the best model, like Fable. You know, I've recently started using that. Sometimes it's Astra. Sometimes it's the other ones. But you have one producer design. You have the other one being like, hey, my colleague produces. Can you kind of tear it up, you know?
Starting point is 01:19:36 Aversarial review it. And it's not always perfect, right? You know, but I think there is utility, especially for something that's super, super critical. We're pushing towards the launch really soon and kind of rushing towards the finish line. I'm telling folks on the team, it's like, you know, especially our very senior folks, like, you should use the best model right now. Yeah. It's worthwhile to do that.
Starting point is 01:19:52 I'm generally just using the best model. And part of the reason is I don't actually know that I'm getting a lot more intelligence from it, but I don't want to have the cognitive overhead of deciding on a case-by-case basis. Yeah. Should I use sonnet? Should I use opus?
Starting point is 01:20:06 Should I use fable? Should I use soul or astra? And I think, you know, maybe if you, like, I really need the speed, I would make that decision. But oftentimes I'm like, I'm doing things in parallel. So you're asking how many agents I'm spinning up? Well, it depends. Like, oftentimes there's like a certain number of sessions.
Starting point is 01:20:22 you might be using. I find my kind of cognitive overhead is about five to ten sessions concurrently, but sometimes those sessions will have many sub-agents doing things. Yeah, they'll be long running. Yeah, it'll depend on what I'm doing. So as I was coming here, I was on the train, I spun up something to do a little bit of research, and that was just probably using three sub-agents. Other times, it might be 20 sub-agents, you know, with Anthropic has dynamic workflows. I can't remember what the Codex thing is called, but they have various ways of spinning up kind of graphs of agents and, you know, prosecuting them. I mean, I think it's sometimes I'm probably using 100 sub-agents. Others, it's only like five. It depends on what it's happening. And sometimes
Starting point is 01:21:01 I'm not doing any of that agentic stuff because I'm often thinking about something else. Since you were an Emacs user for 20 years, you know, using the terminal, right? Well, E-Max, but how come you move to graphical interface with agents? A lot of a lot of people still use terminal UIs? Yeah, yeah. Well, I, so I moved from, EMAX to VS code. I was just like a colleague was like, hey, these IDs are really good. You need to try them out. I was like, EMAX is an ID.
Starting point is 01:21:25 And then I kind of moved and like, you know, some of the stuff was just more seamless. And like Emax keeps on catching up. It always feels a little bit, you know, six months through a year behind. And like every time they have a new version, my setup would break. And I just got tired of that. So I moved to Vs code. But then I was like, I'd have many terminals open in Vs code. My setup would always be like terminals in VS code tabs.
Starting point is 01:21:46 The terminals and tabs, you know, have like eight of them open. and then they started taking over. It used to be like just Claude code in the terminals. Yeah. Yeah. And then, you know, just some point recently is like six weeks ago, colleague was like, well, if you tried out the cloud desktop app recently, it's really good. I was like, yeah, no, I'm very, very happy.
Starting point is 01:22:02 And I made the switch and it was a little bit awkward at first. And then suddenly I'm like, holy crap, this is awesome. You can like manage the agents a bit better easier sessions. You have the sessions there, the sessions are down the side, you know, and it's like they're kind of like tabs in some regard. And shit, but the whole integration that Anthropics have been doing in. Open the eye is doing the same thing now. By the way, like, I'm not a dissing cursor and factory and cognition.
Starting point is 01:22:23 They're all pushing the same thing. It's incredible how fast these systems are innovating right now and evolving. What do you think good software engineering looks today compared to like, you know, four years ago before we had a, has it changed, has it changed? Has it changed? Yeah. I think the ambition has to increase, you know. I think quality has to be higher. Security has to be higher.
Starting point is 01:22:46 I'm glad you mentioned quality. Can we talk about that? because I'm seeing across the industry, just quality declines, which you cannot fully put your finger on AI, but oftentimes it is people pushing out more and more and just not paying attention to just small regressions here and there. Again, like you're building a database. Have you noticed any or have you gotten feedback of any quality regressions or if not, like, how come? Because like when, like this is just basic, you know, like we talk about the law of business. This is an observed thing that when you start to have more output, you increase your deployment
Starting point is 01:23:24 frequency, you often, not always, but you often have like more regressions and more bugs. If you produce a certain number of lines of code, you're probably going to have a certain number of defects per line of code, and now you can produce more lines of code, so you probably would have more defects, right? But the thing that pushes against that is that you can be telling these agents, and you have to give them kind of firm hand at this. This is a thing that I hope the model providers are listening to. But you need to give them quite a furred hand.
Starting point is 01:23:51 They kind of get a little bit lazy on the testing side. And you have to make sure the tests are comprehensive, but also they're using all the testing techniques. And there's a lot of testing techniques out there. We have a lot of knowledge about how to do testing well. And you know what? The agents are lazy. Humans are a little bit lazy.
Starting point is 01:24:08 So getting the humans to actually be very disciplined about their testing is also challenging. And I think it's actually easier with agents. You know, you can kind of instill that. You can get them set up. You have to give them the guidance. Use property-based testing. Use metamorphic testing. Use, like, these advanced testing techniques, deterministic simulation testing.
Starting point is 01:24:25 You know, there's, like, technique after technique you can use. And humans have always, like, I'm lazy myself. Like, I'm pretty good about being disciplined about testing. At some point, you're just like, okay, that was enough. We got a ship. And, like, now you can get a little bit, like, you know, a little bit stricter and firmer. I think the same thing applies over on the security side, the performance side. I mean, we've always had at Cockroach and in the industry, we've known security coding practices.
Starting point is 01:24:49 Yep. You can literally have every single line of code on every commit reviewed by a security expert. Doing like an adversarial security review by an agent or multiple agents. Or multiple agents, right? And we're going through a tough time in the industry right now with like kind of hacks and security leaks and whatnot. There's only a limited number of bugs that can be in software. I think we'll get it, you know, out it on the security side and on the quality side. And then also just on the, when I say quality, it's not just bugs, but it's also like, you know, just a little thing.
Starting point is 01:25:18 It's like, oh, that UX element is wrong. And there's no excuse for that right now. Like, fixing it is so, so easy. And what we're also starting to see evolve is like, you know, designers using Figma. Well, I think that should be a thing of the past right now. Our designer just deals with HTML and CSS and JavaScript directly. And sometimes is even just producing pull requests, you know, PRs, which she loves. We all love.
Starting point is 01:25:41 Everybody's happy with this. there's none of any of this waterfall handoff. So she, she's producing poor requests for the production code base. Yeah. Yeah. And you know, this is on the UX side, not on the core database. Of course, but this is the area that, like, she owns, right? Yeah.
Starting point is 01:25:55 Yeah. It's just, I mean, she loves it. Everybody loves it. I mean, there's no downside. This is not new. Yeah. For sure. Yeah.
Starting point is 01:26:02 What about code review? What's your take on code review? I think it's pretty, it's starting to get a bit controversial, like, is it going to stay or not? Because there's been a practice that's been around, like, think about, I mean, you worked at Google. Google has been so big on code review. As I understand, there are two code reviews, right?
Starting point is 01:26:18 The domain expert reviews it, and then there's a language expert reviews it. So you went through that. How did you think of it and how are you thinking of it now? I used to be part of the, I was one of the early members of the C++ coding. I'm asking the right. Okay, you will have strong opinions. Tell me. Yeah, yeah.
Starting point is 01:26:35 I mean, I see that code review, I think, was fairly essential before the age of AI. but now we're getting to the point where I'm reviewing and looking at code. I look at most of the code that the agent's producing still. And yet it's like you're not giving it the same level of scrutiny, right? And this has always been the case. If you get a pull request from a junior engineer, you have to give it more scrutiny than if you get it from your most senior engineer. And you ask any, you know, tech lead, any engineer manager, any software agent,
Starting point is 01:27:02 they'll see, yeah, yeah. So the junior engineer or someone new to the code base, you know, you just have to get more scrutiny. And what I'm finding is the agents are getting better, you're having to give less and less scrutiny. And you do so have to give scrutiny to some things. It's like I said, it's like the testing will be incomplete. And maybe they didn't follow the security stuff.
Starting point is 01:27:18 Maybe, you know, like there was a performance regression. And I think it's like, you know, assuring there's kind of constraints on the system so it's hard for them to do the wrong thing. I suspect that, you know, I don't know if it's going to be this year or next, that we're materially going to stop looking at the code in the same way that we don't look at assembly anymore.
Starting point is 01:27:36 Yeah. Like you will... We trust the compiler to produce pretty good assemblies. We trust the compiler to produce good assembly. Unless you're one of those people where it's really important, you're maybe a game developer or someone and you look at it, but there's fewer and fewer of those folks. Yeah.
Starting point is 01:27:51 And even then, I think what will be migrating to is why does the human have to look at the assembly? The AI should look at the assembly. Like I'm doing something right now where I want to have a zero overhead abstraction, you know, for doing something that is done at test time and what it's in production is compiled out. Babel set up for me a little system where he actually looks at the decompiled code to verify, like, that there's only a few extra instructions put in place. I never would have put that in place, but now it has this guardrail, that every time it changes this, it can verify that there's no regression.
Starting point is 01:28:23 And I think you're going to see more and more of this where you kind of put guardrails in place. I think the model kind of likes it because it can work within that guardrail. Now, for what, like 20 plus years, you were writing so much code. Like I'm sure you were in the zone. Do you remember being in the zone and just turned? And now you wrote a lot of like production ready code. Now that you're coding with AI, like, do you get into the zone? Yeah, absolutely.
Starting point is 01:28:48 And how is that zone? Is it the same? Is it different? It definitely feels a little different. It's maybe a little bit less intense, but you're managing more things cognitively. Like, can you describe me like what is it right now when you're in the zone? Yeah. Well, I'm thinking of ideas that normally would have taken me a week to experiment with.
Starting point is 01:29:08 and I think multiple of these experiments, and then I fired them off all simultaneously, and then I'm kind of like reviewing, like, what else should I do while that's, what's being complete? Sometimes I'm kind of reviewing, like, they'll be giving me progress updates of like, oh, hey, this is coming in,
Starting point is 01:29:22 we're seeing this stuff, and I'm being like, well, that doesn't sound right. Hey, what about this? Or do we do this, you know, correctly? You know, maybe I had design, and it's not implementing the design quite perfectly. But I have this little feeling that, you know, I haven't been a college professor,
Starting point is 01:29:34 but maybe I was a, you know, if I was a college professor, I had a whole swarm of research, research assistants and they're all off doing things. And it's coming back, but it's coming back just really rapidly. Really rapidly. You're not waiting months or weeks. I'm not waiting months or weeks.
Starting point is 01:29:46 And then I'm iterating. I'm like, oh, that one failed. That's fine. You know, you just got to let go. And this is the nature of the software I build. I think there's other pieces where it's like you just whip out a website. You can whip out something that doesn't have this level of kind of scrutiny, you know, very quickly.
Starting point is 01:30:00 But I talk about this in a, I have a whole bunch of analogies about like the feeling to be. Let's talk about analogies. What analogies do you have about using AI or? I mean, the one I was, you know, advocating for, I had kind of two that I was advocating for, like, late last year and this year, which is, you know, AI is coming for us. It's here, right?
Starting point is 01:30:19 And it's like, you know, you're producing software. You're like walking down the road. And sometimes, you know, someone will pass you. They're running. They got an efficient gate and whatnot. But everyone's under their own locomotion. And these agent-e coding agents came. And it was like a car pulled up next to you.
Starting point is 01:30:35 You get into the car. You don't know how to drive. You don't understand the controls. But you got to get in. you start, you might figure out the gas pedal, and it takes off and crashes into a tree. But you have to learn how to drive. You know, I think using all these coding tools, it just isn't, doesn't just happen naturally. It's learning how to drive.
Starting point is 01:30:50 I think we might be in the era of the F1 driver right now, which is like the really good people can drive these systems a lot harder, a lot faster than the people who are just picking them up. If you've never used an genetic coding tool, there's a vast difference between someone who's like really expert in them, knows where they break, can pay attention. that versus someone who's just picking out for the first time. One analogy I've heard is we used to talk about the 10x engineer. You remember?
Starting point is 01:31:16 Yeah. This used to be a debate for a very long time. Is it or is it not? But now what I'm hearing is the 100x engineer. Yeah. And so you're saying that you are seeing some folks who maybe let's not use the 100x engineer, but like this like F1 driver who was just really good at it. Like how would you describe a person who we've seen this?
Starting point is 01:31:34 Is it just like rock salt engineering basics and they picked up. they lean into using these tools or what do they like? Yeah, yeah. I mean, there is quite a bit of a, it feels like, you know, it's directional that if they're good software engineering before. I mean, I sometimes think it's like, you know, everyone's in this kind of spectrum of capability of software engineering. And this is just like, you know, taking that line and spread it out.
Starting point is 01:31:56 And it's not quite true, you know, I think it's helped some people more than others, but, you know, it feels like it's just stretched it out. So, you know, your ability before is now amplified. I'm always interested to learn that, for example, Boris Churny, Tebow at Open AI, they both have been really, really good software engineers. Boris wrote one of the first TypeScript books, the first TypeScript book for O'Reilly. He built some massive system, same with Tebow, who built it. And now you're kind of seeing, oh, these people who are building all these tools and innovated.
Starting point is 01:32:25 And like, yeah, they don't really good before. Yeah, yeah, yeah. Well, I mean, this is a, I mean, I have two other analogies to give you about, like, what it feels like an AI. You've undoubtedly heard about the term paraprogramming. Yeah, pair programming. Yeah, parprogramming. You know, and the idea behind paraps programming is it's good to just like, you know, have one keyboard, one monitor and two engineers at it, one of the keyboard and the other one sitting beside them, kind of like looking over the shoulder and giving guidance.
Starting point is 01:32:48 And I think there's an aspect of that feeling where I actually got chat to do a little image of this where, you know, it's like the Android is at the computer typing and you're just there giving instructions. But that was maybe the way it felt like a year ago. I think it feels a little bit different now. The one I've just started recently saying is that I feel like the domain experts, the people who are, who. were really strong before are now massively amplified. Have you seen all this mathematical stuff coming out like the crazy proofs? Well, I don't understand it, but I don't understand them either. Yeah. Yeah. So I told you I wasn't very good at math. You know, it's not complete. I just was like, I stopped in the freshman year of college, right? Yeah. But, you know, I kind of like watch along
Starting point is 01:33:25 with these advancements in, you know, Terence Tao. He's like probably the most famous living mathematician, you know, super genius. He actually posted this session. There's this recent breakthrough or it's called the Jacobian conjecture. I don't even know what it meant. But he posted this chat GPT session where he's interacting. I think it was chat GPT. And you can see him interacting
Starting point is 01:33:44 with this intelligence. And it was crazy because he's talking to it as a peer colleague. It's responding. And it honestly looks like, I'd encourage everyone to go look this up. It looks like, you know, almost a foreign language.
Starting point is 01:33:56 It's like his domain expertise is getting amplified by the system he's interacting with. And you can see how he's like kind of learning and exploring ideas just really rapidly. Now, there's all this controversy about AI mathematics,
Starting point is 01:34:07 but I feel like the analogy that comes to mind to me, though, is the domain experts, they're a little bit like sorcerers in their particular domain. You know, you have the Earth sorcerers, the Earth wizards,
Starting point is 01:34:19 the water ones, whatnot. And if you know the magic incantations, the right words to say the right order, you actually get something kind of magical. And if you don't, you just get sparkable stuff that doesn't have anything there behind it.
Starting point is 01:34:30 You know, you would probably admit, you don't know much about distributed databases. If you're going to ask, Fable or Astra, build me a distributed database like Cocker, TB. You will get something out, but it'll kind of be ultimately hollow inside. If you're an expert in databases, you ask to build a distributed database, and you can point out all the various things you have to know about a distributed database,
Starting point is 01:34:47 here's what you have to worry about the storage layer, the networking layer, here's the various data structures, runtime inside. You can actually get something quite magical very, very rapidly out of it. So what would your advice before, advice before engineers with, like, mid-level to senior level who, you know, who have been figured out how to do coding to become strong engineers in this, I guess,
Starting point is 01:35:10 age of AI. Well, the first off is you have the most amazing tutor kind of readily at hand. And I mean, one of the things that I would always do, you know, throughout my career, and now I've kind of stopped doing it, but it's like the reason why it's going to become obvious, which is like, I would
Starting point is 01:35:26 always look at other people's code. So I was at, you know, Google early on. You probably heard of Jeff Jeff Dean was amazing coder. His colleague, Sanjay Gamowat, also just an incredible coder. And, you know, I would be looking at their poll requests. I'd be looking at their changes.
Starting point is 01:35:42 It wasn't called poll requests at Google. There's a different name for it. But I'd be looking at my... C.L. Right? Yeah. CL. This is P4. It's a different version control system. But I'd be looking at their changes. Be like, how'd they do what they did, right? I'd be looking at their code. Oh, my goodness. Sanjay's code is really always very elegant. You know, Jeff's his high performance. How's he doing that?
Starting point is 01:36:00 You know, like, how's he going about it? And it's almost just like you acquire via osmosis. But now what you can do is not just acquire via osmosis, but you can literally like, I mean, I would be encouraging if you're a junior engineer and you know there's a senior engineer nearby, like you could ask them how they're doing what they're doing, but you could just ask the AI to dissect what they've done and explain it to you and explain to you at various different levels. Like, how does this code work?
Starting point is 01:36:23 What is it doing? Give me a diagram. Explain it to me like I'm five. Explain it to me like I'm 10. Explain it to me in French, whatever like you want. Yeah. Like, I mean, fundamentally, to some degree, AI is a translation tool. They're translated from whatever, you know, kind of language or understanding it's in,
Starting point is 01:36:38 and keep on interrogating it until it gets it, you know, increases your understanding. You were saying how domain experts are very much amplified. I guess one strategy as a software engineer is like obviously become a great software engineer and use it as a tool. You can get a lot faster. You can get at distributed systems. Like, I'm not a distributed databases expert at all. But I use AI to explain a few.
Starting point is 01:37:00 things for me to understand upfront, which was very helpful. And it would have taken me a lot longer time beforehand. So I can use this. But I wonder if there's another part of like as a software engineer, you can use it to become more of a domain expert wherever you're working. If it's a payments company, I mean, use it to learn about payments as well. So you can help the business, you can help your team. And honestly, you'll just learn more, right? Yeah, yeah. I think, you know, I would encourage everyone. You have to have a little curiosity, right? Don't, don't be bound in by, you know, the area you're working on, explore outside of it, you know. I was working on Gmail, but I was fascinated about how, like, the main Google search engine
Starting point is 01:37:34 worked. I was fascinated by, like, how the internals of Big Table worked, even though I wasn't directly working on a Big Table, just explore in. Look at those things. And now it's so much easier because you have this super advanced patient intelligence there to explain to it. Like, why do you think it was done this way? And then, like, I mean, once you become an expert, you can be integrating.
Starting point is 01:37:53 I see it's done that way. Once we change this, would this be helpful? And, you know, that's where you, you know, that's where you? you go from just kind of learning to actually contributing back. I think everyone has to have the personal agency to do this. You know, if you're just sitting there waiting for someone to educate you on how to do this, it's going to be really hard right now because anyone who's coming in explaining how to use AI or explaining how to be a better software engineer, they're going to be out of date, right?
Starting point is 01:38:14 You just got to get in there, be using this tools all the time yourself. And using it like, use it to learn. I feel like I've learned more in the past probably even year than the previous five years combined. It's weird. Given your trajectory, given the environment you were working, right? I mean, like, everybody's been in the industry for a while.
Starting point is 01:38:34 Like, I'm definitely a better coder. I was a better coder 10 years ago than when I first got in the industry. It's like I could look back every decade and realize, like, I got a lot better. And I feel like I just got a lot better over this past year. And this was also one of the reasons I was really excited to talk to because when we started to just exchange your messages,
Starting point is 01:38:50 the first thing you wrote to me and I asked to like, hey, you know, how are things going? You said, like, you wouldn't believe. leave, but my coding output is insane and it's high quality and its database quality level. And those were the things I don't really usually see it. I usually see, okay, I'm not producing more code, but it's slop. But again, like, to me, this is a bit of an inspiration. Like, look, like, you can use these tools to just, like, amplify yourself as a software engineer.
Starting point is 01:39:13 Like, you are one example, right? Yeah, yeah. Yeah, yeah. No, I'm not the only one in Cockrookrish Labs. We have other people doing this as well. I find it very exciting. You know, it's a little bit exhausting right now, but it's very exciting. Like, I got into software engineering because I like building stuff.
Starting point is 01:39:28 I can build stuff faster. You know, the stuff you might have had to compromise in the past, and you can take away some of those compromises. I mean, you see this in the U.S. of software coming out. I think the U.S. is a lot higher. You see all the fancy web animations and whatnot. That's only just like the surface level. It just extends way, way below that.
Starting point is 01:39:45 Peter, this was awesome. Thanks for coming on a podcast. Yeah, this is wonderful. Thanks for having me. One reason I was excited to talk to Peter is because he's been a very high profile and productive engineer pre-AI, building some of the most resilient distributed systems in production. Cockroach DB is known for its resilience and how even if several nodes are destroyed, the database still operates with their data loss. Basically, it's as hard to get rid of as cockroaches are,
Starting point is 01:40:08 hence the name. One interesting part of our conversation was how Peter built more efficient data structures than the standard coding libraries had, thanks to him and colleagues paying attention to parts of the library that seemed slow. He did it for the C++-STL map and then go for the switch table implementation. I found, both stories a good reminder that you can improve the existing library or even the language, especially if you measure which parts feel slow. Another part of his conversation that I liked was how Peter came a bit of a full circle. He used to write 100,000 lines of code per year, being a very productive engineer and CTO. He didn't stop writing code aiming to coach engineers
Starting point is 01:40:42 between 2022 and 2024. And then he started to code again because with AI tools, he wanted to coach his engineers better, but it's hard to do if you don't use the tools yourself. And now we've finds himself being extremely productive, and this time the team around him is productive as well. And we're not talking about vibe-coded software, but database-worthy, high-quality code-generated and committed to production. Peter is convinced that AI amplifies existing expertise, and this is one reason why he probably learned more this last year, building with AI, than the previous five years combined. And I find it a valuable reminder that learning and building deep expertise in software engineering, this is very valuable. And as closing,
Starting point is 01:41:20 I appreciate it that Peter said that not only is he excited, but he's also exhausted. There's a lot to learn, but it's tiring, and neither him nor anyone I know is immune to this. So if you're also exhausted with all of the things going on with AI, know that you're not alone. Check the show on us for more of the pragmatic engineering deep dives on Google's engineering culture and on distributed systems. If you liked this episode, please make sure you're subscribing your podcast player and a special thank you if you leave a rating. Thanks, and I'll see you in the next one.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.