The Pragmatic Engineer - Distributed databases with Peter Mattis
Episode Date: September 30, 2026Brought to You By:• turbopuffer – a vector and full-text search engine built on object storage. It’s fast, cheap, and extremely scalable• Linear – the product development system for teams an...d agents• WorkOS – everything you need to make your app enterprise ready.—How is it that a software veteran who regularly shipped ~100K of database-grade code to production each year, pre-AI, feels like he’s even more productive today, with no drop in quality? Peter Mattis is co-founder and CTO of Cockroach Labs, and an original creator of GIMP. He also worked on Gmail and distributed storage at Google.In this episode, Peter reflects on his journey from open source to Google to founding a database company, and we explore how to keep systems fast, reliable, and correct at scale, from Gmail’s early storage challenges to the tradeoffs in building distributed databases.Peter tells us how AI has brought him back to writing code after his work shifted toward management, and why he believes AI can improve quality and multiply the impact of domain experts. We also consider the future of code review, and Peter has some advice about how to level up our engineering skills.Timestamps00:00 Intro02:42 Peter’s path into tech04:00 Building GIMP09:30 Working on Gmail at Google14:51 Google’s infra: google3, build files, Bazel, and Colossus21:30 Distributed storage bottlenecks23:59 Latency, throughput, and availability30:04 Contributing to libraries41:52 Google Spanner46:10 CockroachDB52:00 Manual vs. automatic sharding55:28 Consistency models and strong consistency1:00:03 Raft consensus1:06:15 How AI brought Peter back to coding1:19:12 Peter’s tools and agentic workflows1:23:08 How AI can improve quality1:26:39 Code reviews: are they done?1:29:17 100x engineers1:35:33 Peter’s advice for leveling up your engineering skills—The Pragmatic Engineer deepdives relevant for this episode:• Inside Google’s Engineering Culture• Resiliency in distributed systems• How to debug large, distributed systems: Antithesis• Pushing software engineering limits with “napkin math”• Designing Data-intensive Applications with Martin Kleppmann• Formal methods with Hillel Wayne—Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com. Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe
Transcript
Discussion (0)
He built a Gimp image editor while at college, designed the original storage system behind Gmail,
and built many more large, complex, and widely used systems.
This is Peter Mattis, co-founder and CTO of Cockroach Labs.
Before we sat down to talk, Peter told me,
for the last 30 years, I've always been a prolific coder,
but my current output is a bit insane.
And this isn't vibe-coded junk,
but database worthy, high-quality, high-performance code
thanks to working with strong coding models.
Today we cover why bee trees are so important when building databases
and why Peter kept breaching for the data structure over and over throughout his career.
How he wrote 100,000 lines of production code per year pre-AI,
stop coding between 2022 and 2024 and why he is back now.
Why he thinks AI agents are lazy about testing and how this is an easier fix than it looks.
And many more.
If you're interested in distributed databases, distributed source systems,
or knowing how Peter manages to use AI to produce unusually high quality and production
ready code, this episode is for you.
This episode is presented by Turbo Buffer.
A ridiculously scalable, fast and cheap hybrid search engine
built on top of object storage by an engineering team
I've really gone to like after spending time with them.
The turbo buffer engineering team is doing something really, really cool.
They're completely redesigning their storage architecture
from first principles to make storage faster, cheaper, and more reliable at scale.
If you follow Turbo Buffer, you know that their storage architecture
was a massive part of their early success.
Redesigning a winning architecture is a big deal.
It's one thing for your query plans to pass unit tests.
It's another to bring each query plan to performance parity or better while maintaining correctness and reliably in production.
Here's a cool part.
TurboPuffer is documenting the whole thing.
Their new storage architecture, which they are calling T-Puff V3, demand some hardcore systems engineering, and they're building it in public,
sharing design decisions and benchmark results as they ship.
They're keeping a work log of their journey and the first post just dropped today.
Follow along at turbopuffer.com slash V3.
That is turbopuffer.com slash v3.
Peter, welcome to the podcast.
Oh, I'm happy to be here.
This is awesome.
I wanted to get into.
How did you get into tech originally?
When did you figure out computers are interesting?
I figured that out in kind of elementary school, high school.
Gaming was a little bit of a gateway drug for me as a software engineer as many other people.
I remember, like, early on, my mom did programming at IBM at some point.
I'm never even quite sure what she did, but we had computers always around our house,
like Apple 2 Plus, Apple 2GS.
I'm from that era on up.
And, you know, go to the bookstore.
I'd find a book on basic or magazine, type in programs.
No idea what it was doing.
But, you know, just kind of like I was addicted.
Like, you can produce, put stuff into these computers.
You'd get interesting stuff out.
And then I got to college and I was like, I didn't think there was any money in computers.
I didn't know anything about it.
I started a mechanical engineer following the footsteps of my dad.
Oh, you started mechanical engineering as your socialization.
Yeah, as my major.
Yeah.
And I got in there and I was like doing, like, I'd done some.
computer stuff before. It was a real foolish move to do this, but first semester of doing these
homework assignments and mechanical engineering, they were god-awful, like six pages for a single
problem. And I happened to take a CS course at the same time. And it was so easy. And then
everyone was failing it. I'm like, I'm in the wrong field. Let me switch. And then you switched.
No, I switched, yeah. What was the first software that you built, either at college, it must have been
a college, that you were like, all right, this is a piece of software that I'm kind of proud of. That's a
complete piece of software. Well, I mean, the big thing that I did, along with my roommate in college,
we had this course, was it a compiler's course? I'm quite sure anymore. This is like 30 years ago.
And we were kind of bored with it. So we wanted to do something like kind of fun on the side.
And I'd done journalism in high school in my senior and senior year. I knew stuff about kind of
computer graphics and wanted to do something like Adobe Photoshop. So started just kicking the tires,
and we built up this program that a lot of people know of called the Gimp. And along with the GIMP, I did a lot of the
the graphics library, GTK.
This has since evolved just massively since then.
It's kind of interesting because after college, I kind of stepped away from it.
Didn't really stay involved much past my first year out of college, but definitely people
still know me and it led to other some interesting events in my career.
And it was like starting Gimp, it was literally just you saying, all right, I wanted to do
something like Photoshop, how hard could it be?
And then you just, this wasn't through what you learned in college, right?
This was like you figuring out of how to build, you know, like a graphical.
I guess engine, rendering, drawing, data structures, all of that stuff, right?
All that stuff.
Figured it all out.
I remember trying to look at some papers back then.
My roommate was looking at papers.
We're just fingering it all out.
And it's one of these things.
You almost kind of need to be a little bit naive to start anything like that.
Because if you know how hard is going to be, you would never start it.
So I think there's this like, you have to have this level of like, if you're really
truly wise about the effort, you like, you never get into it.
But then when you get going, you just keep on snowballing.
It gets far further and further along.
And then at some point, you're like, wow, this is really.
awesome. But there's an interesting little tidbit associated with our work on the Gimp, which is kind of
a little bit known. I think we've talked about this before, but it's just kind of fascinating where
we were getting to the point where we're like, we should release this to the public. And back then
there was Usenet news groups where people would post stuff and there's one on graphics. And remember, like,
just a couple weeks before we were going to release our first version of the Gimp, someone else came
on there like, I've been working on this graphics program. And it did everything the GIMP did,
every last thing. And then some. And we're just like,
well that sucks.
I guess we'll just keep on working on.
It's been fun.
And then we release the gimp.
Never heard from this other guy again.
And I just took a little lesson with that.
There's always going to be someone else working on your idea.
You can't get dissuaded if they pre-announce it.
Nothing ever comes of it.
And a lot of the marketing behind some of that stuff that you might hear is like,
I mean, I don't know if we, like, we stole his thunder.
I'm not even sure what happened.
Never traced down what happened.
But there's a real lesson there.
So there's a possible future where you read this
announcement of someone saying, I'm going to build all of this thing, and you go like, uh, you know,
someone did it and you kind of go back and just do something else and the game never happens.
That's right. Wow. Yeah. I guess especially with today with startups, you know,
sounds like just do your thing. Put it out there at the very least, right? The advice I give people
is like you might have a unique idea, but most likely there's like a dozen people out in the world
who've had the same idea. And there might be a couple of them working on it, but a lot of people
don't even work on it. There is like whatever they can't get going on their idea. So, you know,
I wouldn't be concerned at all.
If you hear someone else working on your same idea, that's probably the case.
You know, we're working on some cool stuff now at my current company.
I guarantee there's other competitors out there working on the same thing.
And just, you know, it's like it's a competition.
You got to enjoy that aspect of it, not be afraid of it.
And what were some fun things that the Gimp led to?
Well, you know, I got kind of tired with graphics.
And that's why I kind of moved away from it after college.
I got into storage systems and I bounced around, you know,
I worked at one of the early search engines thinking to me.
And then I got over to another startup.
And I met through that time period in that first startup I was doing,
met Lergan Sergei from Google.
Because it turns out the learning, I think it was maybe Sergey.
I'm not quite sure.
The very first version of the Google logo was done in the Gimp.
So they knew of us.
Somehow one of them, you know, looked me up, asked me to come interview with Google.
I did back in 2001.
This is when Google was three years old.
Yeah, three years old.
And I said no.
You did not.
I said no.
Yeah.
No, here's the calculation in my mind.
I was like, Google's down in Mountain View.
I was living up in San Francisco, and I didn't want to do the commute.
So I went and worked at another startup for a year.
And that one, after a year, like, I saw it wasn't going anywhere.
And they called me back up and they said, hey, do you want to interview again?
I was like, sure.
Like, we're not going to be able to offer you the same stock options you got before.
I'm like, okay, I can't remember what it was.
And I still can't remember what they offered me the first time.
But it probably would have made a lot more money if I'd take that first.
But I did all right.
I'm not like complaining, but it's one of these things like, yeah.
Now I think back about that I'm like, oh, a lot of good stuff came out of it.
At some point along the line, Red Hat was going public.
They actually offered friends and family stock during their IPO to a lot of people.
We got offered friends.
I made a little bit of money off that, not a lot, but it was like just after college.
It was actually quite significant at the time, you know.
So, I mean, there's a number of good things that kind of came out of that.
And I also used the gimp as a, you know, I was looking for like, oh, I can't support
Photoshop and here's the gimp.
And it did so many things.
So like, I'm sure there's so many people like you had a positive impact of being able to use
this thing. And the nice thing that I, what I really liked about it is it was free, but I
wasn't like stealing any software from anyone. You see what I mean. This was at a time where
free software was not as common as today. Free and open source was not as mainstream as it is
today. Yeah. And then the other cool thing I've heard from a lot of people, because people still,
like I mentioned this, are like, oh, I've used to it. And then some people, you know, software
engineers are like, I learned how to program by looking at your code. And I'm just like, wow,
that was the code I wrote 30 years ago. And I wasn't nearly as good as software engineers I am
And then you got into Google second time around.
You said yes.
What did you start to work on?
Yeah, yeah.
Well, I got in there.
They said like, hey, you know, this was early day, 2002.
It was April 1st, auspicious date.
Your start date?
Yeah, April 1st, 2002.
And I got in there and they're like, well, you know, we're going to actually build email,
Google email.
And it wasn't called Gmail at the time.
There's a code name internally.
It's called Caribou.
Cariboo.
Yeah.
And got in there.
I started working on it, and I was kind of tasked with working on the back end threading and message storage and indexing system.
And worked on that, you know, pretty hardcore for the first year and a half.
Actually, I was on it, I think it was three years on it through the launch, which actually happened to be April 1st, 2004.
Do you remember this?
There was a, it was like considering April's full joke.
Yeah, I remember.
So correct me if I'm wrong, but the launch said Gmail with one gigabyte of storage or something or infinite storage.
and like star one gigabyte. I'm not sure what it was, but back then, most email providers
would give you about 10 megabyte of storage, the free email providers. And then you could pay
for maybe 50 megabytes or 100 megabytes, but that was really expensive. This launch, it looked
like an April Fool's joke because who could possibly give you a free email service with
20 to 50 times or 100 times more storage? We were talking about this internally. It was like kind of
a shock and awe campaign for the industry. I think it was actually four megabytes for like
hotmail or, yeah.
our email and then you got this. And not only that we had so much more stores, but it was indexed really
fast. So you know, you could do a search, you know, come back almost instantaneously. But can you
tell me internally? Like when the project started, okay, we're going to do email for Google. How did
you and the team arrive to the point of like, okay, we will offer all this storage and we'll do it fast?
Because this was at a time, if you can take us back. But what I remember is hard drives were still
expensive. They were relatively slow. We're talking HDD. We're not talking necessarily.
about SSD, if I remember.
But if you can take us,
can you take us back of what,
what it was like,
what the constraints were like,
and then how you innovated,
like to actually, like,
do something that's never been done before.
Yeah, yeah.
So at the time,
Google internally had this large distributed file system
called GFS, Google file system.
Yeah.
And we were like,
basically looking at numbers and being like,
yeah, we think we can build on top of this.
They had a lot of knowledge
about how to do search and retrieval.
It ended up being that, like,
some of the existing systems
they had for search and retrieval,
they started building the prototype on there and then that was completely written.
That was actually what I got involved in because I got in there and there's already a prototype
and then it was completely rewritten.
So I was on that, you know, threading side.
We decided to have message threading right from the get-go, which is also kind of innovative.
That was in common in email systems.
And spreading, do you have the message shred, which I assume needed data structures on the back
that it needed like, you know, storage and figuring out the read intensity of those kind of things.
Yeah, you know, there's some bee trees involved in that.
I think that might be the second time I implement B-D-D-Sorture.
trees and I think I've implemented them like a dozen times now. How do B trees relate to email threads?
It's just like you have to have some storage there where you have a thread ID, a message comes in,
you have to look up. It was using the search index to match actually the subject or the message
ID into the thread. And there are some other things taking place in the B tree as well because we were
keeping track of the unread counts of threads and whatnot. So, you know, I can't even remember all the
details now. This is ancient history. This is 2004. What are we in? 2026? 22 years ago. So it's like, left my
memory, but there was definitely B trees. There was also this, you know, inverted into next taking
place there. All this code has since been completely rear in. What were the economics? Like, in the
team, you must have done the economics of being able to offer free email, which of course, I'm sure
you did some mats of like how it could be subsidized, but to like not make a terrible, terrible loss.
Yeah, yeah, yeah. No, I mean, we were kind of a kicking around ideas. And, you know, one of the ideas that
had come up at some point was, hey, maybe we can put ads on this. And it was one of these crazy things.
wasn't involved in this. This is Paul Bukite,
went on to do some cool stuff
at a friend feed and Facebook
and ended up, I think, a partner
white combinator at some point. But
he just like one night, he's like, no, I think I can
just take out of some of our existing ad
functionality and incorporated in there.
And it just went like gangbusters. And there's
a huge business that, you know, kind of grew up from that,
which is kind of incredible. So, I mean,
I think there's a lot of things that's just like, you
kind of have a sense of what can be alone. Like, oh,
this is going to cost, you know, per user
a couple of dollars per year. How do we
monetize that and we didn't want to charge for it. Eventually they did charge for it because
you have the whole Google workspace stuff. But at the time, it was more like, can we do this for free?
Oh, it looks like it can't be economical. And it took a little bit of a leap of faith.
And do you remember the launch? The reason I'm asking, because I remember that it was an invite-based
system. Like, not everyone could get in. And I assumed that must have been to control the expected
demand. Because again, you were, you're offering something for free that was paid before. And it was
kind of pretty obvious that there would be massive demand.
Yeah.
How did you think about it, kind of monitored demand, decide how many people to onboard?
Honestly, this is a team effort.
And I wasn't involved in that.
I mean, I do remember being worried about like the load that would happen.
And then someone came up the idea of like, hey, we should do an invite base system.
And this also had like, it played dual role.
So it kind of constrained the growth.
But it also kind of drummed up this excitement.
Oh, can get me a Gmail invite.
I remember people asking me at the time was like, yeah, I can get you on this
my Gmail invites as you want.
And then after after you built the system, where did you move on to?
What was your next project?
Was it the built system?
Well, there was the build system.
I mean, that was kind of muddled in my mind because I was kind of doing that part time
when I was doing Gmail as well.
At some point, you know, Google actually had this large monor repo.
I think they actually still have the monor repo.
They still have the monor repo.
And they started out with the Google One repo.
That was before my time.
Then they moved on to Google 2.
That was what was there when I got to Google.
And at some point, we saw the strains of Google 2.
And Google 2 was just one single monolithic make file.
I think it actually had some sub-make files,
but it was this really large unwieldy make file.
Someone came to me and was like,
I think we could do something.
I'd had some interest in the build systems.
And I kind of put the foundation in for Google 3.
There was a bunch of people involved.
I was kind of doing like kind of the initial work on Google 3.
And Google 3, the initial insight was like,
hey, make files kind of sucked a right.
We introduced this thing called build files.
And I decided to do it as just this stripped down kind of Python language.
But it was still Python at that point.
And what G-Config spit out at the end was a monolithic make file, but one you didn't have to write.
Over time, this evolved.
It became Blaze internally.
There's some other systems that are associated with now.
I don't even know how complicated it's gone.
I haven't seen it for a long, long time.
Then that became basal externally.
It became Buck.
There's some other people went to Facebook and they like, that was awesome at Google.
What was the reason that make files or make didn't really work?
Was it, were we talking about build performance?
Are we talking about maintainability or readability?
Yeah.
So, like, I think Make itself is like, it's okay in declaring dependencies,
but it's a little bit like kind of assembly language.
And you didn't want to write, you know, all your dependencies in assembly language.
And if you didn't know what you were doing, it was easy to make a mistake and miss dependencies and whatnot.
And so the kind of my thought behind the build files is like, hey, you need to express these dependencies,
but at a higher level and just like kind of cleaner semantics associated with it.
And from that, you can compile down into the assembly.
complete language. And then eventually folks were like, oh, you don't have to compile down to that.
We can just kind of implement, you know, the dependency kind of update engine directly.
And then that's where the performance improvements were able to come in. So like,
because at Google scale, like when you have a large repo, in general, like, my understanding
is that the reason Basil and Buck are so popular for large code basis is it can help you
improve your build performance. It gives you a lot more levers to play around with from
caching, from being smart about cache generation to obviously just raw performance.
Yeah, that's exactly.
So you kind of just dabbled and like, okay, I'll make this build file.
Did that people take that over?
But what was your next main focus?
Oh, well, so I mentioned GFS earlier, Google File System.
And at some point, we realized that there were some limitations in GFS, scalability bottlenecks.
Because I've been working on the storage system for Gmail, I actually dabbled at another storage system, you know, kind of a research thing.
They never went anywhere.
But because we're working on that, I got invited to participate in, like, the founding team of Colossus, which was the successor to Giac.
And as far as I know, Colossus still exists.
It's gone through multiple iterations at Google,
but it's like the second generation, you know, distributed file system.
So what is Colossus?
Yeah.
So when I say distributed file system, externally, you might think of something like S3.
Kind of blob storage, it had a flat name space, kind of like S3, you give names.
And there's like a minor hierarchy there, but it's like very limited.
But it's not like a Puzzix file system.
So you don't have the full directory hierarchy.
You didn't even have the full like kind of permission system.
A lot of that stuff got it added later.
But the files are not.
stored on your local machine. There is a fleet, you know, kind of a service out there that has all the
files. They're writing it down to their hard drives and now SSDs and your client can access that.
And it's all replicated. So if there's any crashes and whatnot, you're not losing your data.
I don't know when at some point S3 got erasure coding. We did a ratio coding in classes. That was
kind of a big breakthrough. Which coding? We used Reese Solomon. So for the audience who's not
familiar with erasure coding, you might think of like, I want to have replicas. And there's
this thing in hard disk is called Raid, where it's like, I don't actually have to have
full replicas. I can actually, you know, if you take A plus A and B, you can X-Werm together,
and then you can have this kind of third kind of version. And there's more and more complicated
versions of that. Reed Solomon is kind of, you know, I think there's actually better codes now,
but it's like one of the known ways to do this. One of the known ways, yeah. And we had to,
you know, kind of pioneer internally like, oh, how are we going to actually make this work
in distributed file system? Where you focused on latency, on being able to store data more
efficiently. It's storing it more efficiently. So in GFS, and there's varying costs where it's
like, well, you're not storing two copies or one copy of data or two copies is storing three.
Triplication. Three times as much storage you're having to use. And with Reed Solomon, you can get that
down quite a bit lower. I can't remember offhand exactly what are, what we used for Reed Solomon.
I think it was like essentially two X. But you get that with also the redundancy too. So it's like
it's smaller and the redundancy is higher. So because I guess the naive thing, if you're saying,
all right, I want my data to be replicated at three places.
You take three notes, three machines, physical machines,
and you say like copy one, copy one, copy one.
I have it three places.
Great.
If one explodes, I still have two, wonderful.
And then you're saying that the algorithm here is you could take not three X
the data, but two X of data, split it smartly across machines,
or maybe you could take it and take it lower.
And you still have the thing where like, oh, one of them explodes.
I still have all my data because it's split enough.
That's exactly right.
And like kind of the mental model, you know,
you just want to understand at a high level, which is essentially you might want to say,
like, I want to have eight replicas of this data.
I think, I can't remember offhand.
I think S3 might use nine.
They've actually talked about this publicly.
You have kind of nine chunks of data, but any five of those chunks can be used to reconstruct it.
And what this means is you can lose any four copies and you can still reconstruct your data.
And oftentimes it's more like, you know, the first five chunks are exact replicas and the other
four are kind of parity ones.
I've kind of forgot some of the details that's escape my mind.
But it's along this.
Yeah, but when you come up with an algorithm, you can then prove that this algorithm will work, right?
Like, this is a little bit like, I know in software engineering, like, maths and algorithms a bit out of fashion.
But in this case, this is really important because once you, once you can prove it that this algorithm works, it will work.
It will work.
And, you know, it's like the math behind here is like a Gaw fields over like GF2, something like that.
I don't even know that.
I never actually understood the full math behind it.
I always regretted not doing more math in college.
but you didn't have to you.
Like read and solve them and proved how this work.
I think it was like back in the 1970s,
something associated with like communication network.
So you just take that and, you know, kind of use that expertise,
but leverage that and have to do all the engineering behind it
to make it work in a storage system.
And then when building a distributed storage system like Colossus,
what were other things beyond, okay,
you want to store data in a resilient but efficient way?
What were other things that problems that you needed to solve?
I'm thinking things potentially like sharding and resharding,
or metadata being important, those kind of things.
Particularly for the scale that Google wanted to operate at,
GFS kind of had this scalability limit.
I believe it was like, you know,
you could have a thousand machines in a GFS cluster
which I guess sounds big until like today,
it's kind of really small, right?
It sounds big for many people.
And then Google was like, no, we need to have this scale up to 10,000 machines.
And there's some bottlenecks.
There's this GFS master, it's a single node.
It was a bit of a bottlenecks.
We're like, oh, we need to have a distributed
and master to store the metadata for all the objects.
And Google at the time happened to have this system called Bigtable.
And so Colossist stored its metadata inside Bigtable.
And one of the things I'm kind of proud and like, kind of like also, you know,
a little bit embarrassed by one of the design choices I went down,
but actually worked out is we wanted to use Bigtable for the metadata for Clossus.
And the metadata, just for those of us not as into this reason, what is the metadata
and the distributed file system?
Yeah, it's like the names of the five.
And for each of the files, the files have broken into chunks.
What were they, 64 megabyte chunks.
And then you have to have the list of chunks for each file.
Yeah.
And then you have to periodically, you know, the master has to be scanning over this and
doing repair, repair work.
And, you know, but there's more metadata in that, but that's in a nutshell.
So we wanted Colosses to have this, you know, kind of scalable, you know, service, Big
Table to store its metadata.
The biggest user of GFS is Big Table.
So we want Big Table to work on top of Closses.
So you have the circlic defendants.
Yeah, well, there's a bootstrapping thing, right?
Boostrapping, yeah.
Oh, yes.
Yeah.
Which one starts up?
Like, you need to mock something somewhere, right?
Yeah, no.
I mean, the way it actually worked at the time, and they've since replaced this, you know,
because this is like the way to get started and leverage what you have and eventually you kind of
get rid of it.
But there is the, uh, kind of a foundational big table.
That big table didn't use classes.
There's the classes using the big table.
And then there's normal big table sitting on top of classes.
And it all worked for years.
So I don't even know when they got rid of that.
got rid of it at some point, but that was...
But I guess sounds like you can make, like, hacks that go really long, knowing that they're hacks and they get you off the ground, right?
They get you off the ground, right?
Because if we had to, like, you know, implement that kind of big table layer from the get go,
or it just delayed how long it took to get, you know, Colossus built.
One of the things that strike me about a system like Colossus is it promises, or this was internal to Google,
but even distributed file system that are external, they will promise high things.
throughput, high availability, and low latency.
And to me, it's always a bit conflicting of like, well, it's pretty easy to, I guess,
build a distributed file system with my limited knowledge.
I could probably do something where I have either high throughput but high latency,
because whenever I write something, I read it out to all the replicas.
You already, like, mentioned one technique of doing it, but how did you kind of reconcile?
Like, how did you get, like, low latency while you have high throughput, while you also have
a replication going on on the file system?
Well, I mean, these distributed file systems, I mean, the latency isn't super low.
In particular, when they're running on hard disks, which Colossus was doing at the time, which S3 does, you actually notice the latency.
So, like, S3 is a high performance system.
Google GCS, the competitor from Google, which is built on top of Clossus.
It's high performance has incredible throughput, but the latency is like 20 to 30 milliseconds for first read.
And that is bounded by your hard disk latency.
If you put it on SSD, it gets down to closer to SSD latencies.
But not actually kind of the state of art.
SSD latencies, which is kind of this crazy thing that's been happening in our industry.
It's just like how much faster the hardware has been getting.
Reading from a hard drive, maybe five to 10 milliseconds nowadays, and we never touch hard drives.
Reading from an SSD over NVME, 30 microseconds, 50 microseconds.
So that's a microseconds.
There's a thousand microseconds in one millisecond.
So we're talking like a huge, huge difference.
Well, there are now startups for infrastructure companies that are starting to take advantage
of the fact that they can have an NVME layer.
and they pull things up either predictively or not.
But as you say, like, when the physical reality changes,
you can build systems on top of it,
that should take advantage of it.
Yeah, absolutely.
You know, stuff.
The disks are so much faster with SSDs,
the networks are so much faster.
I mean, just kind of crazy how fast,
like the intra-zone leitzees are to Google or Amazon Center.
And it wasn't like that, but when we were building colossus,
you know, I can't even remember what the numbers were,
but milliseconds to do a network round trip.
And now it's down in, you know,
100 microseconds within a zone.
I mean, I just look at these things and I'm like, holy crap, you know,
the hardware guys have really done a good job.
Yeah, sometimes I feel that our software should feel way more snappy.
And there are some snappy software, but sometimes I almost wonder if we're getting too
complacent with all the abstractions or not even doing this like napkin maths.
Simon Erickson at Turbo Buffer talks about this napkin math word light like you.
It was like, all right, here's the theoretical limitation of the hardware.
reach from SSD might be, I don't, 30 microseconds.
And then, like, how can I build a system that is as close to this as possible as opposed
to the other way around saying, okay, like, you know, human will notice like 20 milliseconds, like 100
milliseconds.
Let's build around that.
Yeah, yeah.
And sometimes, like, when you're architecting something, you need to think about the human kind
of perceptible latencies.
But oftentimes when you're dealing at the storage system layer, you know, I know Simon working
on TurboPuffer, they're doing great stuff over there.
You have to think about the machine scale and the machine speed, which is a lot.
a lot faster than human perception.
Like, a human can tolerate maybe 100 milliseconds of delay, or if you're playing a game,
maybe you need to have, like, you know, frame rates of like every four milliseconds.
But the machine wants much, much faster in that.
Simon says, napkin math, I call this speed of light numbers.
And, like, sometimes it's literally the speed of light bottlenecking you.
The cross-zone latencies between zones, cross-region latencies, is speed of light and fiber.
You want to hear a kind of crazy fact?
I love hearing crazy facts.
The fastest way to send active data across the globe is to send in a descent.
space. Is it because the speed of light is faster in vacuum? Quite a bit faster.
Or, well, it's not full vacuum. No way. So you cover a higher distance. Yeah, well, you actually,
I believe the way to do this, you send it straight up. And then with Starlink, you send it straight up,
you bounce around between Starlink, you send it down to the other side. So you want to get out of
the atmosphere as quickly as possible. But now, if we're effort on my speed of light, this will also go.
There's like a digital transformation happening. So you would need to calculate how long it takes for
that system to process and do it.
But you're saying that even if you do this like really well, it will be faster than
beaming it through an optical cable.
And an optical cable slows down the speed of light, right?
Yeah, yeah.
Yeah.
The speed of light is only the speed of light and vacuum in every other medium.
It's slower.
I was talking.
I did a deep dive on the hedge fund industry.
And they didn't tell me exactly.
They said that they do use satellites and microwaves and some of these things.
They will not tell you because, you know, this is their thing.
But I had a suspicion that they might have found a faster way.
And I think this is like somewhat well known, but the details they're not going to get into because, again, that.
But like, yeah, so they're probably balancing stuff in space.
Yeah, yeah.
So, I mean, one of the things that they do in the high frequency train is like between New York and Chicago,
it's actually not far enough distance-wise to make it worthwhile to send it in space.
So they were doing microwave beams there.
Yeah.
But, I mean, if you really want to get faster, you need to build like a vacuum to you between them and it's like send a, and maybe they're doing that.
Yeah.
But, I mean, the thing that I think is like just kind of awesome about performance nowadays is, I mean,
There's so many layers you have to be paid attention to in terms of performance.
The rabbit hole just goes so deep there.
You're almost certainly running on multi-threaded systems, right?
And you're like, oh, I need to have a multi-threaded program.
I need to have synchronization in there.
Well, you get the best performance if you just kind of avoid the synchronization.
And part of this is lock-free programming, but part of it is arranging so you don't need locks
at all.
You need to carry about your processor caches.
And there's this whole kind of setup of caches above the CPU.
Like you can think it registers of the cache, and you have L1, L2, L3.
even your memory and onto disk,
and there's just you have to pay attention to all those levels.
And if you do, your performance gets way better.
And if you ignore it and you're like, oh, I'm not worrying about kind of,
kind of the cash access is.
I'm just accessing data all over.
Your program will just be way, way slower.
And some folks pay a ton of attention to this.
The high frequency trading, I mean, they do this all day long
and in so many other places like no one pays attention to it
and you get this kind of gradual degradation of the performance.
of the hardware or the software.
I did want to talk a bit more about low-level stuff,
but not about the speed of light,
but low-level data structures and programming language features.
You made some contributions to the standard library, right?
Well, I've done a couple.
Not quite standard libraries.
So, I mean, I just, like, have been always fascinated by data structures.
It's kind of awesome.
I mean, I think just algorithms in general,
and you're just, like, sorting algorithms, just kind of awesome.
You could probably explain, you know, insertion and, you know,
to kind of every...
Could you explain?
quicksort as well.
Quick sort's a little bit harder.
Right, that's the thing.
Yeah.
No, no, but it's a smart one.
Yeah, and then you get to these levels of like, you know, it's like, oh, wait,
someone really smart came up with this.
So one of the things I worked on, just as a little bit of a side at Google at some point,
was a colleague came to me, and he was like, you know what?
We're using the STL map structure all of the place.
And the STL map structure is a balanced binary tree.
I can't remember if it was red black trees or one of these other balancing algorithms
that any CS college students implement.
And he came to me, he's like,
I think we could do better because, you know,
there's actually a cash problem here.
Every time you're every node you're traversing down,
you're going to a different cache line.
And he was thinking about using something else,
a skip list to do this,
which that's another awesome data structure
everybody should kind of look at.
But at some point I was like,
actually, this feels more like a bee tree.
So I implemented bee trees a couple times before
and figured out like how to implement a bee tree
that implemented almost all the semantics
of the STL map.
It couldn't quite do it perfectly.
And the reason is when you insert into a B-Tree node,
you have to shift stuff around
so you don't get pointer stability.
This is just kind of fundamental.
But if you can, you know,
you don't need that for your use case,
you actually can pack more data in.
So the thing about a B-tree,
like the real easy way to describe this,
is like you just have a small list of items,
like eight items.
The best way to store that,
if you want to kind of fast access
in sorted order is just to sort the items,
right?
Literally not to have a tree at all.
Yeah.
For like eight items.
So I have it in the very simple list.
Very, very, just an array, sort the array, and then you can either do a linear scan over it.
You can do a binary search.
And oftentimes, the linear scan is faster.
And then you think about that I'm like, well, if I want to store worn eight items,
I can just have one node that has eight items.
And then I have another node.
And then you have a parent node that connects them together.
And that is essentially like the, you know, you build it, you think about building
it bottom up.
You start with just one node of eight items.
Oh, I need to insert the ninth thing.
I'll split into two chunks.
And the two chunks, one will have four, the other will have.
five and then you have an apparent node that points to them. And then you just kind of recourse on that.
That's the B-tree algorithm in a nutshell. Everybody go implement it. Actually, nobody should
implement this anymore because nowadays we have something else that we'll implement this in all the
optimizations because there's a crap ton of optimizations that you can do on a B-tree.
But I just want to go back to this. Like there was already an existing implementation for maps.
And then so your your colleague looked at the code and said, I think we can do better.
What I want to figure out is like in my mind, in someone sitting outside,
of, you know, I'm not involved in how some of these libraries are, or data structures are built.
I always thought, and again, this might be naive, but really smart people sit down,
they kind of look at the state of the art, they implement it, and there's no way it can be
faster.
In fact, I've had arguments in the past saying, like, oh, let's write a faster sorting thing.
Like, it's surely it is the fastest.
But if you could bring us a little bit of, like, how, like, you've been inside of how it
actually how it happens and how other people like yourself and and your colleague can say like,
oh, what, what if we, what if we try something else? Yeah, yeah. So I mean, my recollection here is he was
working on this kind of the big internal system. I think it was called Gaia that actually had
the mapping from, you know, you log in, you have your user ID and you have to look this up. And
they were storing, you know, all like this, the map from user ID and email to, you know, whatnot to the
metadata about the user in STL maps. And you just know, it's like, well,
there's a lot of memory usage here and it shows up on profiles.
And then we're like, well, what can we do to do better?
And that was kind of the genesis of it.
And he happened to be working on it and he happened and they were working with me.
Like, we just started kind of noodling on this problem.
Like, oh, can we do something better?
And it's not one of these things like, I think now with Google,
they have a whole team working out of their kind of internal libraries.
At the time, it was more of like, you know, everybody working on their own systems
and contributing to a shared base.
But I guess it still goes back to what you were just saying of like,
just go down the layers, try to understand.
And if something just doesn't add up, like, suddenly like, oh, there's this big exposure memory usage, like, you know, just ask the questions, why is this?
And if you're able to or you happen to be like, oh, can we do something about it, right?
And one of the things that, you know, he observed earlier on.
I think part of one of the things was it was like a map from integer ID to something else.
And you know, like if you look at red black tree, every node, you have your kind of value that you're storing the map.
And then you have two pointers.
You might have an energy ID that's like four or eight pipes.
And then two.
It's a waste.
Yeah.
And you look at it and you're like, oh, that seems like a lot of overhead.
And you're like, you could just a bit.
Well, the bee tree actually has a lot better.
It has better spatial locality and that's what made it faster.
But it was actually smaller as well at the same time because you had less pointers involved.
You also contributed to go, right?
Yeah, yeah.
Well, that came later.
Yeah, it came a lot later.
But can we talk about that?
Yeah.
I mean, just one of these other things, you know, I pay attention to like, you know,
when there's research papers coming out about new data structures and like hash tables.
Hash tables are like the, one of the earliest things you.
learn about in college and data structures. Like, how do I map keys to values where the ordering
is unimportant? That's when hash tables come in. And there is like, you know, the very earliest
ways to do this. I implement hash tables multiple times is like, you take your key and it might be a
string and you put through a function and it spits out in integer and then you map that into an array of
buckets. And if multiple things map to the same bucket, you have to have a link. You have a linkless.
Yeah, this is a naive implementation.
naive implementation, used quite frequently.
And over time, people discovered, like, a lot better ways to do hash tables.
There's very, like, that's called a chaining of your hash.
There's another technique called open addressing, where instead of actually having a linked list,
you just kind of hash it again and move on or kind of walk down to subsequent buckets to find out,
like, or subsequent slots to find out where you should be.
And I remember reading about this new technique and it came out of some folks at Google.
I believe it came out of their Swiss office, because it was called Swiss tables.
I believe that's where the naming came from.
I'm not 100% sure about that,
but I remember reading about it.
And then I was working on Go for a long period of time,
and Go has this built-in map structure.
And it's a hash table.
It's a very highly optimized hash table
because the Go team is very competent,
the Go Run team.
And various folks have taken an attempt at like, you know,
putting together a Swiss table implementation for Go.
And I tested some of them,
and I was like, this is kind of fascinating what Swiss tables do,
and I'll explain how it works.
in just a second.
But I looked at it.
It's like, well, it's really hard to beat the performance of the runtime.
The runtime was really good.
And I kept on, I noodled on this for a little while.
And eventually I ended up having to take this business trip to India, to Bangalore.
And so I was on a long flight.
No.
Yeah.
It always starts like this.
And I'm just like, I'm just going to try to pull on this.
I pulled on it sufficiently that I can get some of the benchmarks to be faster.
Wow.
And then I'm like, you know, that like is kind of like catnip for an engineer.
Like kind of make it all faster.
figure all the rest of it, you know.
Got some help from the runtime folks.
There's an issue on the Go issue tracker that, you know,
where other people have been attempting this.
Because people propose like, hey, let's use Swiss tables.
And like, like, oh, first is like, well, you know,
we don't quite know all the details.
You're going to have to navigate this and that.
And there are some ideas there that combined them together and got to the point
where it had kind of a complete implementation that was faster on,
on most benchmarks, not quite all of them, but most of them.
And then the Go folks eventually picked this up and push it over the finish line.
And then so you, you know, like,
You came with the idea you got to the point where you were able to show an implementation
that showed how some of the benchmarks were faster. And then you started to work with some
folks on the goal team too. Well, I didn't, it wasn't quite that. I came up with an implementation
that we ended up using at my company. It was good for our use case, but actually putting it into
the runtime as a whole other, you know, kind of level. But then you just showed like here is this
implementation and then they, they kind of took the inspiration and the ideas. Yeah, they're like,
well, this is great. We want to make all. They always are looking for ways.
to make it run time faster and there's like, you know.
Oh, wait. And then so, so you did this in this,
this is a lot of years after you left Google, right?
Yeah, yeah.
So this was from the outside.
This is from the outside.
That's awesome.
Yeah, and other people contribute stuff to the outside as well.
You know, we had another colleague, um, he contributed one of the CRC implementations,
you know, adapting some stuff.
CRC, uh, cyclic redundancy checksum.
Mm-hmm.
You know, Intel published some papers about here's how to do the CRC and assembly
very fast and he contributed one of the implementations.
You see a number of those things where, you know, people just like are contributing
externally. It's not a lot, honestly. I mean, I actually don't know the full details, but, you know,
people are regularly contributing to these things. Peter just described how the Swiss table work
came together using an issue tracker in the GoT tracker with different people contributing to the
work and then the Go team pushing all of this over the finish line. This is where I need to mention
our season sponsor, Linear, which is a place to coordinate work between humans as well as agents.
One thing I've noticed about how most of us work with agents is how it's a pretty single-player thing.
You open a terminal UI, go back and forth with an agent, and it usually produces a PR.
But the rest of your team has no idea what happened in that chat unless you tell them or copy the whole history.
And when everyone in the team works like this, a lot of work happens that's invisible to the rest of the team.
Linear's take is that agent work should be teamwork.
Even today, teams already use linear to define the work to be done.
Now, Linear can already delegate an issue to a coding agent.
This agent could be an AI agent that linear integrates with like Codex or Cursor or Linear's own agent or a custom agent.
Either way, the engineer and delegating stays responsible for the outcome.
What I really like about how linear works is how the work stays visible.
Your teammates can follow the session of the agent, check out the PRF producers, and join their review.
We've gone from single-player work to multiplayer engineering work with agents.
Oh, and one more thing I like about linear, a focus on costs.
Linear agents' auto-riding chooses a model that is the best suited for the task.
Teams can also inspect usage and set limits, track usage, so you can use capable agents
without having cost balloon out of control.
Hop on board at linear. app slash pragmatic.
Peter also previously mentioned Gaia,
Google's internal system that map logins to users.
It's not surprising that Google custom-built all systems,
including this one,
but most of us won't build our internal GAIA.
This brings us to our season sponsor, WorkOS.
You can think of WorkOS as something like Gaia for the rest of us,
identity infrastructure you've otherwise spent quarters building yourself.
Workerless includes single sign-on,
skim, directory, sync, audit logs, role-based access,
control basically everything a big customer security team asks for, delivered as a handful
of clean APIs.
It's how companies go from, we have a login, to we can sell to a Fortune 500 without
standing up their own internal identity platform.
And WorkOS is already building for the next version of the agentic authorization problem.
Their newest product is Airlock, the authorization layer for AI agents.
Think about what happens when you had an agent attached like clean up the sale opportunities
in our pipeline.
The last thing you want is for this thing to have standing permissions to delete what
whatever it likes. Airlock sits between your agent and the tools they call. It evaluates every
request against the agent's intent and your rules, and then it allows it, denies it, or routes
it to human for approval. You write down the policies in plain language, the agent never sees
your credentials, and every call and verdict gets logged. It works with coding agents like Cloud
Code and Codex and with MCP gateways. So if you're working out how to let agents do real work
in production without over-permissioning them, take a look at
at WorkOS Airlock at WorkoS.com slash airlock.
And with this, let's get back to Peter and why he left Google after building Colossus.
You're at Google, you're building Colossus distributed phall systems.
You're at this point probably working on probably the larger system on the planet, honestly.
Why did you even consider leaving?
Yeah, yeah.
Well, after Colossus, I kind of dabbled in this other project called Google Goggles for a little while.
Remember the glass holes?
Yeah.
Keeps coming back, the idea, by the way.
Yeah, yeah, no, it's still here present.
Seems like Google was early.
Yeah, yeah.
And, you know, I think that was a technology before its time.
I don't think it was ready to do at that point.
It looks like the actually doing the glasses is quite a bit harder.
You know, the Android phones we were trying to power it on were, you know, not powerful enough.
And then, you know, I just kind of got wanderlust, you know, like, you know,
am I just kind of stagnating here at Google, which is a strange thing to say, but, you know,
some other people feel it as well.
And decided to go off and try my hand in another startup.
That didn't work out.
We got Aqua hired by Square.
So I want to pause for a second.
So this company, what was the company name?
The company that we found is called Viewfinder.
Viewfoundure.
Yeah.
It was in the mobile photo sharing space, which should sound familiar.
This is like Instagram.
This is like Snapchat.
This was in 2012.
Yeah.
Right as Instagram and all the more we're taking off.
Yeah, yeah.
We were right there in the play.
And we just didn't have the right go to market kind of like how to track the users,
how to get viral growth kind of thing.
Because from the outside, like what I read, when I, when I check, you know,
the story just like, oh, you know, like you, you co-founded a startup. It got acquired by Square.
Hooray, like it sounds like you had bigger ambitions. And this was a decent outcome, but not the
dream, right? Yeah, yeah. No, it wasn't dream at all. So the term I used was aquired. So sometimes
a company will get acquired, get bought for, you know, their IP, for their product, for their business.
And other times, they get bought just for the talent. The people. The people. And we got bought just for the talent.
They acquired the IP, but I don't think wherever did anything with it.
It wasn't kind of like where they were working.
But we built up a kind of strong technical team and that's what we were hired for.
And, you know, like, I can't remember the details.
We'd raised a small amount of money.
We were able to pay our investors back, make them whole.
Maybe they got a little bit of a haircut, but maybe they actually got a little bit of a,
but it was essentially, they got their money back, which is like, you like, you know,
as a founder, you, uh, you know, investors are big boys.
They're used to losing their money, but you kind of feel bad if you lose a lot of money for them.
So, you know, getting them paid back kind of makes you feel a little bit better.
Yeah.
Yeah.
And then you, you spent it a little time at the company that acquired you were just square.
Yeah.
And then you started itching that a little bit again.
Yeah, yeah.
Yeah.
Because, you know, we've been working on these, you know, distributed file systems and storage systems, you know, Colossus.
One of the kind of sister projects to Clossus is Spanner.
And how is Spanner different to Colossus?
Well, Spanner is essentially a distributed database.
Colossus is a distributed storage system.
And the way I think about the difference between a distributed storage system and distribute database, you might think, oh, they're both storing data.
I was about to ask because a distributed database will at some point be a storage system.
Right, right.
So what's the difference?
Yeah.
So for Colossus, you know, it was targeting large files, large append only files.
You can't update the append only.
Yeah.
Yeah.
Large append only file.
So, you know, 64 megabytes, maybe up to gigabytes in size.
But if you're a database, you want to be storing like kind of small, like, you know, kind of
You know, if you're using SQL or like the relational data, you might have table with, you know,
billions of billions of, you know, rows.
Those rows are broken up into columns.
The columns are typed.
That just has a very different nature to the engineering challenge for database than it does
for a distributed storage system.
And usually distributed databases are implemented on top of some kind of distributed storage system.
And that was the relationship.
So Spanner was implemented on top of Colossus.
And some of the design decisions in Spanner kind of directly fell out of the appendage.
only nature of the files in classes. You can't update a file in place, so you have to, you know,
make the files immutable in your database. And this is where, like, you know, log-structured merge trees,
you know, come into play. And they weren't invented at Google, but Google really popularized
them with Level DB, which emerged out of the work on Big Table and Spanner. That got popularized
into RocksDB. I subsequently re-implemented one of these things, and this is what we use at Cockroach
Labs. It's called Pebble. So I'm very familiar with the internals of that. But it's kind of all
based on this idea that the data is kind of stored in the mutable files.
So how did you decide to found Cockroch Labs?
While we were working at Square, my co-founder and I, there's actually three of us.
We were all at Square.
And one of them is Spencer, I mentioned earlier.
He was working on the Gimp.
He's my college roommate.
And he was also at Google.
He was also at Google.
This other guy, Ben Darnell, who was also at Google.
He joined us at Viewfinder and ended up at Square.
We, you know, like, we're kind of just noodling on a project to do.
And we'd actually had this design.
back in Viewfinder were like, ah, we didn't, we looked around for a data space to be using
in order to build viewfinder on top of. We didn't really like the things that were out there.
The technologies that side Google looked better. We had the big table. We had Spanner and whatnot.
We were looking around and, you know, like H-base existed, but I wasn't quite happy.
And there's some other systems like React and others. And, you know, in one point, we're just kind
of like, you know, came up with the design for Cockroch to be an initial design. And I was like,
no, no, guys, we're doing a mobile photo sharing site. We shouldn't build a distributed
database. So we put it on the back burner, which I think was absolutely the right thing,
maybe, or maybe we should just pivoted away from doing the mobile photo sharing site,
given the way things worked out. And then we got to square and we saw some of the same problems
that they were experiencing with data storage systems. And Spencer is very convincing.
My here's convinced some of the management that like, hey, can you just work on this part-time
and see if it had life behind the design? And then kind of conscripted Ben and I into it.
And eventually it started gaining attention externally. And we're like, hey, can we go and
spin this off into a company? And that's what ended up happening.
And so you started a company, but I understand you didn't raise VC funding initially, right?
That was the case at Viewfinder.
We did it differently.
At Viewfinder, we kind of eschewed the VC money.
And, you know, in hindsight, I wouldn't recommend that.
You wouldn't recommend?
Yeah, I would recommend taking the VC money because my experience, the VCs are very, very intelligent.
They can help you navigate a lot of challenges.
You know, I think sometimes there's this perception.
you know, the VCs will push you into various areas.
And maybe there's some bad ones out there that do.
The VCs I've had experience with, just like some of the sharpest people, you know, I've ever met.
And so you kind of have like an extra like person helping you on the team pretty much.
Mentoring you, giving you guidance, telling you what they're seeing.
They give you advice seeing what they're seeing in the market where things are going that is very hard for sometimes for a founder.
Especially as a technical founder, right?
That you're, your focus on the engineering part.
Yeah, yeah.
Yeah, no, we actually took money right away for Cockroach Labs.
it was almost like as soon as we left
we got and did a little road show
you know kind of in the Bay Area
and got some interest and you know
got an investor right away.
I have to ask about the name though.
Yeah.
How did the name of cockroach.
Yeah.
So, you know, we named the Gimp.
Yes.
That was, that was mine.
Pulp fiction had come out in college
and like, oh, which we named this thing.
Oh, the new image manipulation program.
I think we're thinking image manipulation program initially
and I'm like, oh, Gimp, it's obvious.
It just stuck.
And at some point, you know,
we're newly on this new database, and you kind of want to give things a name.
You can't just say, oh, we're working on this distributed database.
You kind of need something.
And the sponsor was like, oh, Cockroach DB.
Like, cockroaches are unkillable.
I want these things, this database to be unkillable.
You know, Cockroach is going to survive the nuclear apocalypse.
So that was where the Genesis was and just stuck.
Yeah, we're, one of our bunch of our nodes goes down.
This thing will still be up.
Yeah, yeah.
And, you know, this is where we're at today with Cockroach TV.
It's like one of the things that I point out on like, holy crap, this is awesome.
You kill a node.
We did this whole campaign last year, which was really just to prove out something that already been present.
The campaign was performance under adversity.
But just like you can run a workload against it, you can kill a node, you can sometimes kill a whole region, and the system keeps on going.
And it's like stories like that, you know, like what we did on the marketing side there, but also we hear this from our customers too.
They've had fires and data centers and all the other data systems go down and Cockroach TV keeps on going.
I'm like, that's awesome.
When you started out, who were companies, startups that who wanted to use Cockroachers DB?
be and how has it changed since? Because, you know, like just making the case like, okay, I'm starting a
startup. It's a small startup. Like, I will need a database and I'll, I don't know, I'll typically choose
a Postgres, right? It's free. Everyone's using it. I'm running it on Node. At what point
did you see that typically tech companies are like, okay, like, this is not enough for me that it's
running on a node either because it can go down or because I'm outgrowing. What was it the outgrowing?
I'm trying to get a sense of, like, at what point that companies say, like, tell themselves, like, we need something distributed in a database.
Yeah.
I mean, oftentimes we have companies calling us up after they've had disaster.
So, like, a node went up or a hard drive fail, that kind of stuff.
You know, it's not quite like we're ambulance chasers, but if you see an outage, like a big outage from some, you know, company, you know, like we'll sometimes be trying to knock them up.
But also, they will call us, you know.
I mean, like, there's a very big bank who's now a customer.
Don't think I can name them, but you can go read.
They had a very serious outage due to a weather event.
And after that, there-
Which probably knocked down, I'm assuming a region or database or a networking cable
or tree fell on something.
I think it knocked down a whole region.
You know, it was a region-wide power out.
You knocked down the region.
And there's a mandate from the CEO.
It's like, no, we just have to, you know, be able to survive these things.
And that gets pushed down all the way.
And you see this in other places where, you know,
one of our early customers, they were running on AWS.
and they just got to the maximum size you can run in an aurora instance on.
And then what would typically happen at that point is then you have to charge your database.
This is a very standard practice.
You take your single note database, you create 10 or 20 or 100 charts.
And this is what Google has done for some period of time.
And that's a heavy burden on the application developer.
And the way we always phrase this is like, I mean, the application developer is becoming
a database developer at that point, and they're doing it poorly.
You know, they're trying to implement distributed transactions or indexes and whatnot.
And we felt the burden for that belongs on the database developer.
Can we talk about automatic sharding?
I think it's safe to some, most of us will know what sharding is when you're,
but actually, let's start from like manual sharding and then how you can implement automatic sharding.
And if you can tell us like, you know, tactics that a database like HockwoodDB can do to actually just take that load off of you.
Yeah, yeah.
So I think the very basic form of sharding is a little bit like the hash table.
Let's say, you know, you have a fixed number of shards.
Like, let's say it's just a hundred shards.
your data model is a user with a lot of data that's said with the user.
You just take the user and you say like, oh, they map them to one of the shards.
And, you know, you kind of just rely on the hash function to get like fairly even distribution.
The problem with this is at some point, you know, one of your shards will get full and you have to kind of reshard.
And that's a very, very onerous process.
And reshart, I guess, simple way to do is like if it's just a hard drive, I don't know, per node,
where you write the user data, it gets full and you're like, okay, well, I now need to split it somehow.
I need to move it.
I need to remap it.
I need to reject my metadata,
which knows where the data lives,
that kind of stuff.
Yeah, yeah.
And depends on exactly how you're doing that mapping
from like,
you know,
the user ID or whatever your shard key is to the shard.
You might have to remap them all, right?
This is very typical.
Exactly.
So, I mean,
this happens in hash tables where,
you know, oftentimes in order to grow the hash table,
you just have to essentially create a new hash table,
double the size and copy all the data over.
Now, that's kind of like the very basic,
straightforward way.
And there's various levels of complexity you out on it.
One of them is called consistent,
And there's various techniques to do this.
It's kind of fascinating, like how they all work.
But in consistent hashing, you can add an additional node,
and then it only moves a fraction of the data from each chart over there.
There's various systems to do that.
And I believe this is like when it lies Cassandra.
The way Cockeridge DB does it is more into big table,
more into span, or more into H-base, where instead of actually
hashing, we actually take, you know, you can imagine all your keys in a system.
And this is always true in a system.
You can imagine just in one big, contiguous key space,
And then you kind of partition contiguous spans of that.
And then you have to build up an index on top of those contiguous spans.
And what I just described there actually sounds a lot like a B tree.
So there's this index on top that is like that maps you from, you know, like, I need to have this range, which node is it on?
And this is a little bit like a B tree.
You know, he kind of squint, you know, it's like, I think you squint and everything's either B tree or it's a hash table.
And, you know, that index structure.
But it's now I'm starting to make sense because when I remember when I read, it might have been the Wikipedia article on B,
trees, it's said, this is a data structure that is frequent to use in databases.
It's now coming back to me because I didn't think too much of it.
I'm not a date.
I'm not someone who builds databases.
But now that we're talking, we just like organically keep touching on trees again and again.
And the other place that it comes up in databases.
So this is where kind of comes up in distributed databases.
And no one ever really calls it a bee tree.
I just kind of squint sometimes.
And I see like actually kind of a bee tree.
But the other place that comes up in databases is for your indexes.
So if you have a table and you have like an index,
on your email address.
And you want to be able to scan over those email addresses in order.
That's a B tree under the hood in a database.
And, you know, like any kind of index you have,
it usually provides sorted order.
There are hash indexes, but oftentimes it's the B tree index.
And they're ubiquitous in databases.
There's actually a paper called the ubiquitous B tree.
And, you know, basically, like just identified that.
I think that paper was written back in the 80s.
And they're still ubiquitous today.
They are the foundations of single node databases.
And pretty much every data system I worked on has had B trees at some point,
it placing them. I want to ask about strong consistency. So CockroarsDB offers strong consistency.
Now for people who are a bit more newcomers to distribute a systems, can we talk about
the consistency models and then why strong consistency is important and why it's hard to implement it?
So, I mean, there's multiple ways to kind of approach this, but I mean, if you've used a database,
you've probably heard of transactions. And transactions are a way to perform a whole bunch
of mutations atomically. So databases talk about atomicity, consistency,
isolation, durability. The durability is really easy. It's like when I write you the database,
it has to be durability written. So if anything crashes, it comes back. The animicity is just referring
to the fact I want to do a whole bunch of changes. I want them all committed or all aborted at
the same time. I don't want to have like some kind of partial operation. And why is this important?
Why is the adamicity important? Well, the adamicity is what gets you to the point where it's like,
as an application, I can do a bunch of operations. And if there's an error in, an error occurs,
it all kind of gets rolled back, and it's a much simpler development model to work within.
And then there's the consistency in isolation, which kind of get, you know, kind of muddled.
The isolation is referring to isolation between transactions.
I don't just want to run one transaction at a time.
That's easy to do, right?
I want to run a lot in parallel.
Oh, yeah.
And make it so that when they're running in parallel, they are running as concurrently as possible.
But you want to have the appearance when they're running as concurrently as possible,
that there is kind of a serial order to them.
So this is like kind of the whole track, the kind of the gold standard for isolation is called linearizeability.
Don't worry about that.
The step down from that is called serializability.
And that literally refers to having a serial order of your transactions.
But you're having to construct that in a way that you're doing everything as concurrently as possible.
And the benefit of this, the serializability, is again, it's a very simple model for the application program.
They don't have to worry about weird kind of defects occurring in their program.
And some of the ones that can occur, it's like the classic description.
is of a bank, right, where I might want to read, you know, like have an operation that reads and says,
like, do I have $100 in my bank account to transfer somewhere else? And you could like arrange
for lesser isolation levels that you might be able to subtract that $100 twice. And that's bad,
right? You know, we want to keep accurate, you know, track of your bank account or, you know,
what's in your shopping cart or, you know, it's kind of like the use cases are endless there.
And you do this wrong and you have very egregious bugs. But now going, going,
back to weak consistency and strong consistency.
Anything less than linearizability or serializability might be considered kind of weak
consistency, but there's also like, you know, kind of strong consistency and eventual
consistency.
So the eventual consistency is like sometimes I can do an operation and it might not be
immediately, I might not be able to see all the updates.
The read result will not necessarily give me the current update.
But they'll come back, you know, at some point, you know, I've written part of it and
the rest of it will show up at some point.
And oftentimes when you're doing-
It's easiest thing is your credit card balance.
right?
Yeah, yeah.
Yeah, and it's faster to do it that way,
faster in terms of just what the performance
you can get out of the system.
But again, it's a little bit harder
for the application to deal with.
One of the places this often comes up
in distributed databases or databases
that have any sort of replication
is that I could write to the primary replica
and then I read from the secondary
and it's not there yet.
That's eventual consistency.
Yeah, that's eventual consistency.
And you can often work around this,
but you just like puts bigger burden
on the application of Elk because you have to pay attention
to that.
And then with Carcores
DB, you have strong consistency, meaning as soon as you're writing it when you're reading from
the database, you already get the written value back. You get the written value back. You know, it's like
you read whatever you wrote. It doesn't matter if you're reading from the same, you know,
node you wrote it to. If you read it from another node, you still actually get the data you just
wrote. Is the trade-off logically not that you would have higher latency? Because clearly,
to implement strong consistency, you would somehow need to, in the naive approach, you would
need to write all replicas, right? Yeah. What are you doing?
inside Cockworth TV.
Well, we are writing to all the replicas.
Well, yeah.
Yeah, yeah.
You're just doing it fast.
You're just doing it fast.
You're making that efficient.
I think this is one of the things that kind of also fascinates me about the software industry
is we keep on finding ways to be more and more sophisticated in order to provide, you know,
like do things that make it easier to write the applications, but do it at a very high performance way.
And we've gotten, you know, very, very good at this over the years.
And this is the area I know about the databases.
This is happening everywhere.
Like, I'm just fascinated by how fast graphics have gotten where when I should enter the industry,
you were literally writing out each individual pixel,
and now you have these GPUs that are doing like billions of triangles per second,
whatever the current numbers are,
and just like kind of astounded, like,
how much sophistication has gotten into every area of computer science,
wherever you look at it.
One more thing on Cockroach DB, I want to ask about this is RAF consensus.
What is the RAF consensus?
Yeah, I mean, consensus protocols,
the original consensus protocol is called Paxos,
and it was famously hard to implement.
Raft was, you might think of it as a variant of Paxos.
It was kind of like an alternative to Paxos,
But in my mind today.
And then because this is algorithm being that you have like a number of noes, like three to
a lot more.
And then how do you get them to agree on?
What do you typically agree on?
Yeah, yeah.
So you agree that the right occurred.
So like the way to think about consensus.
So you might think I want to replicate data and I write it to a primary and write it to a
secondary.
And you can't actually have consensus when you only have two replicas.
And the reason you can't have consensus is if there's a crash and I come up.
Like if I'm on the secondary, how do I know?
if something was written to the primary.
If I'm on the primary,
how do I know it was written on the secondary?
Right?
And you're always going to be in this kind of confusing place
where it's like you either have to roll back a little bit
or like, you know, you lose some data.
And consensus requires at least three,
but you can have consensus across more than three replicas.
And the idea with consensus is I'm going to write to three places.
And normally you don't actually,
when you're doing a read,
you don't read from multiple of them.
But if there's a crash, I have to do recovery,
then I'm reading, oh, I can read from any two of the three
and I know I can kind of determine what had happened previously.
And it's usually just on the recovery time
that you're actually doing that consensus read.
So reads are typically just happening from one replica.
You have to write to all three.
And on a crash during that kind of failure
is when the consensus read occurs.
And inside Cockroach, DB,
how many replicas do you choose for either consensus?
It's typically three.
For some system tables, it can be five,
and then customers also have control of this.
At the database level, you can write to five, seven.
Five is like, you know,
if you're really concerned,
about the data durability, you might use five.
But there's a slowdown.
You know, the more you're writing to you, it's like the more storage space that.
And obviously, like, we're talking like a slowdown with nodes.
But of course, if they're like between regions, there's now you have a lot more resilience for,
let's say, an earthquake or power hours or whatever.
But now you will have additional latency.
It's just speed of light, basics, right?
Speed of light latency, right?
And it's, you know, tens, you know, or up to hundreds of milliseconds or even higher if you're going
across the globe.
So, you know, you have to be very careful with that.
in terms of how you architect your queries.
And one of the things that it's just a general truth of some of distributed databases
is you don't want to have a lot of back and forth.
You want to kind of do all your reads in one kind of parallel read set, get them back,
then do your rights, right?
But if you're doing like kind of serial operations where I read a row, I write a row,
I read a row, I read a row, I read a row, I read a row, I mean, the latency's just add up.
We talked about founding CockroachDB, but how is the company grown and where are you today?
Yeah, yeah.
I mean, we're being used in like we're powering mission critical applications.
That's our bread and butter.
But by the way, can you elaborate a mission critical?
Because it's like, if you're not, you are in the industry where you know what this means,
but from the outside it can feel hard to put a thumb on.
What is mission critical?
Is my SaaS that is like showing as a mission critical?
Probably not.
Yeah.
So mission critical in my mind is like these kind of the other term of art is tier zero applications.
The ones that are like just the core crown jewels of what's running a company, you know,
like a trading system.
You know, your trading system can't go down.
If the trading system goes down, this is.
kind of a critical problem or the firm that is running the training system, you know, banking systems.
But also, you know, like we work with DoorDash, you know, like some people might think that
delivering your burrito is a kind of a mission critical system. It certainly is for DoorDash, right?
You know, if that goes down, it's problematic. We power shopping carts, you know, other stuff like
this where it's like, well, if the shopping cart goes down, you know, that company is losing, you know,
hundreds of thousands, millions of dollars per hour. So that's kind of the criticality you think about.
Yeah, I guess of course that they're losing, but this is like when, yeah, their customers are also like they're used to this just working like running water.
And then when it's not the same thing as when your utility breaks, right?
Your water or electricity is out, you'll survive, but it's not what you expected.
Yeah, yeah, yeah.
And everybody's like, what age are we living in?
The electricity goes out.
We were kind of had this ingrained into our heads at Google.
It's like Gmail cannot get down.
People are, you know, relying on it.
Search cannot go down, you know.
If it goes down too long, people are going to move to other systems.
And it's not like, you know, like in some way search isn't his mission critical, except, oh, wait, every single search is ad dollars behind it.
And you can actually notice the blip in the revenue.
And it's not just the blip in the revenue.
It's the blip in reputation as well.
I mean, I think that's the one that really poisons companies is like, if your bank is down for a serious amount of time, the reputational damage there will be horrifically.
And, you know, like we often talk about like, oh, and then grandma won't be able to pay her rent and she'll get evicted.
You have to take this like responsibility really, really seriously.
No, but also like just Gmail being mission critical.
just on the way here, we only exchanged numbers later, but we were communicating over email,
like, oh, I was telling you that, that you were telling me that you're here, I emailed,
and I never for a second thought that it could go down. And I think we were like responding
within 30 seconds, right? And it's just, I just know it's there. Like I didn't like bother
setting up a secondary communication line. So. Yeah, yeah. It's like when people just like,
when you have that kind of level of trust with your users, you got to maintain it and invest in it.
But then it leads to this kind of freedom for the user as well,
where you just don't think about it, I don't have to worry about it.
It's just going to work.
And then in terms of the company, like, how many engineers do you have roughly?
We have some hundreds.
I don't actually know the precise engineering number 110,
but there might be 150 in R&D overall.
There's other folks besides just engineers.
I'm clear as engineering managers.
Those are engineers as well.
And then, you know, like we've been grown steadily.
It takes quite a while to build a distributed database.
Not for the faint of heart.
Yeah.
So it took a couple of years.
You've done it a couple times.
Yeah.
Yeah.
Well, I did distribute a storage system.
I did a distributed database.
It's not for the faint of heart, right?
So there's a lot of work getting it to a level of stability, then a level of kind of quality
beyond that level of stability, getting all the bugs out.
And then continue to innovate and put more performance into the system, adding functionality
to integrate better within enterprises.
A revenue has been, you know, kind of steadily growing over the years.
And, you know, it's at this place now where we see a path to future successes.
as well. And I want to ask about your coding habits. So when you co-founded the company,
how much code did you write for the first few years? I wrote a lot. So I've always been a very
prolific coder. There were a lot of code early days. And early days, I mean, I was a kind of,
we were all technical co founders, Ben Spencer and I. And we were already a lot of code. And I was
no exception. But, you know, I look back at my GitHub output. And, you know, it's like kind of
peak years, maybe 100,000.
lines of code in a year, which is a lot.
Yeah.
Yeah.
No.
So, I mean, rule.
We're talking pre-AI.
Pre-AI, right?
This is when you're back doing this manually, right?
You know, at some point, you know, we started out using a system called RocksDB, which is an L-SM.
At some point, I think it's back in 2019, you ran to limitations with it.
I decided we wanted to rewrite it and did a big push to, you know, right, that might have been like 40, 50,000 lines of code.
And then a bunch of other people have come up and helped and, like, you kind of look at that
output, I'm just like, oh my goodness, that was a lot to keep in your head. It's a lot just to type,
you know, 100,000 lines of code. The average kind of like that the industry talks about is
3,000 lines of code from an engineer in a month. And so if you multiply that out, maybe 36,000
in a year, that's good, right? So I was doing a lot. I kind of look at that. It's like,
there's kind of a max that you can hold in your head at a time. The tools have gotten a lot better
since I first entered the industry. We've gotten better debugging techniques, better testing techniques,
but still quite significant.
And you were CTO from the beginning, co-founder of CTO,
but there was a time,
sometime around like 2002,
when you decided to kind of be a bit more hands off, right?
Yeah, yeah.
Can you tell me about that?
I mean, the general rule of thumb for engineering leaders is,
well, you got to have your team,
you know, to manage your team,
and we have a VP of engineering,
but I was kind of getting to the point of like,
okay, is my, or my coding days done, you know,
like kind of just direct from a higher level.
And, you know, I got this advice for a long period of time and I pushed back on it, but, you know, I kind of acquiesced at some point.
And I think it was the right advice.
I'm not saying that the advice was wrong at the time.
But there was time period from about 2022 to 2024, whereas, like, my output declined.
I think I actually did the Swiss tables thing in that time period, but I wasn't doing much on the core.
The business.
Yeah, the core business.
You know, I would get in there and do some work, but like, it's really hard that if you're in meetings all day,
to also do coding.
I think this is the fundamental.
attention.
If like, so you kind of took on the kind of the meeting burden, the coordination burn and
the stuff that was, if I'm reading correctly, before you spent a lot of your head in
the code and now you're spending a lot of your head like above the code, the business, the engineering,
the org, the whatever, customers, that kind of stuff.
The customers and just being an executive as well.
Yeah.
So very hard to wear all those hats simultaneously.
And then I got back into it because AI started to emerge.
So how, when did you get, when did you start using AI?
when you start to find it useful in terms of coding.
Yeah.
Well, it's interesting because, you know, those initial versions of like kind of glorified
auto complete came out.
And we're talking about the GitHub copilot, the cursor, the early version.
That was the one I had the first exposure to you.
We tabled with cursor at the time.
But they're all like kind of glorified auto complete.
And it was kind of crazy that you could just start typing something.
You know, like, don't the rest of the function.
You look at me.
I'm like, wait, you kind of got that right.
This is crazy.
Right.
And, you know, we're trying to encourage our engineers to use this.
And at some point, you know, I can't remember if this is my idea or my co-founders or someone
basically like, you know, like in order to like really guide people about how to use it,
you have to be a user yourself.
You know, I think this is true of like engineering management in general.
Like you want to like manage engineers.
You have to know how to be an engineer.
Like if you if you don't know how to be a good engineer, it's like really hard to manage
other engineers.
I feel like you'll have a hard time.
I'm like just relating to them at the very least.
Exactly.
Exactly.
So kind of took it on me like no.
I mean like, I mean, this is like, it was.
clear very early on, like this is probably going to go somewhere, but it wasn't quite clear how far, how fast it would go.
And you start dabbling this and I was like, oh, okay, well, it's not quite good enough.
It's not quite good enough.
But, you know, let me start getting back into the coding very rapidly, you know, you started seeing the signs of life, like, you know, kind of the opus models coming out.
It was sonnet first, then the opus.
And you're like, you're looking at these and I'm like, oh, wow, okay.
Well, they seem to be able to do quite a lot, but, you know, the code is still not great.
But then it was just like just this cadence of continuing improvements.
And, you know, I was starting to do a lot with Sonnet and then starting to use Opus.
And, you know, I had that same moment everybody else did.
And this was last year, last November.
November, December, winter break, right?
Yeah, it was Thanksgiving.
I distinctly remember it because, you know.
You were not chilling.
You were coding, weren't you?
You were agenting.
I was agenting, just like everyone else.
I had this thing I'd wanted to do for a long.
period of time on CockroachDB, which is like, I wanted to like test like all these
configurations of Cockeridge DB across like, you know, different vertical scaling, like how many
CPUs you have on a node, how many stores you, discs you have on a node, how many nodes you
have in the system and tested across all this huge matrix of it.
This is one of these things they couldn't ever quite get prioritized appropriately because
it never seemed like kind of critical, but I've like, I always had this intuition there was
something there.
And then over this like four day span, the code just like materialized as I was using, I
that was Opus 4-7, or is it 4-5, whatever the number was.
Like, the recollection I had is, you know, like, a pretty fast typeer.
And then I just remember having this feeling of, like,
well, the code is just, like, materializing before my eyes.
You know, you would ask for these things.
I got out of the habit of actually typing,
and you just kind of, like, ask for something materializes.
If you've ever seen, like, some of those, you know, a movie where they had this kind of
archetype of a software engineer gets in front of the keyboard,
and they start typing, show the screen, it's just like,
it's going way beyond.
on human speed.
It was gone that, right?
Like, it was, I think a software engineer is like, we used to laugh at these.
Like, I still remember Swartfish, the thing when they're visualizing, the things are
going for, or the code appearing, you know, you're seeing that the person's typing and
then it's a big line.
And a software engineer is like, we're laughing because it's not how it is.
But it's crazy that, like that effect, right?
Yeah, yeah.
But now it's.
It's even better than that effect, right?
Because it actually works this time.
And it's, that's slow in comparison.
It materializes faster than that.
Like, you can literally go.
Like, we've been talking about bee trees a whole bunch.
I implemented another bee tree in the past month.
Of course he did.
Yeah.
And it took about 30 minutes to implement probably 10,000 lines of highly optimized rest.
I mean, it's just like it boggles the mind.
I mean, we should look up later at 10,000 lines.
It's like a crazy amount.
You physically cannot type that fast.
You started getting back to coding.
Like, are we talking about kind of vibe coding prototyping,
or are we actually talking you started to contribute like
proper production ready code that is up the level of what you're doing a whole
to be.
Well, it started with this tool that was like this kind of benchmarking tool that tested this
matrix.
I have the CTO.
You have an office of the CTO.
The office of the CTO's mandate is to innovate.
And I was looking for places like we can have innovation.
And one of the things that we want to innovate in was, you know, better auto scaling
of a Cockroarstabee cluster.
and kind of in the January time frame,
I came out with, like, what I would think is, like,
kind of a research breakthrough, you might say.
And it came about, because I was dabbling this area
and just working with the models trying to understand it.
And they're not just good at coding.
They're also good at, like, helping you explore design ideas.
And I think this is kind of the fascinating thing
where it's, you have to get out of the mindset of, like,
I know exactly what I'm going to build,
but more like, hey, we have this problem, talk through it,
be a partner with me.
And it would be a sparring partner.
You know, like, there's this advice you might have heard
that like, you know, if you're stuck on a problem, you should say, go yellow duck it,
go rubber duck it.
I think I've heard it as yellow duck, you know, just go talk to something.
You don't even need it to respond.
Now you can talk to this system, this intelligence, and it will give you back stuff.
And it's not always right.
I mean, this is the thing.
Even today, these models are fantastic, fables fantastic, Astros is fantastic.
And they're not always right, but they kind of like, they know so much.
It's just encyclopedia.
And you can explore ideas, super, super, super.
fast and then be pointing out, well, that doesn't sound right to me and it'll be like,
oh, yeah, you're absolutely right. You know, like, I hate that sick fancy. It's like it kills me.
Or you're right to push on back on. You were right to push the back on that. That's a new absolutely
right. Yeah. And yet, like, just able to make such fast progress. And I, I find it absolutely
incredible. So it quickly moved from just doing kind of side things, building up some tools and whatnot
to building, you know, essentially over the last eight months since January. We've been building
towards a, you know, a new launch. And, you know, like a lot of that has been produced
a gently powered engineers. And like the entire company has gone on board with this now,
where I think it's like probably 100% of engineers are using it. To a greater or lesser degree,
and a lot of code is being produced. I think it's very high quality code. These models
not test things adequately. They don't look quite intently enough about performance. But if you're like,
you can guide them in the right way. And I think there's actually a huge advantage.
to anybody who's done management before.
It's a little bit like being a manager of people
where you're a manager of a large group.
You're not looking to every line of code,
but you're definitely kind of helping architect the system.
I think there's a very strong analogy there.
I sometimes push back on this analogy.
The reason being I was an engineering manager.
My take is that working with these agents,
it's not like management because management has so much of the human stuff.
Like as a manager, when I think of all this,
a bunch of stuff I dealt with,
which was the people side of things,
the conflicts between people,
the performance reviews, the meeting, et cetera,
and you have none of that.
You do have the orchestration.
Like, again, this is like almost like,
I guess a really like naive way of management where like,
you know, they don't push back.
They start to do it.
Sometimes they're like unreliable, but you know,
I almost like to use like orchestration a bit more because I feel management is
so much more involved.
Like I think this is saying where like a tech lead who has no management's
responsibilities, but they have a group of interns,
but they don't need to do with their performance,
with their anything.
Like it's a lot closer to that if you know what I mean.
Yeah, I do agree. I use the engineering manager shorthand, but it's really about being a tech lead for like a 30 or 40 person organization, you know, or being an architect for one. I think the term architect gives me a little bit of a distaste, but if being an architect, he's also on the ground. A hands-on architect. And you don't have any of the management stuff, which is a blessing and curse. But it's kind of remarkable that you can spin these things up. If they make a mistake, you can keep on correcting them until they get it right. We've always been able to do that on the human side.
I can do it faster.
I think the ultimate result for me is you just have to be more ambitious about everything
you do.
You can produce more, higher performance, higher quality, more secure.
So our ambitions have to raise up.
You also mention that with AI, like you're ever using and you're building stuff,
but also with CockwoodsDB, you're now building something that is also related to AI.
Can you talk about that?
Yeah, yeah.
We're building multiple things.
I mean, AI is the future.
It's here to say, right?
Yeah, it's here to say.
Yeah, I mean, like, you know, one of our thesis, which is not crazy, every application in the future is going to be written by AI.
You know, I think there will be some kind of bespoke software, you know, handcrafted software.
I think we'll see that continue to exist, just as people write assembly still, you know, but it's going to be diminishing in size.
So, you know, asymptotically approaching 100% of software will be written by AI.
Generated, right?
Yeah.
It's hard to say if, like, when it's going to end where the humans are the agent, you know, kind of like supplying the agency and the vision.
behind it. I think that might exist for many, many years, but I think the code will fundamentally
be written by the AI. And I think we're going to just see this explosion of applications. And we're
seeing that inside Cockrookrookers lives. We've seen it elsewhere. Earlier this year, we kind of rolled out
this internal platform where non-engineers could write kind of many applications. I mean,
this is like, you're hearing this at other companies. We did the same thing. And over the course of
just a couple months, you know, 500 applications, a thousand applications.
By non-engineers.
By non-engineers, primarily by non-engineers.
And I thought it was awesome.
Like our HR team is building like these little applications.
Like, this is the thing I've always dreamed about doing.
And like, they were never serviced.
Like, like your CFOs all over the place.
And ours is no exception.
Producing dashboards that they can never create it before.
And I think it's very empowering.
I'm married.
I have a wife.
She needs software.
She cannot produce that software in her own.
And I have never actually helped her produce the software, which is my own failing.
But I think there's a world in the future where she gets, like,
everyone's game custom software.
software built for them. And you're just going to also see greater and greater applications and
higher quality systems being produced by every company as well.
Today, what is your stack? What do you work with in terms of harness model, how you run
agents? Is it one agent? Is it multiple? What can terminal do you use? Yeah, yeah. It's evolved
over time. So, you know, like when I first started dabbling the AI stuff again, it was, you know,
GitHub co-pilot. It was an Emacs user for like 20-something years. I got convinced to move to VS code.
but that's all gone now.
At some point, I moved to using cloud code.
That was the thing that, whatever reason I just got to start using it,
it was cloud code in the terminal.
Of late, I use a mixture.
So I'll just grab my current setup.
I use the cloud desktop app, cloud code via the cloud desktop app.
It's fantastic.
Good job, Anthropic.
I also sometimes use the Kodak's desktop app,
just so I have an alternative model to turn to.
For some very critical things we're working on,
I will get, you know, one of those models,
sometimes like the best model,
Oftentimes I'm using the best model, like Fable.
You know, I've recently started using that.
Sometimes it's Astra.
Sometimes it's the other ones.
But you have one producer design.
You have the other one being like, hey, my colleague produces.
Can you kind of tear it up, you know?
Aversarial review it.
And it's not always perfect, right?
You know, but I think there is utility, especially for something that's super, super critical.
We're pushing towards the launch really soon and kind of rushing towards the finish line.
I'm telling folks on the team, it's like, you know, especially our very senior folks,
like, you should use the best model right now.
Yeah.
It's worthwhile to do that.
I'm generally just using the best model.
And part of the reason is I don't actually know
that I'm getting a lot more intelligence
from it, but I don't want to have the cognitive overhead
of deciding on a case-by-case basis.
Yeah.
Should I use sonnet?
Should I use opus?
Should I use fable?
Should I use soul or astra?
And I think, you know, maybe if you, like,
I really need the speed, I would make that decision.
But oftentimes I'm like, I'm doing things in parallel.
So you're asking how many agents I'm spinning up?
Well, it depends.
Like, oftentimes there's like a certain number of sessions.
you might be using. I find my kind of cognitive overhead is about five to ten sessions
concurrently, but sometimes those sessions will have many sub-agents doing things. Yeah, they'll be
long running. Yeah, it'll depend on what I'm doing. So as I was coming here, I was on the train,
I spun up something to do a little bit of research, and that was just probably using three sub-agents.
Other times, it might be 20 sub-agents, you know, with Anthropic has dynamic workflows. I can't
remember what the Codex thing is called, but they have various ways of spinning up kind of
graphs of agents and, you know, prosecuting them. I mean, I think it's sometimes I'm probably
using 100 sub-agents. Others, it's only like five. It depends on what it's happening. And sometimes
I'm not doing any of that agentic stuff because I'm often thinking about something else.
Since you were an Emacs user for 20 years, you know, using the terminal, right?
Well, E-Max, but how come you move to graphical interface with agents? A lot of a lot of people
still use terminal UIs? Yeah, yeah. Well, I, so I moved from,
EMAX to VS code.
I was just like a colleague was like, hey, these IDs are really good.
You need to try them out.
I was like, EMAX is an ID.
And then I kind of moved and like, you know, some of the stuff was just more seamless.
And like Emax keeps on catching up.
It always feels a little bit, you know, six months through a year behind.
And like every time they have a new version, my setup would break.
And I just got tired of that.
So I moved to Vs code.
But then I was like, I'd have many terminals open in Vs code.
My setup would always be like terminals in VS code tabs.
The terminals and tabs, you know, have like eight of them open.
and then they started taking over.
It used to be like just Claude code in the terminals.
Yeah.
Yeah.
And then, you know, just some point recently is like six weeks ago,
colleague was like, well, if you tried out the cloud desktop app recently, it's really good.
I was like, yeah, no, I'm very, very happy.
And I made the switch and it was a little bit awkward at first.
And then suddenly I'm like, holy crap, this is awesome.
You can like manage the agents a bit better easier sessions.
You have the sessions there, the sessions are down the side, you know,
and it's like they're kind of like tabs in some regard.
And shit, but the whole integration that Anthropics have been doing in.
Open the eye is doing the same thing now.
By the way, like, I'm not a dissing cursor and factory and cognition.
They're all pushing the same thing.
It's incredible how fast these systems are innovating right now and evolving.
What do you think good software engineering looks today compared to like, you know, four years ago before we had a, has it changed, has it changed?
Has it changed?
Yeah.
I think the ambition has to increase, you know.
I think quality has to be higher.
Security has to be higher.
I'm glad you mentioned quality.
Can we talk about that?
because I'm seeing across the industry, just quality declines, which you cannot fully put your finger on AI,
but oftentimes it is people pushing out more and more and just not paying attention to just small regressions here and there.
Again, like you're building a database.
Have you noticed any or have you gotten feedback of any quality regressions or if not, like, how come?
Because like when, like this is just basic, you know, like we talk about the law of business.
This is an observed thing that when you start to have more output, you increase your deployment
frequency, you often, not always, but you often have like more regressions and more bugs.
If you produce a certain number of lines of code, you're probably going to have a certain
number of defects per line of code, and now you can produce more lines of code, so you probably
would have more defects, right?
But the thing that pushes against that is that you can be telling these agents, and you
have to give them kind of firm hand at this.
This is a thing that I hope the model providers are listening to.
But you need to give them quite a furred hand.
They kind of get a little bit lazy on the testing side.
And you have to make sure the tests are comprehensive, but also they're using all the
testing techniques.
And there's a lot of testing techniques out there.
We have a lot of knowledge about how to do testing well.
And you know what?
The agents are lazy.
Humans are a little bit lazy.
So getting the humans to actually be very disciplined about their testing is also challenging.
And I think it's actually easier with agents.
You know, you can kind of instill that.
You can get them set up.
You have to give them the guidance.
Use property-based testing.
Use metamorphic testing.
Use, like, these advanced testing techniques, deterministic simulation testing.
You know, there's, like, technique after technique you can use.
And humans have always, like, I'm lazy myself.
Like, I'm pretty good about being disciplined about testing.
At some point, you're just like, okay, that was enough.
We got a ship.
And, like, now you can get a little bit, like, you know, a little bit stricter and firmer.
I think the same thing applies over on the security side, the performance side.
I mean, we've always had at Cockroach and in the industry, we've known security coding practices.
Yep.
You can literally have every single line of code on every commit reviewed by a security expert.
Doing like an adversarial security review by an agent or multiple agents.
Or multiple agents, right?
And we're going through a tough time in the industry right now with like kind of hacks and security leaks and whatnot.
There's only a limited number of bugs that can be in software.
I think we'll get it, you know, out it on the security side and on the quality side.
And then also just on the, when I say quality, it's not just bugs, but it's also like, you know, just a little thing.
It's like, oh, that UX element is wrong.
And there's no excuse for that right now.
Like, fixing it is so, so easy.
And what we're also starting to see evolve is like, you know, designers using Figma.
Well, I think that should be a thing of the past right now.
Our designer just deals with HTML and CSS and JavaScript directly.
And sometimes is even just producing pull requests, you know, PRs, which she loves.
We all love.
Everybody's happy with this.
there's none of any of this waterfall handoff.
So she, she's producing poor requests for the production code base.
Yeah.
Yeah.
And you know, this is on the UX side, not on the core database.
Of course, but this is the area that, like, she owns, right?
Yeah.
Yeah.
It's just, I mean, she loves it.
Everybody loves it.
I mean, there's no downside.
This is not new.
Yeah.
For sure.
Yeah.
What about code review?
What's your take on code review?
I think it's pretty, it's starting to get a bit controversial, like, is it going to
stay or not?
Because there's been a practice that's been around, like, think about, I mean,
you worked at Google.
Google has been so big on code review.
As I understand, there are two code reviews, right?
The domain expert reviews it, and then there's a language expert reviews it.
So you went through that.
How did you think of it and how are you thinking of it now?
I used to be part of the, I was one of the early members of the C++ coding.
I'm asking the right.
Okay, you will have strong opinions.
Tell me.
Yeah, yeah.
I mean, I see that code review, I think, was fairly essential before the age of AI.
but now we're getting to the point where I'm reviewing and looking at code.
I look at most of the code that the agent's producing still.
And yet it's like you're not giving it the same level of scrutiny, right?
And this has always been the case.
If you get a pull request from a junior engineer, you have to give it more scrutiny
than if you get it from your most senior engineer.
And you ask any, you know, tech lead, any engineer manager, any software agent,
they'll see, yeah, yeah.
So the junior engineer or someone new to the code base, you know,
you just have to get more scrutiny.
And what I'm finding is the agents are getting better,
you're having to give less and less scrutiny.
And you do so have to give scrutiny to some things.
It's like I said, it's like the testing will be incomplete.
And maybe they didn't follow the security stuff.
Maybe, you know, like there was a performance regression.
And I think it's like, you know,
assuring there's kind of constraints on the system
so it's hard for them to do the wrong thing.
I suspect that, you know,
I don't know if it's going to be this year or next,
that we're materially going to stop looking at the code
in the same way that we don't look at assembly anymore.
Yeah.
Like you will...
We trust the compiler to produce pretty good assemblies.
We trust the compiler to produce good assembly.
Unless you're one of those people where it's really important,
you're maybe a game developer or someone and you look at it,
but there's fewer and fewer of those folks.
Yeah.
And even then, I think what will be migrating to is why does the human have to look at the assembly?
The AI should look at the assembly.
Like I'm doing something right now where I want to have a zero overhead abstraction,
you know, for doing something that is done at test time and what it's in production is compiled out.
Babel set up for me a little system where he actually looks at the decompiled code to verify,
like, that there's only a few extra instructions put in place.
I never would have put that in place, but now it has this guardrail,
that every time it changes this, it can verify that there's no regression.
And I think you're going to see more and more of this where you kind of put guardrails in place.
I think the model kind of likes it because it can work within that guardrail.
Now, for what, like 20 plus years, you were writing so much code.
Like I'm sure you were in the zone.
Do you remember being in the zone and just turned?
And now you wrote a lot of like production ready code.
Now that you're coding with AI, like, do you get into the zone?
Yeah, absolutely.
And how is that zone?
Is it the same?
Is it different?
It definitely feels a little different.
It's maybe a little bit less intense, but you're managing more things cognitively.
Like, can you describe me like what is it right now when you're in the zone?
Yeah.
Well, I'm thinking of ideas that normally would have taken me a week to experiment with.
and I think multiple of these experiments,
and then I fired them off all simultaneously,
and then I'm kind of like reviewing, like,
what else should I do while that's,
what's being complete?
Sometimes I'm kind of reviewing,
like, they'll be giving me progress updates of like,
oh, hey, this is coming in,
we're seeing this stuff,
and I'm being like, well, that doesn't sound right.
Hey, what about this?
Or do we do this, you know, correctly?
You know, maybe I had design,
and it's not implementing the design quite perfectly.
But I have this little feeling that, you know,
I haven't been a college professor,
but maybe I was a, you know,
if I was a college professor,
I had a whole swarm of research,
research assistants and they're all off doing things.
And it's coming back, but it's coming back just really rapidly.
Really rapidly.
You're not waiting months or weeks.
I'm not waiting months or weeks.
And then I'm iterating.
I'm like, oh, that one failed.
That's fine.
You know, you just got to let go.
And this is the nature of the software I build.
I think there's other pieces where it's like you just whip out a website.
You can whip out something that doesn't have this level of kind of scrutiny, you know, very
quickly.
But I talk about this in a, I have a whole bunch of analogies about like the feeling
to be.
Let's talk about analogies.
What analogies do you have about using AI or?
I mean, the one I was, you know, advocating for,
I had kind of two that I was advocating for, like, late last year and this year,
which is, you know, AI is coming for us.
It's here, right?
And it's like, you know, you're producing software.
You're like walking down the road.
And sometimes, you know, someone will pass you.
They're running.
They got an efficient gate and whatnot.
But everyone's under their own locomotion.
And these agent-e coding agents came.
And it was like a car pulled up next to you.
You get into the car.
You don't know how to drive.
You don't understand the controls.
But you got to get in.
you start, you might figure out the gas pedal, and it takes off and crashes into a tree.
But you have to learn how to drive.
You know, I think using all these coding tools, it just isn't, doesn't just happen naturally.
It's learning how to drive.
I think we might be in the era of the F1 driver right now, which is like the really good people
can drive these systems a lot harder, a lot faster than the people who are just picking
them up.
If you've never used an genetic coding tool, there's a vast difference between someone who's like
really expert in them, knows where they break, can pay attention.
that versus someone who's just picking out for the first time.
One analogy I've heard is we used to talk about the 10x engineer.
You remember?
Yeah.
This used to be a debate for a very long time.
Is it or is it not?
But now what I'm hearing is the 100x engineer.
Yeah.
And so you're saying that you are seeing some folks who maybe let's not use the 100x engineer,
but like this like F1 driver who was just really good at it.
Like how would you describe a person who we've seen this?
Is it just like rock salt engineering basics and they picked up.
they lean into using these tools or what do they like?
Yeah, yeah.
I mean, there is quite a bit of a, it feels like, you know, it's directional that if they're good
software engineering before.
I mean, I sometimes think it's like, you know, everyone's in this kind of spectrum of
capability of software engineering.
And this is just like, you know, taking that line and spread it out.
And it's not quite true, you know, I think it's helped some people more than others,
but, you know, it feels like it's just stretched it out.
So, you know, your ability before is now amplified.
I'm always interested to learn that, for example,
Boris Churny, Tebow at Open AI, they both have been really, really good software engineers.
Boris wrote one of the first TypeScript books, the first TypeScript book for O'Reilly.
He built some massive system, same with Tebow, who built it.
And now you're kind of seeing, oh, these people who are building all these tools and innovated.
And like, yeah, they don't really good before.
Yeah, yeah, yeah.
Well, I mean, this is a, I mean, I have two other analogies to give you about, like,
what it feels like an AI.
You've undoubtedly heard about the term paraprogramming.
Yeah, pair programming.
Yeah, parprogramming.
You know, and the idea behind paraps programming is it's good to just like, you know, have one keyboard, one monitor and two engineers at it, one of the keyboard and the other one sitting beside them, kind of like looking over the shoulder and giving guidance.
And I think there's an aspect of that feeling where I actually got chat to do a little image of this where, you know, it's like the Android is at the computer typing and you're just there giving instructions.
But that was maybe the way it felt like a year ago.
I think it feels a little bit different now.
The one I've just started recently saying is that I feel like the domain experts, the people who are, who.
were really strong before are now massively amplified. Have you seen all this mathematical stuff
coming out like the crazy proofs? Well, I don't understand it, but I don't understand them either.
Yeah. Yeah. So I told you I wasn't very good at math. You know, it's not complete. I just was like,
I stopped in the freshman year of college, right? Yeah. But, you know, I kind of like watch along
with these advancements in, you know, Terence Tao. He's like probably the most famous living
mathematician, you know, super genius. He actually posted this session. There's this recent breakthrough
or it's called the Jacobian conjecture.
I don't even know what it meant.
But he posted this chat GPT session
where he's interacting.
I think it was chat GPT.
And you can see him interacting
with this intelligence.
And it was crazy because he's talking to it
as a peer colleague.
It's responding.
And it honestly looks like,
I'd encourage everyone to go look this up.
It looks like, you know,
almost a foreign language.
It's like his domain expertise
is getting amplified
by the system he's interacting with.
And you can see how he's like
kind of learning and exploring ideas
just really rapidly.
Now, there's all this controversy
about AI mathematics,
but I feel like the analogy
that comes to mind to me, though,
is the domain experts,
they're a little bit like sorcerers
in their particular domain.
You know,
you have the Earth sorcerers,
the Earth wizards,
the water ones,
whatnot.
And if you know the magic incantations,
the right words to say the right order,
you actually get something kind of magical.
And if you don't,
you just get sparkable stuff
that doesn't have anything there behind it.
You know,
you would probably admit,
you don't know much about distributed databases.
If you're going to ask,
Fable or Astra, build me a distributed database like Cocker, TB.
You will get something out, but it'll kind of be ultimately hollow inside.
If you're an expert in databases, you ask to build a distributed database,
and you can point out all the various things you have to know about a distributed database,
here's what you have to worry about the storage layer, the networking layer,
here's the various data structures, runtime inside.
You can actually get something quite magical very, very rapidly out of it.
So what would your advice before, advice before engineers with, like, mid-level to senior
level who, you know, who have been
figured out how to do coding
to become
strong engineers in this, I guess,
age of AI. Well, the first off
is you have the most amazing tutor
kind of readily at hand.
And I mean, one of the
things that I would always do, you know,
throughout my career, and now I've kind of
stopped doing it, but it's like the reason why
it's going to become obvious, which is like, I would
always look at other people's code.
So I was at, you know, Google
early on. You probably heard of Jeff
Jeff Dean was amazing coder.
His colleague, Sanjay Gamowat,
also just an incredible coder.
And, you know, I would be looking at their
poll requests. I'd be looking at their changes.
It wasn't called poll requests at Google. There's a different name
for it. But I'd be looking at my... C.L.
Right? Yeah. CL.
This is P4. It's a different version control system.
But I'd be looking at their changes. Be like, how'd they do what they did, right?
I'd be looking at their code. Oh, my goodness.
Sanjay's code is really always very elegant.
You know, Jeff's his high performance. How's he doing that?
You know, like, how's he going about it?
And it's almost just like you acquire via osmosis.
But now what you can do is not just acquire via osmosis, but you can literally like, I mean,
I would be encouraging if you're a junior engineer and you know there's a senior engineer nearby,
like you could ask them how they're doing what they're doing,
but you could just ask the AI to dissect what they've done and explain it to you
and explain to you at various different levels.
Like, how does this code work?
What is it doing?
Give me a diagram.
Explain it to me like I'm five.
Explain it to me like I'm 10.
Explain it to me in French, whatever like you want.
Yeah.
Like, I mean, fundamentally, to some degree, AI is a translation tool.
They're translated from whatever, you know, kind of language or understanding it's in,
and keep on interrogating it until it gets it, you know, increases your understanding.
You were saying how domain experts are very much amplified.
I guess one strategy as a software engineer is like obviously become a great software engineer
and use it as a tool.
You can get a lot faster.
You can get at distributed systems.
Like, I'm not a distributed databases expert at all.
But I use AI to explain a few.
things for me to understand upfront, which was very helpful. And it would have taken me a lot
longer time beforehand. So I can use this. But I wonder if there's another part of like as a software
engineer, you can use it to become more of a domain expert wherever you're working. If it's a
payments company, I mean, use it to learn about payments as well. So you can help the business,
you can help your team. And honestly, you'll just learn more, right? Yeah, yeah. I think, you know,
I would encourage everyone. You have to have a little curiosity, right? Don't, don't be bound in by,
you know, the area you're working on, explore outside of it, you know.
I was working on Gmail, but I was fascinated about how, like, the main Google search engine
worked.
I was fascinated by, like, how the internals of Big Table worked, even though I wasn't directly
working on a Big Table, just explore in.
Look at those things.
And now it's so much easier because you have this super advanced patient intelligence there
to explain to it.
Like, why do you think it was done this way?
And then, like, I mean, once you become an expert, you can be integrating.
I see it's done that way.
Once we change this, would this be helpful?
And, you know, that's where you, you know, that's where you?
you go from just kind of learning to actually contributing back.
I think everyone has to have the personal agency to do this.
You know, if you're just sitting there waiting for someone to educate you on how to do this,
it's going to be really hard right now because anyone who's coming in explaining how to use
AI or explaining how to be a better software engineer, they're going to be out of date, right?
You just got to get in there, be using this tools all the time yourself.
And using it like, use it to learn.
I feel like I've learned more in the past probably even year than the previous five years combined.
It's weird.
Given your trajectory,
given the environment you were working, right?
I mean, like,
everybody's been in the industry for a while.
Like,
I'm definitely a better coder.
I was a better coder 10 years ago than when I first got in the industry.
It's like I could look back every decade and realize, like,
I got a lot better.
And I feel like I just got a lot better over this past year.
And this was also one of the reasons I was really excited to talk to
because when we started to just exchange your messages,
the first thing you wrote to me and I asked to like,
hey, you know, how are things going?
You said, like, you wouldn't believe.
leave, but my coding output is insane and it's high quality and its database quality level.
And those were the things I don't really usually see it.
I usually see, okay, I'm not producing more code, but it's slop.
But again, like, to me, this is a bit of an inspiration.
Like, look, like, you can use these tools to just, like, amplify yourself as a software engineer.
Like, you are one example, right?
Yeah, yeah.
Yeah, yeah.
No, I'm not the only one in Cockrookrish Labs.
We have other people doing this as well.
I find it very exciting.
You know, it's a little bit exhausting right now, but it's very exciting.
Like, I got into software engineering because I like building stuff.
I can build stuff faster.
You know, the stuff you might have had to compromise in the past, and you can take away
some of those compromises.
I mean, you see this in the U.S. of software coming out.
I think the U.S. is a lot higher.
You see all the fancy web animations and whatnot.
That's only just like the surface level.
It just extends way, way below that.
Peter, this was awesome.
Thanks for coming on a podcast.
Yeah, this is wonderful.
Thanks for having me.
One reason I was excited to talk to Peter is because he's been a very high profile and
productive engineer pre-AI, building some of the most resilient distributed systems in production.
Cockroach DB is known for its resilience and how even if several nodes are destroyed, the database
still operates with their data loss. Basically, it's as hard to get rid of as cockroaches are,
hence the name. One interesting part of our conversation was how Peter built more efficient data
structures than the standard coding libraries had, thanks to him and colleagues paying attention
to parts of the library that seemed slow. He did it for the C++-STL map and then go for the
switch table implementation. I found,
both stories a good reminder that you can improve the existing library or even the language,
especially if you measure which parts feel slow. Another part of his conversation that I liked
was how Peter came a bit of a full circle. He used to write 100,000 lines of code per year,
being a very productive engineer and CTO. He didn't stop writing code aiming to coach engineers
between 2022 and 2024. And then he started to code again because with AI tools, he wanted to coach
his engineers better, but it's hard to do if you don't use the tools yourself. And now we've
finds himself being extremely productive, and this time the team around him is productive as well.
And we're not talking about vibe-coded software, but database-worthy, high-quality code-generated
and committed to production. Peter is convinced that AI amplifies existing expertise,
and this is one reason why he probably learned more this last year, building with AI,
than the previous five years combined. And I find it a valuable reminder that learning
and building deep expertise in software engineering, this is very valuable. And as closing,
I appreciate it that Peter said that not only is he excited, but he's also exhausted.
There's a lot to learn, but it's tiring, and neither him nor anyone I know is immune to this.
So if you're also exhausted with all of the things going on with AI, know that you're not alone.
Check the show on us for more of the pragmatic engineering deep dives on Google's engineering culture and on distributed systems.
If you liked this episode, please make sure you're subscribing your podcast player and a special thank you if you leave a rating.
Thanks, and I'll see you in the next one.
