The Peterman Pod - Dropbox’s Former Most Senior Eng: Building Great Systems and Advice for the AI Era | James Cowling

Episode Date: May 25, 2026

James Cowling is the CTO at Convex and was previously the most senior engineer at Dropbox. We discussed technical details of his past projects, simplicity vs complexity, and career advice given where ...AI is today.• My ergonomic keyboard project I mentioned, you can follow along here: https://read.compose.llc/Podcast links:• YouTube: https://youtu.be/3XkmNSuHFmY• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835• Transcript: https://www.developing.dev/p/dropboxs-former-most-senior-eng-buildingThank you to this episode's sponsor for supporting my work:• WorkOS: makes your app Enterprise Ready with easy to use APIs to add SSO, SCIM, RBAC, and more in just a few lines of code, check them out at https://workos.com/Timestamps:00:00:00 Intro00:00:53 Systems work during his PhD00:13:05 Dropbox technical deep dive00:21:57 Why Dropbox migrated from AWS00:36:40 How to do massive migrations00:44:31 Simplicity vs complexity in promos00:49:23 What technical teams should be focused on01:00:25 Doing the right thing vs promo hypothetical01:08:13 Why he dipped into management sometimes01:11:36 Why you should not lead by example01:23:23 How to mentor Senior Staff engineers01:27:30 Career advice for the AI era01:37:21 Why he started his own company01:46:05 The most technically challenging work of his career01:48:10 How he got involved in Silicon Valley01:52:16 Career regrets01:55:54 Top technical book recommendation01:56:36 Younger self and permanent underclass adviceWhere to find James:• LinkedIn: https://www.linkedin.com/in/jcowling/• Twitter/X: https://x.com/jamesacowling• His company: https://www.convex.dev/Where to find Ryan:• Newsletter: https://www.developing.dev/• X/Twitter: https://x.com/ryanlpeterman• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/• Threads: https://www.threads.com/@ryanlpeterman• Instagram: https://www.instagram.com/ryanlpeterman• TikTok: https://www.tiktok.com/@ryanlpetermanReferenced in this episode:• His PhD Thesis: https://www.usenix.org/system/files/conference/atc12/atc12-final118.pdf• Masters paper: https://www.cs.princeton.edu/courses/archive/fall19/cos418/papers/vr-revisited.pdf• Papercuts writing he mentioned: https://medium.com/@jamesacowling/embracing-papercuts-e6390055dfc4• "Don't lead by example": https://medium.com/@jamesacowling/dont-lead-by-example-4f86b1174e64• His writing about orienting teams around missions: https://medium.com/@jamesacowling/your-system-is-not-a-sports-team-e17f9eb16b94

Transcript
Discussion (0)
Starting point is 00:00:00 The best engineering comes from a deep understanding of why. This is James Cowling, formerly the most senior engineer at Dropbox and now CTO at Convex, and we dived into all the technical work of his career. It's not about getting a system to work. It's what do you do when it doesn't work? Simple systems are way harder to design than complex systems. He also had some interesting takes on technical leadership. Really, a team should be oriented around what problem do they solve.
Starting point is 00:00:29 They should not care about the system that survives. I've had many friends whose promotions were rejected because their work wasn't complex enough. Yeah, I mean, it almost angers me. I just like it so much. Is there any career advice that majorly changed in the last five years? Someone will argue back, no, what you can do is get really good at using Claude. Guess what? It's not very hard to use Claude. Here's the full episode.
Starting point is 00:00:59 Well, first off, your PhD thesis was huge. It's 156 pages. I didn't know I thought papers, you know, maybe 10 pages or something like that. It's its own book. It's very similar to Spanner, ultimately. And it was, what's funny is I did that work. It came out a little bit before Spanner,
Starting point is 00:01:19 so I got a citation in the Spanner paper. But then Spanner came out and then everyone kind of, everyone forgot about my paper. You know about the research? Like, if you could kind of describe the problem, that granola was solving and how it solved it, those types of things. We can talk about that. Yeah, absolutely.
Starting point is 00:01:39 So, I mean, my interest in my career has been two parallel threads. One has been just abstractions in general, the idea about how to build simple models for complex problems. And I find that it's very, very difficult and very intellectually stimulating, like design exercise, how to design APIs, basically. And the other has been large-scale transactions. systems. Big fan of transactions.
Starting point is 00:02:06 And transaction, you know, is do a bunch of things at once. And, and again, I think transactions are just one of the most incredible abstractions we've invented because it allows us to manage probably the most difficult problem in computer science, which is concurrency. And so granola is a, you know, that was a long time ago in my life now. But it was, it was an algorithm for how to do distributed transaction quarterly. nation. And in particular, transactions that I described at the time as one-shot transactions. I'm not actually sure if that was a standard term at the time or whether I made it up.
Starting point is 00:02:43 But a one-shot transaction, meaning you kind of send some code to the server and say, run this function, all these reads, all these rights on two, three, whatever nodes at once, and how to make sure this commits atomically across multiple shards in a distributed system. But yeah, that was my PhD thesis and my master's thesis was on a Byzantine fault tolerance. So a consensus protocol in the presence of malicious nodes. So how to achieve agreement in state across multiple parties when there's malicious entities. And I guess the most well-known work I did as part of grad school was work on this paper called View Stamp Replication Revisited. and that was a paper on
Starting point is 00:03:29 redefining a protocol called Vuesnep Replication, which was a predated Paxos a little bit. Very similar algorithm. Paxos, Raft, VR, they're all basic virtual synchrony, all basically the same thing. And that was a paper I wrote at the time, which ended up being, I guess, influential to a few really great companies like Tiger Beetle.
Starting point is 00:03:53 You mentioned transactions, and I saw on the paper, there's this idea of an independent transaction. And if I'm understanding correctly, a lot of the efficiency in a distributed system is lost by needing to get consensus and to vote for consensus. And I saw that in your research, you found out a way to avoid needing to vote to reach consensus. And how did you do that? Yes.
Starting point is 00:04:20 I mean, a lot of people, I think, think about performance maybe from the wrong angle. because I think of performance as a factor of just of this raw horsepower, you know, how fast are your disks, how fast is your network, etc. And no one has particularly faster discs or memory than anybody else. What really matters to performance in a large-scale system is eliminating points of coordination. So it's how to allow systems to progress without having contention between parties, you know, where like you're basically reducing parallel throughput to serial throughput across large numbers of transactions. And so in Granola, we had this idea called a independent transaction. Again, that was the terminology at the time, where there were, you know, two entities are basically processing pure functions.
Starting point is 00:05:11 So they have independent state and they will both come to the same conclusion as a result. And all they need to do then is serialize those transactions. So they have to decide if it was to happen atomically across multiple. nodes, what time stamp should it get? And in Grinola, the paper was mostly about how to exchange these timestamps very efficiently. So how to how to have multiple parties each propose a timestamp and then choose, you know, the maximum of these, basically. So you could safely serialize the transaction. And there's a lot more complexity that goes into this. But really the focus there was about how to maximize throughput in a large distributive system without resulting in the alternative,
Starting point is 00:05:51 which is two-phase commit and two-phase locking. And two-phase commit with two-phase locking is basically the standard approach for when you want to have multiple nodes agreeing on the same thing, where they both agree to lock their state, not process any other data, and then commit a transaction. Two-phase commit can be quite, one,
Starting point is 00:06:13 it can be low performance because you're blocking basically the systems for the duration of the transaction. And it can be high risk, too, because you are taking a dependency on another node. You basically blocked waiting for another node to return. And so that was what granola was. Now, I've kind of, you know, it's funny you asked about granola because I forgot about it.
Starting point is 00:06:33 You know, that's so far in my history, I kind of forgot about that work. And I even forgot the phrase independent transactions, to be honest. So it's great to hear you bring it up again. But I guess all this stuff does feed into all the work you do later on in interesting ways. When I think about distributed systems, I think about you're out a big company and you got a bunch of machines. But when you were at MIT building this thing, how did you build and test it? Did you, were there spare machines that the college had or was this in the cloud? Yeah.
Starting point is 00:07:06 This was a long time ago now. I'm showing my age. This was pretty early in the days of AWS. And so we had a rack of servers in our office that we could use. and I was fortunate enough to be at MIT where we could afford Iraq. And there was also a service called Planet Lab. And Planet Lab was like a big communal set of nodes that academics could use to run tests on. But Planet Lab was a communal system.
Starting point is 00:07:34 And so it was continually having problems. I mean, just a free-for-all. And so for the longest time, I would just sleep next to my desk and wake up whenever Planet Lab was free. because, you know, you need to run some benchmarks. These days you just spin something up on AWS, but then I would, yeah, I would sleep next to my desk, and I would wake up in the middle of the night or random times and check Planet Lab status
Starting point is 00:07:56 and then kick off a benchmarking job. Now, this was right at the cusp at when, I think, in some respects, Google ruined systems research. And I say that with affection and respect to Google. Because before then, most of the research was coming out of, academia. And a lot of distributed systems research was done on smaller scales, and a lot of the
Starting point is 00:08:19 value prop was the ideas. So, like, hey, I wrote a paper, here's an interesting new idea, and yeah, it's all in theory. And so the value that came out of it was the idea. I ran about this time you started seeing papers out of Google, Amazon, and these various companies, where it wasn't necessarily a paper about an idea. It was a paper about a system. Hey, I built this giant system. It has whole bunch of features, some of which are interesting, some of which are not, and by the way, it powers Gmail or whatever. And that kicked off a pretty interesting transition because at that point then, you know, program committees reviewing papers started to expect to see realistic benchmarks. They're frankly grad students that are not able to, at least at the time, produce. I mean,
Starting point is 00:09:08 I wasn't running Gmail on my system. And I think it did in some obscure intellectual ideas, right? Because I'm more for like papers about systems, but I think there's also value in a paper about an idea. It's not we built this big thing and it works and up to you to figure out what's interesting about it, but instead here's a new thing. There's a paper, you know, an old paper just called leases, right? When someone invented the idea of a time-based lock and it's just a paper called leases. And though that was a cool era of systems research. I think a lot of it now has shifted towards industrial research, which people are building stuff for practical purposes, whereas I think it's hard for academia to compete on pragmatism. And academia is really a great place to do impractical work.
Starting point is 00:09:59 When you look back on you getting the PhD in academia, being enabled to completely explore an idea versus going into industry, but maybe doing, you know, maybe you worked on Spanner at Google. or something like that. You know, when you look back, which path do you think would be better and why? I think that's a bit of a misconception a lot of folks have about PhD programs. I think a lot of students are like high achieving students in college and they think that the PhD program
Starting point is 00:10:27 is like college plus plus. It's like, hey, I like learning things. And so I'm gonna go do a PhD, but it's not. PhD is training to be a researcher. And for most people, they shouldn't do that. If someone wants to be a professional software engineer, they probably shouldn't spend their time trained to be a researcher. But with a really important caveat,
Starting point is 00:10:49 I think there's a really interesting and challenging developmental experience you go through in the PhD program at a top university at least, which is you get to a certain point in time when you have a problem you're facing that no one else in the world knows the answer to. You have a problem that you are the world expert on. And you can't, you know, you can chat about it with folks, but you can't ask your advisor because your advisor doesn't, know either, right? So you face with these difficult challenges that there's, you can't read a book,
Starting point is 00:11:17 you can't, you know, you can't ask chat GPT, right? And so you're forced to go through a quite difficult and frankly emotionally challenging experience of being unsure, being uncertain and learning to think for yourself. I think that is really valuable for all engineers. I think that was really, really valuable in my career. I do see a lot of folks early in their career, think that all knowledge comes from reading or, you know, absorbing it from someone else or that there's a right way to do things. But I think being in a PhD program or being faced with really demanding open questions does train your mind to be comfortable with that discomfort. I mean, look, I feel I feel a little bit lucky that I went through grad school without LLM's existing.
Starting point is 00:12:11 because I had to deal with this uncertainty. There was no crutch to help me out, which I think was very valuable. So, I mean, a lot of, you know, if you want to be a researcher, go do a PhD. If you want to build stuff, go build stuff, you know. And, you know, since leaving academia, I've done a lot of work that I guess one could call innovative, like it advanced the industry in certain ways and did novel stuff. But our goal was never research. Our goal was just to solve problems.
Starting point is 00:12:48 And I think that's the difference. I mean, in academia, your goal is to advance knowledge. In industry, your goal is to solve problems. And I gravitate more towards just solving problems. And I find that a more comfortable environment within which to work. Going to your time in industry, you know, Dropbox, you became the most senior engineer at the company. And looking through all the projects that you did, I had a series of just technical curiosities.
Starting point is 00:13:19 I saw this idea early in your career about multi-homing, and I wasn't familiar with that concept. What is multi-homing? What's the problem that solves? Yeah, I mean, multi-homing is the ability to have data in two locations, two homes. And multi-homing can be valuable for a variety of reasons. One is called, you know, primary secondary or, you know, lazily replicated multi-homing, whereby all rights commit authoritatively in one region and get replicated to a secondary region. And so this is normally used for, you know, business continuity.
Starting point is 00:13:55 You know, one thing we did at Dropbox, we made sure that if the entire West Coast blew up, you know, which hopefully wouldn't happen, Dropbox would keep running because there would be enough, the data will be replicated in other regions. but with a window of vulnerability, with a window of time where there may be some data lost. So that is, you know, that's kind of primary, secondary replication or multi-homing. There's something else called active, active multi-homing, where there is truly an authoritative copy of the data in multiple locations. So basically, when you, for example, write to a system,
Starting point is 00:14:28 you don't externalize that right as having succeeded until it has landed in all the regions. And so, for example, in the storage system at Dropbox, the block storage system, we truly had a multi-region replicator system where we could take down an entire region, say take down a region in Ashburn, Virginia, where everyone's data centers are. And there'd be zero downtime for the company, and the data was still safe in multiple regions. I think multi-homing is something a lot of engineers aspire to work on, because it seems like the real. right thing to do. But I think the reality is for most companies, it is not. For most companies
Starting point is 00:15:09 not because there's very high cost, there's very high latency costs for this. Very high, because the speed of light is just not getting faster. The speed of light is fixed. And if you have to, you know, synchronously write data across multiple regions in the United States, you've going to have, say, 60 milliseconds in the commit path of your protocol, which for most applications is not tenable. So there is a lot of desire. A lot of engineering teams reach out for kind of advice on how to adopt active, active multi-homing for their company. I would normally say, don't do it. I would normally say, frankly, if US East is down, if Amazon is down that day, that's okay.
Starting point is 00:15:49 If Amazon's down that day, your company will be down. That's a shame, right? But by avoiding that complexity, you're going to be able to move much fast and build a much better product. And so in my mind, systems is all about tradeoffs and making the right ones. I'll generally recommend most people do not make the tradeoff to have partition tolerance or multi-region availability, even though for a company like Dropbox, yes, that does make sense. When your job is storing data and you have several exabytes of it and hundreds of millions of customers, I think that's when it starts to make sense.
Starting point is 00:16:26 So it's just another term for, I guess, data replication. and having it available in other regions. Okay. I imagine also, I mean, the cost of storage is a concern. How many replicas would you keep for something? Like, let's just say I stored something in Dropbox. It's my document. Is that on the order of one or two, or is there multiple?
Starting point is 00:16:47 No, it's on the order of many, many. And so if you were to store a file in Dropbox, I modeled the storage to be, as we advertise at least 12 nines of durability, Internally, the models look around 24-9s of durability. And so that means the data is secure with 99.999% where there's 24-9s, right? Which means, at least according to the model, you know, the universe will be extinct before any data is lost. And the way that is done is by a combination of what's called erasure coding. So erasure coding is how you take several blocks of data and combine them together with an encoding scheme.
Starting point is 00:17:30 and spread them around in different locations. At that scale, at Dropbox scale, you're taking things into consideration like putting data in different racks, different rows of a data center because they're on different power feeds. You're taking into consideration different eras of hard drives and different manufacturers of drives
Starting point is 00:17:46 because they could have correlated failure patterns, and then replication across regions. So you could be looking at 27 fragments, for example. That doesn't mean you're storing 27 times the data. It's kind of encoded in many regions. But it's really, the replication schemes get really quite sophisticated at that point. And we actually had our own custom encoding matrix we had developed. It's called a Van de Mont matrix where you kind of take a bunch of data and you combine it together to produce outputs.
Starting point is 00:18:18 And we will plug in some variables like how much do disks cost? How much does network bandwidth cost? Because there's a trade-off. You can either store, if you don't want to lose your data, you can store more copies. on more disks, or you can store fewer copies, and any time a disk fails, you re-replicate it really fast. And re-replicating really fast costs network bandwidth. So these kind of variables go into this equation. It ends up being an extremely complex field of endeavor, but it's kind of abstracted away into a part of the system that doesn't leak into anywhere else.
Starting point is 00:18:52 So with erasure and coding, my document in Dropbox is fragmented into a bunch of different chunks of data and loaded potentially from many different machines? Yes, absolutely. It, I mean, the first thought I have is now there's, I might be waiting there and one of the 27 machines is slow and I can't look at the whole doc. So how do you prevent against that? It's actually faster than not replicated. Because if you imagine, I'll pick a simplified example. Imagine, this is not the encoding scheme, Dropbox users, but imagine to
Starting point is 00:19:29 reconstruct a file, you have to read six out of nine fragments. So if you read six out of nine fragments, you can just ask all nine and reconstruct and return the data as soon as you've heard from the first six. It's actually faster and you can construct these encoding matrices. It's actually faster to have a ratio coded data than then not. I see. Okay. So you can like over subscribe, over request, and then you complete on a portion of them being received. Yes. Now, in practice, it was a bit more complex in that.
Starting point is 00:20:05 We'd often have a copy and a single disk to provide fast access. We'd often try to make sure that you could serve your data out of a region close to your home region. So we'd make sure that your data was mostly served with low latency. But if that region had failed, you could reconstruct it from the remaining regions. So there's a lot of, a lot of people, you know, there's a lot of talk about building your own infrastructure, and you can save money by moving off the cloud, almost definitely you can't. Unless you, either you have very small requirements
Starting point is 00:20:39 or very fixed requirements or very, very, very heavy investment. Because if you want to compete with Amazon, if you want to build a more efficient storage system than Amazon, you have to have a supply chain team that's working with Western Digital and Seagate constantly negotiating on prices of disks and buying shipments at certain time and capacity teams and data center teams.
Starting point is 00:21:00 There's a lot of work that goes into optimizing this because ultimately our desire was to use the disks to 90, 95% of their disk size to maximize storage efficiency. To put it this way, I mean, that was, you know, I guess it was a ballpark figure, a billion dollar project. And it was, you know, and I think at the time, as far as I know,
Starting point is 00:21:24 was the largest ever data migration in history. I think at the time. And so extremely large engineering project with very high technical investment. So yes, if you have that scale and you have the engineering team to do it and you're willing to keep innovating, if you're willing to keep optimizing
Starting point is 00:21:44 and investing effort in, then yes, you can do it. But I think the cloud has been an incredible innovation, right? Most people are not experts at this, and most people should not be experts at this. Most people should focus on their applications. When you say that most people shouldn't, my immediate thought was, but one of the big projects you worked on at Dropbox was migrating away from S3. Yes, yes. So, you know, why did Dropbox migrate away from S3?
Starting point is 00:22:12 Yeah, that was a desire at the company for a long time. You know, even before I was there, so I started a Dropbox in 2012. And I spoke to Drew, the Dropbox founder, I think, in 2010 about this project. And so I think there was a desire to control the destiny of the company from a strategic perspective. I mean, at the time, this is before Dropbox kind of reshaped itself as being more about collaboration. You know, at the time it was a file sink and share category. That was that was the market sector, you know. And so, and owning the file system was really valuable to the company.
Starting point is 00:22:47 Ultimately, we saved a huge amount of money. I mean, we, and this is before the public company went public, we really drove, massive cost efficiencies through the project. But it was hard, you know, in a way that I think it'd be very difficult to emulate without a huge investment. And I do think that there is, there is a benefit to an organisation from having hard problems to solve. Because if you have a company with extremely hard technical challenges, you can attract
Starting point is 00:23:17 engineers who like working on those hard technical problems. And when they've solved those problems, they cycle off and work on different parts. parts of the system. So after we all worked, we had such a great team. It was a very small engineering team. And after we kind of shipped the story system reliably, we all went off and, you know, Jamie went and redesigned the sync protocol, the desktop client. And I worked on the, on the file system and the distributed databases. And so, yeah, there's value to a business to have that level of technical investment. But it is, you know, it's like, it's like having a baby. And then you have to, you've got to raise the baby.
Starting point is 00:23:55 You can't just build a system like this and be, that's it, we're done. You own it and you have to keep investing in it. Did S3 do any counter-negotiation before you set out to leave them? They say, like, oh, you know, we'll cut you a deal. I guess I'm allowed to talk about this now. It was a long time ago. Yeah, for the longest time there, I don't think they were particularly aware that this was happening. But, you know, the data center folks talk, and certainly it was noticed that Dropbox was buying up a lot of data center space.
Starting point is 00:24:30 So, yeah, we had, obviously, at the scales that we were at, I mean, we were negotiating very good rates with Amazon. You know, we weren't paying sticker price. We're paying very, very, very good discounted rates. But yeah, at a certain point, they weren't able to meet our cost efficiency. Because when we launched the system, it really was more efficient. an S3. And that's for a variety of reasons. One was that we were using kind of new experimental disks called shingled magnetic recording. We were the first ones, I think, to use these disks at scale. And two, we had a very tight understanding about workloads. So we were able to design
Starting point is 00:25:06 the system specifically optimized for our workloads, whereas S3 has to design the system for everybody. So it got to the point where Amazon would not have been able to offer us a more competitive deal because we had a more efficient system. I wouldn't recommend another company do this right now. But I think at the time, it certainly made sense for us as a company. Can you give an example of a tight understanding of your workload leading to, like something you could do that S3 couldn't? Yeah, absolutely.
Starting point is 00:25:37 So, for example, I know that, well, I'll try not to leak any confidential data, right? But when you upload a file to Dropbox, there is a pattern of access, right? So typically people access the file very quickly shortly afterwards because you're sharing it with someone or maybe Dropbox is processing that file to generate an image preview. And then it decays at a certain rate. And so we understand in general the average block size and we understand also the access pattern. So we could do things like at a certain point we had these two clusters. One was designed for contemporary storage
Starting point is 00:26:20 that was kind of storage inefficient, but access efficient. So it was very cheap to read and write to, but it was inefficient to store. And data would get written to there first, and then in the background, it would get moved in bulk to this colder storage system. And this colder storage system was far more static.
Starting point is 00:26:40 And so it was able to have kind of more efficient algorithms and be written to in bulk. And if it went down for rights, that was no problems because it wasn't in the live path. And so we're able to trade off, again, like systems are all about tradeoffs. So we're able to trade off the live data right path from the long-term read path. That was one of many examples where knowing the size of your data, where it's accessed from, how frequently it's access, how long it takes to delete that data, you can really tune a system to your workload. You know, stuff like looking at even things down to knowing how much power to put in a rack.
Starting point is 00:27:20 You know, you have a rack of hardware. There's a power distribution unit, a PDU at the top of that rack. It has a circuit breaker which can handle a certain number of amps. We would have to figure out, you know, how many amps are required for that rack based on access patterns. And there were times where we got it slightly wrong. There was times where we got maybe a bad batch of hardware and disks were failing to, frequently. And as a result, we were re-replicating the data more regularly. And I was getting, you know, messages from the data center team saying, hey, we're running the racks really hot right now.
Starting point is 00:27:51 You know, so that's a level of optimizations you can make when you get to that scale. But again, that's multi-exabyte scale, you know, million hard drive scale. Yeah, when I was reading about this migration, I saw somewhere in the migration, you initially started with go. And then the racks or something that you're running on the hardware itself was ooming too much or requesting too much memory. And then you migrated to Rust. So initially, the prototype was in Python, if you can believe that. And to be fair, Python is actually pretty efficient for I.O.
Starting point is 00:28:26 I think people give Python a bad rap for I.O. bound workloads. It's pretty good at I.O. But obviously, not great for concurrency, not great for memory management, and very hard to refactor. And it was critical the system was correct. And so we migrated everything to Go, and we built most of the storage system in Go. This is before Go was in general availability, I think. It's before Goa was GA. And we built the system in Go.
Starting point is 00:28:54 Go is a great language for concurrency, a great language for proxies. You know, it's really well designed for like servers that moved out of one place or another place. At a certain point, though, we would have, you know, let's just pick a number. Let's say a million. Let's say we have a million nodes in the system. and every node has some amount of memory, some amount of disk, some amount of sheet metal in the chassis, right? And we would itemize all these things.
Starting point is 00:29:20 You'd have a pie chart of how much money is spent on all the things. And it would have stuff like, yeah, sheet metal and screws and stuff. And so you're trying to optimize the storage. And a big problem for us was the amount of memory these nodes were using, not just the memory they were using, but the arm. predictability of it, with Go having a runtime and, you know, in a story system and out of memory error is pretty bad. Because if a node runs at a memory and restarts, that looks like a disc failure. So that looks a lot like a disc has failed and has to be re-replicated. And so a batch of nodes
Starting point is 00:29:59 ooming can lead to cascading failures throughout the system. Because maybe I remember a time, it was a band. It might have been Dila Sol. I can't know. There was a band released an album on Dropbox. So it was a big spike in load to a few files. So it was very, very high bandwidth. And it caused the machines to oom. So those machines oomed, those disks oomed. And so as a result, no problems.
Starting point is 00:30:30 The system went to try to recover that data from a whole bunch of other replicas. Now all of a sudden you've taken one amount, one fire hose worth of load coming. and you've turned this into seven fire hoses worth of load coming, because now you have to do a more expensive reconstruction operation. So now you've seven extra load. And these are the kind of cyclical kind of behaviors that can lead to something called congestion collapse. Congestion collapses when workload to a system crosses a threshold
Starting point is 00:30:57 where it all kind of collapses. So designing against congestion collapse is really the hardest part, or one of the hardest parts is Magic Pocket, which is the name of the storage system. So ultimately we switched to Rust for the storage nodes themselves. And this again, this was before Rust was in GA, so that was a bit of a risky move. But we rewrote the switch to Rust
Starting point is 00:31:21 coincided with getting rid of the file system entirely on the disks and directly addressing the disk heads. So there was an instruction set, I think it's called ZBC, zone-based block control, something of that. There's an instruction set for accessing disks that we were using. The disc manufacturers gave us the draft specs of these new disks, and we were operating off the draft specs and directly controlling the disks.
Starting point is 00:31:45 And so all that product was tied up together into a product called DiscoTech, it was the Disc Technology Project. And ultimately, if you look at that pie chart, one congestion collapse stopped and reliability improved, but if you looked at the pie chart of where all the money was going, it really shifted to be almost all discs. And what we wanted to do is get that pie chart to be almost all money spent on disks and it adds little money spent on RAM,
Starting point is 00:32:14 on compute, on network, on power, on sheet metal, et cetera. The cascading failure you mentioned, was there, so that happened and then there was a postmortem? I don't think we needed a postmortem, but I think we knew as it was happening. I think when I was getting paid in the middle of the night on these things, you'd be pretty evident.
Starting point is 00:32:33 And so, look, this is a, again, this is an argument in favor of the cloud. And imagine you've spent hundreds of millions of dollars on a storage system and it's out in production and it's on physical hardware. You can't just go and put new memory chips in every one of them. I mean, you can, but you have to pay people to come in and swap them out, you know. And so that's a tricky place. And so as a, and this is what I love about industry, you know, because in like, Some of the more challenging moments on Magic Pocket were stuff like in one week, just a weird coincidence, two trucks crashed that were delivering service. So there's two trucks showing up to deliver racks.
Starting point is 00:33:16 And they're both crashed. I don't know, you know, the drivers were okay. So we lost capacity for two weeks. So we lost capacity for more than two weeks for probably six weeks. What do you do? What happens, you know, is the equivalent of your disc filling up on your lap? except it's a million disks and you can't tell the customers to go away. You can't delete their files.
Starting point is 00:33:39 So a lot of tricky capacity work. And ultimately, what that led to was trying to build in all these protections against the unknowns, making sure we had the right amount of buffer planned out for anything bad that could happen, making sure we design the system so they couldn't be congestion collapse, so there wouldn't be memory spikes. And at one point, there's a process called, I think it's called FMEA, which is like a threat modeling process where you had a big spreadsheet and you kind of write down every bad thing that could possibly happen
Starting point is 00:34:17 and then how bad it would be if it happened. Existential risk, does someone die? You know, if there's a fire in the data center. All these kind of things get put into the spreadsheet. And you kind of do a bit of an almost pre-mortem kind of work to figure out all the potential failure modes and then design around them. And I love that work. I really, I don't know, I do like the firefight.
Starting point is 00:34:42 I like the, I can't say I like getting paged because I've spent my whole life on call. But I do like that rubber hits the road stuff. I like that, wow, there's congestion collapse and there's no one that can help you. And so you've got to think through this problem. When I worked on infrastructure, Instagram, we had this concept of defecon knobs, which are basically these configs that you could flip that would gracefully degrade your system, where you can still operate it, but maybe in the case of Dropbox, you store less replicas. So you take on a temporary increase of risk in losing data because you need to. Did you have something like that and did you flip it in that type of case? Never when it came to durability.
Starting point is 00:35:31 So, I mean, we just had just absolute zero non-negotiable, like there was no room for negotiation on users' durability. And so we had those knobs for background processes for CPU and memory, for example. So if there was like a spike in load, you could turn off background processes and turn off the test load, for example, on the system. Eventually, we built this system called trampoline. And what trampoline did was, which actually was that it saved us a ton of money, right? When if we ever got too close to the threshold, we would just start writing data to S3 because this is right there, right?
Starting point is 00:36:11 So you can run your capacity way closer to the edge if you're willing under worst case scenarios, just to dump 30 petabytes on S3 and then move it back when it's done. And now that didn't happen very often. We would do it to test it. We would do that just to make sure the system worked. But being able to have an escape hatch for worst-case scenario was really nice. I see. So S3 is kind of, it's like elastic storage. Exactly.
Starting point is 00:36:40 When I looked at this project, it's such a massive migration. And my first thought is, how do you coordinate this whole project without breaking the system as it's running? Yeah. What are your thoughts on doing such a large migration? without breaking things. Yeah. Now, I mean, in terms of the engineers that built the initial version of the system,
Starting point is 00:37:04 it was a handful, three, four, five, six engineers. It wasn't a team of a thousand, you know. It was a very, very small team. And so how do you build such a large system with a small number of people? And then how do you do a high-risk migration? One of these, one thing you do is you keep things simple. You try so, so, so hard to,
Starting point is 00:37:26 build very cleanly abstracted simple systems that their failure modes are very understandable and so they and they're decoupled so they don't have congestion collapse etc so focusing on simplicity it gets tricky by the way with people doing agentic development they're not the best at building simple systems simplicity is still the domain of human beings for now but a big focus on simplicity the other was you know this very thick layer of valid validation checks during this migration. And in fact, we had, when we did the migration off of S3, we had this something called the Dark Launch where we would be moving data off of S3, but we would keep it in both locations.
Starting point is 00:38:11 And we had to demonstrate to, we had this kind of contract with the Dropbox founders. We'd have to kind of keep this system running with no incidents, no downtime, no data loss, whatever, for six months before we would delete any of the data from S3. So we have double-rided. And there was one point in time halfway through this process where there was a bug got through to production. And it didn't, nothing bad happened with the bug.
Starting point is 00:38:38 But it was like a, it was like bugs slipped through our multiple layers of release process, et cetera. And so I went to the VP and I said, hey, a bug made it through to production and we're going to reset the launch clock. And as a result, it's going to launch later and it's going to cost us some amount of money. Let's say double-digit millions, you know? And that were like, great.
Starting point is 00:39:07 Thank you. That's good. I trust you. And that was a cool thing. I mean, Dropbox had a lot of incredible cultural values. But that was like, okay, cool. If you're prioritizing user safety, that's the right thing to do. And no one was mad about that.
Starting point is 00:39:20 It was almost like they were proud almost. You know, it was just like, it was like, yes, you're operating in accordance to the principles of this company. Okay, so it was double writing both systems for a while. Some subset of data, yeah. Okay, and then you switched over reads. Once we were sure the system was durable, and we had all as validators running in production 24-7, there was just migrating as fast as we could. I think at some point we got up to 700 gigabits per second, 764 gigabits per second of peering.
Starting point is 00:39:53 bandwidth between Amazon servers and ours. Certainly someone on the network team over there noticed that there was that much data moving out. At one point, I got a slightly nasty email from someone saying it was super weird. I hadn't the phrase super weird. It's like it's super weird that you're doing so many reads and not that many rights.
Starting point is 00:40:15 And I didn't respond to that email. But to be fair, like Amazon were a great partner for Dropbox. So Dropbox still uses AWS. They always, there was no concern that they would do the wrong thing by us as a company. I mean, like, we only had excellent experience with AWS. But yeah, it was a moment for us.
Starting point is 00:40:36 You know, it was a, it probably strained the relationship somewhat. You mentioned simplicity and intuition-wise, it makes sense. Do you have a concrete example, though? Yeah, I've got a concrete example for you. And maybe it shows the difference between, academia and industry. So the storage system is a giant distributive system with a file stored in various locations. And so you need a mapping from the file to where it lives on these disks. And so all we did was had a cluster of a thousand my SQL nodes, right? Big giant database,
Starting point is 00:41:16 and it was indexed by the block ID, and it said this block is on these disks. And that's a pretty like simple. It's not sophisticated, you know? And it, and every time we'd hire someone out of academia or maybe from other companies, they say, oh, this is not very sophisticated, because like you could use a Patricia try or you could use a distributed hash table. And that would map a block to a set of, of locations. And I think that's optimizing for the wrong thing. because the really nice thing about dumping a list of files and the locations in a giant database, it is written in one location. If I want to validate what happened, if I want to check all the data is where it's meant to be,
Starting point is 00:42:03 I just walk over the table and check. And we did. We had services constantly walking over the table and checking. Whereas if it was a distributed hash table or some giant complex data structure, it's very hard to validate. So, I mean, designing for validation is very important. Designing for understanding is very important. It's not about getting a system to work. It's what do you do when it doesn't work, right?
Starting point is 00:42:25 And so having a very simple boundary. That's a very basic example. You know, there's more sophisticated examples that take more time to explain. But like something like that is a lot of engineers will, you know, a lot of engineers will want to do interesting work, will want to advance in their career. They want to be seen as an intellectual problem solver.
Starting point is 00:42:49 And so the tendency can be to design complex systems. And my argument is always that like simple systems are way harder to design than complex systems. Like simplicity is so hard. And I think to like maybe the untrained eye, a simple system can seem like obvious. And the best compliment you could ever get about anything you design is people say like, oh, isn't that the obvious way of doing it? It's like the same as convex. You know, people say, oh, isn't that.
Starting point is 00:43:19 what's that? That's just like the obvious way of structure. I'm like, great. Because it wasn't obvious when we did it. No one else was doing it, right? Everyone thought we were idiots. If after the fact, people think it's obvious, then you really nailed it. I think, but I think that's a, it requires an understanding that simplicity is the hardest thing in systems. And because simplicity is, is scalable. And I don't, yes, simplicity is scalable in terms of numbers of queries per second, right? But what I really mean about scalability is you can take a simple system and have it run for five years and that people work in it for five years. And I have all sorts of features added to it and have requirements changed because the company realized the product didn't work the way it wanted to work and it wants to change things. And it still stands the test of time.
Starting point is 00:44:08 Whereas the complex over-optimized system will not. And I think that's the tough thing about distributed systems design, especially LLM, organizations. the distributor systems design is just because something works doesn't mean it's maintainable over a long period of time, doesn't mean it's understandable, doesn't mean it's clearly architect and abstracted. That stuff's really very hard. Absolutely. And I agree with you. I think it's the long-term beneficial thing to do. One unusual thing, though, in the industry that I've seen is the incentive system for engineers is actually, I mean, you mentioned the desire for an engineer to want to be seen that they can do something difficult.
Starting point is 00:44:50 There's that, but there's also the incentive system of promotions. And I've had many friends whose promotions were rejected because their work wasn't complex enough. And so that kind of forces, it's a force is complexity, which is kind of unusual. I wanted to know what you thought about that. Yeah. I mean, it almost angers me.
Starting point is 00:45:12 I just like it so much. Partly why I started my own company, you know. I think the ideal for anyone is to be doing work where you're being appreciated for solving the problem. Like, you know, and this is, if we get philosophical, this is what it was like going back to the farming days, right? There was no incentive to make it really complicated to milk a cow because the goal is to like milk the cow and then the reward is you got milk, right? And I think it sounds so silly, but at a startup, that's the same thing. The startup the goal is to build the system, have it work, have the users like it, have it grow, and everyone gets rewarded and celebrated for solving the problem.
Starting point is 00:45:57 It gets hard to scale that. So at large companies, you end up with so many layers of organization that people end up building alternative incentive structures. It's like I'm so far away from whatever the hell we're trying to do over here, that my goal now is to get all green checkmarks on my OKs. plan. But who cares about your okay.R plan unless it solves the problem. And so the thing that really drives me, really drives me insane is when people cry to chase artificial goals, right? And I understand that if you're in a company with like this, then you may have no choice in the matter. But what I
Starting point is 00:46:39 want to tell people, there is a better way. And that, you know, and that better way may not be available to you may not have job opportunities near where you are, for example. But if you do have the ability to go work at a company where you are being appreciative for problem solving, that will make you so much better as an engineer. And I see this when I interview people, you know, and if I, if I, you know, do a, I mean, everyone knows Google has tremendous engineering and tremendous engineers.
Starting point is 00:47:13 A lot of folks there, though, I'm not, that interested in hiring because, you know, if I'll do a deep dive with them and they'll say they build a system and I'll say, well, why did you build it? And they're like, I don't know, the VP told me to. And like, oh, how's the system used? And they're like, I don't really, I think ads uses it. I'm not sure. And this is a caricature. But I think it's very, very hard to do good engineering in that environment. You can do competent engineering. But the best engineering comes from a deep understanding of why. And this is something we just drill into, you know, the team here at convex or, you know,
Starting point is 00:47:49 the team embodies so strongly at convex is like everything exists for the why. Do, like, don't build a fancy low balancer unless it's not needed. Turns out we do need a fancy low balancer. We're building it right now. But, but, you know, it should always start with why are we doing this. What's the point? And I feel for people stuck in environments that are not like this. But you know what?
Starting point is 00:48:11 Like, I don't know. Try to fight the system a little bit. I think I do see a lot of maybe nihilism, a lot of defeatedness sometimes amongst junior engineers. A lot of this cynicism. Like, you know, what does it matter? Like, who cares? It's just a big organization and nothing matters. But I think it does matter.
Starting point is 00:48:32 Like, I've just, if I think of, like, the happiest times in my life, it's been, like, just dedicating myself to a cause. and trying really hard and trying to do the right thing. And I felt good when I went home and not trying to get promoted, just trying to do the right thing. And then assuming I'm going to get promoted, and if not go somewhere else. I know it does sound quaint when I'm saying this,
Starting point is 00:48:56 but I think it's possible to do this, and especially possible if you surround yourself with people like this. And if someone isn't a big company and they're feeling frustrated by politics, look around and see if there's a team of folks who just seem to want, to do the right thing. Just seem to want to do good stuff. I don't think that's selling out. I think that's being true to yourself. That's what real engineering is. Not trying to make a complicated, fancy thing to get promoted. Just build the coolest thing that solves the problem. This really reminds me of something you had written. I thought it was really good writing. And in the writing, there was this
Starting point is 00:49:28 idea of system bias. And you have this quote, your writing. It says, here's some examples. It says, the team is spending six months to improve performance by 10% when it was completely fine to begin with. Or, you know, the team is trying desperately to force their tooling on clients who don't need it. Or, you know, the team is riding their outdated system to the grave, like the captain going down on the Titanic. I've definitely seen examples of all those types of things in industry. And so, yeah, I think it was in the... It was in the context of your writing about what you should orient your team around, not systems, but actually missions.
Starting point is 00:50:16 And maybe that's a way to fight system bias. Yeah, I mean, one of my jobs at Dropbox wasn't the most fun job, but it might have been one of the most impactful jobs, is shutting down projects, you know, looking around and being like, huh, that thing over there that has had 60 people working on it for two years doesn't seem to make a lot of sense to me. And then I, you know, it wasn't like a hostile thing, but I'd go and chat with a team and I'd say, hey, what are you all doing? Do you believe in what you're doing? Does this make sense? And the team in private would say, oh, I don't really know. I don't really. But inertia is so strong. You know, this whole desire to not get in trouble, you know, to just keep doing what you were previously doing is so strong. And talent people can end up doing things that don't make a lot of sense. One of the things I said in that article is it, it's kind of a cheesy story. But when we started building the storage system at Dropbox, the team was called the Magic Pocket Team,
Starting point is 00:51:13 because that was a silly code name for the system we built. And so the teams oriented around building that system. But as soon as we shipped it, I renamed the team to the storage team. And that actually took a bit of work because you had to rename all the email addresses and the channels and the repos and the things. And so it seems like a waste of time. And but my argument to the team was that the responsibility of the storage team is not to advocate for magic pocket the storage system. It's to solve the needs of storage for the organization. Because who else in the company knows more about storage than the storage team, right? And if there was a point in time where S3 was a better idea, it would make sense to move
Starting point is 00:52:02 back. Or maybe there was a different kind of storage system that we meant to use, right? Is the job of the storage team to advocate for moving back, right? And so I've seen this before, you have like a team called the puppet team, you know, for people don't really use puppet that much anymore, but like for, you know, the puppet, the job manager, or puppet versus chef, and then they'd be kind of advocating for their team's thing,
Starting point is 00:52:24 when really a team should be oriented around what problem do they solve. They should not care about the system that survives. Because if you are on the Magic Pocket team and someone says we should move back to S3, that's pretty threatening to your identity and to your career, right? But if you're on the storage team and it turns out it makes sense to move back,
Starting point is 00:52:49 I don't think that's the case, but then that's an exciting new product for you to own. And so I think it seems like such silly management philosophy, but I think it's really, really important to orient a team and an identity around solving a problem and not owning and defending a system. Because you just see this in big companies, it's just inertia is so strong.
Starting point is 00:53:10 Yeah, and you see people doing things they don't believe in because that's just what they do. If inertia is so strong, how did you fight it and close down all those projects? I think I got lucky insofar as I was there pretty early on.
Starting point is 00:53:25 And I worked hard enough, you know, I was putting in probably 16-hour days at the start. I'm not advocating for that, but I was like, I was dedicating my life. to the company. And I think it became pretty obvious to people that I cared. You know, here's a guy over here that really wants to do the right thing and cares about the company. At that point, you build up enough confidence, capital, you know, that you feel comfortable
Starting point is 00:53:49 saying things. You know, I wasn't, at that point, I wasn't afraid for my career. I just, I was afraid for the wrong decisions getting made. So I, so, so I felt, you know, psychologically comfortable making's observations. And then at a certain point, you, um, you know, there's another, you know, a lot of engineers, so at one point I would mentor, you know, a lot of the staff plus engineers at the company. And people would sometimes get grumpy, you know, because they were like, oh, we're not doing the right thing over here or this is inefficient. And every engineer listening to this has a story like this.
Starting point is 00:54:25 They're annoyed about some inefficiency at the company. And my response was generally like, do you think we should solve this problem right now? because if we should, let me know the team to take some engineers off and the product is shut down and I can redirect resources and we can solve this problem right now. And they'd be like, oh, well, we shouldn't shut down
Starting point is 00:54:45 any other stuff. I'm like, cool, well, we just have this many engineers right now. And so if there's anything higher, lower priority, let's stop doing the low priority thing and do the higher priority thing, this thing. And oftentimes the answer was, oh, no, nothing else is lower priority. and then the answer is we just have to accept.
Starting point is 00:55:05 You just have to accept, right? There's no point in being angry or upset that we're not doing the right thing all the time. So I think there's a dimension to you break it, you board it. I don't think, like when I said part of my job was shutting products down, it wasn't going around just causing problems, right? It's ideally like solving problems.
Starting point is 00:55:22 Like, oh, this product is not going the right direction. Let's redirect it and do this alternative thing. And so I think the thing that helped me, I guess, was having a sense of ownership that instead of complaining, I just wanted to go fix problems. But that was part of the culture.
Starting point is 00:55:40 I mean, there was, when I started a Dropbox and I remember being a Dropbox, and Infra was, I don't know, seven, eight, nine people. And, you know, we'd have someone join from Google, for example, and they'd say, well, someone should go build,
Starting point is 00:55:56 I can't do anything without this logging framework. Couldn't possibly do anything with that's logging framework. I'm like, well, we're going okay without it. So, and then like, someone needs to build this thing. And everyone would be like, well, who is someone, right? Because it's just us. It's just us. We build it or we don't build it. And they'd pretty quickly come to understand, oh, wait, it's just us. There's no other idiots out there. We're the idiots, right? So that was, I don't know, I loved, I loved at that time because it was just a time of accountability. Life gets easy.
Starting point is 00:56:30 and harder when you realize that everyone else is not an idiot. When you realize that everyone else is just dealing with their own stuff, right? And so I think, I do not think someone will have good luck going around complaining about stuff and just saying this is a dumb idea and being negative. I think people will have, everyone wants problem solved, though. So if you're someone in an organization who is willing to put their head up and say, you know what, I think this thing over here is a bad idea, but here's a different idea, and I'm willing to own it and put the effort behind it. I think that's a recipe for success.
Starting point is 00:57:11 From that article, you had a great quote, or I guess a great question to think through this. It said, you went to everyone or a bunch of people, and you would say, if we could be spending these resources working on any project at the company right now, would this still be the best use of time? And I feel like it frames exactly what you just described really cleanly. Yeah, I mean, I guess this is like maybe a trick for being a tech leader or a manager is almost nothing's a yes or no question. It's a like it's a prioritization question. Yeah, it's not like should we redesign the database? I don't know, maybe, I guess, is it the most important thing to do right now?
Starting point is 00:57:48 No, cool. Let's not do it. I think it's much easy to have those conversations than to think about it as a yes or no. I often hear, oh, my VP won't let me do blah. Okay, well, it could be that your VP is dumb, but it's probably not, right? It could be your VP has a different set of priorities. Maybe their VP knows that you need to ship these features, and if you don't, then the company's going to be struggling.
Starting point is 00:58:10 I don't know, but to really frame it around prioritization, and the way I think about the career ladder for engineers is, I think some people, maybe it's not common, but think that, like, Be becoming a senior engineer means you get better at programming. But I don't know, I think my program abilities went down from like level four onwards. You know, I think once they got to like level four, that was peak programmer for me. And then I probably went downhill from there. And I got wiser, whatever. But I think the real thing that happened is the scope that I cared about increased.
Starting point is 00:58:46 So at a certain point, you're just thinking about, cool, what matters most at the company for the next five years? and I don't think a junior engineer on their first year of the job should try to do this because you probably don't yet have the wisdom, insight, knowledge to be able to make a good assessment. And I think it probably would be a mistake to be like to try to come up with like redirecting company strategy. But as you grow, I think the real key part of it growing is the IC six, seven, eight engineers are not necessarily the best programmers.
Starting point is 00:59:22 They're just getting better at having broad perspective in decisions within a company. Open AI, Anthropic, Cursor, and Versal all use this product to make their lives better. And the problem it solves is when you're building SaaS or an AI product and you want to sell to other companies, there's all these requirements you need to meet.
Starting point is 00:59:44 There's SSO, there's SCM, there's Rback, there's audit logs. These are all things that take time. to integrate but aren't the main focus of your app. WorkOS is an API layer that lets you meet all of these requirements in just a few lines of code. So let's say you have a new SaaS product and you want to sell to other companies. WorkOS will solve all of these critical feature gaps for you. You can check them out at WorkOS.com to learn more and get started. And I appreciate them for supporting my work and sponsoring this podcast. You mentioned this idea of do the right thing
Starting point is 01:00:18 and then, you know, the byproduct, you also get promoted. And I think that's the dream. You know, do the right thing, get promoted. But in reality, oftentimes people would have to make the trade-off. Imagine a two-by-two matrix of doing the right thing, doing the wrong thing, getting promoted, not getting promoted. Obviously, you know, do the right thing, get promoted, great. Do the wrong thing, don't get promoted, obviously bad. But I'm curious about the other two quadrants.
Starting point is 01:00:46 Which one would you have picked when you were earlier in your career? Let's say I came to you, I said, hey, you can do the wrong thing, but you're going to get promoted, or you can do the right thing, and I guarantee you're not going to get promoted. Absolutely, the second one. So I'm going to say a really tacky thing, right? There are a set of engineers who at a certain point just make infinity money. Like the amount of money they can make, you know, going and whatever, going to work at Anthropic is a huge amount of money, right?
Starting point is 01:01:18 And so there's a certain point where it just doesn't matter anymore. Like the money is whatever, right? But so if you really want to, if you really want to maximize long-term income, I don't think you should try to do that. But if you wanted to maximize long-term income, if you get to a certain level of experience and skill and seniority, you've made it. That's it.
Starting point is 01:01:39 You're done. Money's fine, right? And so I think there's this desire amongst junior engineers early in their career to be kind of over-optimizing for promotion and salary, etc. As opposed to investing in themselves. And I guess they probably don't want to hear me say this because it's easier for me as an old guy to say this, right?
Starting point is 01:02:04 But I think, you know, for the longest time at Dropbox, there was a time at Dropbox where like we just didn't have our level system right. I was tech leading the team and I was making the least amount of money of the team. Because at some point, at the time, I was technically manager. I saw everyone's salaries. I was making the least money. But whatever, right? It all worked out in the end, right?
Starting point is 01:02:22 And that was a happy story. But I do really think I benefited so much from working with the best people. Like, if you have two choices, like making 20% more money now or working with like the best people in the world, right? Just be around the best people because like you will be, that's going to set you. on the ship to success, right? And now, not everyone wants to do that. Not everyone wants to go on that. Like, there's no shame in just wanting to be, have a regular job and just be chilling
Starting point is 01:02:56 and just be getting paid and fine. There's, that's nothing wrong with that. But if you do want to maximize, if you want to maximize your career growth, then the way to maximize that is to maximize your skills, right? Like, if, and hopefully you're not so cynical to. I think that there is no correlation between talent and compensation. I don't, if you think that's the case, okay, I don't know what to say, but there is a correlation for talent and compensation and growth.
Starting point is 01:03:29 And so my advice to people early in their career is land at the best company with the best people doing the most important problems. And your life will be great because we're so lucky as engineers. Like, where else can you work solving problems for a job, like solving puzzles for a job and getting paid so well? Like engineers get paid so well compared to most jobs. The privilege, the real privilege that we get is to work on something cool. And that's what I advocate for. I love the idea you mentioned earlier that's kind of counter to the common opinion I often hear.
Starting point is 01:04:10 Like the common opinion I hear Someone working at a big company They're in this machine It's kind of defeatist They're going I shift this thing But I hate this or I don't believe in this And I really liked your perspective on Doing the right thing
Starting point is 01:04:26 And like basically giving a damn That you know someone Outside of you is doing the right thing too And expanding your sense of ownership And I want to know What is your motivation for that? Because there's so many other people who are faced with the same inputs and they get to a very different conclusion.
Starting point is 01:04:44 So what motivates you to actually do the right thing when the machine doesn't necessarily incentivize you to do so? Yeah. Well, I think the machine does incentivize you to do so, but I don't think it's visible immediately. Like I think people, for example, who do less job hopping really do grow the most. Because I would say very strongly, if you're not in a job for three years, you're not going to see whether you're not going to see whether you're not.
Starting point is 01:05:10 decisions were good. Like you can get more money as a junior engineer, but you cannot become a senior, a very calendar senior engineer without being around for long enough to own the consequences of your decisions. It's like, it's like playing a basketball game and leaving before the game's over. Like you're just not learning, right? So, so I do think there's an actual structural incentive towards staying in a job. Now, if you're in a bad job, leave the bad job, right? But there's a structural incentive to be able to stay long enough to have an impact. But also, I don't know, like, again, I guess it's easy for me to say. And also, I joined the tech industry when it wasn't like a very lucrative field. Like, you know, it wasn't like it is now, you know. But I just like
Starting point is 01:05:54 it. The reason I left academia was because I wasn't confident I was making the right. So what would happen how I would, how I was in academia. I'm not saying everyone was like this. But you, so, firstly you don't know what to do. So you have to make up a problem to solve. And you make up a problem. And then you make up a solution, you know, hopefully a good one.
Starting point is 01:06:18 And then you write a paper where you try to convince everyone that was a really good idea. And I hated it. Because I didn't want to convince people. I just wanted to build it and see if it was good. Like I just wanted to build it and ship it and see.
Starting point is 01:06:33 I didn't want to play pretends and argue about it. was good. That's what makes me feel good. I don't know. And like, obviously, if you're an engineer who's living paycheck to paycheck, you should maybe ignore what I'm saying. Like, I don't want to come across as unempathetic to anyone who's really struggling financially. But if you are not struggling financially, the best thing you can do for your quality of life is, you know, in be enjoying what you do every day. You know, like, I don't know, you could have a fancier car, or you can enjoy what you do
Starting point is 01:07:13 every day. And I love cars. Like, I'm a motorhead. But I promise that enjoying what you do every day is going to have a much bigger impact in your life. And personally, I don't enjoy, like, going to work and working on pros that I'm believe in. Like, I can only operate.
Starting point is 01:07:30 And say, Jamie, my co-foundland is the same way. Like, we really get along really well in this respect. we can only we can only put up with doing things we actually care about and believe in and I think that's the luxury
Starting point is 01:07:45 that's more luxury than taking a first class flight it's more luxury than go into a three-miltern-star restaurant the luxury is you don't get up and have to do a terrible job you know you get up and go to a place that you like your co-workers and you solve cool problems and you go home and feel proud of yourself
Starting point is 01:08:04 If you can construct your career that way, and it's not easy. It requires very active effort. Then you're going to have a good life. I saw in your career journey, it said that you're occasional manager. And you seem like someone who really enjoys the technical aspects of things. So how did you decide to occasionally dip into management? I feel like most people should not want to be managers. And I sometimes think it's a bit of a red flag if someone wants.
Starting point is 01:08:34 to be a manager too much, right? Because being a manager's hard job. And so I think most people should go into management because they need to, because it's necessary. And so what happened every time I went into management, I'm in a management role now as well because you have to, right? But someone needed to manage the team. And so I do genuinely enjoy accountability.
Starting point is 01:08:58 I like responsibility. I like having weight on my shoulders, I suppose. And so I went into management several times. But as soon as I had the opportunity to get out, as soon as someone else would come in and manage the team, I would kind of bounce out of that role back into engineering. Now, going to management is tremendously educational. Being a manager really lets you see the world in a different way.
Starting point is 01:09:25 You realize companies are more complicated than you thought. You realize the engineers are more complicated than you thought. You realize everyone on the team is going through something. And you realize, oh, wow, now I understand. why that thing happened. So being a manager is very educational. One thing I would say, I would strongly caution people against is going into management too early in their career. And I see this happen a lot with well-intentioned people who want to push people into management as a means of career advancement. I think it's doing people a disservice. I think you should let folks
Starting point is 01:09:58 take their time in an in a career journey. I don't think you should be going into management under most circumstances in the first three years of your career. I think you should get to
Starting point is 01:10:09 ideally get to staff engineer before you do that. That's not going to happen for everybody. But ideally you take the time because if you go into management before you're an excellent technician, it will limit your ability
Starting point is 01:10:22 to influence strategy later in your career. How does that play out? It plays out with people who, who stuck or have roles as people managers where they see their job as maybe making a team happy or maybe coordinating a team or dealing with all the day-to-day challenges of people on a team
Starting point is 01:10:46 but they're not organizational leaders. They're not like pushing the team towards excellence. They're not like, hey, how can we reframe what this team is about? how can we, you know, help influence technical strategy. And again, people management is a fine job. But I think the best companies are ones where all the managers are very technical. And so they're able to make sure the company is doing. Like a certain point, like if you're not a particularly technical manager,
Starting point is 01:11:16 you're not going to be able to evaluate the work of your team. You're not going to know whether your team is even doing well. And I guess maybe you might contribute towards some of the stuff you were saying about, you know, cynical, organization. organizational attitudes where like my manager doesn't understand me. Maybe a manager doesn't understand you if they haven't spent enough time developing technical steals and and struggling. Yeah, there's another piece that you wrote. I thought it was really good about, you know, leading by example. And you actually say it's bad to lead by example, but there's a quote in there. It says, modern tech workers
Starting point is 01:11:51 can be an anti-authoritarian bunch at the best of times. Let's say, you're not. Let's say, you you're a tech lead or tech lead manager or manager, how do you strike that balance so you don't lose credibility as a tech lead? Yeah, there are command and control companies that like the manager just tells people to do things and they just do them. Those typically are not excellent companies. And not every company needs to be excellent. But if you want to have a company where your team is really innovating,
Starting point is 01:12:22 like the engineers feel very personally responsible of making really high quality decisions, you have to have them believe in what you're doing. And so, you know, at convex, I'm the founder and CTO. I can't really go to someone's desk and say, do this thing. Now, they might do it just because they like me. It's a good chance they would do it, but not because I'm the boss. They'd be like, well, why? And that's because that's our culture. I mean, we don't just do things because we're told, like if I want someone to do something, I'll spend time with the team talking about why we're doing it. Where the market's trending? Where's the gap in our product, what, you know, and then once we're really well articulated, um, why something
Starting point is 01:13:03 matters, people are just going to do it anyway. Now, there's a, sometimes you have to have hard conversations, but I think, um, the way I think about it in terms, certainly in terms of conflict in an organization, there's like this kind of hierarchy of, of the values you have, and then why, what and how. And engineers are very often debating the how. They're very often debating, like, what algorithm should we use for this and should we use this container service or that container service and these are kind of like the implementation
Starting point is 01:13:30 details but most times when I see organizational conflict is because well-intentioned people are debating as best they can about how to do something but they don't agree on why we're doing it so if one team thinks the most important thing we can do right now is get more features out
Starting point is 01:13:47 to expand our customer base and the other team thinks the most important thing we can do is increase reliability because there's a risk to the business. They're both very valid perspectives, but it's going to lead to them doing very different things. And within an organization, I'm a strong belief that everyone needs to have 100% Y alignment. So the stuff we argue about or we debate or we talk about ad nauseum is why. And then largely, I just trust the team to do the right thing. And I think that's the case for any tech lead. If you want to have credibility on a team,
Starting point is 01:14:21 if you think that you can just tell someone to do something and they'll listen to me because I'm the senior guy and no, they won't. They won't. If they listen to me it's because I've come in and explained it in a way that resonates with them. And again, one of my jobs at Dropbox was to resolve these situations.
Starting point is 01:14:38 You know, someone would say, oh, that team's an idiot. They weren't do blah. And then I would go talk to that team. That team was not an idiot. And we talked through it and they would decide to do the project. And it wasn't because they were scared of me. I hope they weren't at least. It was because, you know,
Starting point is 01:14:54 take the time to figure out what their motivations are. And that's something you really have to learn. If you're a tech lead, you don't really have that much authority over people and really you have to kind of encourage them and get them to believe in what you're doing. And if you are leading through authority,
Starting point is 01:15:11 you're not going to have a culture where good ideas arise from within the org. People are just going to do what, they're told. Influence without authority is huge. And I think there's, I mean, even big companies like meta, they technically don't have titles. Everyone's just a software engineer. Is that also how convex is run? I mean, we're a very flat organization. There are people in tech lead roles. I don't completely buy the everyone's a software engineer thing. It's like, you know, at, is, you know, anthropic, everyone's a member of technical stuff. Because it's kind of like a wink,
Starting point is 01:15:48 wink thing. You kind of know. It's like no one's mentioned the title, but you kind of know if like the most senior person of the company comes to your desk, you're probably going to notice, you know? And so I think there's no point in, you know, playing pretends. People, people know, but I will say that you should have an organizational culture where you don't do something just because a senior person says something. Now, I do think that reputation matters. Like I certainly think that, you know, I don't know, if a very experienced person at a company comes and tries
Starting point is 01:16:22 to explain something to you, you probably should listen to them because they probably have some wisdom. There might be something to be learned there. You should be open to it. But ultimately, even the most senior person, I don't think should be leading through authority. They should use the benefit of their
Starting point is 01:16:40 experience to be able to win the hearts and minds. And that's how you, in engineering culture, It's so important. And if you want a culture of ownership and innovation and drive and enthusiasm and people trying to do the right thing not get promoted, right? You have to have a culture where everyone believes in what they're doing. And no one's going to believe in what they're doing if the senior principal engineer comes to the desk and says, I'm not going to tell you why, but you have to delete this database and do this other thing.
Starting point is 01:17:08 That's not an empowering statement. What the empowering statement is, hey, let's spend some time together to talk about where this product's trending and how this is probably not going to work out and how there might be a different way we can solve this problem. On the topic of tech leadership, I mean, in that article I mentioned, the title was don't lead by example. And I think that kind of might be confusing for people. Can you explain why you think you shouldn't lead by example? Yeah, leadership by example is a very passive thing to do. And engineers are passive people at the best of times, you know.
Starting point is 01:17:40 I think there's stereotypically a little bit part of that personality. So, I mean, the very concrete example is, you know, when I first started becoming an engineer leader, I was trying to lead by example. So I wanted to kind of demonstrate the behaviors I want everyone else to have. And very specifically, like with regards to on-call and people getting paged, I want people, I want people to have high ownership. I want people to jump on issues as soon as it happened. So I would do it. I'd be always the first one to respond to a page. I would always be like, you know, writing up the reports.
Starting point is 01:18:14 I would always be jumping on all the bugs and stuff and like really falling over myself to kind of show how I want people to do, to be. But from their perspective, all they see is that the leads just doing all these jobs and they don't know, they're like, oh, maybe that's James's job. Or maybe James knows how to do it and I don't know how to do it. Turn it, I didn't know, I was just kind of figuring out.
Starting point is 01:18:39 or maybe he likes doing those things. And I think at a certain point, being a leader is about understanding human psychology. And I don't think you can just act a certain way in front of people and wait for them to copy you. I think you should, obviously, you should act with integrity and values and you should, you know, you should own the values of the team. But you have to explain stuff.
Starting point is 01:19:02 Sometimes you have to tell people, hey, very specifically, there's a trajectory almost every leader goes through, every high-achieving leader, where they become a tech lead and they care so much that they become a micromanager. Like they review every line of code, they're involved in every decision. And at a certain point, they are the bottleneck for the team. So this person is so overwhelmed, you might have gone through this yourself. They're so overwhelmed and they're like, wait, my team doesn't even seem busy right now. And I'm so busy.
Starting point is 01:19:35 I'm reviewing all this code. doing the strategy, what's going on? And then a manager will come along and say, hey, like, you're micromanaging. You got to let your team have more ownership. And so the tech lead says, okay, sure, whatever, I'll let them own stuff. And they'll just take their hands off the wheel. And then the team falls apart, right? Because you can't just stop doing the things you're doing, right? You have to go and have a conversation with people. And so I see it, I see leadership as this kind of slider between oversight and accountability. So when someone's new to the team, when they're very junior,
Starting point is 01:20:14 they're in a mode of oversight. Like you're checking their work, right? But at a certain point, you have to dial down the oversight and very importantly, dial up the accountability. So instead of saying, I'm not going to look at what you're doing anymore, you say, okay, cool, you've got this project. Let me know when it's going to get done. Next Thursday, cool, all right, what's the thing?
Starting point is 01:20:36 the plan. This is going to happen? Okay, how are you going to know it's correct? Great, great, great. It's on you. I expect you to do that. Great. Let me know if there's any issues, right? And then the ownership relationship is explicitly on them. Because what you want to do is basically encourage ownership within teams. You can't go from owning something yourself to not owning something yourself and expect that to develop. You have to go have those conversations. You have to go and give people accountability. Now, what I have found, even though that can feel like an awkward conversation, most people genuinely like accountability.
Starting point is 01:21:12 Most people like to own their work. Most people like to say, hey, this is on you. We're all going down with the ship, right? I'm not going to leave you high and dry. Like, you know, I'm the tech lead. I still take, you know, responsibility to. You're accountable to this project and go for it. And that's how you can develop people than your team.
Starting point is 01:21:34 You know, a tech lead shouldn't be about you. as the boss and the team is like the people who, you know, do the work or something, right? It's about this kind of flow where you're the more experienced person, typically, maybe the more organized person, the more strategic person, and you're working on developing your team members so they can take your job. And then you can do something else. That slider you mentioned, what if you give someone accountability and they blow it? Do they go back down the slider to, you know, where you start microvenging? recognize it, right? So this is growth. I mean, this is growth, right? It doesn't always work.
Starting point is 01:22:12 And so I think I had another article about paper cuts or something. So, you know, we'd call these paper cuts, you know, so some decisions don't matter that much. You know, you get them wrong and maybe it sets you back a week. You get them wrong and maybe something's a bit suboptimal. And these are great candidates to give people accountability for because they do it. And if it doesn't work, they get to experience it and it will feel bad and I guess that's good it's good to feel bad you know you shouldn't be demoralized but it's good to try something it doesn't work and you're like oh wow that's that that didn't work and that and then you feed that back into your little you know l-lm in your brain and you get better next time right that's that's a growth experience I think what you don't want to let people do is like
Starting point is 01:22:59 lose their arm you know like like paper cuts are fine but I will not let someone a junior person especially design the replication system at convex, whatever. Like, you know, you have to have safeguards. And one of the arts as a CTO, a leader, a tech lead, is figuring out the right level of altitude for how to know whether you can trust someone to take on ownership or something. You mentioned, because you were the most senior engineer at Dropbox at some point, that you had to mentor other very senior engineers. And, you know, when I think about mentoring a junior engineer, it's relatively straightforward patterns. But when I think about, let's say I need to mentor a senior staff engineer, maybe even a principal engineer, how do you mentor someone like that
Starting point is 01:23:46 that's already so polished? They already know how to take ownership. Yeah, how do you mentor someone so senior? Yeah, I mean, there's in two ways. One is there is still just a whole bunch of commonality. Everyone goes through the same problems. They don't know how to deliver harsh feedback. They They don't know, you know, there is a bunch of the standard stuff people working on. But I do think that at a certain point, everyone needs to become the best version of themselves. Sounds so cheesy, right? But there is not an archetype for what, according to me, there's not an archetype for a senior principal engineer. You've got to be your own brand of engineer.
Starting point is 01:24:26 So you might be, for me, I'm like the strategic kind of collaboration, simplicity, abstraction, engineer and my coding just fell off a cliff you know some engineers are like the the deep science fiction you know the super hardcore science you know problem solving engineer so everyone's going to find their own brand so oftentimes what i'd be working on them is is a combination of two things one finding their strengths and helping them be more spiky helping them to like really excel in the area that makes them special at the same time I'm almost always working on the personal side of the engineering. Like the organizational personal understanding people, understanding the why.
Starting point is 01:25:10 It's just, I don't know, we're so mathy, you know, as an industry. We think there's somehow, like, even silly things. Like, for example, I have to tell so many senior engineers that you can't change someone's mind in a meeting. You just can't do it. Like, you can disagree with someone and they're going to be mad or whatever. but to change someone's mind, involves going through a complex series of neurological processes that don't happen live in front of 10 people in a meeting, right?
Starting point is 01:25:44 And so if you force someone to agree with you about them being wrong, they're just going to go along with it and the ego is going to get bruised and whatever. So one thing, you know, firstly, meetings generally aren't for decision-making. most of the time when you identify a disagreement, I would point out the disagreement and why and why I don't agree or where the problems are, provide enough information, and then let us sit and let them go and reflect on it
Starting point is 01:26:12 and come back and have a conversation a week later after they've gone back and float about the why. And so this is kind of very psychological, but it makes no sense to be like, well, they should be able to change their mind. I mean, well, they won't. That team should do this. this, well, they didn't, right? You know, people should like my API. Well, they don't, right? And I think
Starting point is 01:26:32 it's like that whole, like, it's almost, um, it's almost like a capitalist attitude towards like interpersonal behaviors, right? It's like doesn't like, you know, capitalism rewards success, you know, or, or impact. It doesn't matter how well-intentioned you were if you didn't manage to convince them. You could be the smartest person in the world, but if you can't, change someone's mind, then you are kind of useless, right? And so I think a lot of senior unions still struggle with this. It doesn't matter how smart you are. It doesn't matter how right you are.
Starting point is 01:27:06 Are you effective? And a lot of being effective at that level is about, you know, uncertainty and project management and simplicity, blah, blah, blah, blah, blah, blah, but also kind of interpersonal dynamics, because most hard problems happen with a team. Not always, but generally when you're at that level, you're going to have to have 10, 20 people working with you to get something done. And that's a whole different ballgame. On the topic of career advice, the industry's changed a lot in the last five years because of all these agentic tools. And I wanted to know if you thought, is there any career advice that
Starting point is 01:27:43 majorly changed in the last five years? Something that you used to say five years ago, you don't say anymore or vice versa. No, I don't know if I've changed my perspective much, but I think the industry has changed this perspective dramatically. I think, I mean, let's just be honest about it. There is incredible demand in Silicon Valley for senior engineers, and it is getting harder for junior engineers to succeed and grow for a variety of reasons. Why is their demand for senior engineers? Well, because LMs can't do everything. Well, you know, the architecture and simplicity and design, they're still the domain of human beings, despite what you might hear on Twitter, right? and every company, including the labs,
Starting point is 01:28:26 are desperately hiring senior engineers. But junior tasks are getting a little bit commoditized, and that worries me. Because I do think that learning, it's very, you know, wisdom is kind of facts put into practice and then synthesized, you know? And so a lot of people would argue that, oh, well, you know, it's easy to learn now because of chat, GPT,
Starting point is 01:28:51 because you can just go ask it a question about, how does two-phase commit work or what's the difference between snapshot isolation and serializability? And it will give you probably a pretty good answer. But I think growth as an engineer does require wisdom. And wisdom only really happens when you synthesize it, in my opinion. And so I think I'm still very bullish on young people. We're hiring junior engineers at Convex. I'm very excited about, I mean, I love working with junior engineers who are really hungry to grow.
Starting point is 01:29:28 But what I would say is train your mind. Like, do not listen to anyone who tells you that there is an advantage to having less knowledge. I'm not sure if you've seen people say these ludicrous things that like, oh, maybe in the future, like, not knowing engineering will be an advantage because you won't have biases and you'll just use clawed. I think these are like ludicrous statements. Like, software engineering is an intellectual discipline that helps you think. Like, like the software engineering, the best software engineering is not about knowing syntax. And it's not about knowing an algorithm as being really good at conceptualizing problems and being able to break them down to building blocks and coming up with clean solutions to them.
Starting point is 01:30:11 And that requires experience. It requires like, it's like doing weights with your mind. And just like if you went to the gym and you do. just like it never hurt. Like if you went to the gym and just picked up really lightweights, you're not growing. And also, if you went to the gym and picked up a heavy weight and like, okay, cool, I think I can do it. And you let the robot pick up the weight for you the right. You're also not really growing, right? You have to do the reps. And so I would say, and it's tough, but find a way to stress your brain every day. Find a way to avoid. Now, obviously,
Starting point is 01:30:50 agentic coding is here, right? Obviously, that's, you know, I could never tell someone to, like, never use a coding agent because that'd be silly, you know, but I can say that it is easy to fall into a passivity trap, like where you're just being passive about your learning. And I do think you need to spend some time in the intellectual wilderness of not being able to solve a problem and struggling. I would say if you're running into a new, a new, um, a new, problem, try to think of a solution yourself and then go check with an LLM, it's going to be hard for you. It's almost like, for example, I'm not very good at reading anymore. I'll be honest. I would find it hard. I mean, I can read obviously, but I find it hard to sit down and read a book
Starting point is 01:31:40 because my brain has been fried by the stimulation economy, right? Like, like I find it, I go home and I'm tired and I watch a YouTube video. I don't tend to go home and read a novel. I should be better at that. But it takes discipline to do that. I would say similarly, it's getting harder to solve a difficult problem without reaching for the help.
Starting point is 01:32:03 It's getting harder and harder every day to be faced with a very difficult intellectual problem and not being like, well, I'll just do a Google search or I'll just ask Claude, you know, whatever. I'm still doing the work. I don't know. I would really encourage people
Starting point is 01:32:16 to practice using your brain every day. If I was just devil's advocate or just thinking from the junior engineer perspective, I might think, well, the proof that Claude can do this work today means that I don't need to know it today or tomorrow or in the future. So why even build that skill in the first place? And also these agentic tools, they have positive trajectories too. Yes, so firstly, the agents are not particularly good at a lot of parts of engineering currently. And so probably everyone should agree there's, you know,
Starting point is 01:32:55 Claude is not good at designing distributors systems protocols right now, for example, you know, or managing a three million line code base, you know. And so there are parts of engineering where engineers are valuable. And like I said, you know this because the labs are all hiring engineers. No matter what they say, there's still hiring engineers, right? desperately hiring engineers, really, really aggressively hiring engineers. So engineers still have value. And maybe they weren't in the future, but I doubt it. I really do think there's a role for human ingenuity in engineering. And so here's the trade-off I would pose to people.
Starting point is 01:33:32 I mean, there's kind of three paths. One, get out. Get out of engineering, go mow lawns and do whatever, if you want. That's the most nihilistic attitude. I don't believe in that. I believe, I really love engineering. I believe that there's promising roles to human beings in engineering. So the second is to realize that there's value right now for humans in engineering, and maybe one day AGI will be here. Let's say in the years time, AGI will be here. And there's two people. one person just gave up on problem solving right now and they're just like feeding the machine and they're running 17 coding agents in parallel
Starting point is 01:34:11 and they're just accepting everything Claude says and they've given up right and one person said you know what I still want to have an active role in my learning I want to understand what it's doing I want to think hard about problem solving right played out one year AGI arrives who is it going to be a better place for the future
Starting point is 01:34:28 right the person who's been training their mind right like there was engineering is not a means to an end. It's a mechanism for improving your mental processes. I still do whiteboard and coding interviews, and by the way, Anthropics still does whiteboard coding interviews, just in case you were wondering.
Starting point is 01:34:49 They don't say just use Claude, right? So I still do whiteboard coding interviews with candidates. Candidates are getting worse at coding. That's absolutely true. But I don't know any better vehicle for evaluating someone's intellectual capacity for problem solving, then seeing them solve an engineering problem. It's a great mechanism for doing so.
Starting point is 01:35:10 It's like, you know, if you were doing a math degree, right, a big part of doing advanced mathematics is proving theorems. And by the way, the theorem's already all proven, right? Part of the, you know, the exams and is proving a theorem that has already been proven. And so you might say, well, what's the point of proving that theorem? it's already been done. Or the point is that the act of proven that theorem improves your mind, right? And then you can go and do more innovative stuff.
Starting point is 01:35:40 So I would say, sure, you may not, I still do think that there is a very, very promising role for human beings in engineering, maybe not in coding, but coding and engineering are very different things, right? But even if you don't believe me, you can either give up now, but or you can just keep trying to improve your brain and you'll be better off anyway. Imagine if you're wrong. Imagine if you take the defeat of you. Imagine if you say, well, AGI is coming next week.
Starting point is 01:36:05 Why bother doing anything? And imagine you're wrong. Oh, my God. You just got off the, you just got off the ride. You go off the ride. It's so much cool stuff. I mean, this is a cool time for engineering, right? This is a really cool time.
Starting point is 01:36:20 There's so much cool stuff happening. And I don't, and I hate this talk about like, AI means humans don't have a role anymore. AI means no one's going to have a job. AI means we're going to do nothing else new and there's no one's going to. What I like is, hey, look at this, all this new cool stuff we can build. Like, look at how these ways we can make people's lives better. And to be honest, I really wish my peers and my cohort would stop it with the real Duma, like, human elimination kind of narrative. And because I think there is a little bit of a, like, you know, a psychological anchor.
Starting point is 01:37:00 You know, I, and as part of convex, we didn't start convex to eliminate jobs. We started comments, we want to make it easy people to build cool stuff. And I think the more we use in industry get behind, let's do everything we can to make it possible for people to do more cool things. I think it's a really exciting future we have ahead of us. I wanted to talk about what you're working on now, or convex, and why did you quit Dropbox to build convex? Dropbox is a great place. A certain point I was there. eight years, whether I, I don't know if I said I outgrew the company, but I'd been the
Starting point is 01:37:35 most senior year for a while there, you know, a certain point, I wanted to grow and do my own stuff, you know, and so, so I left the company, you know, very amicably, and started convex. Now, why convex? Convex started pre-agentic era, because my observation was that the real differentiator in the success of a project, especially a large project, is the quality of the abstractions, the quality of the design, the quality of the architecture, right?
Starting point is 01:38:08 Systems that thrive over time and are extensible are systems that are architects well. And in my mind, the most difficult challenge in engineering, according to me, is distributed state management. How do we store state reliably and reason about it,
Starting point is 01:38:23 you know, modifying it concurrently with other users? And so we designed a platform for application building based on our experiences building large-scale systems. So Combex is a, is a transactional database where the transactions have written in TypeScript. They run as stored procedures. TypeScript, store procedures. They're serializable. There's, you know, automatic reactivity. So what the client sees is a consistent view of what's on the server. And I would say convex is a very designed
Starting point is 01:38:52 platform because convex is designed to be very composable and fit together well. And so we designed convex for developers to use. In particular, we wanted to make it so that application developers were able to build complex full-stack applications. That was the goal of convex. Now, all of a sudden, a gigantic development came along. And that's been really interesting for us, because it turns out that what humans find hard is also what agents find hard. You know, you know, the coding agents are not particularly good with large code bases. They're not particularly good with reasoning about action at a distance, you know, race conditions across services.
Starting point is 01:39:32 They're not particularly good at simple architectures over time. And these are the things that convex gives you as a developer. So the idea now, and almost everyone using Combex is using Comex because they have the coding agent doing their front end, but they need a back end abstraction, a higher level abstraction is something like AWS or something like posted Postgres, which makes their problems go away. And that's the company. That was one thing I wanted to ask, because immediately when I think, oh, I just need a back end or something like that, just go to AWS or just host something like that. So this is a layer of abstraction on kind of on top of those types of primitives.
Starting point is 01:40:08 Yes. That makes it easier for an application developer. It ties into a lot of the stuff I said earlier about making problems go away. And frankly, I mean, I watched the interview you did with Barbara Luskov and Barbara was my advisor in grad school and we worked a lot together and a lot on abstraction and you know the value in clean designs that minimize complexity and so AWS is a fine tool postgres is a fine tool although none of the mainstream databases are that great frankly but they're fine tools but they're not they don't make problems go away right and so the idea of convex is a higher level set of abstractions that you can use and not reason about state management not reason about concurrency, not reason about scheduling, not reason about transactions,
Starting point is 01:40:53 not reason about polling and data sync and type safety and all those things. So convex is a, if you think about, you know, the history of engineering, over time the abstraction floor raises. You know, when Barbara was first starting, she was using punch cards, you know. I don't know if she mentioned it to you, but when she started as a programmer, she never heard the word programmer before, you know. that was the first time she heard the word, right? And then you went from punch cards to like, you know,
Starting point is 01:41:22 having proper operating systems and, and, you know, then languages like C and then higher level languages and then you had cloud computing. And over time, the abstraction floor raises and you largely forget about what's going on beneath the surfaces. Most people don't think about how S3 is implemented. I do, but that's what I used to work on. But like most people just use it and it just stores your data
Starting point is 01:41:44 and it gives it back and that's great. That's a successful abstraction. But I do strongly believe that the world is and has been overdue for a new abstraction, at one level up the stack. And especially now that people are doing agent development, largely they don't want to own a Postgres instance. Largely they don't want to think about Kafka versus Rabbit MQ. They don't want to think about, you know, what set of tools to use. They want it just to work so they can focus on building their application.
Starting point is 01:42:11 When you talk about, you know, the abstraction, there's obviously a lot of stuff going on behind the scenes and Convex and the technical side. And what is it that Convex is building behind the scenes that you're most excited about and why? Basically, Convex is a new operating system in some respects. So we have the primitives, queries, mutations, actions, subscriptions. What I think is kind of cool is how we built this. We have our own database that we built.
Starting point is 01:42:42 We have our own distributed database that does, you know, tracks read, ranges and right ranges and does very efficient subscriptions of a web sockets etc. So that's the current operating system set of primitives. But Combex is getting much larger workloads now and much more interesting workloads and more high performance workloads. And so we're in the process of developing a slightly lower level API for doing very efficient kind of background processes, singletons, APIs like fork, like you know, like, you know, like operating system primitives.
Starting point is 01:43:17 And I'm pretty excited about launching these and how much is faster is going to make various convex components like the workflow system. And to be honest, the thing I find exciting every day, challenging every day, I still find convex very hard.
Starting point is 01:43:35 Like, to be honest, like, I struggle every day. I don't find my job easy. I mean, I feel confident at my job, but it's not easy. like designing the new API for this is super hard. I can't just go ask chat GBT. It's not going to give a good answer, right?
Starting point is 01:43:51 And because it's innovation, it's new ideas. I really enjoy it. I find it stressful sometimes. I find it challenging and tiring. But I also find it exciting. And I would encourage engineers to try to find this kind of stuff to work on, where it's like you're on that edge of like, I'm really liking this, but also, it's a bit tricky.
Starting point is 01:44:14 You know, it's a bit tough. You mentioned fork, and in operating systems, I'm familiar. You know, you just take the existing process and kind of split it. What's the idea of fork in a distributed system? So convex almost never has scale issues with regards to live traffic. You know, live traffic is like typically bound by user-facing interactions, people clicking on stuff, running website, you know, acting a website. Every now and then, someone will come to convex
Starting point is 01:44:42 and want to kick off a million background jobs to do something. background processing. It's a big workload. You programmatically, you can trigger huge workloads, right? And so one of the things we have to scale is kind of these background workloads, and a lot of them involve things like scheduling. And there are a lot of workloads in convex that would be very efficient if you had a background singleton process to perform things like aggregates. I'll give a very silly example, right? If you have a, let's say you're building an election on convex, a voting system, and every vote is a new row in the table. And you want to show a tally of the votes.
Starting point is 01:45:22 One way of doing this is having a bunch of background processes or crons adding these things up. One way is doing a table scan, which is the obvious way to use Postgres, which doesn't scale. The other is to have a background job, which if there's new votes, it adds them all up, keeps a tally. If there's no new votes, it goes to sleep and waits on like a condition variable to wake up, to wake up again was a new job to perform. And so these are the kind of primitives that we're working on right now. Most people won't even know they exist,
Starting point is 01:45:50 but allow us to build this very high performance, primitives for scheduling, aggregates, you know, background aggregations, et cetera. And I'm pretty excited about like the next generation of work clubs we can support as a result. When you reflect on your career, and it sounds like you've done a lot of gnarly technical work across your PhD, Dropbox seemed like pretty intense systems work and Convex is also doing a lot of
Starting point is 01:46:19 cool stuff. When you look back on your career, what was the most technically stimulating work you've ever done and why was it hard and what did you learn from it? There were certainly times in grad school where we were like formally modeling consensus protocols and stuff and I'd be on the phone with Barbara and weekends and talking through trying to reason about this in our heads. That was pretty intellectually stimulating and fun. But I think the stuff I found most stimulating was stuff like working on very large storage system with a team where, you know, things are going wrong. You know, where the rubber hits the road, that's where I find. And this is every day at convex. You know, the rubber hits the road like, you know, hey, we have a compaction process that runs in the
Starting point is 01:47:02 background, but it's running into issues. We might have to redesign it using partitioning, etc. I feel most intellectually stimulated where there's a really clear constraint in front of me and that to me is engineering. I don't actually know what the definition of engineering is, but I'm just going to make it up. In my mind, engineering is science with constraints. It's like how do you solve problems
Starting point is 01:47:32 in the presence of resource constraints? I'm not particularly interested in constraint-free environment. That's art. I like craft and engineering, and the more visceral and difficult the constraints, the more fun that is for me. And I've been lucky enough to, whether it's luck or intention, I don't know.
Starting point is 01:47:54 But I've always placed myself in those environments. You know, like, let's go get on the hardest team and own the hardest problem, and then put the effort in to, to survive. This question might be a little bit off topic, but I know you were a consultant for the show, the TV show, Silicon Valley.
Starting point is 01:48:18 I love that show, and I got to hear, how'd you get involved with that? Yeah, that was a lot of fun. So a lot of folks might not know this. I had nothing to do with season one. So a lot of TV shows, they don't know whether they're going to survive as a TV show. So Mike Judge, who started,
Starting point is 01:48:36 who wrote Silicon Valley also was of Beavis, Budhead and Office Space fame. He started his career as a software engineer at I think Lockheed or something. So he actually was a software engineer that a lot of people don't realize. And so Silicon Valley was like a throwback
Starting point is 01:48:51 to the kind of work in it. And if anyone's seen the movie Office Space, you would get this. That's really a dystopian cubicle era tech industry film. So they did season one of Silicon Valley and then it was very popular. and they got picked up.
Starting point is 01:49:07 And so they had to figure out what to do for season two, but they didn't know what to do because they've written a storyline that gets to the point where there's a compression algorithm and then what happens. And so they needed to find an expert on compression. And I guess ostensibly that was me. And I don't know whether I was an expert on compression.
Starting point is 01:49:27 I guess I was an expert on storage at least. And so they came to the office and we just chatted. It was so much fun. you know and so I was involved in um you're pretty heavily involved in the show um a lot of it was you know storyline and and so first like yeah sure what would you do with the compression algorithm what would could you design a storage system and so coming out with story ideas that that are technically accurate they were really you'd be surprised to know how much they care about accuracy a lot of um people I know can't watch that show because it's just so creepily accurate they find it so cringy
Starting point is 01:50:05 And partly why Silicon Valley can, the TV show can be so cringy, is because it's real. Like, those stories are almost, almost, I don't even everyone. So many of the stories in Silicon Valley are just real stories. They just, they went to a bunch of companies, just farmed everyone for like stories of crazy things that happened in the tech industry, and they wove them into the series.
Starting point is 01:50:29 And all the characters are based on real people and real archetypes. But they also cared very much about technical accuracy. So I would also do technical consulting. And then they'd say stuff like, oh, we're building a data center in a house. What should the rack look like? And, you know, what should the diagram on the wall look like? And partly, you know, I wanted to be like, oh, well, it doesn't really matter.
Starting point is 01:50:50 No one's going to care. And they're like, no, no, no, it matters. They really cared to get it right. So, yeah, I love the show. But it can be hard to watch just because of, oh, my God, how real it can feel. Are there any Easter eggs where you look at it and go that's unusually accurate or that system diagram actually is like very spotting. I can't, I don't think I can even say them because there were stories of like early Dropbox.
Starting point is 01:51:21 There were stories of having ideas ripped off by other companies and being tricked into having meetings with folks only to, you know, it may be the competitors team there to steal the information. a lot of those stories are real. And so there's people who watch Silicon Valley and like, oh, wow, that was something I went through. And probably it was because it was about that situation. That's such a cool experience. Did you get paid for that or it was just? Yeah, that's a complicated question. I got paid because I had to get paid because it was like a Hollywood union thing.
Starting point is 01:52:00 I didn't want to get paid because it made my visa more complicated. I've got a green card now. I'm all good. But yeah, but there was something where they had to pay me $400. Anyway, I made a grand sum of $400 off that show. Looking back on your career, is there any regret that comes to mind that maybe other people can learn from? Yeah. I mean, ooh, I think I underinvested in my personal life, to be honest. I don't probably say that much, you know, because I think you can be all about, like, growth. And you're sure I could have grown more. I could have dropped out of grad school, say, three years in or four years in.
Starting point is 01:52:40 I probably would have learned just as much. I could have taken a job at Dropbox back, you know, two, three years earlier and made a lot more money. Everyone has been in the industry long enough has been offered to co-found several billion dollar companies. So, like, everyone has a story about the times there could have been a billionaire several times over. I don't really regret those. I think, yes, it's been a bit. a lot of sacrifice to be to be blunt. I've been on call my whole career. You know, I've been, I've carried a laptop almost every day. There's many dinners and parties and events I've had to
Starting point is 01:53:17 skip and there's people in my personal life who have suffered as a result and I really appreciate this people and they, you know, they love and care about me and they know that I have a passion for this. And so, you know, they accept me for who I am. I think it's, um, it's, it's, it's, it's, This is like, this is, you know, this is old person talk. But yeah, I think everyone has to decide, like, how much they want to really drive their career. Because there's a trade-off. Like, absolutely, like, I made a tremendous amount of sacrifices in my career. And I've really prioritized building as probably the number one.
Starting point is 01:53:52 I mean, values first and then building. But if, yeah, if we went back in time, yeah, I would have had a good life, but I probably would have done more vacations and, you know, I just had a bit more of a balanced life. I really do think, I mean, I see that with all the 9-96 stuff and this kind of performative, like, you know, photos of being in a bar and a laptop and stuff. And I'm like, that's not, that's not real. Like, that's not, I mean, sure, I was working more than 996 back then. But, like, I probably still do work more than 996. But that's, but I don't do it like, as a checkbox.
Starting point is 01:54:26 You know, I do it because I just really want to be doing stuff. And so I think just, I would, I would caution people against that. hustle culture. I'll, that's not, firstly, like, you know,
Starting point is 01:54:38 you're only young once. You should, you should have fun. You know, I've got quite a few grays in here, you know, but, but also,
Starting point is 01:54:48 that's, that's acting. I mean, focus on solving problems, work hard, be passionate, yeah. Yeah,
Starting point is 01:54:55 when I studied your career, I mean, there's, there's mentions of you working a 16 hours a week and various, but, you know,
Starting point is 01:55:03 you, you like on call you like fires um you like ownership like all those things are recipe for working obscene hours yeah and you know i go home and i'm tired and i wind down by building stuff now like i i'm lucky enough to have a little workshop at home and so i go home and i you know make things with my hands uh that's just as you know that's just you become an infra person um you become an you become an engineer like in all aspects of your life um but yeah I would just say like the cool thing is doing cool stuff that humans use. The cool thing's not working long hours.
Starting point is 01:55:43 The cool thing's not like showing off about that you were running a coding agent all night. Who cares? The cool stuff's, the cool things enjoying doing important things. Do you have a best technical book recommendation for people? I mean, I have to be honest. I haven't read almost any technical book. I mean, obviously I was in academia for a long time, so I read a lot of papers, read a lot of papers.
Starting point is 01:56:13 Learning is awesome, reading is great, but balance it, right? Read something and then go put it into practice. And most of my career, I have, I mean, sure, I did a PhD, so I guess that that is like the academic side. But after that, most of my learning's been by doing. Because there's nothing like being faced with a real problem in your face. to like to really to really develop as an engineer yeah i think if i answer the question i'd probably say the same i think learning by doing is that's that's really where where it matters most
Starting point is 01:56:44 um and then last question for you is if you could go back to the beginning of your career and give yourself some advice or what would you say i'd say you know what it'll it'll be okay like don't sweat the small stuff as much it's hard because like my my whole like engineering brand is about caring about details and like I I love like Dieter Rams and like design you know and and and so so being obsessive is a little bit part of my DNA but I think I would go back and say careers are long I mean there's just any story anyone reads about 22 year old billionaire blah blah blah just ignore that story that's not that's not real that's not repeatable that's not normal and it's not that healthy and it's not that good for the people either, right?
Starting point is 01:57:36 You probably won't, ideally won't max out your growth for 20 plus years as an engineer. I'm still learning all the time. I've been an engineer for several decades. So my advice would be, don't sweat it. There's time to grow. And I do think, again, it's a little bit of a modern phenomenon, but there is a feeling right now,
Starting point is 01:57:59 oh my god, AGS come and better max out my growth in the next three months. Well, guess what? You ain't going to do it. It's not going to happen, right? You can't do it. You can't max your growth out in the next three months. It won't happen. Don't have to be running 17 agents at the same time.
Starting point is 01:58:15 All you got to do is orient your career around learning every day and getting better at what you're doing and trying to solve things in the most simple ways. Yeah, I mean, on Twitter, I see this take all the time of this permanent. underclass idea where if you don't make it in time for AGI, then you're going to be part of this permanent underclass. Yeah. I mean, look, it's a bit challenging economic times for a lot of folks. And I don't want to be, I don't want to be unsympathetic to people who are, you know, having financial difficulties. At the same time, it's just not an instructive attitude. There's not much you can do with that information other than feel bad about about yourself. and I don't and I and someone will
Starting point is 01:58:59 someone will argue back no what you can do is get really good at using Claude well guess what it's not very hard to use Claude as well like grab some like sometimes people tell me like oh my God I'm finding it hard to keep up with all the new models the new model drops I don't even know what the new models are like I use them but I forget the latest model because it doesn't doesn't matter right like like there was this thing called Ralph right I guess it's still Ralph, it's like a loop thing.
Starting point is 01:59:26 I don't really know what Ralph is, and I haven't heard anyone mention Ralph from the past few weeks, but it was the biggest thing in Twitter for like a month. And it just doesn't, this is noise. Sometimes I feel like, it's like, it's like tech tabloids. It's like people think that they're learning by somehow knowing, like listening to what Jensen said today or like, oh my God, Boris said that Claude writes itself,
Starting point is 01:59:53 I don't think people You're like that's like it's just like reading about Beyonce But you're a nerd And so you're reading about Jensen Right But like it doesn't matter You're not growing like you don't have to know It's okay
Starting point is 02:00:07 The new coding agent could come out And you could miss it And then next year If it turned out as the big one You'll just use it like it's not hard But I haven't seen any skill so far I guess there's some skill But it's not a hard skill
Starting point is 02:00:22 if you're good at engineering you can figure out how to use Claude Code right or open code. So just watch out for tech tabloism. It doesn't matter. Just be building stuff. Just just do real crap. I love that mindset. And yeah, well, thank you for your time. I'm part of it.
Starting point is 02:00:43 I'm on Twitter too. Like I'm part of the thing. But like, but just ignore me too. Like, oh God. All right. Well, thank you so much for your time, James. I really appreciate it. This is a lot of fun. Brian, it was great. Thank you. Hey, thank you for watching this podcast. If you liked it and you want to see the show grow, please support with a comment or a like. Also, if you have any recommendations for people you want me to bring on,
Starting point is 02:01:08 please drop a comment. Guests like Barbara Liskov, Mike Stonebreaker, Mark Brooker, these were all people that I brought on because someone left a comment. On another note, aside from the podcast, I'm working on building the ergonomic keyboard that I wish existed. Here's a glance at the prototype. It's a split keyboard. So there's two sides. This is in the case.
Starting point is 02:01:29 But yeah, we launched on Kickstarter and we hit our goal within eight hours of launching. I really appreciate it if you were one of the people who grabbed one of the early units. We're now working on the long journey of building the tooling now. And so if you still want to pick one up, I've left the late pledges open on Kickstarter.
Starting point is 02:01:46 So you can grab one there. I'll put a link in the description. Thank you again for watching the podcast. and I'll see you in the next episode.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.