The Pragmatic Engineer - Building Pi, and what makes self-modifying software so fascinating

Episode Date: April 29, 2026

Brought to You By:• Statsig — ⁠ The unified platform for flags, analytics, experiments, and more.• Sonar – The makers of SonarQube, the industry standard for automated code review• WorkOS ...– Everything you need to make your app enterprise ready.—Mario Zechner is the creator of Pi, a minimalist, self-modifying AI coding agent, that is the foundation upon which OpenClaw (created by Peter Steinberger) is built. Meanwhile, Armin Ronacher is the creator of Flask, and a longtime user of Pi. The pair are also friends.I sat down with Mario and Armin for the latest episode of the Pragmatic Engineer Podcast for an interesting conversation about AI and their reservations about it – even though both are heavily invested in building AI-powered tools.Mario explains why he built Pi, and gives his take on why it has become so popular. Armin walks us through how he uses AI tools, including building a game with Pi, and why he always puts human judgment firmly at the heart of his approach.We cover the risks of over-automation, the limits of agentic workflows, and why strong engineers with informed judgment still matter. We also get into the challenges of working with code written by non-engineers, and whether open source can withstand a tidal wave of agent-generated code.—Timestamps(00:00) Intro(07:30) How Mario, Armin, and Peter Steinberger met(15:15) How 30 dev teams use AI agents: learnings(21:50) The importance of judgment(24:26) Challenges when non-engineers write code(28:30) Downsides of over-automation(32:18) Pi(48:09) OpenClaw + Pi(50:54) “Clankers”(57:32) Open source and AI(1:00:22) Complexity as the enemy(1:02:50) Building an AI-native startup(1:11:52) “Slow the F down”(1:16:40) MCPs vs. CLI(1:25:03) Predictions and staying up to date—The Pragmatic Engineer deepdives relevant for this episode:• The impact of AI on software engineers in 2026: key trends• Cycles of disruption in the tech industry• The AI engineering stack• The creator of OpenClaw: "I ship code that I don't read"• What is inference engineering? Deepdive—Production and marketing by ⁠⁠⁠⁠⁠⁠⁠⁠https://penname.co/⁠⁠⁠⁠⁠⁠⁠⁠. For inquiries about sponsoring the podcast, email podcast@pragmaticengineer.com. Get full access to The Pragmatic Engineer at newsletter.pragmaticengineer.com/subscribe

Transcript
Discussion (0)
Starting point is 00:00:00 What if I told you that one of the most influential AI coding agents of 2026 was built by a single developer in Austria who got frustrated with existing AI coding agents? This is Pye, a minimalist self-modifiable coding agent, which has quietly become the engine behind the wildly popular personal AI assistant OpenClaw. Mario Zekner is the creator of Pye and joining him today's Arminoraneranercher, the creator of Flask, and now an early adopter and contributor to Pye. In today's episode we cover the backstory of Pye and why self-modifying software is much easy. easier to do with AI agents. What Armin learned interviewing 30 plus engineering teams about how AI agents are changing how they work and why software quality feels like it's trending down. The case against MCP and why CLIs are becoming so popular, and many more.
Starting point is 00:00:44 If you want to hear from two very grounded voices in the industry honestly talk about what's working on what isn't and why we need to slow down as an industry, this episode is for you. This episode is presented by Statsig, the Unified Platform for Flags, Analytics experiments, and more. This episode is brought to you by WorkOS. Engineers love to build. Today's episode will be a great example of this. We'll get into why and how Pi was built from the ground up.
Starting point is 00:01:08 But when you're shipping a product, some problems are better to solve with trusted infrastructure built for scale. Enterprise features like Samo, directory sync, and audit logs are some of those. WorkOS gives you APIs to add them in days, not in months. Shift faster without reinventing the wheel. And now let's get into the episode. Mario and Armin, it's so good to have you here on the podcast. Thanks for having us.
Starting point is 00:01:30 Thank you. So as a kickoff, Mario, how did you get into tech and eventually into building AI stuff? Oh, well, that's a long story. How much time do we have? So I'm a case of the 90s, actually. And I got my first PC at 96. And the trigger for that was that I loved computer games. We were kind of working poor, so we couldn't afford any of the Game Boy and N.S. Super NIA stuff. But I had an uncle with an Amiga 500.
Starting point is 00:01:57 and I would go to his place every second day and just play games there. And eventually my parents told me if you work, you can save up and buy yourself a computer. And in reality, my dad would do what's he called? Schwarzenberg. Well, you're not necessarily pingotaxes on your work. Yeah, so he would do his normal job. And after his normal stroke, he would go fix cars and work at construction sites. Yeah, it's very common in Europe.
Starting point is 00:02:23 Like, I know everyone's about it. And after two or three years or so, they just said, it's time. and took me to a computer shop in the nearby big city and bought me of 486 and that's how it started basically antium 486 yeah an Intel 486 dx 40 megahertz with turbo button and that's where i started and i've always been into games a lot which also led to graphics programming and through shiriluk i got a job while i was studying at university at the applied science organization who was doing LLP stuff, machine learning, applied machine learning, basically taking research results and trying to stuff them in the industry applications. And that's where I learned the ropes of machine learning.
Starting point is 00:03:09 That was all before deep learning became a thing. And I actually quit that kind of domain in 2010-11-ish, because I joined a startup in San Francisco. And then later came back and joined another start up with two friends in Sweden where we did an ahead of time compiler for a Java bytecode to iOS that got sold. And since they might have a little bit more time. And I've always kept up with machine learning stuff because obviously is super interesting. And yeah, and then GPT happened. That's the story. Yeah, and here we are.
Starting point is 00:03:44 And Armin, where are your roots? My roots are definitely not working poor, but because my parents run an architectural office. where they kind of adopted computers for cat's drawing. My first computer was like old computers that they recycled. So my first computer, even though I'm younger, was in 386. I'm so sorry for you. And so basically none of the computers that I ever had were capable of playing computer games properly. Because one, they used Windows NT, which at the time didn't do anything.
Starting point is 00:04:18 So you had to sort of like build their way through it. And like the only way next you could actually get them to run was because before, it didn't know yet how to get the Windows 95 or like Windows 311 that was like before it booted into either one of those you could put it into DOS like really old to DOS games at a time
Starting point is 00:04:34 when you could already get better stuff but because it was sort of this kind of thing I started throwing around with QuickBasic a lot with Tau Pascal I bought a bunch of books on that and I that that was my roots of learning how these things work and it just I wasn't ever really good at
Starting point is 00:04:53 this, but I found it really interesting. This idea of like, no, for sure. Like I, like, I, like, I, I, I, I, I swear to you, like, I was, when I, when I started dabbling of this, I just really sucked. But, like, over time, like, if we keep doing this, you get better. And then in, uh, 2000, two or three, I, I used, I used to use, uh, Delphi a lot, which was like a visual version of, uh, of, uh, of Tuval Pascal.
Starting point is 00:05:25 And in 2002 or 2003, someone also showed me, because I forgot this idea, like I want to use Linux, and then I thought, Delphi didn't work on Linux, and then I found Python. And through that, I started doing some PIF programming, and there was a Ubuntu just came out in 2004. And that was a venture-packed vehicle, but they created all this local community.
Starting point is 00:05:45 So there was, like, Ubuntu Association. So together with a bunch of friends, we started the German Ubuntu Foundation, not a foundation association and we ran this online community called Ubuntu users for four or five years because Ubuntu was popular
Starting point is 00:06:01 the community grew and then the skating problems came so like that's how I got into web development and then for building this I just wanted to build a templating engine a web library all of this and then eventually I bundled that together and made this flask framework
Starting point is 00:06:17 which was very popular and even nowadays still is a thing that Clank has like to spit out. It's hilarious. But I left that and then in 2013, 14, so I worked on computer games for a couple of years in London, but then afterwards I went back to open source and I worked on Century for 10 years and then left in April last year to try something. So both of you are originally from Austria. In fact, you right now live in Austria as well, right? you were doing games,
Starting point is 00:06:48 you were working at Century, you also did games before, and then the third person who's not in the room, but was on this podcast just before us, Peter Steinberger, also from Austria.
Starting point is 00:06:56 Where did the two of you meet? Where did the three of you meet? Because I've recently seen a bunch of photos, especially before OpenClawe and Pye started, you hanging out, the three of you experimenting, playing with AI. I think the two of us met on the internet, right?
Starting point is 00:07:13 I think we, it depends. because I definitely met you once when I was at a university. All right. So, but you didn't recognize me at a time and I was useless. I was already famous. But yeah, we sort of abstractly met on the internet. But eventually we met up in Vienna.
Starting point is 00:07:30 We were screaming a lot at each other, but on the internet. But in a very cute kind of way, in a very non-confrontational kind of way. And even though we might not think alike in all areas of our lives, it was a cultured exchange, I would say. So that was nice. And Peter, I, like six degrees of Peter Steinberger, basically. I was working at an office in my town. And the company that gave me free office space in exchange for being like a mentor to the CEO,
Starting point is 00:08:01 had some kind of business dealings with Peter's company, PSPDF Kit. And eventually he came to the office in Gratz. And I think that's where we met the first time. And then also the same year we met at the conference in Istanbul. I just hung out for an entire night, and that's basically where it all started. Nice. And then how did the both of you go from being skeptical about AI when these tools came out? And both of you have, at that point, by 2022, you've been doing a decade plus of building
Starting point is 00:08:31 complex software in different domains. What was your first reaction to it? And then eventually, how did you kind of come across the side of like, well, this thing is actually really interesting? So for me, it was, I think, in 2022. I think copilot, GitHub co-pilot came out before GPT. Yes, in 2021. Yeah, and through my previous startup stuff,
Starting point is 00:08:52 I was working with Nat Friedman and Miguel de Kaza from Xammerin. Oh, with Xamarin. Yeah, they acquired the company. I talked about the Java compiler thing. I knew Net Friedman from our early startup stuff and eventually moved to GitHub and then was in my DMs in 2022, I think, and asked if I wanted to have access to it up co-pilot the tap tap tap all the complete thingy and it was like i i don't really care i don't
Starting point is 00:09:18 think it is going anywhere and he's like no man it's the future got to try it it's the future so i tried it and it was absolutely horrible but yeah after after when gpt came out and especially when when they started providing API access i did a lot of projects just figuring out what works and what doesn't work not necessarily in a coding space but eventually once they had tool calling that's when they became very interesting or function calling as OpenAI called it back then. But it took until 2000 and I would say 24, end of 24 October or so for that to actually be useful. And that's where the coding agents also became kind of interesting. And then 2025, the CloudCode team came out with CloudCode,
Starting point is 00:10:00 and that introduced the Gentic search. So basically just give the agent a way to plow through your file system and read all your files. and then made the whole difference, actually. Like all the things that came before, like cursor with indexing and any AST-based stuff and all of that, that just went away. And I know that the CEO of Croma
Starting point is 00:10:20 is probably mad at me for saying these. That was the difference. That it didn't, it wasn't like a dense and sparse search thing that the agent could go through. It was just give it access to your files. That was it for me. That's where it clicked for me.
Starting point is 00:10:36 I think my path was kind of, similar because I think Copilot came out quite a bit earlier but I know that there was a program at GitHub that gave you early access to Copilot at the time I think it was like this maintainers group or something where it still was in I got the feeling for co-pilot that this will actually be really interesting but not in any way in which it is now because I felt like oh I am in an open source for such a long time and now they're doing like training in open source data it's like there is something At the very least, this will be controversial.
Starting point is 00:11:10 I didn't think about, like, it being productive. I felt like, oh, this is going to be, it's going to be, like, a controversial thing. I was, like, bringing up source data. And I was, I remember for, like, I was trying to probe it, like, really. Whether there's flask in there? No, no, I was trying to probe it, like, really adversarial. So one of the things that I probed on is, like, I probing on, like, will it retell GPL code? And I remember at one point, I got it to spit out the...
Starting point is 00:11:36 Carmax inverse. Carmex inverse square root function, which is very easy, because also it was had a very specific name. So, like, it was very easy to get to recall. But I also found that, like, you can sort of tap in a certain way, then it would then continue putting, like, license text on top of it. It was completely wrong.
Starting point is 00:11:49 So it came from an open source, GPL drop of, of Doom originally, I think. And so it was like, it would have been GPL code if it would have done that. But it actually attributed, like, MIT license from a random dude. And it is like, oh, like, Mr. Copeland, that's the wrong thing. And that tweet at a time, got really, really popular, and then sort of people started, like, sharing with me. Because I was at a time not really exposed to how much actual AI progress was being made in those labs. Like, I didn't come from this AI space or a mail space.
Starting point is 00:12:23 So, like, I learned about a university and I was like, oh, there's AI winter and then nothing happens. But through this tweet and some other things, I like, I recognize that there was something there. Like there's actually CEOs in certain companies are convinced this will get off. And that's how I started paying attention to it. And I was essentially I was trying all kinds of stuff with the API. Like can you do like buck fixing things? So I got really interested in it, but it didn't at all feel like the world is going to change until, quote, quote. And you also changed your stance on the whole, oh my God, this is spitting out open source code.
Starting point is 00:12:58 It memorized. So because like my like my schick for many. years now has been that I really I'm like a I want people to share stuff. Like I think like human progress comes from like building on top of each other and I'm a huge supporter of the fact
Starting point is 00:13:15 that in the US you basically take knowledge from one company and another company that then no competes. Like I like this pirate kind of approach to sharing. Yeah, spread of knowledge. Yeah. And so like I was like my optimal versions like copyrights don't exist in a way
Starting point is 00:13:31 or like very very like a limited kind of version of this. I was like, I really didn't care that spits out GPL code and doesn't attribute. Like, I was like, oh, maybe this will just completely destroy copyrights. And like, for me, I was like, oh, this is, like, if that's the outcome, I'm fine with it. But it was, it was an interesting kind of thing
Starting point is 00:13:49 in the beginning that it sort of like, it sort of creates this license violation. Like, I want to see, like, what chaos will emerge from it. And so far, I think mostly what has emerged from it is like a strong belief now that, like, the system in place for copyrights, has some assumptions in the US about how it's supposed to work
Starting point is 00:14:07 and we're all kind of like ignoring that right now because we want to create the mess first and then re-regulated probably because like at least in theory a lot of the things that we're producing right now are probably by historic readings of the copyright interpretation actually not copyrightable.
Starting point is 00:14:25 Yeah, that's an interesting one. But speaking of jumping to today, so an interesting thing that you did recently, we talked about it just before is as part of your new startup is building things on top of agents and you talked to about 30 different engineering teams saying, hey
Starting point is 00:14:40 how are you using agents inside of your company inside of your team? What did you learn from large companies to startups? I think that a bunch of learnings, I'm highly unsurprising is that whenever people had vacation, there was more time spent on
Starting point is 00:14:57 trying these tools. And just to be clear, like, talk with like folks at the likes of like meta startups yeah so like a bunch of different people right so a bunch of different people from like different like European dinosaurs like are you pointing at me well I mean like the European dinosaur would be someone like seamets yeah or I also talk to two companies which are sort of in a critical space and what I mean like when adoption happens when people have vacation is that like when you're a CEO or the tech it comes and says like you've got to use cursor now, you've got to use cloud code now,
Starting point is 00:15:33 is actually you don't get it in a way because you need to actually spend some time. Like there's a, it's like a two to three week kind of thing until it really clicks on you. And so I always felt like with the people that I knew, like I had a lot of free time. Like I left the company in April until October. I was like, I can dive into this. And I, I feel like, this is like, how does nobody get this? It's like catnip for all. It was crazy catnob.
Starting point is 00:16:00 I didn't sleep much, all of this. But what happened within the company seemingly is that when there was like Thanksgiving, there was for the Europeans, a lot of it was over summer. And then at Christmas, a lot of people sort of, and they also get free credits during those times. And so like more and more people get them. Oh, you mean the AI companies often give you generous research credits? More and more people went into this.
Starting point is 00:16:21 And especially after Christmas, I would guess like in more than half the companies I talked to, after Christmas, it really exploded. and it exploded in all the ways and we'd expect it where like all of a sudden the quality drops. And it doesn't necessarily drop because like people want to make worse code but because it actually takes some effort
Starting point is 00:16:43 to stay within this. And we have seen this in the startup ecosystem already in the summer last year. Like if you pay attention to like the YC startups, a lot of them, some of them have their stuff on GitHub or for some period of time on GitHub
Starting point is 00:17:01 and you can look at it and like at the time because like plan MD files checked in and like everything attributed to Claude so like that vibe coding kind of thing was for like prototypes and whatever and like that built it out. It was already out there to see.
Starting point is 00:17:15 But then gradually a small version of this has like been code basis with a little bit of vibe slop on top. And an interesting sort of part of this was like how engineering teams and companies are now responding to that. With all kinds of different findings,
Starting point is 00:17:32 but a lot of it has been challenged to review PRs. They're getting large and larger, and they're becoming more psychological taxing. So engineers specifically are having a hard time keeping up with the longer PRs, they're more frequent. Yeah, and they're also, a lot of the code in those PRs is how an engineer wouldn't do it.
Starting point is 00:17:52 Because as an engineer, you sort of get a really bad feeling, committing certain code because you think of your future self and the agent really does not care I will retell this story over and over but like I worked for an Xbox one game at the time right around the Xbox one launch
Starting point is 00:18:10 so that was like a fixed day it has to release on that day so I worked on the Halo Masterchief collection and there was a game where you had like a matchmaking component and you had to like store this thing and whatever and it was like it was an all hands on deck kind of situation where people had to go in and unslop the human-made slop that was the matchmaker. And it was like, it was a system with like way too many states. We call it an emergent state machine because it was like 16 bulls on one massive thing
Starting point is 00:18:38 and like in theory, they were only six valid states. But in reality, it was a dramatic explosion of possible states. And that's how a Chentee code feels like, where it really should only be like a very clearly defined system. But in all reality, they're like, oh, we can, config doesn't load. let's catch it down and load the default config. So instead of actually failing, it now recovers. But now your code is way more complex than it should be,
Starting point is 00:19:03 because instead of failing properly, it is now recovering and entering these many more failure states. And that makes it much harder to work with this code because you can also not really ask the agent to reflect. So it because it's like, oh, yeah, this could be possible. So we need to maintain this invariant. I think it's kind of even worse, what you described about your human-made complex system
Starting point is 00:19:24 because there are moments of brilliance in agents where they spit out perfectly fine, simple code, exactly the amount and type of code you didn't need for that specific thing. And the USB steering engineer looking at better, like, wow, this is amazing, I can just sit back and not care because it's obviously doing this thing.
Starting point is 00:19:42 Like, two minutes later, you have another agent running in this window and it spits out the worst horrible garbage, because, but you might not notice because now you have fallen. into automation bias. I think your your agent is doing the job well. Do you think this might be a bit of a human bias because, because you know, like typically like onboarding a new engineer, you have a new join a new grad, you review their code, and if it's terrible code, you will review the next one thoroughly,
Starting point is 00:20:10 until they get to the point that, oh, it writes the code that I do, and then it typically takes, you know, six months or a year, or something like that, but then, you know, I can trust this person. Yes, but you don't have anything like that with agents. Like agents don't learn. You can put as much stuff in the agentism D. You build a memory system, but that's not the same type of learning than a human does. Obviously, humans are failable as well, no matter, but they have some capability of learning.
Starting point is 00:20:37 And retaining that learning, right? Yes, and they also feel pain. I think that's one of the defining things about humans. It kind of ties back to what you said. Eventually, if the pain gets too big, you as a human are incentivized to, fix the cause of your pain. And in the codebase, the cause is usually terrible interfaces, terrible complexity that you want to get rid of because you can't longer maintain that system. Isn't this why just calling on to it? You know, like senior engineers are always in demand
Starting point is 00:21:05 because the CEO sees a senior engineer as like they just get it done. But in reality, as a senior engineer or most senior engineers who are effective, they've had battle scars. They've been burned. They felt the pain. They saw what happened when they left tech dev spiral. So they now make all these decisions that they know they will help avoid. And of course, through this, progress goes faster. I personally think, and your mileage may vary, but a good engineer is an engineer that says no a lot, and I don't need this a lot.
Starting point is 00:21:35 Because that keeps complexity down. If you're using agents, the exact opposite happens. You say, yes, I want this and that they want this and I want this because I don't have to type it myself. I don't have to think about it. I just give the little machine a prompt, and it will spit out something that kind of looks like the thing I want. good enough.
Starting point is 00:21:51 And that's where all the problems start. And one thing that I also think is like, good engineering is all about knowing the tradeoffs that you have to make. And there is sometimes the right solution is actually, if you were to sort of like sit at university and learn about it, you kind of learn that you shouldn't be doing this in a way. I think Cal Henderson had this once where he said, like, you do the dumbest solution first until it doesn't work anymore.
Starting point is 00:22:16 because the actual problem is there's so much stuff that you need to do that if you actually do the right solution and the correct solutions and all of this you're creating the kind of complexity that kills you at scale and the engineer learns that but also like if you
Starting point is 00:22:32 if you don't have that battle scar it's actually very hard for you to argue correctly because it is this learning process that gives it the authority to then convince other engineers in the engineering org that you should be doing it this way. That is part of the learning.
Starting point is 00:22:46 of it you learned that. But the other thing is also that the agents give you now world knowledge access. And one of the other things that I learned through interviewing engineering teams now is that the senior person says no, knowing something. And then 48 hours later, the junior comes by and said, like, I talk to the agent. And I already had this inkling, but now you have all the evidence of why we shouldn't be doing it this way. Because like previously, you really didn't have that ready-made access to... Someone who can tell you a senior of. Yeah.
Starting point is 00:23:20 And this creates other stresses now that were previously... Not every team has that because they have a really good engineer. It's like people going to the doctor with a chat chitp-t printout and saying, this is what the machine said. You better do that. Is it fair to say that we are based on what you're seeing and talking, we might face a thing where it's very hard for experience engineers to it's harder just for them to say no in spite of the product manager or a junior engineer saying.
Starting point is 00:23:51 It's much worse because the product management now comes in and sends pool requests and all the mortgage shoots them. Yeah, that's another thing. It comes. Like non-engineers participating in engineering process is a thing now. Ask Arvin how that works. Ask him, how does it work, Aramon? Well, it's hard because if, because on one hand, like, it's well intended, right? If someone who is like not an engineer.
Starting point is 00:24:12 What is your experience? Is this your company talking with other people? So first of all, like, we have a little bit of this errandover. Like, we're small. And so, like, like, my covana, for instance, sometimes sounds like a poor worker's on the website. I talk to people that have that at scale, where, like, the marketing team all of a sudden does stuff on a website. And the sales team, like, creates ever more elaborate, like, sales demos that sort of land up on a GitHub org. And partially at this one of the most.
Starting point is 00:24:42 funny as one was like, where the sales demo built a feature that didn't exist, but nobody noticed. Right. So this, this is all, like, this is new, right? Because like, previously, none of that happened. But I think it's empowering. Like, if you're entire- It's like, there's a good thing to it in, too.
Starting point is 00:25:00 If your entire org, if everybody in your org can participate in, in the creation of software, in some form, right? Previously, people couldn't do that. Like, you had a designer who could figure something out in Figma, but they were, you might not be able to kind of put it into a clickable dummy demo, whatever. You might have a PM who wants to try out a feature without kind of wasting time of an engineer. Now you can do that. The problem is that people are now so focused on everybody can do everything now
Starting point is 00:25:29 that they forget that you still need a process to kind of guardrail all of that. And the integration part is the hard thing. It's like that Peter gave this idea of like the prompt request, but I'm actually really warming up to this idea. Like once you have demonstrated, it. I no longer need their code. And just to just to recap, the prompt request was him saying that he doesn't like to get poll requests and said he would rather see the prompt because he will run the prompt
Starting point is 00:25:53 or he will tweak it and it will generate it in the style that... For me, it's less about like, I want to see the prompt as it like, what is it supposed to be doing? And now that we understand... Because like, actually, in many ways, I think, like, the interesting part is, like, often you don't really fully know what you wanted to do in the first place. And so, like, the act of creating clarifies what you really want to do. And so like that part is highly valuable. Often the approach and the code that comes out of it is not what an engineer with sufficient seniority would have done.
Starting point is 00:26:21 So it's not like I want your prompt so that I can reclank my clanker so that it does it slightly better. But more than like, now that we know what we wanted to build, probably faster for me to start. Yeah. And I also kind of disagree with Peter and I just need your prompt. I actually value seeing a terrible implementation of something. Like if I get a pool request and most of the pool requests,
Starting point is 00:26:42 we get that on the Pi repository are made by agents without a lot of human touch, let's say. Then I immediately know, okay, this is going to be garbage. But it's valuable garbage because someone has put in at least a minimum amount of thought instructing their agent to create this pull request. And I get to see how a shitty implementation of what they wanted to build looks like. And I get to, I don't need to waste my own time on trying that out. So somebody else tried it out already, that the naive dumb agent, do the thing, do no mistakes version.
Starting point is 00:27:14 And that saves me time. I'm not saying I like pull requests by agents because they're terrible and they're all to close them now. But they have value. It's not just a prompt. It's on an exponential, right? I mean, it's a sigmoid eventually always because of thermodynamics. But I think we're going to find out way earlier than in previous cycles that this is
Starting point is 00:27:34 a bad idea. That's a good news. What I think is going to be interesting. And I don't know the answer to this. But I write this fascinating retelling of the British industrial revolution and how it changed the textile industry. The industry revolution, yeah. Yeah, and so the general thesis under the article was like every time something ahead of the pipeline got optimized,
Starting point is 00:27:55 it created an incentive downstream of the whole thing to create something, right? So like in the beginning, like if you can weave the thing faster, then eventually you need to have garn that can be weaved at faster speeds, then eventually you need to, everything sort of turned the bottleneck all the way down. And like ultimately the biggest bottleneck in the entire thing had turned out to be what I think is actually the next bottleneck we're hitting in engineering, which is like at one point you made a shirt
Starting point is 00:28:20 and if you didn't like the shirt, you went back to the person that made it and they fixed it up for you. And so the actual thing was like if the shirt is bad, nobody cares about any more who destroyed the shirt in the process. Is it just going to get a new one? Right. Like the responsibility actually went from anyone in this chain to the entire factory as a whole
Starting point is 00:28:37 doesn't have to carry responsibility anymore because we have commoditized the whole thing so much that you don't have to do this. And take the engineering approach of it. It's like a pretty significant part of running a company and running a service is like running it reliably. And so you have these postmortems on incidents to figure out like what went wrong in the process.
Starting point is 00:28:57 And you go back and fix the shirt. Yeah. And the thing is like we are running all on this idea that every engineer that sort of is in this creation process that ultimate led up carries some responsibility. and that we're going to that person and not saying to blame that person, but to figure out like, why did you do wrong here?
Starting point is 00:29:17 And so like if you do, if it like the machine now produces stuff at like 10 times of speed, the responsibility thing does not scale in the same way because the machine cannot yet be responsible. And I don't actually know if there is a future where you can abstract away human failure so much in how we run engineering
Starting point is 00:29:34 that now the entire company now no longer cares about who signed off on a poll request or something like that we automated in the same way I think as we are sort of automating t-shirt creation. I just don't yet see that. So here's the thing.
Starting point is 00:29:50 I think one thing we software engineers or IT people underestimate is just how freaking complex the world is and how much human squishiness is in each little nook and granny and corner, right? So we were thinking, oh, we were now able to automate that thing.
Starting point is 00:30:07 Now we can automate everything. like every bit of knowledge work. But we as software engineers are so bad at becoming domain experts that we don't see all the non-machine parts that go into a workflow. And we're running to the same fallacy here again. We're seeing models doing incredible things. I'm not disputing that.
Starting point is 00:30:25 Like this is, for me, this is like, whoa. Basically all my research in the 2000s is null, null and void because transformers can do all the things. But we are overextending that to everything. like we always do in software. Like we did in at tech. Yeah, we have tablets in classrooms now. Sure, now it's soft.
Starting point is 00:30:42 Education is soft because we have now computers. Well, in fact, I've heard, I don't know which country it was, but they're now rolling back. Sweden, they're taking the tablets out from the classroom. It turns out if you do some scientific investigations into the tactics and effects on pupils, if you do, just throw a bunch of tablets into a classroom, close it and hope for the best.
Starting point is 00:31:02 Turns out the best is terrible. So, yeah, for me, I think the biggest takeaway in the past two to three years is the hype is terrible because it dehumanizes everything. And I want to not be part of that circus. Well, speaking of not wanting to be part of the circus, let's talk about Pi, which is a very popular. Let me get my clown nose. And also minimalist coding agent. Can we start with the backstory of why you decided to build Pi at a time where they were already. already agent harnesses around, right?
Starting point is 00:31:39 Because they were suboptimal. Tell me more. Yeah, sure. So I was a believer in cloud code just because they kind of created that whole genre through the invention of a genetic search. I mean invention. There were precursors to that and shoals of giants and so on, but they were the first that packaged it up in a really compelling package.
Starting point is 00:32:02 And at the time, that fit my workflow really well. It was simple. It was predictive, saw the LLM juristic nature or stochastic nature of being kind of unpredictable. But everything around the LLM
Starting point is 00:32:16 was kind of nice and tidy and easy to understand. You were a happy user of Claw code, right? I was super happy. I was proselytizing it. But eventually the team started dog fooding and getting more and more tokens,
Starting point is 00:32:28 I guess. And kind of increased velocity and team size. And with that came more features and much, much, much more bugs and I personally like simple tools that are stable that I can rely on even if they have non-deterministic parts but all the deterministic parts should be as stable as possible and that was just not the experience with Claude code around summer 2025. So I kind of soured on that real hard.
Starting point is 00:32:52 Was it bugs? Was it unexpected behaviors? So they take away your control of the context. They would inject stuff behind your back which is bad and then your workflows that used to work stop working because there's now a system remorse. that you don't even see in the UI that will modify the behavior of the model. They would also do this to the system prompt. I would reverse engineered, I mean, I wouldn't call opening an obfuscated JavaScript file and unobfuscating it reverse engineering coming from a more low-level background. But I reverse-engineered cloud code during the summer of 2025 and build a little service
Starting point is 00:33:26 where I can track the progression or evolution of the system prompt and tool definitions in cloud code. And it's like every release it was like messing with stuff. cchistory.marsechner.80 if you want to see that. And yeah, that just messed with my workflows and I don't appreciate that. If I commit to a development tool, I wanted to be a stable, reliable thing, like a hammer. I don't want my hammer to break a different spot every day. Yeah, that's terrible. So that's what happened with Claude.
Starting point is 00:33:53 But again, I'm not, like, I'm not roasting the team, I think there, some of them are really nice people I got to know on the internet. They're just dog fooding and that's perfectly fine. We need somebody who like goes to the full velocity kind of way. But I don't want to work with a tool like that. Yep. Because again, get work done. It sounds like the move fast and break things.
Starting point is 00:34:11 The break things was not for you. No. And then I looked into alternatives and Amp and droid came out around that time, I think. Pretty early in Tucson, 25. I don't remember. Amp was earlier, right? I was very early. I think they sort of spun off from the same experience of taking,
Starting point is 00:34:31 because I think Amp was around when Claude Cod came out. I'm pretty sure. Around that time, yeah. In any case, I looked into those harnesses and they were super good. They were just super expensive as well. Because none of them could basically use what made CloudCode enticing on top of it being a cool tool on the subscription. And that works in an enterprise setting where you're paying by token anyways. But it doesn't work for the small tinkerer in the garage.
Starting point is 00:34:57 While I'm not a small tinkerer in the garage in the financial sense anymore, I kind of still relate to that community. and I would like to use my subscription with something, so I looked into open source alternatives and found open code. But while that kind of wipes me with my OSS roots, it too did stuff to the context that I didn't appreciate behind my back.
Starting point is 00:35:17 Pruning tool results after a certain amount of tool result token output or asking a let's be server after every single edit the model makes if there is an error. Yes, there will be an error because the model isn't done yet with its work, so the code doesn't compile, so the LSP server will...
Starting point is 00:35:37 So like reaching out LSP, the language... Language server protocol server, yes. So when you go into BS code and you type some TypeScript, you have in the bottom some error diagnostics, and that comes from an LSP server for TypeScript, and open code runs an LSP server on your behalf in the background and feeds the model with diagnostics from that server
Starting point is 00:35:58 on every edit. We as programmers, how do we work, right? We go into one of more files, we edit line after line after line, and only then look at the errors that resulted from that. In open code's case or in other harnesses cases that also support LSP, the model calls an edit tool to change lines. And they would inject the diagnostics after every edit call. And that's just not smart, because now you're confusing the model with,
Starting point is 00:36:23 you have an error, you have an error, you have an error on the mouse, like, yeah, I know, I know, I'm not done yet, oh. It's not great. Anyways, TLDR, open code wasn't for me. either. It was also, I had to fork it to modify it, which I don't think should be necessary. So then I just thought, how hard can it be? I built my own little thing. And then your own little thing is pretty minimalistic. What does it use? What's the basics of pie? The basics of pie are my own abstraction over all the LLM provider APIs, because I didn't like the Bursale SDK, the Bursale
Starting point is 00:36:55 ISDK for various reasons. Armin kind of wrote a blog post eventually about that as well. It's obviously Good to use. Lots of people use it. It just didn't fit my old man sense of abstraction. This is the beauty of software, and especially open source, you can build your own always. Yeah. And now with agents, you can even do it faster and produce terrible complex software. Now, so I built an abstraction with that, and I built a little abstraction for a generalized agent with tool calling and streaming all of that.
Starting point is 00:37:25 I built a bespoke little tool that doesn't flicker or not a lot. And then I tied it all together into a coding agent that looks like. like clot code or codex or whatever we have. That's it. And the extensibility comes from the fact that this minimal core has so many hook points that you can basically hook into with a simple TypeScript module. That gets loaded into the same node process. And that allows you to do things like provide the LLM with custom tools,
Starting point is 00:37:51 do your own compaction implementation, fully revant the toi itself. You can modify everything in the TIE. So if you have a special... The terminal UI. Yes, exactly. If you want the to behave differently for a specific workflow you have, like say you're non-techie, you can change the to be whatever you need as a non-techie.
Starting point is 00:38:12 I have a couple of non-techie friends that did that because they don't need to know how to build this. They can just ask Pi to build it, and Pi will modify itself. Oh, so this is a thing, right? So you can ask Pi to modify itself because of the extension points and it can write code that extends itself. And it's trivial, but it's, a big unlock.
Starting point is 00:38:33 Is this what you meant when you said that? Or open code, you needed to fork it to modify it. It doesn't have this. It does have a plugin system, but there's not a lot of extension points and it was very rigid. I think they changed it recently. I think it's much more open now. I haven't kept up with it, but it might be better now. So I guess PY stars has this very minimalistic thing.
Starting point is 00:38:52 As I understand the tools it has is read, write, edit bash. It's all you need. That's it. And then you can actually start to make it your own like, okay, like, what are examples of people would add? Pi doesn't have MCP, people just ask Pi to build MCP support into Pi. Pi doesn't have a plan mode, arming goes, and my plan mode must be fantastic bespoke and super special.
Starting point is 00:39:13 I don't have a plan most. Yeah. But he has like five implementations of a plan mode until he realized plan mode is entirely useless. Other people just like messing with the UI and making it their own, like a different visual style of the editor box where you enter your prompt stuff, like trivial stuff, more cosmetic stuff. Other people
Starting point is 00:39:31 have rejiggered it for a full-blown RL environment for open weights models where they use PIE as the agent that does part of the RL execution environment. So it's, you can do anything really. What drew me to it beyond like actually
Starting point is 00:39:47 using the library abstraction was in fact the custom tools part. One moment for me was over Christmas again like many people had some time and I tried to build other things and I and Peter was talking to me in in November that he's like vibing without looking at code more or less I don't know exactly how he said like he's like he can do this now like okay I want to build a thing where I don't look at the code I wanted it's not look like
Starting point is 00:40:13 slop I thought like I want I wanted a version of it where like afterwards like even though I don't really look at the code it should look like what I would have written and so and I want to make a game and so then I basically started the whole experience of like a just basic pie I was like We want to build a game, but actually, before we build a game, I want you to set up the code base in a way that you can validate the changes that you're making, but also I can see them. Like a two-prong kind of approach. Like I want to be in the loop, but also have the agent be able to validate itself.
Starting point is 00:40:46 And what sort of emerged out of that was, well, first of all, like it built itself some debugging tools into the game so you can make screenshots and run a simulation and sort of dump out state and read it again. But also Pi can show image. images in the Tui and and and I added so a bunch of like I talked with the twanker to figure out like what would be interesting things to do but we we ended up having like a all these screenshots I can tap through quickly in the UI or I can pile also this great feature I can reverse to an earlier state in the conversation and then it can branch within the conversation to build a bunch
Starting point is 00:41:19 of stuff around that and because like these these sessions especially with screenshots and it became very token inefficient very quickly was actually one of the other things that pie was rather quickly rather good at was having a lot of screenshots in it. Because OpenClaw, people had a lot of screenshots in the chats and OpenClaw is using Pye. Yeah. So we had to do it. But having this, like it felt really
Starting point is 00:41:39 magical for me to actually treat the problem as I don't know what the right way of engineering here is, but very clearly part of it is like I should be in a loop so we can figure out like how to specifically for the problem at hand, do that. And it turned out like for web
Starting point is 00:41:55 project and computer games and some of the other I tried, they're kind of different. But very many of them are sort of come down to a similar thing where like the agent interacts now with my program and should do the most optimal way. And I want to interact with it
Starting point is 00:42:11 in conjunction with it interacting with the program. And the entire experience should be as little confusing as possible to both me as a human and to the agent. And I found it very, very fascinating just to see how that emerges. Or like your tool all of a sudden
Starting point is 00:42:27 when you launch it in this program, looks and feels different than if you launch it in the other program. I realize this point Armament made just a few seconds ago that AI works best when the engineer stays in the loop and the system can actually validate what changed. And this is a great time to mention our season sponsor, Sonar. AI can now generate code faster than you can verify it. Sonar, the makers of Sonar Cube, ceases leading to a serious gap in verification. With the rise of coding agents autonomously writing code, verification is no longer nice to have. While the latest coding models are extremely intelligent, they also are error prone, and they don't fully understand your code base and your context or your objectives. This is why verification must be mandatory and agenic
Starting point is 00:43:10 workflows. SonarQ provides a zero-trust, multilayered approach to code verification that is consistent and repeatable. It analyzes semantic syntax, data flows, and architectural boundaries at agent speed. Acting as a critical trust and verification layer before any code reaches production. Covering 40-Bust languages and 7,500 issue types, SonarCube is the most comprehensive code verification platform available. And with easy integration via MCP, CLI, and hooks, it fits right into your existing AI tool chain. Let agents move fast and have SonarCube as the independent, multi-layered verification for safe,
Starting point is 00:43:45 reliable, and auditable agentic development. Head to SonarSource.com slash pragmatic to start verifying your agentic workflow today. I'd also like to talk about our presenting sponsor, Statsig. Statsig built a unified platform that enables both experimentation and continuous shipping. Built in experimentation means that every role that automatically becomes a learning opportunity with proper statistical analysis showing you exactly how features impact your metrics. Feature flags let you ship continuously with confidence. And because it's all in one platform with the same product data, teams across your organization can collaborate and make data-driven decisions. To learn more, head to
Starting point is 00:44:22 Static.com slash pragmatic. With this, let's get back to the episode into the topic of general versus purpose-made tools. Yeah, I mean, I spend a lot of my youth on construction sites to earn money, and you don't use a hammer for all your problems at a construction site. You have a screwdriver, you have your hammer, you have your drill, you have whatever.
Starting point is 00:44:42 And I think in engineering, it's kind of the same. I'm not using the same tool for every task I do as an engineer. So now if I use an agent, I don't want a general agent for every, task per se. I want a specialized thing where I know the performance will be top notch for that specific task because we built the harness in the way that the agent can be most affected at this task just because of the construction of the way the harness is constructed. And that's what I wanted to enable with Pi. That said, I'm probably the person that has the least amount of modifications
Starting point is 00:45:12 in Pi. I have like two extensions that I use and they're trivial. They're basically just, if you see a URL that looks like a GitHub issue or pull request thing, pull down the details via the GitHub a and display me a small little widget on top of the editor that gives me the issue title, the author account and a link to the issue. That's basically all I do. Well, it might work for you as a minimalist. Yeah, I mean, that's how I work on the Pi Mono repository because I might have two or three of sessions open in which I process an issue or pull request.
Starting point is 00:45:43 That way I remember what the session was about. But sounds like you also made your Pi for that for working on the Pi Mono, a specific one And if you were working on a, if you went back to building games, you'd probably have a different. I never thought of the fact that you might want a different harness for a different task. I guess we just kind of assume that most developers, you work on your main thing at work. You might have a site project and just experiment, experiment, whatever. But this point in this, I wonder if this is a new thing that we could never have. We could never have custom tools for a project.
Starting point is 00:46:18 That just sounds crazy, you know. Here's his stuff. Like, my intuition is this. I think where we are going is software that modifies itself on behalf of the user's wishes and needs. And the agents can do that now if you give them enough rope to modify themselves. And I think with Pi, that is my first foray
Starting point is 00:46:35 into this kind of self-modifiable, mailable thing, just for the coding agent sector, but I think this actually can be extended to all kind of knowledge work. For specific tasks within the broader set of knowledge work, obviously dehumanization. on, you know. But yeah, the next plan here is actually to have an alternative user interface to the toi, because the two is obviously limited. And the best alternative stack is obviously
Starting point is 00:47:02 the web because it works everywhere and can do anything. So once I have that built out, then it really becomes interesting because then you're not limited anymore to the line waste or rendering of a terminal. Now you can do really, really interesting stuff. And so yeah, we'll see how that works out. And one reason that I learned about Pi. Before I knew that it was as minimalist interfaces, how OpenClaw is using Pi. How did that come? From November, hanging out and reviewing each other's blog posts
Starting point is 00:47:33 and just throwing ideas at each other. And in October, I started building out Pi and Peter started billing out VA Relay, his little WhatsApp assistant. Oh, that's how it started. Yeah. And he was in search of a genticore he could reuse or copy. I think it started out by him taking Pi
Starting point is 00:47:51 and cloning it and calling it tau and then modifying it, but eventually he got tired of having to maintain that, so he just said, I'm going to use your stuff. That's how it ended up being. I wouldn't have compaction if it weren't for open claw. I specifically built that because Peter was crying in chat, and, I need a compaction. Okay, you get compaction.
Starting point is 00:48:12 But I'm going to tell all my users, don't use compaction. It's bad for you. Yeah, but that's, I guess, the beauty of a building on top of open software one another, right? I mean, it has pros and cons, yes. I now get to enjoy all the open claw instances that think bugs in OpenClaw are actually PiBugs. So they autonomously send me a gazillion issues and pull requests
Starting point is 00:48:32 without their users probably even knowing, and I get to deal with that in my open source. So that's a negative side effect. Well, so you're really on the receiving end of this, I guess. I mean, just like OpenClau itself is, which is much more exposed to this problem. I mean, there are tens of thousands of issues now, and there's no way they can get a good,
Starting point is 00:48:50 creep on that. But how are you dealing with the fact that you now have OpenClaw just AI autonomously opening things on your repo as a maintainer? Do you build tools to battle this and try to close them out? Build a tool for OpenClauans which embeds issue and pull requests into a 3D
Starting point is 00:49:06 space so I can see the clusters of similar things that agents would have sent to the repository and then I can bulk select things and close them out in Oh really? So you actually have a really like visualization. OpenClau for context at I think it's less crazy now,
Starting point is 00:49:23 but end of December to I think mid-February. I mean, it was exploding, obviously, but like this explosion almost like directly translated to, I was on this repo refreshing pull request and the number went up. Yeah. We actually tried to contribute and help out Peter a little bit, but I immediately gave up. I didn't know how to do anything useful.
Starting point is 00:49:46 I was looking at this and I was like, this is a type of software engineing I'm just not used to. I would fix two things and spend an hour on them and then five minutes after I committed and pushed it. Some Clanker comes along and just reverts my fixes. And this is not how I... Can we talk about the name of the name Clanker?
Starting point is 00:50:04 Oh, sure. So Clone Wars, Star Wars. I actually never watched it. But kids of friends of mine watched it a lot while we were visiting them. So I kind of, through osmosis, got the lore. and there is an army of robots and the Jettas would call them clankers or people who would call them clankers because when they move they clank, clank, clank, yeah, that's the origin of that. So an AI, a droid. Yeah, exactly, yeah.
Starting point is 00:50:32 But coming back to the, how do you deal with the influx of agentic pull requests and issues, I just auto-close every pull request. A human agent doesn't matter. What I do is, if I haven't had contact with you previously, my GitHub workflow knows about this because if you had, you're in a file in my Git repository, your account name. So if you're not in there and you send me a pool request, your pool request gets auto-closed.
Starting point is 00:50:58 And then my little workflow posts a comment under your pool request that says, hey, thanks so much for contributing, really appreciate it. Could you please open an issue in a human voice, no longer than a screen's worth of text? And if I like it, I type, looks good to me. and then that account name gets put into the file. And the next time they send a pool request, they pass. And it turns out agents don't see the comment my GitHub Workflow
Starting point is 00:51:24 posts underneath their pool requests. So this is a great filter for filtering out agents and keeping the humans safe more or less from. It's interesting. I wonder if this might be an unavoidable future where we just need a way to separate. is this coming from a human with an intent or an AI? I don't necessarily care.
Starting point is 00:51:48 If it were actually a good PR, then if it came from a machine, it's actually fine-ish. I think what's interesting in Pai is like, an open-closure more so, it accumulates pull requests. Well, actually, there was no intentionality behind it at all. And so the person that dispatched the machine
Starting point is 00:52:07 didn't actually care that much about it. But didn't even know about it. or I didn't even know about it. And I've done open source for many years. And there was also, there was a big difference between someone that's any pull request up or like an issue. And it's like, hey, please fix this. But actually didn't care enough to even reply to questions anymore. Like, there's not uncommon.
Starting point is 00:52:29 And then you don't actually have to fix that. But you have to close it out because, like maybe it's still useful input. But like, it clearly that person wasn't caring enough. And if the pull request is even worse now because they come in so quickly, that many of them cannot be merged anyways without manual resolution of the conflict. And they said there's a lack of back
Starting point is 00:52:48 pressure mechanism because even I as a human if I see there's like 500 pull requests open, I probably will not contribute to this thing now. Because at worst I will make it worse. And I think previously in open source you had the people who would just send issues
Starting point is 00:53:04 and be very entitled and say you're the worst person on the planet if you don't fix my little issue. But that's fine. be handled and poor requests were kind of special because it needed a human to invest quite a bit of time to reduce them and you don't have that anymore you just have people oh this this should be easy agent please do something make no mistakes send it to this repository and that's just not going to happen so basically what we need are bottlenecks i'm not necessarily i don't necessarily need human verification or a verification that you're human i just need a bottleneck that allows me to
Starting point is 00:53:35 process the amount of incoming things as a human because in order for Pi to not deteriorate into a pile of garbage, I still believe that it needs me and other capable people reviewing at least the important code. And for that, I need bottlenecks because otherwise I can't deal with. It's a second law of thermodynamics, right? It's like everything degrades towards chaos and you have to put extra energy into,
Starting point is 00:54:00 to keep it away from this outcome. And we don't see and feel like the pain of the code base anymore if we stop looking at it. And people don't feel the pain or like they feel no restraint anymore. And it's, the issues are also interesting because on the one hand,
Starting point is 00:54:20 it is something great about someone doing an investigation and sending you a description of that. That can be good and can be bad, but it look very similar. Like, it takes quite a bit of energy to tell apart a good and a bad AI generated issue request. And unfortunately, like, most of them are not great.
Starting point is 00:54:41 But some of them are actually good. And that's also kind of, it's weird. Like, all of it is weird. I really don't know what the feature of open sources in many ways because, like, a lot of open source really worked because people piled out on hard problems and so they congregated around it
Starting point is 00:54:56 and said, like, now we need to have a good database. So we're going to put all this energy on building a good database. And so the value of open source came from, there's some hard problems, and we're going to throw our energy together and we're trying to figure out how to solve it. And now it feels like open source is all about like throwing stuff up.
Starting point is 00:55:14 What really grinded me so mad was people, particularly, like a lot of agentic engineering right now is like building more stuff for agentic engineering. So it's like it's Euberus or Euberus or I call it. And I see this tweet and it's like, oh, I solved problem X where I see and here is my solution for it. And you click on this thing as like it's 48 hours old. That person probably never used the thing that they built. I would like to suggest to the viewership to look at Arvin's GitHub account over the last year and what happened there.
Starting point is 00:55:44 Yeah, I built a lot of this stuff but I don't then go on Twitter and say like, hey, I solved the problem. It's like I have a shit ton of vibe slop on my GitHub account and I wish I could market differently because maybe there's some utility in it but unless you're going to actually
Starting point is 00:55:59 have that code base still be there a year, a year and a half from now and someone is still using it. The utility of that is actually not very validated in that way. And there's so many markers and metrics you can look at now for GitHub that really demonstrate this explosive growth of it. But if you were to then maybe find somewhat a number to see, like, how many of the things
Starting point is 00:56:23 that are being created are actually turning into like really fundamental pieces. It can sustain open source communities that can actually deliver this value that scales amazingly. We have actually created many Vipe engine. engineered projects that have become that. But I like how you mention energy and how open source always worked if we just think pre-AI, let's say Linux, the most successful or widely used open source project. It has both an energy and a structure.
Starting point is 00:56:54 You know, people come in with intent that they want to add something. They have a process where it goes through. There's human trust at every level. There's a little pyramid. And in the end, it all goes back. Exchange request goes up one level. And in the end, Linus does the cut. But there's a lot of energy, there's a lot of intent.
Starting point is 00:57:11 There's a lot of humans. There's a lot of humans. And it was always about human energy. And now we suddenly have this AI, which it's just tokens. Right now, who knows how much they're subsidized or not or it's just machines doing. And then suddenly, you know, they create plausible things that look like human energy and it's hard to differentiate. And suddenly, just like throws this wrench. Actually, disagree.
Starting point is 00:57:34 I don't think a lot has changed for open source. Okay. the volume has changed. No, yes, but that's just a number. The amount of, as you said, the amount of actually useful and maintained projects probably not changed a lot. So you're saying that the ones that were there, they're so useful and maintained? Not even the ones that were there, there might, I mean, there's a specific rate of new
Starting point is 00:57:54 open source project that survive longer than two weeks. That's always been the case, right? So now we just have more projects that die after two days than before. But we still have the same amount of, projects that will have a long-term viability just because there are humans that actually care to maintain the thing over a long time, build a community
Starting point is 00:58:14 of humans that support the entire thing, build an ecosystem around the entire open source project. You're saying you're not believer in the moldbook. No. I mean, good job, matter. Putting that up. Super useful. No, I think
Starting point is 00:58:30 at the end of the day, we're kind of freaking out when we don't actually need to, because apart from the fact that I personally cannot generate code faster to a speed of light. For me building an open source project, and that entails not just the code, but the community around it, the spirit around it, the ecosystem around it, nothing changed.
Starting point is 00:58:47 What changed is mechanical parts. I need the bottlenecks to deal with the influx of exponentially growing agent's pull request, whatever. GitHub itself is under immense pressure because now it's not just humans hammering their infra. It's now billions or millions of open claw instances hammering their infra. Everybody complains about GitHub going down.
Starting point is 00:59:08 I actually think they're doing a pretty good job. Like, that's a lot of traffic that's coming their way since basically Christmas. It's basically open call. So, yeah, I would be a little bit more optimistic. We're just indeed messing around and finding out stage at the moment and everybody wants tokens to be a KPI, just like lines of code used to be a KPI. We've seen this. Speaking around of things that don't change and messing around and finding out,
Starting point is 00:59:35 you wrote a tweet or you wrote somewhere that your biggest enemy is complexity, it's also your agent's biggest enemy. Can we talk about that? Very simple. If I have 600 lines of code, code BIS, and my agent can at best be effective up to a context window size of around 200,000 tokens, how much of the code can the agency? A third, right? Right.
Starting point is 00:59:58 if you manage to get all the relevant code for a task into that context window you're probably okay although that is a separate project an information retrieval problem which is not solved and which agenting search also doesn't solve that is are you sure that the agent finds all the relevant code it needs to find to fulfill a thing that's also where all the garbage code comes from because it doesn't see all the thing in this case, let's assume the best case, information retrieval is solved, everything fits into a context, agent does a good job, okay? That's not the reality we're living in, because now the agents spit out so much code that they themselves cannot possibly read into their context on a new
Starting point is 01:00:42 task anymore. You know what I mean? Yeah, they fill up their own context window. Yeah, exactly. The complexity they add is their own worst enemy because eventually the code will be so big and so complicated and so interconnected, that the agent has absolutely no way on a technical level to ingest all the context it needs to do the new task. And I would like to point out that the agent has learned all of this garbage from the internet and from us, because on the internet there's all our old code. While there are some pearls, there's also a lot of swine. Because we have a gazillion GitHub projects from the old days where we just try it out
Starting point is 01:01:20 things and because instances like Linux or any other really well-maintained and well-written open-source project are minuscule compared to all the rest of the garbage and a machine learning model will kind of converge towards not the well simplified to the mean right and what is the mean then it's not the handful comparatively of excellently engineered projects it's all the garbage on the internet all the cargo culting all the trend type of the day kind of stuff And that's what we get when we let the agents do all the things for us. Yeah. So we have this problem of things are getting more complex, which slows agents down,
Starting point is 01:01:58 which will in fact impact quality, which we were just talking about. But Armin, now that you're building your own startup, you two of you're building your startup now, how are you, are you, and you're working with agents, right? And they will have these things. How are you dealing with generating code building products, balancing quality, tagged up, complex. Badly.
Starting point is 01:02:23 I think that... We're coping. We're not dealing. I don't know if I wrote this in a blog. I definitely have it on my slides for the conference here. It was I enjoyed the time from April to about October immensely. Because it felt like I can do so much, but also like there was no heightened expectation. Like the world has not yet gotten used to this idea that everything has.
Starting point is 01:02:50 to now also move at 10 times the speed. And there was a moment of time where I felt like, like we worked in this Vipe Tunnel thing in the beginning and I was like, it felt so much fun because like, I have time now to play with the kids and I just prompted a little bit of my phone.
Starting point is 01:03:05 And like it felt. Vipe tunnel was where you could set up with your phone talking with your agent on the machine where it wasn't as easy. Terminal basically. Yeah. And it's not that we did much with it, but like it had this like happy vibe.
Starting point is 01:03:18 And like I know that I spent too much time on a computer. but I didn't feel any pressure. But now it's like we're collectively feeling like everything has to ship faster. It has to iterate faster. Like the baseline that we want to achieve in terms of fidelity and everything has to be higher. And so now it feels very stressful. Even in your own startup.
Starting point is 01:03:41 Yeah, because to some degree you cannot, like you can be the most stoic person in the world and it's still going to get at you in a way that I'm slowly learning. to work with my own emotions in a way on dealing with this. But I find it very, very hard in a way to, because I was used to things work in a certain way and I knew how I do some stuff. And then I fell a little bit too much in the trap of like giving in to the machine and actually doing things in a way that I normally wouldn't have done things.
Starting point is 01:04:13 That you regret. It's definitely a gentic regret. Jentic regret, yeah. And so, like, quite frankly, the answer is like I feel like, now with a little bit of power of hindsight, learned some things that I wish I would have learned probably November. Tell us. Well, I mean, like, a lot of it is like really the recognition that if you,
Starting point is 01:04:33 there is no back channel to me or to any other engineer. When under normal circumstances, there was a back channel. That was this, this feeling of like things are not quite right in the code base. Like there was this, now the changes harder and like the complexity, like, you sort of see it in the complexity of the pull request getting higher, but like if you rubber stamp it then like what's what's the back channel there and so like this this mechanism this back pressure this friction in the code base you don't feel when you work with the agent i think there's a way to kind of measure it and um like if i scan through my sessions on a project
Starting point is 01:05:07 from start to current date i think the frequency of curse words increases because the agent starts messing up more because it itself cannot deal with the complexity or the edge of the project and i would be actually really interested in whether it is measurable because I feel it in most of my projects now that occurs a lot more. But you mentioned friction in the software. You didn't say tech-de-up, you didn't say complexity. What is this friction? Because I don't remember us talking about this pre-AI at all. So I found this ironically kind of funny. And it's kind of sad. I will not name any names, but there was a, what I assumed was an incident related, at least in parts of the gigantic engineering on a company
Starting point is 01:05:50 where they shipped out a configuration change that ultimately result in a security issue. And look, things happen. But the link that I saw on this had the social preview of that company's tagline. And the tagline was ship without friction. And that really gave me pause because I know as an engineer, we used to talk about like you got to get rid of like all the things in the way so that you feel happy shipping stuff. But there are always where changes where you really wanted to think.
Starting point is 01:06:17 is like, do you want to drop the database? Like, do you want to merge this migration, which might take a table lock that could potentially take you down, right? It's like there's, there's moments every once in a while what you really, you were really supposed to think and people created checklists or people created, like mechanical gates
Starting point is 01:06:36 that would, where you would have to confirm something. There's certain things that we used to put, particularly if you run a SaaS company, did it put stuff in? So to slow things, down or in some of the best engineering teams in order to mature service you have to define an SLO you have to define like expectations and like if your service is supposed to be critical but like there's some other stuff that unlocks on the sort of tree of requirements that you
Starting point is 01:07:03 and and like a lot of engineers feel like oh this is also this bureaucracy but like the reality is like if you do this correctly then it saves you time and it like it makes you're happy you're not waking up at 3 o'clock in the morning like all of this is useful It's like friction injected to deliberately stole things down. I guess the easiest example in any decent size company, you have services based on tier based on criticality. The highest tier software now needs to have, let's say, two or three code reviews or an approval from a director to do a configuration change, which again, all slows down. But it's kind of like, we know this is on purpose. By adding this friction, we want you to think, do I want to push through this friction in terms of time?
Starting point is 01:07:46 I'm invested or effort or having to justify things, etc. It makes you think about, do I really want to add this to the codebase if I know that the end effect will be the test to go through this entire chain of hardware is? So we're coming back to saying no to yourself, to avoid pain, going through that process. And then taking on the pain when you know that you have the conviction, you have the backing, you have the confidence as well, right? So typically when it's a higher friction thing, let's say a tier one service or high-one service or highest your service where a director have to sign off. When you're a new joiner on the first day and you don't know the context, you probably know that that's a pretty large ask and you'll probably socialize, get buy in from an experience and say like, oh, this is the right thing. You'll go
Starting point is 01:08:30 with them, right? Back to human dynamics a little bit. And I think the thing is like there's a, there's a very delicate balance in the whole thing. Because it's like you don't want the friction to be just an accident of having created bad develop experience, right? But some things look the same. But they were deliberate, but they were not sufficiently documented, but there's this feeling now like you'll get rid of all the friction so that the agent can be very autonomous so that they can run many of them simultaneously. A lot of it comes from that.
Starting point is 01:09:02 It's like, I, it's like, these things are actually rather slow and the only real time saving that you get from it is parallelism. and so somewhere there is this trap. I feel like a little bit more experienced now in managing the trap, but I don't have the solution for that either. And I will not say that he's an example of code base where I felt like really, really great
Starting point is 01:09:32 about the stuff that I built, except for pre-existing libraries from before Argentic days where I still feel like a strong emotional attachment to them and they're much more careful about doing them than any of the code that we other than pi, to which I don't have access. Oh, no, there's still no right access.
Starting point is 01:09:53 There's a lot of slopping pie, but I try to avoid it in the bits and pieces where I know that's important code. Like we have an HTML export functionality where it takes the current session and just spits out an HTML file that you can then host and GitHub and whatever. I have not looked at a single line of code for that function.
Starting point is 01:10:11 I don't care if it's broken, if it looks right when it comes out. But then there's the agent loop itself or the extension loading mechanism and all of that stuff and that's important. And the way I deal with ensuring that that has, or at least trying to ensure that has high quality is I re-factor mercilessly because that pulls me into the code base. I need to understand what I want to change structurally, not just line per line and syntactically or whatever, I need to understand what's going on
Starting point is 01:10:42 to do a good refactor. And I'm doing that every now and then, like I'm doing now at the moment, prompted by wanting to add a new feature that's currently not possible with the current architecture. Being in the code is the one thing that keeps the code-based quality high
Starting point is 01:10:55 and the complexity low. But that's against the industry wisdom of burning as many token maxing, basically. Yeah, that's an interesting one happening. But you just recently wrote on the same theme, a blog post called, we all need to slow the, down. Can we rehash some of the thing?
Starting point is 01:11:12 And what triggered you? So let's just put it out there. OK, so the basicist is, OK, your agent can now spit out 10 times more code a day than you can. But it also means it spits out 10 times more boos, errors. Even if it has half your error rate, then, OK, it's not 10 times more. It's 5 times more. And it's still more than you would spit out. So the rate of deterioration in your code base has now increased.
Starting point is 01:11:36 And now go dark factory. Now take 100 agents that do this to your code base. What's the end result of that? So that's the first problem, right? You need some way to review all of that code. It now gets generated to fix all the boo-boos. But you can't as a human, because as a human, you're used to spitting out 1.5K lock a day,
Starting point is 01:11:55 and that's about the limit that you can actually review well, right? If your agent spits out 10 times that, no chance you can review that. And not all of that code by the agent might be important, like the HTML export thing, right? But even if the agent speeds are 3 to 5k a day, you have no way of reviewing that in any meaningful sense. And then if you do the armies, the armies, this is interesting. So you call it the dark factory, the idea being that tens or hundreds or thousands of agents, you give them a spec, they go and they break it up, they organize themselves, like the mayor,
Starting point is 01:12:31 and all that jazzed. They have the quality, the QA agent, they have, you know, you give them roles, you give them context and then you give them enormous amounts of tokens and spend and the idea is or the hope is that your software will be done in oh there will be something will be done definitely something's going to be done first your purse and then no yeah sure more power to the people that make that work i can't make it work and the reason i think i can't make it work is because i still care about the quality of my product and i don't care if it's built by by hand or by agent i just want the quality to be good both in terms of how easy it is to maintain it
Starting point is 01:13:09 and add new stuff to it on an developer side and on the user side. All the companies claiming that all of the code are written by agents, yes, we know. Quality is garbage. We feel it in our bones when we use your product. It's garbage. So I don't want that.
Starting point is 01:13:25 And yeah, basically, I think people need to turn around and say, hey, what are we even doing here? We have these wonderful machines now that can take away so much pain from us by doing the stuff we hate doing and doing that really well. Why don't we start by giving up some more free time
Starting point is 01:13:42 to work on the interesting bits and delegating the stuff we know they can do to them on large, like across the entire organization. Find all the things that annoy the out of you and have the Asians automate that for you. And then you suddenly have time to think about what do we actually want to build?
Starting point is 01:14:04 What do our new? users need. And if we decide to build a thing, then we can pull in the agents again and say, and we're going to polish the sh out of that. Because now we have the time and the means and the tools to do an excellent job. But that's not how we're working. We build an army of agents and install beats and make a big spec that hopefully will result in something crazy. But here's the thing. We talked about where did the agents learn their knowledge from, right? The internet. So garbage to mediocre. Now, if you write a spec, What's the best possible spec you can have?
Starting point is 01:14:37 The best possible spec is, well, you define exactly how it should work, you give it test cases. Best possible spec is the software itself. Oh, I see what you mean. Yes. Okay, you write a spec that's not the software itself. So that means there's a lot of planks that need filling in. Yes.
Starting point is 01:14:53 What do you think is the agent got to fill those planks in? Most likely from his training data. And we already identified what the quality of that training data is, right? garbage to me, okay. Well, and even before AI, don't forget, like, Stack Overflow had a really big criticism because there was this thing of, like, well, you control, C, control V from Stack Overflow, and oftentimes there will be some answers where the first answer was either not correct or not correct in many cases. Reg X for email was a good one. You emailed Reg X for email, first page was Stack Overflow, everyone just copied the first solution, and I think underneath number three,
Starting point is 01:15:28 it was said it missed a bunch of cases. But here's the thing, though, I'm not saying agents or Humans are better. They are clearly not. But agents also don't solve that problem. And if you then don't let just one agent that's already 10 times more productive as you do the thing that it's bad at and that you as a human are bettered, but a hundred of those, what do you think is the outcome? It's just very simple math. Let's talk about another controversial topic, MCP versus CLI. Oh, my God. It's coming up. And right now I'm hearing a lot of people really going for a CLI is the future, and I think I'm sitting with two of them. But also MCPs are also really
Starting point is 01:16:08 popular inside of large companies, especially when you talk with a bunch of people working at large companies. It seems MCPs have found a real product market fit inside of larger enterprises. Despite what people might think, I don't actually hate MCP quite as much. Same. Oh, wait.
Starting point is 01:16:24 We have it on recording. Yeah, no, we don't deal in absolutes. We're in SIF. So my fundamental challenge with MCP is that I think, for first of all, the spec is very complex. I think, for it. But it's like, this is just generally how specs happen to be. So it's a bit like the core bar of its time.
Starting point is 01:16:42 So there's an inherent complexity in it. But if you wait to say like, what is it really doing at the end of the day? It's authentication and it's sort of invoking some stuff. And MCP, even theoretically there's structured responses, but MCP for the most part is run some stuff, put stuff back in the context, and then work with it. So it fills your concept very quickly.
Starting point is 01:17:02 and there's a Cloudflare has this code mod MCP which in principle I really like. I have an MCP for testing, which is a JavaScript interpreter that gives me access to the Google API. And between an MCP like this and a skill, there's not a huge difference because the skill also needs to be
Starting point is 01:17:19 in a system prompt so that defines it. But the agents are just very, very, very, very good at running code. And MCP is not quite running code. It's basically rag. It's like input in and do some stuff. And maybe some state transition at the model also doesn't see. But it is in that sense, it's a hard problem to solve.
Starting point is 01:17:40 But it does solve off. It solves a whole bunch of things. I want it to work. I just still don't get it to work like I wish it could work. And my suspicion is still the glue has to be code execution. But because MCP servers are largely not defined in a way that the model actually understands them, I haven't found ways to compose MCP tools reliably. I found ways to make the MCP itself be composable
Starting point is 01:18:09 by having the MCP be one tool run code, but I haven't found ways to then orchestrate larger ones. I wanted to work, and I think it has found its niche, and I don't think it's going to go away. I think it's just a victim of its own success, really. When the whole thing started, I think it was in October 24, it was more or less, solution to get external services into consumer facing chat apps.
Starting point is 01:18:37 Connect your emails, connect your OneDrive, connect your whatever. Really much. And then IDs also took it over because it was convenient. Yeah. The cursors, the windsurfed. Yeah. But I think the origin was basically the consumer side, not the developer side. And I think that's a totally great use case.
Starting point is 01:18:53 I don't want my mom to having mess around with cogeneration or whatever to invoke some API or call some API and so on. perfectly fine use case. And then developers side also picked it up and thought, oh, this is a great way to provide tools to my LLM. Tools as in the system problem somewhere there is, if you want to call this tool, provide this JSON payload, and you get this thing back, right?
Starting point is 01:19:17 And that kind of felt right at the time, because if you read atropics documentation, they would say, our models can deal with about 30 to 40 tools in the context. And even that wasn't the case like that. 12, 20, they would just break down, but it doesn't matter. But there was still like a, yeah, this can work if you kind of keep it small and contained and very specific to your use case.
Starting point is 01:19:41 And then people started building MCP servers that would just basically map an entire Open API spec into a gazillion tools. And that's where it all fell apart. So that's the first problem. Very bad MCP servers from big corporations that thought, we need this now. What's the fastest thing we can build? I just push the Open API spec of our API. through this thing and make it an MCP.
Starting point is 01:20:03 That's garbage. The second problem is that it's inherently non-composable. If you want to combine the tool out, the MCP tool outputs of two different servers, they need to go through the context. The model itself needs to do the data transformation, the composition of multiple pieces of data fetched through. And then compared to this with a CLI, it's a pipe, right? Exactly. The model only sees the entry salt.
Starting point is 01:20:28 And it is super free in how it must. that data. And that's also the idea behind code mode. Basically, it's a hack. It's basically, okay, we now have MCP, we know it doesn't work for this specific use because we have multiple sources of data and you want to combine them, but don't kind of pull that through the context. So let's build code mode. And code mode is basically, we take all the MCP service. We expose that as functions in TypeScript and then the model can actually just write some code that calls the MCP servers and then does the composition in the code. It's like how many interactions do we want here? we can just let the model write the code
Starting point is 01:21:02 we don't need the MCP server and then the third part is David from Century is a big proponent of MCP because it's off the off thing and honestly that's again for me super valid but the model itself kind of doesn't make sense anymore
Starting point is 01:21:17 I think there's a world for MCP 2 which is ironically maybe based more on I said there's a company called stainless which basically generates SDKs out of open air especially And I'm really warming up to the idea of, like, maybe it is an MCP is entirely based on off plus, like, libraries or, like, directly like, HTTP requests against OOF specs, because if you compose it together there.
Starting point is 01:21:47 And I think, like, one of the things that's also, like, kind of underappreciated and sort of is you see, I think if you see Pi do its stuff, because it's kind of transparent of the tool calls that it does, it's kind of magical at times, like, how creative. agents get at large outputs. Like for instance, Pi, when it runs a program in Bash and it produces too many lines of code, it actually only reads, I don't know what the cutoff is,
Starting point is 01:22:10 but it reads the first couple and it's like, oh, if you want the rest of the file, it's 20 megabes large and it's in this file. And then the agent's like, oh, 20 megagos, that's too much. I'm going to grab on the file.
Starting point is 01:22:20 Right? And they get really ingenious in how they're interacting with it. And like, an MCB takes that away. The question is like, how would you define MCP in a way where it wouldn't take that away,
Starting point is 01:22:31 where it still has all of that magic and capability and I don't really know the answer because I think it's hard, but off-needs-solving and composability need-solving and I think there's a bright future of that kind of stuff. And also like what Mario said,
Starting point is 01:22:47 if coding agents wouldn't have become so popular, then the idea of code generation code running for like non-code-related problems probably wouldn't have taken off quite as much too. But like the most capable personal agents, OpenClob, being a good example of it, they're just coding agents hidden from you. And then that's just naturally,
Starting point is 01:23:11 some random person who is not a programmer is going to say, how am I going to do this? And the model doesn't say, like, install this MCP. The model says like, okay, I can write a Python script that does it. And so you naturally have this, in the sort of the crazy space, you have the adoption of more code execution
Starting point is 01:23:26 and the compliant enterprise space, you don't have that. There's a different path. And I personally don't think that models are going anywhere else other than code generation, going forward for any kind of a gigantic task. I think that's mostly a function of
Starting point is 01:23:42 there being a lot of training data for code generation and code generation being a very easy means to control computers. So I don't see a different paradigm there coming out of the model labs anytime soon. So I think, taking that as the assumption where the future is going, we just need to figure out how to make code generation kind of work
Starting point is 01:24:02 within an enterprise setting with off and all of the other enterprise things that entails. So let's do a fun trying to predict a year out, which is hard, but in 2027, knowing some of these basics, just again from first principles, where do you think these coding agents might be and the software engineering workflow might be? Basically, this is just like, again, speculation,
Starting point is 01:24:26 we know we cannot predict the future, but where do you think that there'll be a lot of focus in the coming year and we might, in an optimistic case, see some results in tools, and how we work and what's working, what's not working? I have no idea. I honestly have no idea. I could make up something that's probably not going to happen. I think the self-materiality thing is obviously something I believe in.
Starting point is 01:24:48 I think we will see more of that. So like self-mutable software. Yeah. Including the tools themselves with which, we built the software and I think that will expand not only to the tech sector but also to non-tech applications of agentic tools my is it dog years which are time seven is that how it works so that that's that's basically the model i have right now of like how this stuff works it's like when you ask me like what's going to be in in a year it's like seven years right and to me that
Starting point is 01:25:22 makes it incredibly hard to have any sort of predictions about the future because like it's still not one year. Maybe now it's a one year from like people starting to use in cloud code, but it feels like it's much, much longer, much more time behind. And my time has passed. And I think like right now that the closest that I can
Starting point is 01:25:40 imagine is going to be like we we know that code execution and code generation and like this harnessing around it. This is going to be it because reinforcement learning gets more of that data. And my strong hypothesis is that as more and more people are starting to wake
Starting point is 01:25:56 up to this, you can do interesting things with agents, there will be a societal recognition also of how much more dependent you are on basically two companies. And I think we'll have a conversation about that part. We should have a conversation about that part, particularly as Europeans,
Starting point is 01:26:12 because we don't really have these labs over here. And so I hope we have that conversation, but like my best guess is that we'll wake up to the fact that we are now, I mean, engineering teams already now telling me that they have code bases, that they think they couldn't maintain anymore without the machine.
Starting point is 01:26:27 My guess is that one of those companies will be public and and it will be expensive and I think that might actually dominate or at least become a conversation that's much bigger than the question of are you using PIE
Starting point is 01:26:44 or using ClaudeCode code or something like this. I also see we've seen this with, was it Mithers, the new Cloudmon? Oh no, Spat. The new GPD model. they will only give this to select partners. So now we are seeing a split in who can get the best intelligence.
Starting point is 01:27:04 Yep. Or the perceived best intelligence. That'll be interesting dynamics. So both of you are working on popular AI tools. You're building a startup that, of course, you're using AI and it's also around agents. How do you both keep up to date? I've just seen things. And it's not as easy to get me on a hype train as it used.
Starting point is 01:27:25 be, but that comes with age. It's definitely easy not being in San Francisco because I think that just drive me crazy. I hear so many things from my peers over there, and that's just like, yeah, I'm not going to go to San Francisco. Thank you. So having a peaceful environment around you where it's not all about tech might be helpful. It helps having a kid. It helps just going outside, climbing trees, going ice skating, and then looking back at what
Starting point is 01:27:48 you did just half an hour ago and be like, why would I do that? That's just stupid. the detriment of maybe people that are trying to stay in contact with me, I got very good at not muting notifications, not reading emails. And that has in part become necessary, I think, over the last year or so. But it actually turns out that passage of time sometimes clarifies stuff a lot because if it was really necessary, it's going to reach you again. Like, I have an unhealthy Twitter addiction, which I'm not particularly proud of. But in, in, in terms of source of like interesting things that is still a thing.
Starting point is 01:28:29 But I try to now sort of consume it in a form of if it's really, really important, it will stay in the discourse for quite a while and I just wait it out. And if it's if it's there into like three weeks after it originally happened, then probably something to it. And I don't need the three week I start necessarily. But it is honestly, it is really hard. It is really hard to deal with. this because there's a genuine excitement in it. And I feel like my more than 20 years of experience
Starting point is 01:29:01 in that space of software engineering doesn't, it tells me a lot of stuff. But at the same time, it hits you in certain ways where you felt like there will be grounding and there will be something to build on a strong foundation. And now it feels like, well, seemingly everybody else doesn't care about that foundation anymore. So maybe you don't need the foundation. And for quite a while, it works and that is sort of weird. I kind of feel like since we've been fun employed in 2025 and all this started that we had like a head start. Like I see all the excitement the two of us and Peter had in April last year.
Starting point is 01:29:40 Has weaned. Nobody else, no, no, but nobody else at the time has kind of shared that excitement that much. And then the Christmas break came. And now everybody else has that excitement that we had in April, right? So now they are learning groups. Now they are catnipping themselves to immeasurable amounts of lost sleep and at terrible codebases.
Starting point is 01:30:02 And I think it will self-correct because it's not sustainable. Yeah, we did see this as well. I did a deep dive at the pragmatic engineer early March when a lot of people who were very excited in January about all and they start to use the new models, what they can do. They went all in at work or on site projects. in about two months time a lot of them were like hang on it introduced all this complexity it has these things i'm not going as fast as i thought i would be etc so i guess there's just a natural thing where you you have a time anything new right a job anything you have a honeymoon period where you've
Starting point is 01:30:38 got the blinders on which you should by the way and then you start to realize and maybe overcorrect but but there's a natural thing where it in general like it just takes time to see the outcome of your decisions So I'm not worried about all the dog factory and all the software is dead and Suss is dead and all that. I'm terrible if this is just part of the hype machine and that will self-correct. As closing, what's a book that you would recommend and why? Code by Patsalt. Classic.
Starting point is 01:31:07 I just love it. It's just such a great read. It's also for non-techism. The first thing I recommend if anybody ask me, what's your job? I'm pointing at that and it's like, it has much less to do with computers than you think. And I read recently Breakneck, which I unfortunately forgot the author of. It sort of goes a little bit into an exploration of like how China works and maybe Europe and the US are different. And I found it at least thought provoking.
Starting point is 01:31:38 Well, Mario and Arben, thanks a lot for this conversation. It was great to have it in person. So having us. Thank you. This was a really fun conversation thanks to Mario and Armin. The idea of self-modify all software really grew up. on me. Mario said how Pi doesn't have MCP support, plan mode, and many other features that devs would want from it, but you can build it into its own code. So far, it's working. Pi is popular
Starting point is 01:32:00 because it modifies itself. I wonder if and when this concept of self-modifying software thanks to AI will spread outside of just the Dev tool. I also liked how we talked about the observation that agents don't feel pain. But humans do. When a code base gets too complex, the human engineer feels the issues this creates. And this tech deft is what pushes, refactures, and rewrite. But agents simply do not do this. They just keep adding to the complexity. And in the codebase where devs regularly feel the pain of the code base and do something about it,
Starting point is 01:32:32 the quality will probably be also better. And finally, the MCP versus the CLI discussion. This was a good one. MCP is more about offering tools for AI through context, and CLIs allow piping one tool after the other. Both Mario and Armin are more of the fans of the CLI, but in all fairness, MCP has its use cases. for example inside larger companies. The right tool for the right job,
Starting point is 01:32:52 do check out the show notes below for related the Pigmatic Enduring Deep Dives that go in deeper into related topics. If you enjoyed the podcast, please do subscribe on your favorite podcast platform and on YouTube. A special thank you if you also leave a rating for the show. Thanks and see you in the next one.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.