Python Bytes - #496 A lake house in Seattle

Episode Date: September 15, 2026

Topics covered in this episode: Pandas Should Go Extinct Pydantic-pint puts real-world units in your Pydantic models How Libraries Run Rust Inside Python (With PyO3) AWS acquires DuckLabs Extras Jo...ke Watch on YouTube Sponsored by Logfire from Pydantic: pythonbytes.fm/logfire Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: Pandas Should Go Extinct Pandas' slowness pushes teams toward "Big Data" tools (Spark, Databricks) they don't actually need — most workloads never hit true Big Data scale Amazon Redshift telemetry: ~95% of tables are under 100GB, ~87% of queries touch 80GB or less — that's "Medium Data," not Big Data Polars and DuckDB fill that gap: single-machine, fast, no cluster required 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory On a real-world NYC taxi dataset (3GB parquet), pure DuckDB ran 2x faster than pure Pandas while using a fraction of the RAM Bonus: Apache Arrow lets you pass data between Pandas/Polars/DuckDB with zero copying, so trying them out doesn't mean a full rewrite Michael #2: Pydantic-pint puts real-world units in your Pydantic models Pydantic-pint bridges Pydantic and Pint so models can validate physical quantities like 4m or 12 meters instead of bare floats. Fields annotated with PydanticPintQuantity parse user input, convert between compatible units, and serialize quantities back out as strings. That closes a real gap for anything consuming API payloads, config files, or sensor data with measurements, letting you enforce units at the validation boundary instead of hoping every caller remembered them. via PyCoder's Weekly newsletter Unit mix-ups have literally crashed spacecraft; now your Pydantic models can refuse them at the door. Annotate a field as Annotated[Quantity, PydanticPintQuantity('km')] and inputs like 12 meters arrive auto-converted to kilometers Validation covers string, numeric, and quantity inputs, and model_dump_json serializes quantities as readable unit strings Installable from PyPI as pydantic-pint, MIT licensed, with docs at pydantic-pint.readthedocs.io Early-stage solo project at version 0.4, so API stability and maintenance are open questions worth discussing Calvin #3: How Libraries Run Rust Inside Python (With PyO3) Pydantic v2's validation core (pydantic-core) is Rust under the hood, built with PyO3 — this post shows how that bridge actually works via a small hand-built JSON parser Four steps to get Rust into Python: write a normal Rust module, annotate with PyO3 macros (#[pyfunction], #[pymodule]), compile/install with maturin, then just import it The parser builds a Rust tree first — Python never touches it until the boundary crossing Key insight: converting the Rust result into Python objects (.into_pyobject) is often the expensive part, not the parsing — 100,000 JSON values means ~100,000 Python objects built after parsing's already done Errors cross the boundary too: Rust's typed errors convert into real Python exceptions (ValueError, FileNotFoundError) via From/?, so callers get clean Python semantics Takeaway for anyone porting Rust in: if you're returning a scalar, don't sweat it; if you're returning a big structure, profile the boundary — that's the real cost, not the algorithm Michael #4: AWS acquires DuckLabs Thank you Dylan McConnell. What does this mean for the DuckDB ecosystem? DuckDB is the open-source in-process analytical SQL engine. MIT licensed. The IP is not owned by any company - it's held by the nonprofit DuckDB Foundation, which was created when the team spun out of CWI Amsterdam. Peter Boncz, the CWI representative on the Foundation board, describes it as the entity that holds all IP of open-source DuckDB. DuckLabs (ducklabs.com) is the company, formerly branded DuckDB Labs. Founded a little over five years ago by Hannes Mühleisen and Mark Raasveldt to give the DuckDB team a stable long-term home, bootstrapped deliberately instead of taking VC, grown to 30+ people in Amsterdam, funded by support and feature-prioritization contracts. It employs the core devs. It does not own DuckDB. DuckLake is one of three projects DuckLabs builds, what they call the Duck Stack: DuckDB, DuckLake, and Quack. DuckLake is the lakehouse format that puts catalog metadata in a SQL database instead of in files on object storage. Quack is newer - an RPC-style protocol that turns DuckDB into a client-server system where both ends are DuckDB instances, slated to stabilize in DuckDB v2.0 in September 2026. MotherDuck is a separate Seattle company, Jordan Tigani's, selling serverless hosted DuckDB. It was started in partnership with DuckDB Labs and has worked closely with Hannes and Mark for four years. It contracted DuckLabs for engineering work and contributes heavily upstream - three of its engineers are among the top 10 outside contributors to DuckDB. It also sells its own DuckLake offering. Customer and collaborator, never owner. What the AWS post changes. Amazon bought the company, not the project. DuckLabs joined AWS effective September 1, with the process concluding August 31, 2026. Hannes and Mark keep leading the team and the project's technical direction, the team stays in Amsterdam, and DuckDB stays MIT under the Foundation. AWS gets the people and a direct line to the roadmap. The license protects your code, not your priorities. Three second-order effects worth tracking: The Foundation board is the real question. It has three directors: Mühleisen, Raasveldt, and Boncz. Two now work for AWS. Commentary on the deal has focused on exactly this - the license protects the code, not the roadmap. The announced counterweight is governance: a technical advisory board on the Foundation, and opening the extension stack so extensions signed by other developers can run in DuckDB. MotherDuck immediately moved into the business DuckLabs vacated. It now sells DuckDB enterprise support, which it had avoided because it didn't want to compete with DuckLabs' business model, and says it has explicit blessing from Hannes and Mark now that they're joining Amazon. It also bought Tower.dev the day before the AWS announcement. Everyone expects an AWS DuckDB service. Tigani says Amazon will likely release one eventually, and welcomes the competition, citing Redshift's failure to slow Snowflake on AWS. The groundwork is already visible: Amazon Quick uses DuckDB to query S3 Tables and has processed over 2.5B queries with it since launching in October 2025. The DuckLake angle is the one to watch. AWS is heavily committed to Iceberg through S3 Tables, and it just acquired the team behind a competing lakehouse format. The stated plan is to use DuckDB, DuckLake, and Quack together to power a new generation of data services, but which format wins internal priority is unannounced. Extras Calvin: astral-sh/uv 0.12.12: code-signed release binaries 🥳 Michael: My MacBook power supply rebooted to install updates (?!?) The Story of VS Code | Official Documentary Amazon/AWS acquires DuckLabs (see recent episode on DuckLake) Joke: We’re agentic now

Transcript
Discussion (0)
Starting point is 00:00:00 Hello and welcome to Python Bytes, where we deliver Python news and headlines directly to your earbuds. This is episode 496, recorded Tuesday, September 15th. I'm Michael Kennedy. And I'm Calvin Hendricks Parker. This episode is brought to you by Logfire from Pydantic. If you want observability for your apps and your AI agents, Logfire is the business. I will be telling you more about them. Later, find the link at the top of the show notes.
Starting point is 00:00:25 Follow us on the socials. All the various things you can think of are there on. the episode page as well. And sign up for the newsletter. I just sent out the last, the most recent one a couple days ago. It was a little bit late. Sorry, folks. But really cool stuff that we add, like, extra information that doesn't even appear in the show
Starting point is 00:00:43 that helps you get a little more out of the show. Yep. I love all the context it adds. I do too. I do too. I'm like, well, that's pretty good. We found some good stuff here. I would say that my, at this newsletter that we're writing here, it shouldn't go extinct.
Starting point is 00:00:56 But something might, something might need, might need, might be. need to go extinct. What's going on here? So I found, so this is a blog post from, what's Eddie's last name? Hold on. It's down here at the bottom of us copyright. Eddie Atkinson. He gave a talk at the most recent, well, latency conference.
Starting point is 00:01:17 So it's actually a talk from last year, but I think he kind of brought it back a little evergreened it into a blog post last week about pandas that should go extinct. And we're not talking about the cute little fluffy things that are. used for international diplomacy, but the Python Data Frame Library. Michael, how many times have you thought you had big data only to find out you were ready to definstrate your laptop because Pandas was the problem. You know what? It's happened. I'm going to get my dictionary real quick and then I'm going to know that that happened. I think the issue is a lot of folks really don't have truly big data problems. I mean, we've done some big data projects in the past, which were
Starting point is 00:01:57 10,000 tables, petabytes of data. That's truly big data. Most folks probably lie in the medium-sized data, but Pandas definitely tops out. He does some interesting benchmarks in here. Gives a couple good code examples. Actually shows a really interesting post from Amazon Redshift team where they were looking at the composition of many of the tables that are out there in the Redshift environment. If anybody's going to have a good view on what the size of data is and what big data could be, that they're probably the ones to look at that. But if you look at this chart, they basically say, on a continuum of data size,
Starting point is 00:02:33 most folks start over here in Excel. You've got like under a gigabyte, around a gigabyte of data, about that point in time, Excel falls over. It's probably time to pick up another tool to handle that. And a lot of people reach for pandas because I think there's just a lot of built up inertia or momentum in the community around the pandas and data frames, and it's an easy UI.
Starting point is 00:02:53 It's been taught in a lot of universities. So there's just not a lot of like need to kind of move out of that space because there's a lot of good code examples. A lot of blog posts have been produced. A lot of data sciences based on pandas. But they're really based on data frames. And there's more than one library out there to handle data frames and probably do it more efficient. So if you actually looked at the chart here, they're basically saying when you get up into like the 10 gigabyte range for data sizes, pandas is probably still pretty good. But then there's a gap.
Starting point is 00:03:21 It falls off somewhere between 10 and 100 gigabytes. of data. And 100 gigabytes of data these days is not infathable. You can easily go find sample datasets that are in that realm and that range and that size. And so the next thing they reach for is typically a commercial tool like Databricks, Snowflake, Dask, or some of these other things that are like Spark. So you're distributing the memory of that data set across many machines or maybe even across one very large machine, but doing it an attributed manner. Most people probably don't need to go that far. Like most people probably are still sitting in the range where you can see on this chart that Polars handles. Polars can handle straight up into 100 gigabytes of data easily
Starting point is 00:04:01 on a single machine. And then DuckDB takes it to the next step, which actually kind of fitting for this episode. I think this will be an interesting episode because there's a lot of information here about DuckDB later on in the show. But he kind of goes on again and shows that basically the average size of our row in Redshift is about a kilobyte. Every redshift cluster has like 10 machines in it. They're capable of guzzling 8 gigabytes per second from S3. But really, in actuality, almost 95% of the tables in Redshift contain fewer than 100 gigabytes of data. Most people are still in the range of just using a single machine with DuckDB
Starting point is 00:04:37 or even just Polars, which is probably similar to maintain and manage, but it's above the reach of Pandas. That's why the post is kind of going on about Pandas needing to go extinct. Another interesting bit, again, kind of good code examples in here. When we get down into some of the tables for the performance, what strikes you here, Michael, on their memory usage? The Pandas Library, we're talking about 30. This is one of his examples. I can't remember which one.
Starting point is 00:05:04 But four minutes, basically for the duration of the processing, 38 gigabytes of RAM. If you get into pollers, that gets halved, 18 gigagram. And if you go into DuckDB to do the same operation, five seconds at 1.93 gigabytes of RAM. So even the... Like nine to 20 times as much, yeah. Yeah, I mean, an iPhone could do this operation against the 100 gigabytes of data. You know, it's fine though,
Starting point is 00:05:30 because you can just get more memory. Memory's cheap these days. Yeah, totally cheap, totally, totally cheap. So I just think folks, you put this one to bed. Pandas was probably a good way to start, but you'll notice in here, like the polar's notation or syntax really, really similar. I like the DuckDB syntax.
Starting point is 00:05:49 I think he's got some examples in here where he reads in and does some more operations. He does a couple against some larger machines, then falls back into an older laptop, like a framework 13, to show that this is still useful as a developer tool. So I think here, Polars versus DuckDB kind of comes down to your workload,
Starting point is 00:06:07 your experience, and your preference. It's a good post. I really liked all the code snippets. He links over into the GitHub, where you can actually, like, try it yourself. It goes against the New York City taxi data set. It's got a ton of data in there. so it's fun to play with.
Starting point is 00:06:22 And you can see that we want to wait minutes or do you want to wait seconds? And would you want to use all your memory for this? And actually, another thing that Polars and DuckDB did much better was utilizing the CPU. Pure Pandas in this case was using like a multi-core machine, 146% CPU, where if you go to PureDB, over 800% CPU.
Starting point is 00:06:42 So obviously eight to 10 cores are being fully utilized as opposed to basically one and one and a half cores. Yeah, that's awesome. Mm-hmm. I feel like this is a pretty data-heavy episode for the data science crew out there. Yeah. Wasn't on purpose, but we ended up that way. Yeah.
Starting point is 00:06:58 I have a little bit of real-time follow-up for you, Calvin. For people who... Yeah. I was going to ask the one last bit in here. He does mention Apache Arrow, and if you've not played with it, it allowed him to switch back and forth between pandas, polar, and DuckDB without having to reload or copy the data into RAM. So you could actually do the same operation with each of the libraries without actually
Starting point is 00:07:18 having to take the data back out of RAM. So check out that Arrow, which is really cool. That's kind of a little bonus side bit that was in the blog post. So data folks who got medium-sized data, this is going to be a godsend for you. Yeah. I think Arrow is the foundation of Panthers 2, if I remember correctly, and also a polar. So that's pretty sweet. Yeah.
Starting point is 00:07:39 My real-time follow-up here. If you were working with one of these and you want to switch to the other, I had Marco Goralie on, really on Talk Python a while ago to talk about Narwhals. And Narwhals is a facade, adaptive layer that speaks native polars, but also talks pandas. So if you want to try, like, ooh, let's see what we're, you know, you could use this as a intermediate layer to kind of swap that out a little more easily than rewrite and everything. Yeah, I think people just need to drop, drop pandas. I mean, it was great. It was great 10 years ago. Yeah. Yeah. Yeah. I have some funny jokes, but let's carry, let's move on.
Starting point is 00:08:21 Let's move on to Piedantic Pint. So Piedentic is really interesting. Do you know Pint? Are you familiar with Pint? I've never used Pint. So Pint, we've covered that on the show back in the day. And Pint is interesting because if you're, I mean, all you got to do is say Marslander, sample return, whatever.
Starting point is 00:08:41 And it's like the $100 million plus fail because somebody used feet and somebody use meters or something like that, right? And so Pint lets you do math in Python with units attached, which is pretty cool, right? So I can say, instead of just having a distance, I have 42, I have 42 kilometers and you can say like two miles to whatever and so on. So it basically forces you to work in units, right? A lot of times we don't do this as regular programmers, but if you do anything scientific, well, there you go, right? So that's the background on Pindt. But what I'm going to talk about is actually not pint. It's called PIDANTIC Pint. Because PIDANIC is an awesome library that lets you validate the inputs and parse them and everything, whenever you read some sort of JSON, right?
Starting point is 00:09:28 Like if it's a FAS API or just a JSON file or whatever, beanie database, SQL model, all those things, yeah. So PIDantic Pint takes this idea and adds units to your data validation libraries. So instead of say, and I have a box that has a length and a width. I could say I have a box that has a length and a width that is a pint quantity. And the validation is to convert it to meters. So even if you parse something that says feet or centimeters or whatever, it will show up correctly. What do you think? That's definitely handy.
Starting point is 00:10:01 Yeah. It's the kind of thing that's like you don't have to, you're not going to use it a lot unless you're really in, you know, some kind of engineer or something. Yeah. Oh my God. This is so good. It's like so perfect. This has got to solve so many, like you said, small mistakes.
Starting point is 00:10:13 that end up in huge damage. 100%. And to be able to validate with it too. Yeah, just automatically, right, just all the Pidentic validation you have a Pidentic base model and it parses over to whatever it is. And if you put, I don't know,
Starting point is 00:10:26 liters into the length. Well, well, leaders, you can't convert leaders to meters. So I don't know. And it kind of fits perfectly under the Pidentic scope of the data validation and serialization. Like, it just, there's natural. Like, this should exist.
Starting point is 00:10:43 and they made it exist. Yeah, it's really cool. So you could have, say, a fast API endpoint that just automatically just takes units and automatically converts units. And yeah, it's a really nice one there. Speaking of really nice, now, this transition here has, this has nothing to do with the sponsorship, the previous one. They just happened to do great stuff in open source code too.
Starting point is 00:11:04 But Pidantic also happens to create LogFire, which I told you about at the beginning. So let me go ahead and tell you about our sponsorship offer, LogFire, not Pythantic Piney which is not even from them, but it's based on Pidantic. So here's the deal. It's 2 a.m. your AI agent failed. Was it the model, a tool call, the database, just the general unreliability of, hey, I think a new model is coming. So the current one starts breaking periodically. So most observability tools, they can't tell you because they only see part of your stack. Pidentic log fire sees all of it. One trace across your agents, LMs,
Starting point is 00:11:37 APIs, and databases down to the infrastructure services Kubernetes hosts. It's It's built on OpenTrematory with SDKs for Python, TypeScript, and Rust, and it works with any O-Tel compatible language. Every prompt, token count, and cost right next to your vector searches and API calls. You query everything with Postgres compatible SQL to understand what your app is doing, and so can your coding agent. It can use the same way because it talks SQL really well. If you connect it to the MCP server, your agent can also figure out what is going on.
Starting point is 00:12:09 So stop guessing, read the trace, Pythanic Logfire. AI, it is still just engineering, even if it's weird engineering these days. So visit Python BySat-FM slash Logfire today and sign up. Get 10 million records free every month, no credit card required. You can even click, and I really like this, there's a little copy of this text, onboard with your agent, click that, and it gives you a prompt. You can drop in the cloud code or codex or whatever, and it automatically knows what to do to set up Log Fire in your app.
Starting point is 00:12:36 So thank you to Pidenta for supporting the show. Calvin, I know you're a big fan of the visibility. into the token. Yeah, but I was just curious now if copy the setup prompt isn't the new pipe to bash, like pipe some curl to bash. Yes. This is the replacing that. That's why I was scheduling. That's why I was getting it. Yeah, no, no, I think it is. And it's amazing. I have some stuff that I'm working on. I'm like, oh, this is like, this idea is perfect. I love it so much. So yeah, pretty cool. Thanks to Pidentic for sponsoring the show. Thank you. And let's jump over to your topic next. see what's going on the wrong order in the other.
Starting point is 00:13:12 So what's next? Well, speaking of Pidentic, this comes from Bob Builderboss, a friend of this show. I know you've had them on numerous times for other events and things. But this one is about how to run, how Rust Code becomes something you can import. I think it's interesting that we can, if people are complaining about performance, with the first news article I had about getting rid of Pandas and bringing back in with polar and Duck TV was about performance, this is similarly veined. Like, if I've got a very computationally intense data structure
Starting point is 00:13:50 or function that's happening in my program, it'd be sure be nice if I could maybe replace it out with a Rust version of that, but have it act natively inside of my Python code. So this blog post from Bob goes over basically what Pidentic V2 does, which is a data validation library that most Python apps are using these days. it's actually a rust extension under the covers. It does the work. Its core, Pidentic core is all built with Pai-O-3, the same tool chain we're going to use here in this example. So it goes over some examples of basically you write a normal Rust module, you annotate it with some
Starting point is 00:14:24 specific macros, and then you'll actually be able to import that into your Python code fairly naturally. I think there's basically the Rust parsers incredibly fast. Python never touches anything it until the boundary crossing. So if you call for some data, that all happens over in Rust. The thing that I think the article covers that it's really important is that if you hold back that data across the boundary, those Rust results get turned into Python objects. And so you're going to want to think carefully about how you bring back parts of that because maybe you're only interested in a small piece of what is coming back and you don't need to populate. Say, for example, 100,000 JSON values means that you're going to get 100,000 Python data.
Starting point is 00:15:08 dictionary objects after the parsing's all done and everything gets passed back. You don't incur the penalty until you cross that threshold back into Python land. Maybe you don't need all 100,000, but there's some other operation you can do to get down to just the pieces you need. So what's nice is errors cross the boundary too. So if Rusts runs into errors, those come back as Python exceptions. So it makes it easy to debug and figure out what's going on. So you get clean Python semantics while still leveraging Rust. So basically for anyone porting Rust, if you're returning a scalar, don't sweat it. If you're returning a big structure, profile the boundary.
Starting point is 00:15:45 And if that's the, you know, if that's a real cost, you want to switch over and maybe do more of the algorithm on the Rust side. So much like you can use C or other languages in, I don't know if you can you use Ruby to do this kind of thing. I don't know if you can or not. But you definitely use Rust. And I'm kind of excited about that. I know you've been doing some coursework on it. And it sounds like Bob has also made some learning materials to lead folks through. It just feels like there's a really nice friendship between the Rust communities and the Python communities
Starting point is 00:16:16 and all the niceties that have been put in place to allow us to use Rust almost natively over in the Python world. So thanks Bob for the awesome post about that. Again, code examples in here kind of explains to the Python folks who have never touched Russ what the function signatures look like, which I appreciate because breaking it down and telling me what each of those pieces means that I can now pretty easily read some Rust code and understand what's going on because it doesn't look terribly, yeah, it doesn't look terribly foreign to me, but it's just different enough, but this goes over a good, a good usage of what each of those pieces mean for you. Yeah, it's pretty surprisingly similar to Python, honestly. Yeah, yeah, and incredibly fast,
Starting point is 00:16:55 but you get to think about things a little differently because the way it manages memory, and I think that's a bit, that's the big differentiator. What was that cartoon? with the guy's like, I'd gladly pay you on Tuesday. It's Popeye. Popeye, that was Popeye, right? Yeah, I mean, I was thinking borrow checker, the borrow. It was Wimpy, who will glad you pay you Tuesday for a hamburger today. Yeah, that's the difference of Rust is you've got the borrow checker.
Starting point is 00:17:19 Always checking. Yep, always checking. So yeah, I don't know. You've been doing a little more with Rust and Python and teaching some folks these things. Yeah, yeah, a little bit, a little bit. I have two followers. up here. So you talked about DuckDB, but the question is, do you have a lake house? You run a lake? I mean, a lake house. So we've heard of data lakes, which is a place you just kind
Starting point is 00:17:46 of dump a ton of like an insane. This is like back to your big data thing, duct DB thing. You just dump a bunch of data into this data lake and you figure it out. Well, that's grown up a little bit. And now there's this thing called DuckDB, but it's an implementation, an example of what's called an open lake format. Who knew? Do you know? I didn't heard that. We've done lakehouse implementations. I didn't know there was an open lake format now. So the story is, what if we could use S3 to scale our data access, right?
Starting point is 00:18:17 S3 scale is pretty large. If you can read stuff off the file system instead of out of memory, you can scale that tremendously large. In the open lake story is, well, if you put file formats in S3 that everything could read, Like maybe JSON files that tell you what the files mean. You could read them first. Here's where the data lives in each piece. And then parquet files.
Starting point is 00:18:38 Yeah. Or zipped CSV. I don't know. Take your pick, right? It could be whatever. So I just had the folks from Duck Lake on, which is a DuckDB implementation story of this open lake format. Well, then I get this message here saying, guess what? AWS, I'm sure you know this as a hero.
Starting point is 00:18:57 I did see this one come by. Yeah. Yeah. in Duck Labs just offered basically the H1 is bad. Let me read the first sense. Today we are announcing that Amazon has signed an agreement to acquire Duck Labs, the Amsterdam-based company behind the open source analytical database DuckDB. I'll link to the announcement.
Starting point is 00:19:18 And I thought, well, hmm, what does this even mean? And I wasn't entirely sure. So I went and did some looking here. I'm like, there's actually a lot of pieces in play. So let me lay it out and I'll tell you what part AWS acquired, what part didn't. Okay. So first of all, thanks to Dylan McConnell who sent this in. What does this mean for the DuckDB ecosystem?
Starting point is 00:19:38 So first of all, DuckDB, which you gave a shout out before, twice really, is the open source in-pros analytical SQL engine, MIT licensed. It's like SQLite, but for columnar data, which if you're data. And way, way, way more. Yeah. Anything you pointed out becomes SQL queryable. It's amazing. Yeah, yeah, yeah.
Starting point is 00:19:58 So you can say pointed out. out of PANDA's data frame and then do SQL queries against your pan. Like there's a bunch of plugins. It's a, it's far beyond just a database. But it's an in process sort of data processing engine, much like that. So the IP of this is not owned by any company. It's held in a DuckDB Foundation.
Starting point is 00:20:17 Oh, good. That sounds good. That is good. It was spent out of CWI Amsterdam from the folks who are mentioned that article. Good, but we'll come back to it. Then there's Duck Labs. And the story.
Starting point is 00:20:29 is, Amazon, AWS has acquired Duck Labs. This is the company formerly branded DuckDB Labs, founded over five years ago by Hannes, Moolisten, and Mark Rosfeldt to give DuckDB a stable home, bootstrapped, grew to 30 people. Now they can go chill on their island, which congrats to them, that's awesome. Because DuckDB really has taken over, right?
Starting point is 00:20:52 Then there's Duck Lake, which I mentioned earlier. It's one of three projects by Duck Lab, and there's DuckDB, there's Duck Lake, and then there's an API for working with this called Quack. Check out the Talk By Thought episode. But it's an open, open lake house format. All the waterfowl puns are great. It is.
Starting point is 00:21:13 And we actually on the podcast had a fun conversation about like, do you need a more serious name? Like RPC for your data lake. Like, no, they're like, we're calling it quack. Come on now. And then also we have Mother Duck, which I think I thought Mother Duck was online version of DuckDB.
Starting point is 00:21:31 But no, that is a separate Seattle company selling serverless hosted DuckDB. And originally it was started in partnership with DuckDB Labs and they worked closely with Hines and Mark for years and even contracted Duck Labs for some of the engineering. So now, what does this all mean? So the foundation owning DuckDB is awesome, but it has three directors, the people who own Duck Labs. So there's a bit of a, how much independence is it really going to have? There are some other folks, other governance and so on there. But, you know, it's cool.
Starting point is 00:22:06 There's a foundation. It's not super independent of DuckDB at the moment. Maybe it will be, though, after this. Mother Duck immediately moved into the business that Duck Labs vacated. They now sell enterprise support for DuckDB and so on. Everyone expects an AWS DuckDB service, probably a Duck Lake as well. Well, it already used S3, right? But maybe just a little more formal.
Starting point is 00:22:27 Yeah. There's S3 query and some adjacent-like technologies that sound like this may, maybe this will augment or replace. What's really nice about Ducklapse is it runs a local DuckDB or a local Postgres server. And a lot of the chatty API that would come from an open table format and the metadata, now all happen in the database. And then it just fetches and reads the files. That's pretty cool. Yeah.
Starting point is 00:22:49 We can't beat physics. If we can keep the data, the bits where they're at physically. and bring the compute to it. That's the win. Yeah. Yeah. So the Duck Lake angle actually is probably the most interesting one because AWS heavily committed to iceberg through S3 tables.
Starting point is 00:23:04 Yeah. Which is a competitor, at least a competitor competing concept to a duck lake. So yeah, check it out. I think I think AWS just got better. We'll see what that means for the rest of the world. What do you think? I mean, you're on the inside of this a little bit. Yeah, a little bit.
Starting point is 00:23:19 But, you know, we've had mixed reviews on their handling of open source. But luckily, they don't have any control over the open source other than they've just bought out the founders who are on the board of the open source foundation. It's still separate enough of an entity. I don't think there's a conflict here. I mean, it's going to be good for the project. The project is already incredible. Like the DuckDB stuff is, it's just like if someone had thought about SQLite and said, I need to grow all these other features that handle all kinds of crazy data and do give me native like JSON access and functions. And it's a really great platform for building cool little utilities or talking.
Starting point is 00:23:53 to giant chunks of data, as we saw in the first article. Yeah, yeah, absolutely. And use barely any memory. I mean, that's why this will run and work. Yeah, the DuckDB part is really interesting on that. And then the Duck Lake is like insane. You could have terabytes of parquet files all broken in a little bits. No big deal.
Starting point is 00:24:10 No big deal. NBD. NBD. Yeah. Well, how about some extras? Well, I will continue on my beating the UV drum. The latest release every week we got some. something new. The latest release from the UV folks, we get code signing on Mac and Windows.
Starting point is 00:24:29 So the binaries are now officially signed, code signed. Again, I think this is all coming together, ensuring we can secure the software supply chain part of this. So I'm excited to see that. That release too. I was hoping I wouldn't see a UV thing this week, but sure enough, it popped up in my feed. And I was like, I have to mention it because they just keep making everything better and better and better. So UVUV is now code signed so you can trust that it came from the right source on your own machine if you're on Mac and Linux. Yeah, that's excellent. Excellent. Excellent. Yeah. I see it's still going. Oh my gosh. The code signing is such a pain these days. Yeah. It used to be you could just build an EXE or a dot app and you could just go here, try my app.
Starting point is 00:25:12 And now how dangerous is that. I know. That's how the world used to be though. Well, we used to have what our login with no password to remote machines. What could go wrong? It's fine. It's fine. Trust. You've got to have a lot of trust. Why would somebody do something mean?
Starting point is 00:25:28 The computers. I don't know. I don't know. I remember in Windows 95, we had a bunch of them at a university I worked at. Plug straight in to Ethernet and the Ethernet. Everyone got its own IP address. And guess what? That thing got taken over pretty quickly.
Starting point is 00:25:44 The university I was at, they all had public IP addresses in the labs too. Yeah, it didn't go well. No. It did not go well. Speaking of things that need to be passed. patched and updated. Check this out. So there's two things that involve restarts. But I have a MacBook Pro in 5 Pro. Very nice. Love it. I got it this summer or earlier, maybe spring. I don't know whenever I got it. And it came with a power supply. One of those power bricks. The power brick
Starting point is 00:26:08 had to reboot the other day to update itself. I was sudden there working and I saw this article come out come by say Apple releases a firmware update for the 141 USBC power adapter, which is the one that runs the MacBook. And it's almost that, that's almost actual size right there. It's so big. Yeah, it's,
Starting point is 00:26:25 it's a beefy boy. It is, uh, yeah, I think this is even a little small, this big old picture of it. It's, it's heavy.
Starting point is 00:26:32 But I was, I read this article and I was sitting there working. And my Mac, when it comes off of power, it dims the monitor. Yeah. So I'm, I'm just,
Starting point is 00:26:40 you noticed. And I was an hour two later, I was sitting there, everything goes dim for a second. The little power just disconnects. And then two, two, three, five seconds later,
Starting point is 00:26:49 something like that. Power comes back, brightness comes back. I'm like, hmm, I think I just, my power brake just rebooted what in the world is going on here. That's crazy. We're a weird world. I replace all my Mac power bricks. I've got, not that I'm trying to be an ad for anchor anything like that, but the anchor prime has 160 watt, like itty-bitty little power brick that, because I travel quite a bit, it has four USBC ports on it. And they can all deliver a, you know, a combined sum of 160 watts.
Starting point is 00:27:18 so I can full bore charge my MacBook Pro M4 Max and my iPad and my phone all at same time. That's beautiful. And it's smaller than that brick. Yeah, I'm also a fan of the anchor stuff. This one since I had it anyway, I just plugged it into the wall, part of the house. And just if I'm in that part of the house, I just grab that cord. But yeah, normally if I travel, I have an anchor that's actually a power brick, a little battery. And it has two USB things.
Starting point is 00:27:43 And it'll do not quite as high as yours, but pretty high. and it's super nice because it's also a power brick, right? So if I need to charge up, it'll even charge the MacBook, but then just get it and just plug it into the wall on it, then it just becomes a power thing. I'm also a fan of this anchor stuff. Okay, a couple more extras really quick here.
Starting point is 00:28:00 We've got, do not play. So the story of VS code, the official documentary is out. Have you watched this? No, I have not watched this. It's an hour and 38 minutes, and I'm here for it. Okay, all right. It's got a lot of people that maybe you didn't see, coming like Eric Gamma for example you know thinking back to the gang of four patterns and all that
Starting point is 00:28:20 kind of stuff because he was apparently involved in the early days um yeah cool we had uh cult repo do the python documentary we had them do the jet brains document or intelligent documentary and here's the vs code one i'm really loving these like high quality production i mean these are these are nice little nice videos they yeah a lot of people behind it i mean there's such there's an audience for all these things, I'm here for it too. I love the fact that the underdogs can feel they are important for us and we can now hear more of the story about how some of these things came about. Yeah, it's really interesting. I mean, VES code has taken over so much and then, yeah, anyway, the origins are way, way more, less ambitious, let's say. It's cool to check out.
Starting point is 00:29:05 So it also is over a half million views. So there is an audience for this, apparently. There is absolutely an audience. Okay, speaking of the things that make you reboot yesterday last night, yesterday, Mac OS Golden Gate, iOS Golden Gate, watch OS 27 Golden Gate. All those things came out. So, uh-oh. Did you upgrade?
Starting point is 00:29:21 I did. Why wouldn't? I mean, I'm like, oh, it's out. Let's go. I'm not yet upgraded on this. I'm usually of that opinion, but lately I may wait a month or until a dot one or 0.1 to come out. I spent one day working with it.
Starting point is 00:29:34 And so far, it's okay. Okay. I'm going to upgrade then on your, on your full recommendation. Well, I've not upgraded my MacBook or my streaming computer. I only recorded my, I mean, desktop. So we'll see.
Starting point is 00:29:45 Ask me next week. Ask me how I feel about it. Honestly, the one thing to be a little careful about is developers is the Rosetta 2. Yeah. That's going away. So your ability to run Intel compiled stuff, you might think, Michael, why would I run an Intel compiled stuff like docker?
Starting point is 00:30:01 Certain Docker things only have Intel versions. So that's going to be a mega hassle. I mean, that's pretty rare. There's people have cross compiled most of this stuff because when was the last time an Intel Mac was released? Yeah, but if you, let's suppose I'm to pull. to an x86 server yeah and i want to test something i think it's gotten a lot better it definitely has gotten better but it used to be the thing is i don't know it used to be certain stuff would only work
Starting point is 00:30:24 in an x86 version of oh i remember this yeah but that was that was like three four or five years ago when i was really dealing with that actually it was when i was dealing with like data bricks and trying to to coordinate that stuff nice so this one still has it but the one after it whatever that's called won't. So this is like your last safe upgrade if you're worried about the reverse the reverse other thing. So did Apple actually deliver some AI features this time? Well, I'll tell you what. The new Siri caught me off guard. I'm like, oh yeah, I did. I did actually upgrade the phone. Like I said, it has the new Siri because it sounded, I asked it something like set a timer and it said something completely different than I'm used to.
Starting point is 00:30:57 And it sounded better. I'm like, oh, wait. I haven't had a chance to test it though. All right. It did set the timer like a champ. Let me tell you. Well done. Way to go. All right. Let's let's talk a joke. Speaking of, you know, the new series supposed to be agentic. So the joke is, we're agentic now. We're an agentic startup. You ready? This is how you, there's certain things you've got to position yourself.
Starting point is 00:31:19 I was just watching an ad because I started watching football yesterday. And normally, ads are excluded from my life, but apparently not a football, American football. And there's some ad for Zoom that Zoom is an AI company. They're not about medians anymore. Nope. They can, they can book the things so that your dry cleaning gets picked. up, they can do, they can do a slideshow, like what? Okay, so everyone's got to be some kind of AI thing now. So here's the joke. I change all of our loading dot, dot, dot states to thinking dot, dot, dot,
Starting point is 00:31:53 we're an agentic startup now. Perfect. I'm going to get right on that. Yeah, get it right in there. Like you can, there's so much VC money to be had from this. Go for it. Discombobulating. No, I'm thinking. Oh, oh, I hate that so much about clogged code. It drives me crazy that it's got all these random little words. Yeah. The reason I don't like it is I don't, it feels like it's made for someone with ADHD who just can't possibly let it just be for like five seconds. And so if I'm like doing something else, I'd look over and like words start, like, oh,
Starting point is 00:32:21 maybe it's, no, it's not done. It's like, oh, maybe it's done. Oh, no, no, it's just still like randomly. Like, could it just have the little icon go? No, no, no, no. It's thinking, it's combulating. It's wording. I don't know, what is it doing?
Starting point is 00:32:34 That's why you need herder. Exactly. That's that's the story for another episode. All right. Sounds good. All right. Well, thanks as always for being here, Calvin, and thank you everyone for listening. We'll talk to you soon. Yeah, bye.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.