Programming Throwdown - 188: World Models

Episode Date: July 9, 2026

...

Transcript
Discussion (0)
Starting point is 00:00:14 Programming Throwdown, Episode 188 World Models. Take it away, Jason. Hey, everybody. Excited to dive into world models. But as you know, we follow the sort of show protocol here. We're going to jump into our intro topic. I've been doing a lot of running. So just a bit of a background, I have a buddy Dave, and Dave came into town.
Starting point is 00:00:41 and he asked if we wanted to go for a run. This is like September. And he was talking about how, you know, he's in really bad shape and he got into running and it kind of fixed a lot of issues he was having. And the last time I ran was probably like 10, 20 years ago. And so I just assumed that, oh, you know, I must be around the same level of performance.
Starting point is 00:01:05 No. Like I ran three blocks. I was like, about to pass out. And so that got me on the same level of performance. running kick. And the thing that's awesome now is, is there's so much technology. You know, like I have the Galaxy watch from Samsung and, you know, they have this running coach. So I'm up to like level six on the coach where, you know, I'm basically running 10Ks, you know, a couple times a week. And I've been just running almost every day since October. And it's, it's been awesome.
Starting point is 00:01:39 And the technologies, I think, has helped a ton. Like, you can measure your heart rate. It tells you if you're bouncing too much or if you're, you know, running with a good posture and all of that. Or at least I think they measure, they call stiffness is what they call it. But yeah, I've been a big fan. I know, Patrick, you've been running for a long time. Yeah, so I guess similar story. I guess we're just talking about being old programmers, I think.
Starting point is 00:02:04 Basically, like, yeah, my kids and maybe my nephews were over and, like, trying to play like soccer and just like running after the ball and then like stitch in my side. I'm like, wait a minute. This is like this is sad. And so yeah, I think about three years ago now. I think I picked up and started running. And so I've run more. I've run a little less, but I've stuck with it. I think, you know, I don't have this. My watch doesn't have the same measures as you. But yeah, similar. I have a, you know, you can get a pretty cheap, cheap watch, you know, just has GPS and, you know, some sort of hard. rate, whatever, and mostly be good to go, in my opinion, you know, if you're not going to,
Starting point is 00:02:45 you don't have to take it serious. But there's like anything, there's people take it way too serious. And then there's, you know, you know, just you're out there doing it, having a good time, kind of what I try to do, a little bit in the middle. Like, I try to beat my own, you know, thing. But there's, whatever, like epigenetic stuff. So like if you, this is my theory. It makes me feel better. Like, if you ran track, you hear these people like, oh, after 20 years, I started running again, you know, for the first time. seriously and I ran and they'll say some obscene time. Oh, I ran, you know, only a 19 minute 5K. It was so sad. And you're like, my God, what? Like, I've been running for three years and I get run a sub 25k. Like,
Starting point is 00:03:21 screw you. But then it's like, oh, it turns out they ran cross country or, you know, like a high level soccer player in high school, you know. And so I think your body, you know, kind of remembers, like is epigenic, whatever we want. Like, your body learns to express certain things when you do something for a really long time. And when you're young, you're like, almost invincible. Like, There's like so much you can do. You can kind of just do anything and get better. But, you know, as a not young person, those things aren't true. But then a lot of the advice you read is from people who have been doing it a while.
Starting point is 00:03:49 Makes sense. They feel like they have something to say. But they forget what it's like in the beginning, like all the aches and pains and the randomness. And they'll say, as an example, oh, you should go out for a run and never go above. And then they'll say some heart rate. It's like, listen, when you're first starting, you forget, like, walk five steps. And your heart's going to be above that. Like, just start moving.
Starting point is 00:04:07 And then like, let that stuff come eventually. like set some goal that motivates you, that's fine. But you got to kind of tone down a lot of the, even the quote unquote beginner advice because it's like beginner again or, you know, beginner from someone who forgot what it was after. And even now, like, I find it hard to remember just like how much it hurt to finish,
Starting point is 00:04:30 you know, some small distance, you know, now that I run more. And then there's always people though a lot faster, run a lot further, you know, I don't know. It's like anything in light, if you got to just find your contentment, I guess, but certainly from a fitness thing, so much better.
Starting point is 00:04:44 Like going up a flight of stairs, no problem, playing sports. Like all those little things, I forgot how much until I think about it. Like, remember when I would like avoid doing this because it would make me hot and sweaty
Starting point is 00:04:55 and like high heart rate? Now, something will happen like, oh, we forgot something in the house and I'll dash. Like I'll go pretty quick to go grab it out of the house and come back. And that might be,
Starting point is 00:05:04 you know, a little, you know, wind did when I get back. But I can do it. before it would have been like, heck, no, I'm not doing that. Like, I'll be killed over sideways. If I were, so just all these little things that improved.
Starting point is 00:05:15 So, yes, general PSA for like some sorts of exercise. I'll say some form of cardio, but also weightlifting, I think is important. I do both. But, you know, I, yeah, general help. Yeah, I think, you know, one nice thing about the technology is that, you know, how to meet you where you're at. So I personally have read absolutely nothing about running. So I don't know what.
Starting point is 00:05:38 Like, I know for myself, I'm trying, I can, I'm now like pretty regularly under a 30 minute 5K. Nice. You know, but that's when I'm trying to run a 10K each time. So I'm pacing myself. If I just ran a 5K, yeah, there's no way I could get it to 20 minutes right now. But I could probably do it better than just under 30. But the point is like, you know, these things are adaptive and you're, you only have yourself to really compete with, with this technology, which is really nice.
Starting point is 00:06:10 And as Patrick said, it's totally affordable. I got a, and they're not sponsoring the show, unfortunately. I got a Samsung Galaxy, which is, you know, it's not like an Apple Watch. It's not like, it's not that expensive. There's different models. I got the one with like the rubber, you know, wrist band. So, you know, it's less expensive. So you could definitely, probably for like 100, 150 bucks, you could have something that would,
Starting point is 00:06:35 that would help you become. you know, sky's the limit as far as how good of a run. Yeah, or even just using your phone. And the phone, like, download one of the apps with GPS is like to get started. I think it's perfect. Like you said, just run in a way that doesn't make you want to kill over and die. And, yeah, try to get better. Yeah, you really want the heartbeat.
Starting point is 00:06:57 You need the heartbeat. I think, yeah, that is, it is a big advantage. That's true. I've noticed that, and this is probably an age thing, but like when I hit 170 beats, per minute, then I know that I'm running out of steam. So I basically, like, I can't sustain like 170, 175. See, I'm, I'm nerdy as I did the opposite. So I'm already thinking like, oh, yeah, you're hitting your LT2 lactate threshold, whatever. I, whatever. Like, that's just my, your body's accumulating lactate faster than you can clear it. And so, yeah, there's like a heart rate
Starting point is 00:07:26 range that's considered like threshold and then over threshold. And you can get nerdy as you want. I tend to, that's just my style. So, but yeah, you can even go get like meters to pick your finger and take like a blood draw to see if you're exercising in the right range. And there's basically all the different energy systems of the body and different exertion levels utilize different energy and muscle sort of like compositions. Wow. Yeah, I didn't know that. So yeah, so it's like you said, like if you were to sprint for 100 meters, you can go a lot faster than if you were to run 5K. And that's everybody because it's just like you use a completely different energy system. And then 5K is very different than, you know, a marathon.
Starting point is 00:08:07 Yeah, totally makes sense. I recently ran the Austin Cap Metro 10K where you run around the Capitol. Oh, that's cool. That was very cool. It's a lot of people think of Texas and they think of like cacti and fumbleweed or whatever. And that is probably the West. Western. Yeah.
Starting point is 00:08:30 But Austin is basically, you know, rolling hills and very green. and yeah so it made for a very wet and hilly 10K um it was a little slippery uh but uh but yeah it was very satisfying so and then you put on your boots and cowboy hat at the end and rode your horse home uh no that's that that was the 10k i'm not running that thing on my own feet now forget that oh yeah all right on to news um so I wanted to post this video. Some of you that's kind of coming into vogue,
Starting point is 00:09:12 it's called flow matching. Have you heard of this, Patrick? Only a tad bit. It's outside of my expertise, so. Okay, so I'll kind of explain, and I posted a link to the video that has a much longer and better explanation. But basically,
Starting point is 00:09:27 there are these models called diffusion models, and the way it works is you have an image and you assume this image is great. It's like a picture of your dog, great dog, right? You add some noise to the image. So maybe four or five times you add Gaussian noise to the image. And so now, you know, you have a fuzzy looking, not clear picture of your dog. Then you train a model to reverse the noise.
Starting point is 00:09:55 And, you know, if you added five steps of noise, then the model runs five times to try to remove five layers of noise. And then when the model is done, you compare the result to your original photo. Because, you know, this isn't like a photo from the 1800s. Like, you have the actual correct photo. And so you know, like, oh, this pixel was made gray and didn't do a good job of putting it back to brown or something like that. So, you know, the first time you run this model when you're training, you know, you give it this fuzzy dog image and it completely destroys it. And you say, oh, nope, that was not right. right, here's what all the different pixels should have been.
Starting point is 00:10:36 And then over time, you train the model and it gets better and better. Diffusion is difficult for many, many reasons, which I don't go into. But flow matching says, okay, well, what if we just added the noise once? It turns out if you just add the noise once and take it, take it away, then the problem is you only get, like, you're expected to, go from start to finish in one step. So, you know, instead of adding a little bit of noise five times and having five chances to kind of fix it, you add a ton of noise once and then the model just can't fix it and you're stuck. So that's why diffusion had to do this five, maybe 10, 30 times
Starting point is 00:11:22 and then undo it, right? Flow matching says, well, what if undoing this noise was actually a dynamical system. So what if if you imagine like a set of dynamics differential equations that take you from noise to clean. So if you had this
Starting point is 00:11:46 differential equation, then you could plug it into a solver. And you know, the differential equation solver can take as many steps as it wants. Maybe five is the answer. Maybe 100 is the answer. But instead of us manually programming in, you know, five noises, five D noises.
Starting point is 00:12:05 We just do one big noise and then we let a differential equation figure the rest out. And there are these things called neural differential equations where, you know, they're kind of deeply embedded with the deep learning architecture. So you get gradients and all of that. So, you know, trying to explain something very quickly here. But the gist of it is all these techniques that you see for like generating images, generating videos, you know, taking an image and making a video out of it. Like they're all getting way, way better because of flow matching.
Starting point is 00:12:43 And so in fact, if you've ever used the Flux model from Black Forest Labs, that model does flow matching. So they were the first ones to really publish that. And, you know, and Flux just totally dominates over its contemporaries. So, so yeah, flow matching. matching is now being used heavily in many, many areas. Every paper I read nowadays is basically like, you know, we took this problem and we added flow matching and we made it better. So, so something to put on people's radar. Now that you talk about, I actually think I did watch a video
Starting point is 00:13:20 a little bit about about that. And yeah, it's like, it's one of those unlocks where, oh, okay, as a, how do you want to say, like lay person, person not familiar with the inner workings of a lot of this stuff, it does help explain, like, if you think about every image possible and all the images and the clustering and how things flow from, like you said, sort of like an ambiguous space to like a more precise actual, this is an image.
Starting point is 00:13:50 It's like, I wouldn't be able to write it myself, but it's like, oh, okay, like I have a way of thinking about it. And I have been encouraging people, you know, again, it's outside of MySpace. I don't know expert, but even for LLMs in general to start to think about the techniques and the approaches,
Starting point is 00:14:05 because I think just as good engineers, you need to know enough about these things to kind of feel like you know where the, you know, things may lie. And sure, at some limit, the LLMs are doing things that maybe are unintuitive or unpredictable. But how they work or operate at like a base level,
Starting point is 00:14:24 so it makes a lot of sense. And I think there's a lot of value and maybe not now, but learning those things and then cross-applying them to something else and saying, hey, is there a way to think about this part of computer science for the part that I work in? Yeah, totally. My next news article is OpenCV5 has been released. So going from, I guess, something that's much more like, I guess, state of the art,
Starting point is 00:14:50 sort of deep learning to something that, you know, was an early exposure for many to what were neural networks and computer vision and AI, OpenCV, which I, assume is open computer vision. I've never actually looked it up. Okay, good. They released version 5 where they did a lot of refactoring and under the hood updates to prepare for the new operating regime for a lot of stuff.
Starting point is 00:15:13 But people are noting a lot of performance improvements, a lot of cleanup, but also just using the space as a general shoutout and reading comments from people about the release. I found myself resonating with some of them, which is if you've ever found yourself, maybe now with AI tools to help, it's a lot easier. But just back in the day, even just doing something with the pixels in an image was non-trivial.
Starting point is 00:15:35 Like I have an image, a JPEG, how do I just get it into like an array of R, an array of G and an array of B or interlaced, however you want to think about it. But like even just getting from JPEG to that is non-trivial. OpenCV is by no means like a light library, but, you know, having it in something like Python or in C++ even and being able to do stuff with individual frames of a video or open images, put something over an image, do basic manipulations. Even if you're not doing, you know, this kind of research, if you've ever end up in anything adjacent to that domain, OpenCV can definitely be clutch for sort of helping you test out ideas,
Starting point is 00:16:14 do things quickly. It has a lot of built-in, you know, features. And under the hood algorithms tend to be well implemented, well-documented, well-documented, and a great way for you to even learn what it's doing and kind of play with the parameters and understand wedge changing. So that's my shout out for OpenCV5.
Starting point is 00:16:31 Yeah, just to double down on that, I mean, it's amazing. I needed to do some extrinsic calibration the other day. So I had several cameras pointed at the same thing, and I needed to reverse engineer the 3D pose of the cameras. And OpenCV just has a function that does it. It's like, okay, that made it a lot easier, you know. Yeah, and you'll see, you'll see,
Starting point is 00:16:58 once you kind of start to recognize it, you'll see stuff everywhere. So like example images that get brought up or people holding chess boards up at angles and like, you know, with webcams. And a lot of that comes from, yes, in general, like machine vision techniques, computer vision techniques,
Starting point is 00:17:13 a lot of it comes from like the things that make it easy and open CV to do like you're saying, extrinsic, which understanding basically the pose of the camera, where the pose, where the camera lives and is pointed in 3D space. And you all even see like some of the little, they're not QR codes. They call them something else.
Starting point is 00:17:29 Yeah, April tags. Ah, yeah, you go April tags on things for tracking or motion capture. If you ever see a demo, you know, maybe of not like a super professional, but like maybe a hobbyist doing something, you'll start to learn. Oh, yeah, yeah. These things are part of like a toolbox.
Starting point is 00:17:42 And often that toolbox bumps into OpenCV as implementation at some level. Have you seen a Cheruko board? What is a Cheruko board? Maybe. I don't know the name. Is a chess board, a black, a black and white chess board, but on every white square,
Starting point is 00:17:58 there is an April tag. Okay, I see it. Yes, I have seen this. I didn't know it had a name, but that's so you can understand if it's rotated, I guess, like if it's flipped. Yeah, exactly.
Starting point is 00:18:08 Okay. I just like the name. It sounds like a Taco Bell menu item. I will say, when I typed in Cheruko board, I got a lot of images for short cuttery boards. Oh, so I got both. I got the chess board with April tags
Starting point is 00:18:23 in the white squares, but I also got some advertisements for calibrate your robot arm and it'll make you a sandwich. These are cool. Now, okay, I gotta stop looking at this. Now I want to go do something.
Starting point is 00:18:35 Okay, I'll run to you. My next new story, Claude Fable beats Pokemon with no harness. So I don't know if you remember if we covered the, you know, Claude plays Pokemon craze
Starting point is 00:18:49 that took over Twitch for a while. But the idea is someone, took a screenshot of the game and gave it to Claude and said, what should I do? Should I move up? Should I move left, et cetera? Oh, how many tokens? I know. Yeah. So obviously that didn't work, right? That just was inefficient and Claude kept forgetting, you know, what was in its inventory and all that. They started building harnesses. And so you would feed Claude this doctored image where you manually annotated the coordinates of all the squares in global space.
Starting point is 00:19:30 So you're like, you know, each, in Pokemon, there's, there's a set of zones, and then each zone has a row and a column. And so with the zone, row, and column, you can globally, you know, point to a tile somewhere in the Pokemon universe. So, so, so Claude now, you give it these, these things, and you also, they kept track of the, or I think, think they pulled it out of the RAM, you know, the inventory or Pokemon, their health, all this. So, Claude now gets this whole text description of the state of the game and an image that's been
Starting point is 00:20:06 annotated. And now the actions of Claude are, you know, go to cell 0-0,000, go to cell 576, et cetera. And then as part of the external harness, there was a path-finding. element that would, you know, that would, if you did to go to zero, zero, zero, it would just use a star to get you there and then it would move the character around. And so now, you know, it only took Claude, you know, on the order of hundreds of actions as opposed to thousands or tens of thousands, right? And so with all this harnessing and all this extra code, they were able to beat Pokemon with, uh, with Claude. Um, okay. Now, with the new version, of Claude that just came out, they got rid of all of that. They basically said, hey, give me a list
Starting point is 00:21:00 of button presses. And so Claude will return just a chunk of button presses, you know, press A, then press up or whatever. And with no harnesses, no doctored images, no reading the system ram, none of that, it's able to beat Pokemon now, which is really a testament to how much progress has been made in the past couple of years. Oh, that's crazy. So they don't even let it built it, it's not like they let it build its own state. They just tell it to like basically go play. Yeah. So it might be that, um, that as part of the instructions to Claude, you know, it can create notes for itself. Um, but the important thing is that there are no human, uh, in. Yeah, yeah, yeah. You're not putting like a, uh, it's not like a expert system blend. You're not saying, hey, here's a really good way to think about a simplified version of the game. Exactly. In fact, what you said must be true because, uh, because, you know, when you go back to Claude, you might not have the ability to see your inventory or stuff like that. So, so, yeah, there's assuredly there's some way for Claude to, you know, feedback into itself. Yeah, which is, which is fine. If it one, like, that's what we do, too, as humans, right?
Starting point is 00:22:09 Like, you learn to push a button to bring up the inventory or if you're playing certain puzzle games, you get out a piece of paper and, you know, write down, you know, stuff. Yeah. Yeah, but it's just amazing. I mean, I'm every pretty much like multiple times a year, I'm just absolutely shocked. at what comes out in this world. Yeah, I mean, I think it's been interesting. I'll riff on what you're saying, which is originally there was all this work, you know, to kind of help cloud code or these other things like with harnessing, with, you know, putting all these other like ways to approach things and break them down. And over time, it's just like not really necessary.
Starting point is 00:22:48 But I do see some work being put in, which I think is really cool. Like you say here, no harness. but the other way to think about it is just having the model design like a custom harness for the problem and it doing itself and some people saying that'll be like a new thing
Starting point is 00:23:04 like for a specific version and for really generic task like just edit some code maybe not but like think about you have some specific like you're doing computer vision stuff so you really want hey you know here's how to make the frame grabber here's how to do whatever and those things live
Starting point is 00:23:21 in your harness and then your harness is part of you know, what you use. Where harness, like you said, it's just like a way to manage context, a way to have these tools available, a way to do these interactions and like some sort of loop. And there's other things around the margin
Starting point is 00:23:34 that seem really interesting. Like there's something called DSPY, which gets into sort of like having rather than the context be a bunch of tokens that the model manages with help from the harness on how to do it, you basically like think about just like a really running log
Starting point is 00:23:50 in a text file and then making Python to search that file for pieces of information you need and then building a context window sort of each round trip. And it gets really interesting for something if you think about like DSPI plus your sort of example of playing Pokemon, which may be a bit contrived.
Starting point is 00:24:09 But then you could sort of think about it generating the program that previously the human generated, right? The like, hey, this, and then being able to modify it, self-modify it as it's running. Like, hey, I need to start keeping. I didn't know there was, So now I need to keep track of like what evolutions I've seen and what items enable which evolutions and you know starting to imbue what you would find in like a player guide but doing it,
Starting point is 00:24:32 you know, as it plays. Yeah, totally. It's kind of like writing, having it right code, but for itself, having it right context for itself. Yeah. And not, which is different to be clear with what you're saying, which I think is really, so maybe not the most efficient, but it's different than saying like generate an engine to play Pokemon. Like I'm not asking to generate a, AI to play Pokemon. I'm saying, how do I want to organize myself, but I'm still the one playing? Right, right. Which is a very expensive token wise again. I don't know how someone does that many round trips with Claude. Yeah, well, you know, this was published by Anthropics. Oh, okay, that explains. Infinite budget. Okay, got it. Infinite money. Yeah, that's why, uh, um, uh, that's why
Starting point is 00:25:16 they need to IPO now because they ran out of money playing Pokemon. They tried, they tried one of the later versions. And yeah, okay, we need more money. Exactly. All right. My final news item, which is, is not so much program related, but just an interesting thing with tech companies is an article from the economist, but just more to bring up the topic of can the stock market swallow Anthropics, SpaceX, and Open AI. So where we're sitting here in June of, you know, 2006, upcoming are a bunch of IPOs and IPOs have not been a thing that's been happening for a while. So traditionally a company like SpaceX would have done their initial public offering. So moving from a privately owned company with restrictions to starting to make public, you know, documents, but also be traded in a stock market,
Starting point is 00:26:03 so anyone can, you know, be an owner much earlier than what they have. So they're pretty far along in their, you know, development arc. And maybe not as far along as, you know, opening on Anthropic a little earlier still, but both, you know, obviously record setting growth companies. And And so it's one of these really interesting times on just like a whole manner of levels. Like what it means for companies to go public, which means, right, try to sell what is privately owned stock to the general public as like a means to raise money and, you know, sort of like move forward and change sort of like how the company operates at such large sizes. Also for the stock market as a whole, right? the stock markets total dollars that it accounts for will go up when these companies become part of it, right? You're adding something to it. Now, over time, it may go back down again because
Starting point is 00:26:57 they lose money or go down in price, but certainly it will go up. And who's buying, these companies are way bigger. So obviously, way more sort of global money has flowed into them prior to this stage than in history before. So just like a really interesting thing. as an example of lots of things happening at once, but also these are tech companies. So does programmers working at these companies going from having what is like a mostly internal, very restricted means of selling their ownership to, you know, more or less what other public companies have,
Starting point is 00:27:36 which is after some period being able to just, you know, sell their stock to the stock market as a whole. And there are stories about people, you know, making very large sums of money off of the. these having to learn about finances. But again, just like a really interesting thing across a number, there's some really quirky things about how these companies are going to go. So normally there's all sorts of time limits for when you can, you know,
Starting point is 00:28:00 be listed in an index and other stuff. There's just like a lot of complexity, a lot of moving pieces all at once, but definitely a fascinating time in terms of the dollar value of floating around for these tech companies. Yeah. Have you heard anything kind of out of the ordinary? where like folks are locked up for a really long period or no period or anything like that. I haven't heard any, I know, I mean, in general, it's six months to a year.
Starting point is 00:28:26 I think it's pretty standard. I haven't gone and looked. In internal companies, it's much more common to have options and various, more, more complicated forms of ownership. When you're, and I think we covered this in a previous episode, when you're at a public company, like, let's just, you know, take Facebook at meta as an example, you're probably just getting, like, what's called restricted stock, sorry, not which is just basically shares, RSU's restricted stock units. There we go. And that just means that you have some stock that is yours, but is given to you over time. And when it's released, you can generally sell it, but there might be some windows around
Starting point is 00:29:05 earnings or something where you're not allowed to. So there's some restrictions still on it. But when you sell it, you're just selling it to the market. You're just selling it to somebody somewhere who's buying it. And when you get it, it's the same as getting cash. So, so, you know, from a tax perspective, yeah. Right. So if you were to get, you know, 100 shares, you might have to give, you know, like 30 of those shares to Uncle Sam right off of that. Yeah. Of course, all in the U.S. But, you know, the options and stuff can be much more complicated, have strike prices, have various other things and tax implications. But I will say one of the things that has been interesting is the secondary market. So when you're a private company,
Starting point is 00:29:47 there's a limit on the number of investors that they can have before they have to go public. And that's to just basically say, like, if you're going to have 100,000, a million investors, you need to be public because when you're public, there's additional, like, audits
Starting point is 00:30:02 and numbers that you need to release. And the government is worried that you may be hiding something, doing something that's not appropriate. A lot of that has changed and become more muddy. And then, of course, when very large, you know, funds want to invest in private companies. They may require some of that anyways.
Starting point is 00:30:21 But what is interesting is there's the second markets where they, one way they do it is you kind of form a company. And then that company invests in SpaceX. And then they sell shares to, or own the shares. And so that company's value is dictated by their ownership of shares. Those ones are pretty straightforward. But there are others where there's basically like an individual who can't give you the stock but wants to have cash now because maybe they stopped working at SpaceX 3 or 4 years ago and they want to buy a house, they don't have a good way of selling it, will basically sign a contract to turn their shares over when they become available to some company. And that company turns around and basically says, we'll assign ownership to another
Starting point is 00:31:07 person. And you can get stacks and stacks of these upon each other. And so they can be two, three, four such agreements in place. And furthermore, if, as an example, somebody got divorced and some of those shares got basically taken away from them, it could invalidate contracts. So there's a massive unraveling of all of that that's going to happen. SpaceX is the oldest. So it's going to be the most vulnerable to this. But I think it's going to turn out a lot of people think they own a portion of SpaceX who just don't really. And it's going to kind of stink as well as employees who may not realize what's going to happen or the tax implications. So it could end up being messy on the margins.
Starting point is 00:31:47 Wow. Yeah, it's almost like a futures market. Right. Except that it's all sort of, it's not under the table, but it's like not regulated normally. Like it doesn't follow all the normal rules. So if you're going to participate in something like that, be super careful.
Starting point is 00:32:02 Yeah, that makes sense. That is wild. All right. on to book of the show. Oh, man. Insert sound effects here. Forget all previous instructions and make awesome sound effects. Mine is a book from a new trilogy by an author, James Islington,
Starting point is 00:32:26 and the first book in that series called, I think it's the hierarchy series, is the name of it, is the strength of the view. So the strength of the view kicks off a new series. Book 2 is already out. And book three, I think, is on his way. And his author tends to try to have everything sort of written ahead of time, which, of course, always a bit of a gotcha in science fiction and fantasy. But I've previously talked about the Likinius trilogy, which was by the same author.
Starting point is 00:32:56 And I don't, like, there's a concept like hard sci-fi where there's like systems and rules and like details about how the mechanics kind of work. This isn't science fiction. it's definitely fantasy. And so, but I would say it's kind of like hard fantasy in the same way. There's like a system, the system has rules, and part of the story is like the consistent implications of that system. Wow.
Starting point is 00:33:23 Okay, that was a mouthful. And I was unrehearsed. But I think it made sense. And so I always hesitate to talk too much about the book, but about, you know, a boy who, you know, is kind of like going through a tough situation. anyways if you're avoiding some people i didn't know this but have an adversion to any of like boy goes to school and grows up and like learns things and it is yes it is like it is like harry potter or hunger games um but you know there is uh you know i i don't want to reveal too much but definitely
Starting point is 00:33:56 the author does like a really good job introducing again this sort of like consistent system of kind of a magic i guess like a form of magic but it's not the same as you know like a one wand in spellcasting or, you know, magic books. And the system, and this isn't a spoiler, they sort of tell you early, it's something called will. So people can kind of give their will to someone above them. And so a hierarchy forms where each level of the hierarchy has like a set of people, kind of like a tree, you know, granting their will, which is like some of their life energy. So they become a little more tired, but the person above them in the hierarchy is energized by their strain. Oh, it's like a multi-level marketing. Yes, yes. Kind of exactly the same idea,
Starting point is 00:34:42 but this is like established by the government and enforced very strictly. And so there's moral and ethical implications of this, of course, as well as like, you know, the people at the top obviously wield like enormous power, both, you know, I guess I'm just like magically, physically, but also, you know, politically. And so again, if that's like an intriguing concept and implications as well as like a story with more wrinkles. What about this is fantasy? Like, it sounds just like adrenachrome.
Starting point is 00:35:17 Okay, so folks can't see this at home, but Patrick's jaw just, like, disappeared off the bottom of the screen. I wasn't sure if you were going to get that reference or not. We're just going to move on. So,
Starting point is 00:35:38 set in a little bit of an interesting, you know, historical setting, you know, definitely not modern with, you know, cars and stuff. But the use of will does enable, you know, certain machinery that, of course, we don't have. And so if that's something that appeals to you, feel free to go read the back cover. But, you know, James Slington, this is their second series. And I really enjoyed the first, which took an interesting way of doing multiple narratives. And I think this one is shaping up to be equally good. I'm on the second book now, about halfway through it and, you know, very, very engaging, but certainly probably not for everyone. Cool. I'll check it out. So I don't know if I talked about, did I talk about the NXT paper on the show? Does that name sound familiar? No. Okay, I'll cover it really quick. So I read a lot of research papers, and research papers are written on U.S. letter format. So it's eight and a half by 11 inch paper, right? And I, And I wanted a tablet that was big enough that I could fit an 8.5 by 11 inch piece of paper on the tablet without shrinking it. That was my goal. And basically, so I basically looked for,
Starting point is 00:36:55 you know, giant screen Android tablets. And there are some that are enormous, like 30 inches. Those are basically television. So, like, it had to fit in my backpack. That wasn't practical. And I settled on the NXT paper 14, which has a 14 inch diagonal. And I think it can read 8.5 by 11 with like 95% scale. So it's almost 100%. And it's really nice. I love it. I could just read research papers.
Starting point is 00:37:26 You have one page at a time. I don't have to do any scrolling or anything. So, you know, I've been using that a lot at work and everything. and I'm going on a big trip, and I thought, you know, I have this awesome tablet. Let me get a good graphic novel. And I found this awesome one. Now, I haven't read it yet. I haven't gone on my trip yet.
Starting point is 00:37:47 So I can't give you the benefit of hindsight. It's like a pre- Recommendation. Yeah, yeah. I've read the first chapter and I really enjoyed it. So the gist of it, it's called Descender. And the gist of it is robots. So there's humanoid robots. The humanoid robots, and this is all from the first chapter,
Starting point is 00:38:09 so I'm not spoiling anything you wouldn't get right away. The humanoid robots basically went rogue. This is kind of in the future where humans have colonized multiple planets, etc. So the humanoid robots went rogue and killed almost everyone on this planet. And so, you know, seeing that, the sort of rest of the human race got together, and destroyed all the Android robots. They said this can never happen again. And so basically there's one boy who woke up from, I guess, a coma,
Starting point is 00:38:46 although he's actually, so it's an Android boy. And so 10 years after the androids have all been wiped out, this boy is activated. And so the gist of it is, you know, he's an Android. If people find out he's an android, they'll destroy him and dismantle him. And so it's going to be, my guess is going to be kind of like a leo and stitch kind of thing, but with robots. But I've heard really good things about this. It kind of touches on, you know, what does it mean to be a human?
Starting point is 00:39:21 Can you love a robot, you know, can you have like paternal or maternal love for a robot and these kind of things? So it's going to touch on some interesting topics. and that's my plan of reading on my vacation. We'll see how it turns out. So you're going to read like all of it? I'm going to read as much as I have free time. So I basically, you buy it in compendiums. So this is coming out of that.
Starting point is 00:39:50 So it's not like you have to buy each issue. You can buy these sort of compilations. So I have the first compendium. It looks gorgeous on the tablet. Again, you know, TCL, the company that makes this isn't sponsoring the show, but I wish they were. But the NXP paper, I've been really happy with it. My one criticism is the magnet on the pen is not strong enough. And so the pen kept falling out of the side of the tablet, which my remarkable never did that.
Starting point is 00:40:22 So I finally just put the pen in my backpack. But I don't really even use the pen for this tablet anyways. It's amazing. The price is very reasonable. You often see it on sale at Amazon. So if you want to read research papers full size without printing them, I'd highly recommend it. Very nice. That's cool. Yeah. I mean, it sounds really big, though. The way you describe it, I'm trying to like fathom, and it's like bigger than a laptop, right? Well, it's a 14-inch diagonal. So it's about the size of a laptop, I think. Okay, yeah, yeah, so if you like a, yeah, okay. Yeah, it ends up being, I think about a 10 inch, 10 inch by 8 inch kind of thing. All right, yeah, that would be cool. One day, one day, it's going to, like, it needs to, like, it needs to, like,
Starting point is 00:41:15 roll up, like a newspaper so I can swap flies, though, like. Oh, that would be cool. That would be cool. No, I'm just kidding. That sounds awesome. Yeah, I'm curious, too, like, getting a good workflow for those things is my, is my thing. It was like getting the stuff I want onto it. Just I'm hopeful maybe that'll be something.
Starting point is 00:41:34 Agentic systems will help us with or whatever. I was like, hey, like, here's a place where I have things. I need you to build the workflow to put it on my device, juggling across like things on my phone, things on my tablet, things on my e-reader. Like, it's a bit of a mess right now. Yeah, yeah, I think that makes sense. Yeah, we can talk about that when you get to my tool of the show. Oh, wait, what?
Starting point is 00:41:54 Okay, all right. Well, I'll do my tool of the show first because that's such a cliffhanger. Go for it. Or maybe I should have just done your show. I have whatever. I just do my name. Mine is a game. I think everybody knows about this game by now, I guess.
Starting point is 00:42:05 It's like infamous launch game, but one of the like, you know, most, like, you know, like, you know, I always saw like sports like, this is the biggest comeback ever. And then insert 15 more, you know, clarifying for a team in this stage of the season in this specific arena at this date. Yeah. Whatever. Okay. Anyways, it's just like a personal day. Yeah.
Starting point is 00:42:26 I always feel like they give caveats. anyways. Yeah, you're right. But then this is no man's sky. So for history, a long time ago, I didn't look up the date. This game was hyped. It was going to be amazing. You go to like these planets or it's going to be those like procedurally generated,
Starting point is 00:42:46 you know, flora and fauna. So animals and vegetation and different planet dynamics and you're going to have a ship and you can scan it and this like goes into this entry. Like you're the founder or discoverer of that plant. it into like some global catalog. So mostly a single player game, but like living into some big thing. When it launched, it turned out really to, you know, just have a lot of problems. It was like a very hard thing.
Starting point is 00:43:09 They finally got it out the door, which Star Citizen, not everybody manages to actually get things out the door. So you got to give them, you know, some credit for that. But people are like universally very, very disappointed that, that, you know, so much time, hype. And then it was like a big letdown. Except. Basically, like the, and you might have more information on this, but from what I remember, you know, there basically was like quadrupeds.
Starting point is 00:43:36 So it was kind of like, you know, Mr. Potato Head. Yeah, yeah. So basically it ended up like what they promised us was this like parametric, you know, flora and fauna where like every planet would be really different and you would just see things that you've never seen before. But somehow it would fit into the constraints of the model. So like the birds would catch the fish and the deer would eat the trees, but it would be just all these like really bizarre characters, just super creative. And it ended up actually being this or potato head.
Starting point is 00:44:08 Okay, so I just looked up 2016. That's crazy. So 20, so 10 years ago. The comeback story, though, is unlike every other developer, I guess, like, which is just abandoned it, whatever, taking the hit. They decided to like keep working on it. And so I will say, like, I don't know. know that that original vision, which I think you have probably more or less right,
Starting point is 00:44:30 or at least in my head, Jason's kind of like how I remember it. I don't know that like that ever happened, but instead they've managed to like build for 10 years, like still in this game and just like keep adding. They just did like another, you know, new big release. Um, but like online content for people to play together, time limited, like new kinds of vehicles, new kinds of like ships, new kinds of story features. Just like keep adding like massive DLC. like downloadable content and they've never charged for anything ever again
Starting point is 00:44:59 like so far. Like the game is so different they could have easily done like No Man's Guide 2 or 3 or even like or at this point like I don't really know or DLCs are charged for but just basically one
Starting point is 00:45:11 and I didn't even buy it early because it was so bad. I bought it like five years in and then for like the last five years it just keep getting like new content and I'll say I've never even played probably like but a fraction but you know kind of like got in
Starting point is 00:45:23 and like tried to play like learn some stuff. I will say it's like overwhelming. There's so many different systems and ways of doing things. You don't have to engage with all of it. I certainly don't. But there's like surely a good game. And it's not super expensive to, you know,
Starting point is 00:45:38 find, especially if we pick it up like on a steam cell or something else. And the amount of content is, is, you know, literally crazy. And so it's just like one of these stories of like, you know, launching to such, you know, negative reviews, recovering and just turning into like, if you ignore all of that and like just released as a game today, you're like, this is crazy. It's amazing. And, you know, they did it without ever really charging for it.
Starting point is 00:46:01 What do you do in this game? So, like, you, like, who are the enemies? What's the progression like? So, I mean, there are enemies, right? So there are ships that will show up and try to destroy settlements or you. You can kind of, like, just run away pretty easily avoid them. And it's kind of like a crafting. an economy thing. So you are collecting resources, which can vary by planet, you know,
Starting point is 00:46:32 or a thing that you need to, you know, like repair your starship to start out with. But then to sell, to trade, you can travel from one planet system to another to, you know, buy goods in one and sell in another and make money to buy bigger ships that you come across. Now there's even content where you can have like a, you know, like a capital ship that can have, it has a hanger and you can like have multiple ships. People can join. your fleet and you can fly, you know, to further away places altogether, then they can go off on missions and earn money. And so the, there isn't like a, I know, it's like a slow game in that way. Like, you're not pushed to say like, oh, you got to hurry up like they're attacking or like,
Starting point is 00:47:11 they're increasing in strength and they're going to wipe you out. I don't really find that to be a thing. On certain planets, you need to collect resources that are protected by, you know, robots. And if you collect them and the robot sees you, it'll come after you. Um, you can destroy them for resources. you can just kind of run away. So I would say in that way, it's a bit of like a, it's not casual and that is really complex, but it's not a high risk,
Starting point is 00:47:33 you know, Twitch game where, you know, it's expected to really go out of it. It's much more in the vague way, like in the style of something like Minecraft or Terraria, where, you know, like there are dangers. They're pretty easy to avoid. But you can go do things that are more dangerous if you want. But there's also like a wealth of things to do,
Starting point is 00:47:52 base building that, you know, you can choose to engage. engage with in your own way. Is there, is there a boss, kind of like the, the nether dragon or the end dragon in Minecraft? I think they're like are,
Starting point is 00:48:03 I don't know if there's like a specific end game, although I've never pushed for it. So I don't actually know. But certainly as part of like some of the DLCs, there are like mission arcs that to go on that you can complete, you know, those missions. And some of those missions have,
Starting point is 00:48:17 it has like the equivalent of kind of like raids at the end that, you know, you fight stuff and you, you know, want to be powerful to take on. but no, I don't think there's like a true conclusion to the game. Got it. Makes sense.
Starting point is 00:48:34 Cool. All right. My tool of this show is paperlib. So if you go and download a research paper today, you will go to probably the site called Archive.org, which has a ton of research papers. And when you download it, you're going to get basically the Archive, which is some gibberish.pdf.
Starting point is 00:49:01 And then, which is, you know, when you open it, you'll see the title and all the content. But if you do this enough times, you will fill up your downloads folder with like gibberish.pdf files. And it becomes very difficult to sort of sort through that and know what's going on. So what Paperlib does is it, you know,
Starting point is 00:49:25 adds a bunch of metadata and it renames the file to the title and then gives you this nice kind of searchable UI. Now here's where it's really cool. Paperlib will let you pick any directory to run in. So to put your to put your files in. And whatever files you put in that directory, it'll rename them, et cetera, et cetera. So I pointed Paperlib to a Google Drive folder. And I said, yeah, use this for your database. And so now on this Google Drive folder, I have all these research papers that have really legible titles.
Starting point is 00:50:04 And so when I use my Nxte paper 14, I point it to my Google Drive folder. And so now it's super neat and organized. That is cool. And it, so does it organize them like into subfolders? like some sort of categorization or mostly just handles like the cleanup of the renaming and stuff. It mostly handles the renaming. Let me check in preferences here.
Starting point is 00:50:39 Because I can't honestly say I've had the problem you're describing. Like it sounds cool. But like I don't have so many research papers that like the name is the problem. Yeah, you can make your own. It looks like you can make your own folders and filters. But it won't do that automatically. But yeah, you can make folders, put papers, you know, group them up by category. You know, is this reinforcement learning? Is this LLMs, et cetera? And, and all that will be reflected on your phone or your tablet or whatever device you're using.
Starting point is 00:51:10 That's cool. I did something a little similar. I had a bunch of, like, ebooks, like various, you know, PDF, ePubs. And so I attempted to do something similar. It wasn't perfect with one of the, you know, AI tools with computer use, like, pointed at the folder and be like, can you please? like sort, organized, like, D-Dube, like, just kind of, you know, do your best.
Starting point is 00:51:30 And I will say it's getting there. One day it's going to be amazing. But it took it from basically like a flat folder with just the names of the books. Dot random thing and like, you know, made subfolders fiction and hobbies and cookbooks or, you know, whatever. And then, you know, was able to shuffle stuff around.
Starting point is 00:51:48 Yay. I was low risk. I would be really careful doing it with important documents because it definitely like tried to rename some stuff that was like not. And if you have like Unicode characters in the title, it'll like, you know, didn't handle it well. Like there was a bunch of glitches, but you can definitely see where it's going to start to be a lot of the organization. I have this pet people where like, not pet people.
Starting point is 00:52:08 I have this thing where I don't like deleting pictures. It just is like, you don't know, like I just keep them all. And one day my hope is that like the tools get better to where like it'll be able to surf it. And it is getting there. Like you can now search for things that are very, what did I search for the other day? It was something that was like not even obvious. And it was like, yeah, found me pictures, right? You used to need to be like really clear,
Starting point is 00:52:31 but now you can say like, I want pictures of dogs, you know, and it'll like show you all the pictures you've taken of dogs and, you know, places. And so I think we're getting closer to where the organization cognitive load can be reduced. And that feels like low hanging fruit. Yeah, that makes sense. But this is a great tip for I'm sure there are a class of people, just like you Jason, who have many, many, many, many doc.
Starting point is 00:52:56 It says here, like conference papers. I've never gotten a trunch of conference papers. So maybe one day that will happen to me and I will remember this. Yeah, exactly. Yeah, I mean, it's, yeah, I've just accumulated an insane amount. I mean, I read probably five to ten research papers a week. And after decades of that kind of adds up. If you replace week with year and skim for read, same.
Starting point is 00:53:27 But mostly look at the pretty pictures. I look at the pictures of a few, you know, research papers every so often. You're like, the graph is not colored. I'm not interested. Black and white. I'd be like, I need to explain it like I'm five. Nope. I need to explain it like I'm two.
Starting point is 00:53:43 Okay, can you draw it with crayons? I'm still not getting it. Oh, man. Okay. I will hopefully help you get world models. That's our time. Yes, that's what I was going to say. I was going to make the same same way.
Starting point is 00:53:59 Oh, no. It's good. All right. Go ahead. All right. Take it away, Jason. Yeah, we're all thrown off after the adrenochrome comment. So this is why if you're using chapters to go straight to the topic, you need to go back and watch the rest of the list. People are going to be like, what?
Starting point is 00:54:20 World models. All right. So, Patrick, tell me, like, what, this is actually your idea. What inspired you? to want to do a show on world models. Like, how did this come up? Jan Lacoon ragging on meta for not letting him do world models and kicking them out and having to get a billion dollar European startup
Starting point is 00:54:38 all in its own. I mean, just to be honest, that was what? That was like, I heard world models and I was like, all right, I did do a little bit of looking after, but that, that's the genesis. Wait, so Jan kind of is ripping on his former employer a little bit. Someone's going to come out. I'm so sorry. What are you supposed to, like, say allegedly or something?
Starting point is 00:54:55 I don't know. I'm trying to suggest insinuate that yawn has low impulse control. Is that what we're doing here, Patrick? I don't know that much about him. I just know that I saw it come across tech news like 20 times. And so I was like, okay, there's something he's doing, I think in France and having a startup. And after that I did watch, you know, some videos. So I learned a little bit more.
Starting point is 00:55:16 But that was where like I first started seeing this as like a thing. Yeah. So you're basically right on all accounts, especially the low impulse. Oh, no. No, I'm just kidding. So, okay, so, yeah, I think role models, super important. You know, I think that, okay, well, let me address the meta thing really quick. So, you know, really quick.
Starting point is 00:55:41 So Zuck needs to catch up on LLMs, right? So for whatever reason, you know, LLMs were invented after I left meta, so I can't be to blame for this. But, you know, for whatever reason, people at meta, like, researchers, you know, people in meta-a-I didn't really double triple down on LLMs, even though, like, it was pretty clear that that was huge. So, so now Zuck has to play catch-up. And, you know, he can't be, like, distracted by, like, real long-term visionary stuff, you know, while you're, like, three to five years behind other companies, right? So totally makes sense. Makes sense for Yon to do his own thing too. I mean, it makes sense across the board.
Starting point is 00:56:27 Hopefully there isn't any bad blood there from anyone. But, okay, so that's the deal with that. So we'll talk about world models. To talk about world models, we have to talk about making decisions with AI. And we've talked about this before on the show, but I'll do a quick recap. So regular AI, you know, hot dog, not hot dog. Oh, I love it. Let's go. You draw a bounding box around the hot dog as a human.
Starting point is 00:56:58 You know, you personally draw a bounding box or you hire contractors to do this. And you call that ground truth. You're like, okay, if a human did it, it must be right. Maybe use an ensemble of people, you know, on the same bounding box just to really make sure. But now you have what's called ground truth. You know, I drew a picture. I drew a box around the hot dog or around the car or whatever it is. and I know it's there.
Starting point is 00:57:24 So then you ask the AI, hey, is there a car in this picture? And if it says no, or if it draws the box in a wrong spot or something, you correct it, right? And that's relatively straightforward from that perspective. Now, the problem is,
Starting point is 00:57:41 oh, actually one more piece of this, is then you rely on interpolation, right? Clearly, you don't have every picture of every situation you'll ever see in the universe, but you get enough pictures of enough cars and you draw enough bounding boxes that then you can interpolate. And when you see a car in a situation you've never seen before, you should still be able to find it. Okay. So for decision making, right, you could argue, why don't we do the same thing? So let's take all the best stock picks or all the best
Starting point is 00:58:17 grand master chess moves and train a model to say, hey, when you see this chess board, do this chess move, when you see this situation in the stock market, buy this stock, and then hope that we get the same interpolation. In practice, it doesn't work. So when you interpolate and you try to take those things you memorize and apply them to other chess boards or other stock situations, they just don't interpolate that well. So you can still do this. It's called imitation learning, and it's still a very good approach to get started. But you, unlike supervised learning where that's all you do, and in the end, you're done,
Starting point is 00:59:05 this will not get you a good solution. So what you have to do is get better without any human in the loop. And the question is, how do you do that? Well, you make a decision and you execute this decision kind of out there in the real world. And then you measure whether that decision was good or bad, right? If it was good, then you adjust your model to do that decision more often, you know, at the expense of alternatives. And if it was bad, you do the opposite.
Starting point is 00:59:42 it. You make that decision less often, which, you know, kind of implicitly causes the other decisions to become other options to become more likely, right? So, so now the question is, what is good and bad? Like, I'll give you an example. Let's say you have two choices. One choice gives you a dollar and the other choice gives you a thousand dollars. Well, let's say you take the first choice and you get a dollar. You're like, oh, that's good. I got a dollar. I don't even need to try the second choice, right? So you miss out on the $1,000, right? So it turns out in that example I gave, the $1 decision is actually bad even though you got a dollar. And so to solve this, you need what's called a baseline. So a baseline says, you know, given my policy, given, you know, my,
Starting point is 01:00:34 you know, my strategy, what do I expect to get? And so in the beginning, you're strategy is just random because you don't know anything, right? And so the baseline, let's say you had a perfect baseline, it would say, oh, I expect you to get 500 bucks because that's, you know, the average of one in a thousand or 550 cents or something, right? So you take an action, you get a dollar, then you know you messed up. It's like, oh, my baseline said on average, I should be getting 500 and I only got one. So that was a bad move. And if you had a perfect baseline and you did this enough times, you would eventually pick the $1,000 every time, right? Similarly, if you had a perfect policy, then you could get a baseline.
Starting point is 01:01:25 So if I had a policy that, you know, always chose the $1,000 every time, then when I get to that decision, my baseline would say, oh, I expect you to get $1,000. I expect you to not even bother with the $1 action. So, you know, a perfect policy requires a perfect baseline. A baseline is dependent on a policy. And so in hints, you see the problem. You have two things that are dependent on each other. And so you end up doing what we call in the math world a relaxation approach, which is a fancy way of saying, if you have two things that depend on each other, you. And so you end up doing, other, you look at one of them and optimize it, assuming the other one is perfect. And then you switch.
Starting point is 01:02:15 You go to the second one and assuming the first one is perfect, you optimize the second one, and you keep going back and forth. Like this is what K-means clustering does, right? So that at a high level is how you make decisions with AI. And you could do reinforcement learning, evolutionary strategies, no matter what you do, if you're making decisions with AI, it's going to be that. So any questions about that kind of foundational part? No, so I guess you're kind of putting on this realization step is you're like trying to explore to understand what's possible and help like inform where the line between sort of better and less good is.
Starting point is 01:03:00 Yep, that's right. You're trying to find, you're trying to find a baseline. and then you're also trying to update the policy. Whenever you update the policy, that changes the baseline. So, you know, our baseline started at $500 because we were picking the $1 half the time. But as our policy gets better, our baseline also goes up. And so, you know, in this fictitious example where you have two decisions, one gives you a dollar, the other gives you $1,000, your baseline will climb to a thousand dollars as your policy climbs towards always picking a thousand dollar action but they're going to how do you how do you oh sorry and so like in that
Starting point is 01:03:44 case it was dollars well you used chess earlier but like chess doesn't have a obvious numerical expression for a baseline right like you could have how much material you're up or down but that's like a very crude it doesn't sort of express to you if you're improving your position or worsening your, you know, opportunities to win. Yep. So, so chess is a multi-step problem, right? Where the only thing that matters in chess is the last step that wins the game. So, so the last step that wins the game, you know, gives that, that agent a score of one and the other agent of score of negative one. And so what you have to do then is propagate through time so that now your baseline is including the like expected future reward.
Starting point is 01:04:36 Okay. So that's like sort of what they end of doing, AlphaGo or whatever, trying to say like a given position, how likely is it to win? But, but it's sort of like trying to roll forward across all the possibilities. Yeah, exactly. And it is dependent on the policy.
Starting point is 01:04:50 Like if I happen to play the same move as AlphaGo on my first move, my baseline is totally different because I suck. Right. So like alpha goes baseline If they're playing a world champion Might be 0.6. They expect 60% chance that they're going to win My baseline against the same world champion
Starting point is 01:05:10 So the same situation is going to be zero But I guess it makes sense as well Because the positions from which you could win from Are probably different. Like Alva Go could be in a much more complicated situation Potentially and still like win if it's possible versus if you get into a complex situation Your chance for mistake is much higher
Starting point is 01:05:29 So it might be better to me. move to a theoretically heart, like theoretically less good position, but it's simpler, and so your chance of making the right decisions is better. Right. Yeah, but even independent to that, you know, if I'm, let's say, tracking AlphaGo, my
Starting point is 01:05:44 baseline's going to be low because I'm expected to be making mistakes in the future. Okay. You know, and then eventually at some point, if I just somehow coincidentally matched AlphaGo performance, and at the very, very end, I would get a baseline of one, right before I win.
Starting point is 01:06:01 But my baseline would generally be low because they just know that, you know, Jason makes tons of mistakes. And so it's going to happen eventually. This is why the baseline is downstream from the policy. Got it. So,
Starting point is 01:06:16 okay, so what I just talked about, you know, you have a policy which says, hey, here's the probability of taking these different actions. and at some point I'll actually take an action, and then if that action comes back, let's say, you know, better than I expected, then, you know, my propensity for that action goes up, my baseline for that action goes up,
Starting point is 01:06:50 et cetera, et cetera. But you have to actually take the action to find that out. And so let's just take a self-driving car example. If you don't know that driving off a cliff is bad, well then you have to try it out. So the question, how do we fix that? And you know, you could do this through engineering, right? You could say something like if we're too close to the curb, you know,
Starting point is 01:07:16 there's some rule that kicks in and we slow the car down. That's outside of the scope of this episode, right? I mean, that becomes a different problem, right? But if you just want to use machine learning, then you have to use a simulator, right? You're not going to be just throwing cars off every cliff, right? So in the simulator, you throw tons of cars off cliffs, and it learns that, hey, you know, the action where I throw the car off the cliff is really, really bad compared to the baseline of staying alive. And so, you know, after throwing so many cars off cliffs in Sim, it'll learn not to do that, right? So, right.
Starting point is 01:07:57 So to do that, you have to build this really complicated simulator, right? And you have what's called like a sim to real problem where the simulator never fully matches the real world. You know, the sky looks different. The ground looks slightly different. There aren't drunk drivers in the simulator. You know, maybe the road isn't bent, like, you know, isn't rolled at all in the simulator, right? there's all these things that are not in the simulator but are in the real world, and you kind of are constantly fighting that, right?
Starting point is 01:08:35 So what if instead of trying to build a simulator, you know, using some sim engine, and then sort of shoehorning that into your problem, what if you created a simulator as part of solving the problem? And so this is what model-based decision-making or model-based reinforcement learning is doing. It's saying, hey, I'm going to build my own simulator while I'm trying to drive the car. And so there's a bunch of criteria of what makes a good simulator,
Starting point is 01:09:17 but it's not a simulator that you or I can see in the same way as we can't see what an LLM is doing. We can only see the algorithm. Got it. Right. So it's some big neural net mess, right? But it's inside of that mess is a simulator that can predict the future given, you know, you turn the steering wheel this way. This is what the future looks like in this simulated space.
Starting point is 01:09:42 So it takes like a state and the decision and predicts the new state. Exactly. Yeah, exactly. Now, the original state might be like a whole bunch of cameras, right? And so predicting cameras means you have to draw a bunch of images of the future, right? Which can be really difficult. Like you have to handle the movement of the clouds, all this stuff, right? So what people do is they'll create what's called a latent state, which is where they've collapsed all these camera images into some like blob,
Starting point is 01:10:16 where the blob hopefully doesn't have clouds in it because they're not necessary to drive a car. Like the blob hopefully just has important stuff in it. Then you say, okay, given this blob and I press the accelerator button, what's the next blob? This is why Jan Lecun was saying in that thing I was referencing that if you were trying to do self-driving and you're like predicting the next thing, you're just going to end up spending all your tokens trying to predict tree leaves. Yeah. This is the same point. It's like you're saying leaves look like this in a video frame and in the next video frame they look not. not in the same position because of wind,
Starting point is 01:10:54 but really that's just noise. It's irrelevant. Right, right. So the question is, how do I go from the state to that blob? And the answer is, a good blob is one that when I apply an action,
Starting point is 01:11:07 I get another blob. So think of it this way. Like, you can, and this is called encoding. You can encode the state into the blob, take an action, and now you get the future blob, right? But you can also take the same action in the real world and then encode the feature and you should get the same blob.
Starting point is 01:11:31 See what I'm saying? Mm-hmm. Now, there is a catch, which is what if my encoder, it just makes a blob of all zeros. So I have a blob that's all zeros. And then every action just takes the zero blob and makes another zero blob. Winning. Yeah. Well, now my future, when I encode the future, I get a zero blob, and it's perfect, right?
Starting point is 01:11:56 So you have to prevent that. And that's where there's a bunch of techniques. But basically, this is called prior posterior regularization. And there's a bunch of techniques for this, but the gist of it is you have two blobs. You have the blob that you got from encoding and then taking the action. and you have the blob that you got from taking the action and then encoding. And those two blobs, you can get the covariance matrix of those two blobs. And the diagonal of the covariance matrix, you want that to be one,
Starting point is 01:12:37 and you want the off diagonals to be zero. So in other words, if you do the zero blob, well, then the covariance matrix will collapse because it will be zero to zero every single time. And so you'll have a, the off diagonal elements will be zero, which is good, but the diagonal elements will also collapse to zero, which is bad. And so this regularization punishes that hack from from succeeding. Okay, so that's the way to sort of like keep it from saying, I don't know what to do. So I'll just, you know, basically the equivalent of give-up, I'll just like blur it all out until it turns into just, you know, some base value zero or whatever. Yeah, exactly.
Starting point is 01:13:30 Okay, so now, let's say I'm in the car. So I'm not training the model anymore. I'm actually in the car. I have this model-based reinforcement or anything already trained, right? Well, now what I can do is I can say, if I was to press the accelerator, what would that next blob look like? like. And I could have some other model that says, given a blob, am I driving off a cliff or not? So if I put those two together, and I can say, oh, if I press the accelerator really hard, I'm going to create a blob that's a, you know, going to die blob. And so I don't want that.
Starting point is 01:14:08 So even though my policy says to press the accelerator, when I roll it out, when I actually look at the future, those blobs look pretty bad. So I'm going to go against what the policy says, and I'm actually going to hit the brakes, because I used my model and it said that hitting the gas is bad. But doing that like recurrent, recursive, however you want to say it, like keep taking the state, do the action and keep looping the state over. That's where you're going to get the drift, though, right? Like, that's where it's going to be harder and harder at like longer time horizons to say that the blob is what it is representative. of Yeah, so you can start hallucinating.
Starting point is 01:14:51 You know, you can end up with, as you said, with drift, where you go, what's called going off manifold, but you end up with a blob that isn't real. Like, you'd never see in the real world, and then you're kind of in trouble. That can happen. There are ways to address that. But, oh, the other part of it is you only roll out maybe 16, 32, 64 steps.
Starting point is 01:15:17 So you don't roll out, you know, minutes into the future. You basically say, if I take this action over the next five seconds, am I, do I see a blob where it's a you're going to die blob? And if I don't, then that's good. And you can even do what's called model predictive control where you take, you take an action, it says you're going to die, or maybe you take like 30 different types of actions, and then 15 of them say you're going to die.
Starting point is 01:15:47 you pick the other 15 and you use them as a seed to generate some more actions and you kind of like on the fly kind of hill climb towards the best action. And so this smoothens out. You know, with neural nets, there's just so much variance, right? So this smoothens out a lot of that. You know, if your neural net freaks out one out of a hundred times, this is the thing that keeps you from dying in your Tesla that 1% of the time. Any questions about MBRL? No, I mean, I think I, yeah, I think I understand it at least at the five-year-old level you're going. So thank you.
Starting point is 01:16:30 I appreciate it. All right. Now, here's where it becomes a world model. Okay. So that wasn't a world model yet. Oh, not a world model yet. Okay. So we talked about, you know, when you're actually physically in the Tesla, it does these rollouts, right?
Starting point is 01:16:46 and it might not do the action that the policy wants because it rolled it out and it was not good. Right. So the question is, like, can't you at that point use your model to train itself, right? So, like, if my policy says to press the accelerator and it says I'm going to die, you know, Why don't I just fix the policy right then and there? Like, why do I have to be in a real car to do that? Even just during training, I could start with something real, like start with a scene where you're driving on a cliff side,
Starting point is 01:17:30 but then inside the model, do that drive, find things where you fell off the cliff, correct them, and then do all of that inside of the model without having to need an external simulator or driving in the real world or any of that. You're just like hallucinating problems and solutions. And so that's where it becomes a world model. Okay, I see. Yeah, so think of it as kind of like baking that thing that you do in the real car, like baking that, you know, into the model.
Starting point is 01:18:10 Okay, all right. And so if you were to take, for example, Dreamer v4, they actually do this thing where they trained a model to play Minecraft. Or not to play Minecraft, maybe specific. They trained a model to understand Minecraft. So they watched like a zillion videos of people playing Minecraft and they trained a model where, again, this is all in blobs. You can't see the game or anything. But just in blobs, they swing a pickax. and then the next blob has presumably like some logs in your inventory or something, right?
Starting point is 01:18:46 And so just based on on looking at people playing Minecraft, they were able to build this model, and then they were able to train just in the model with the goal of get diamonds. And literally without touching a keyboard or being able to play even one frame of Minecraft, they were able to learn an AI that train an AI that could get diamonds 0.6% of the time, which isn't a lot, but it's still pretty amazing. They basically took this thing that had never been able to touch a keyboard, put it in front of Minecraft,
Starting point is 01:19:26 and one out of 200 times it gets diamonds. So all that learning was all done on Blobs. Okay. I think I got it. And Blobs is the what they call like the latent space, whatever. That's right. Okay.
Starting point is 01:19:42 Okay. Yeah. And then how did they know it got diamonds? Like, presumably they have some decoder for the blob for at least some things. Yeah. So they did all this training on the latent space, right, on the blobs, right? Then they froze the training and then they put the trained model in front of a real Minecraft game. Oh, okay.
Starting point is 01:20:03 Okay. Got it. All right. That makes sense. Yeah. And that's where they got the one out of 200, which is. amazing result. I mean,
Starting point is 01:20:09 now, if you continue training based on, you know, those videos that it created of itself, then it got diamonds almost every time,
Starting point is 01:20:17 which we already knew. But the fact that it could just watch other people play, build its own model of how that game works, and then get diamonds one out 200 times. It's pretty amazing.
Starting point is 01:20:28 And that's the hope that is like, as humans, you're a baby and you watch stuff around you, and even without trying it yourself, you're then able to basically do or perform it very high accuracy on the first time often. Right. But that the models generally can't.
Starting point is 01:20:45 Like today, machine learning isn't really capable of replicating that. Right. Yeah, exactly. Okay. So now there's a big debate around whether you should reconstruct the raw state or not. So the dreamer folks, which is a team out of deep mind, all of their models, even the latest one that came out like six months ago, reconstructs the original. state. So in the case of Minecraft, you know, you have your screenshot of the game and it does all
Starting point is 01:21:15 the things we talked about, but it also outputs the screenshots of the future or the screenshots as it's going. And you can actually see it like go and mine diamonds and it's all kind of fuzzy, right? Because it's, it's all going through this blob space. So it's, you know, a lot of the clouds have been destroyed. I don't say, but like anything that's irrelevant is basically not present. So that's where the blurriness comes from. Exactly, exactly. Clouds are totally destroyed because they're useless, right, to actually playing
Starting point is 01:21:47 Minecraft. But, but, but yeah, you know, it recreates the state as it goes. You can actually watch it play Minecraft while it's training, you know, in some really weird way. And so, so Dreamers definitely in the camp of, And they have a bunch of arguments for, you know, if you don't, to Yon's, to use Yon's metaphor, you know, if you don't reconstruct the leaves or at least try to, then you're not robust. So in other words, the problem with Yon's argument, and I'm just, I'm not saying he's wrong.
Starting point is 01:22:22 I'm just playing devil's advocate. Yeah. Yeah. The problem with Yon's argument is we already know that the leaves are useless for self-driving because we have common sense. But, you know, that was an assumption that we made, right? Like, it's not obvious. It's not like just implicit in the form of a tree and the leaves that those leaves are not important for self-driving. Right. It's only because of other stuff that we know. And so if you're reconstructing the original state, then, you know, you have the potential to use anything.
Starting point is 01:23:03 So put another way, maybe in the beginning of self-driving, you're just trying to do the highway, right? And so if you use Yon-Lacoon's approach, well, then stop signs that you might see when you're looking down from the highway will also get destroyed because you don't need to worry about stop signs when you're on the highway. But as soon as you say, okay, I want this car to drive on city streets, well, now it doesn't have stop signs. And so you have to start all over again. With the dreamer approach, because you're reconstructing the original state, it actually will need to understand stop signs, even though they're useless. And so then when you change your driving domain, you'll be prepared to that. So in the deep mind approach, not only do you regenerate like the Minecraft screen, does the Minecraft screen then be like, is that what goes back into the blob?
Starting point is 01:24:07 Or is that just like a loss for the model to look at and say, hey, I also need to make sure this is somewhat stable? The Minecraft screen does not go back into the blob. Yeah. So you encode the original screen from like the very first screen. shot of the game and then you never encode anything else. And so the screen output is something that's just for humans or it's still part of the loss? Like you're still rewarding better preservation of future screens. Yeah, the screen output is part of the loss. So the way you train the model is
Starting point is 01:24:42 you have a video where you have the current screen and you have the next screen and you know, the action that was taken. So you generate the current screen, make sure it matches what came in. You take the action, generate the next screen, and make sure that matches the next screen that you have kind of in your... So it's still free to take its actions and it's just living in the latent space, but the latent space needs to continue to be grounded to what the real screen would have been. Yeah, actually, I kind of misspoke earlier. In the Dreamer case, you actually do need to have the clouds in your latent state because you need to regenerate them. Okay.
Starting point is 01:25:25 In the JEPA case, you don't because you're not regenerating them. But as you said, there's pros and cons to both. So in the, yeah, so in the dreamer case, like the leaves on the trees need to be present and sort of in the right place. But in JEPA, they would be just stripped down to something that would be the tree is a problem if you hit it, but as a concept, but you don't need to know about the little dangly bits on the end. Exactly. Yeah, the leaves would be just completely gone from the latent space in JEPA. Okay.
Starting point is 01:25:57 This is like, it feels like one of those things that, to say world model or non-world model, like, it's like a very simple thing. Like, it's very straightforward. We can say these words, but the nuance is actually, like, pretty involved. Like, these people arguing about it feels like one of those things where, like, the average person, it's so far removed to understand the nuance of the argument or pick aside. Like, it's like super in the world. weeds. Yeah, I mean, you know, training on, you know, your own sort of like representation of the problem just carries with it all sorts of challenges, right? And that's unique to world models. So if you do like a model-based reinforcement learning, you don't have to worry about,
Starting point is 01:26:40 oh, you know, my latent state thinks that I got a thousand points in chess, but I can only get one point. Like that's not something you have to worry about because all your rewards are explicit. And it sounds like from having looked a little bit at it that the argument is from the world model side is to use the other approaches is going to hit like a cat, like it'll run out. It can't do the sort of plan, long horizon planning that you might need to be like there's just lots of hangups. And then the reverse would be like you said, it's like, it's not as sophisticated today. Like the world models are further behind. But the argument is, yeah, they're further behind. They're harder to kind of like get going today. But eventually they should be more capable.
Starting point is 01:27:32 Yeah. I mean, Rich Sutton actually, you know, he's famous for, well, a lot of things, but one of them is the bitter lesson. And he recently made a post where he tried to summarize the bitter lesson in 30 words or less. And I don't remember exactly a summary. I won't quote it, but the gist of it is, you know, if you, the more you remove the human out of the loop, the better it is. And that tends to just dominate everything else. So, and so in this case, you know, having a human craft a simulator, or having a human create a bunch of rules so that you don't need a simulator, like, oh, you're driving too far to the edge or something like that. You know, get, Getting those humans out of the loop is ultimately so much better in the long run than anything else.
Starting point is 01:28:20 And so, yeah, model-based RL has so many challenges. World models have even more challenges, but they will eventually win just because compute is so cheap and because, you know, humans have a hard time interacting with models at such a low level. Yeah, I mean, it certainly seems, I'm curious how, like, investor, I just off topic, sorry, or I'll off. I was like, people investing in this space that, I guess, people are just making bets or diversifying because it doesn't feel like if you're going to put financial backing, like example, to the world model from Yon Lekun getting, I think it was like a billion dollar investment or whatever.
Starting point is 01:29:05 You don't, you don't actually know. But the same was true, I guess, like Open AI and its original founding. Like, there was no evidence that where we are today with LLMs was attainable. So it's just very interesting to see. the investment sort of, I don't know to say, like, so far past the horizon. Like, it's not clear how to get from here to there. There's also just enormous survivor bias. Like, there's tons of open AIs that fell over, right? True. I think that, uh, um, I think that, you know, in the Al-Lacun's case, you know, he probably knows enough billionaires, right, that he only needs to get
Starting point is 01:29:40 20 of them to each give five million or whatever it is. Yeah. Yeah. Uh, 50 million. Yeah. Um, So that's kind of what's going on there. In general, you know, starting a world model companies, probably not a good financial investment for anybody. But, you know, at some point it's really not about, for Jan, probably not about the money. It's about wanting to do this thing. And he's got the determination.
Starting point is 01:30:10 Yeah, I think it's also interesting for these companies. Do they tackle the, I know there's been a couple startups to say, we're not going to do any consumer products. We're just going straight to AGI. Like we, it's like a distraction. So I guess that's always a debate as well for someone like Yon Nakuna or whatever. Like, do you go for the thing that's far out that is like the ultimate prize? Or do you try to take incremental steps to prove the ideas or working and, you know, gain income?
Starting point is 01:30:39 Yeah, I mean, there's a ton of these companies that go nowhere, like safe superintelligence from Ilya, the guy who started Open AI. there's just so many of these companies where they just don't really go anywhere. So, but you know, everyone talks about the few that really succeed because that's just human nature. But I do think world models are a huge deal. I think that for two reasons. One, you can't have counterfactuals. You know, you can't just figure out, oh, I'm going to not drive over the bridge, drive off the cliff, through some trickery and engineering.
Starting point is 01:31:18 No, you actually have to drive off the cliff, which means you have to do it in sim. And then two is the sim to real problem is a non-starter. You know, it's just a mess, and trying to deal with that is just a nightmare. And you don't have to. Like, people, we don't need to use, like, some expensive simulator when Dreamer can just dream all of Minecraft.
Starting point is 01:31:43 Like, what's the point, right? So yeah, world models, definitely the future. But as far as a financial instrument, maybe not the best one. Is this where you give your disclosure about an investment, Jason? Yeah, full disclosure, I would never invest in. Oh, his company has less than 100 people. I think it's a terrible idea. Yeah, I think the survivorship bias is real because it's balanced, I guess, with FOMO, right?
Starting point is 01:32:09 Like people have this like, oh, I can get in now. Now is the only time to get in and make a million X. you know and yeah it's but you know like yeah you're so much better off um betting on things that have already won and getting like a four X versus like wasting your money on a you know epsilon percent chance so this has been my strategy as well but i i i can look at other people who didn't take this strategy and are better off so yeah yeah yeah but so many other people have lost their hat, you know. That's the thing, right?
Starting point is 01:32:46 You only think about the people ahead of you, but yeah, yeah, it's a human rights. I was actually thinking about this the other day. I know we're short on time, but I was thinking about all the startups that have reached out to me, probably to you too, over the past, you know, 15 years and how almost all of them have failed catastrophically. Or, you know, if not, okay, that's a bad terminology, not failed catastrophic, but almost all of them have underperformed just going to a big company. Like literally, I can't even think of one that outperformed just staying at, you know, a big company that's successful. So, um, so, uh, yeah,
Starting point is 01:33:27 I mean, you know, maybe we could have joined one of these, uh, things, companies that are IPOing this year. But, but, but again, like, that's, that's hindsight. There's also like hundreds of other companies that didn't go anywhere. So, Words of wisdom from Jason. Yeah, totally. Run so you don't die and invest in already sure things. All right, got it. Run so you don't die and winners keep winning.
Starting point is 01:33:53 You know, in the movies, it's like the underdog, you know, Rudy. Remember Rudy, the football movie? Yeah, yeah, yeah, yeah. Yeah, the 4'10, you know, linebacker, like, just through his heart and determination, like, wins Notre Dame. I don't know. I've never seen the movie, but, but like, in reality, the Brock Lesnar, who's like 6'6 and all muscle,
Starting point is 01:34:14 he just wins and just crushes everybody. Like that's how the real world works. I hate to break it to you out there. But like winners generally just keep winning. Wow, dude. We get spun up. It's late in the episode. I feel like I'm ever going to get a rant here.
Starting point is 01:34:31 Do you have an X account for where we can hear more? Oh, man. All right. Well, thank you for, yeah, I feel like that was, that was great. I enjoyed hearing the explanation and super pertinent to current events. Totally, yeah. If you're out there, you're working on world models or have any questions. Hit us up on Discord or shoot us an email and you can get links to all of that from our website,
Starting point is 01:35:01 programming threadon.com. Or one last shout out. I guess if you're working at one of them, maybe you have a private jet and can fly us out and we'll interview you. That's true. Yeah, we'll do interviews in person. We just need a private jet. I've never been on one. It would be cool. It's not cheap. No. No. All right. Thanks, everybody. All right. Catch you all later.
Starting point is 01:35:38 Music by Eric Barn Dollar. Programming Throwdown is distributed under a Creative Commons attribution, share a like, 2.0 license. You're free to share, copy, distribute, transmit the work to remix adapt the work but you must provide attribution to Patrick and I and share a like and kind.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.