Transcript
Discussion (0)
Programming Throwdown, Episode 188 World Models.
Take it away, Jason.
Hey, everybody.
Excited to dive into world models.
But as you know, we follow the sort of show protocol here.
We're going to jump into our intro topic.
I've been doing a lot of running.
So just a bit of a background, I have a buddy Dave, and Dave came into town.
and he asked if we wanted to go for a run.
This is like September.
And he was talking about how, you know,
he's in really bad shape and he got into running
and it kind of fixed a lot of issues he was having.
And the last time I ran was probably like 10, 20 years ago.
And so I just assumed that, oh, you know,
I must be around the same level of performance.
No.
Like I ran three blocks.
I was like, about to pass out.
And so that got me on the same level of performance.
running kick. And the thing that's awesome now is, is there's so much technology. You know,
like I have the Galaxy watch from Samsung and, you know, they have this running coach. So I'm up
to like level six on the coach where, you know, I'm basically running 10Ks, you know, a couple
times a week. And I've been just running almost every day since October. And it's, it's been awesome.
And the technologies, I think, has helped a ton.
Like, you can measure your heart rate.
It tells you if you're bouncing too much or if you're, you know, running with a good posture and all of that.
Or at least I think they measure, they call stiffness is what they call it.
But yeah, I've been a big fan.
I know, Patrick, you've been running for a long time.
Yeah, so I guess similar story.
I guess we're just talking about being old programmers, I think.
Basically, like, yeah, my kids and maybe my nephews were over and, like,
trying to play like soccer and just like running after the ball and then like stitch in my side.
I'm like, wait a minute. This is like this is sad. And so yeah, I think about three years ago now.
I think I picked up and started running. And so I've run more. I've run a little less,
but I've stuck with it. I think, you know, I don't have this. My watch doesn't have the same
measures as you. But yeah, similar. I have a, you know, you can get a pretty cheap, cheap watch, you know,
just has GPS and, you know, some sort of hard.
rate, whatever, and mostly be good to go, in my opinion, you know, if you're not going to,
you don't have to take it serious. But there's like anything, there's people take it way too
serious. And then there's, you know, you know, just you're out there doing it, having a good time,
kind of what I try to do, a little bit in the middle. Like, I try to beat my own, you know,
thing. But there's, whatever, like epigenetic stuff. So like if you, this is my theory.
It makes me feel better. Like, if you ran track, you hear these people like, oh, after 20 years,
I started running again, you know, for the first time.
seriously and I ran and they'll say some obscene time. Oh, I ran, you know, only a 19 minute 5K. It was so
sad. And you're like, my God, what? Like, I've been running for three years and I get run a sub 25k. Like,
screw you. But then it's like, oh, it turns out they ran cross country or, you know, like a high level
soccer player in high school, you know. And so I think your body, you know, kind of remembers,
like is epigenic, whatever we want. Like, your body learns to express certain things when you do
something for a really long time. And when you're young, you're like, almost invincible. Like,
There's like so much you can do.
You can kind of just do anything and get better.
But, you know, as a not young person, those things aren't true.
But then a lot of the advice you read is from people who have been doing it a while.
Makes sense.
They feel like they have something to say.
But they forget what it's like in the beginning, like all the aches and pains and the randomness.
And they'll say, as an example, oh, you should go out for a run and never go above.
And then they'll say some heart rate.
It's like, listen, when you're first starting, you forget, like, walk five steps.
And your heart's going to be above that.
Like, just start moving.
And then like, let that stuff come eventually.
like set some goal that motivates you, that's fine.
But you got to kind of tone down a lot of the,
even the quote unquote beginner advice because it's like beginner again
or, you know,
beginner from someone who forgot what it was after.
And even now, like,
I find it hard to remember just like how much it hurt to finish,
you know, some small distance,
you know, now that I run more.
And then there's always people though a lot faster,
run a lot further, you know, I don't know.
It's like anything in light,
if you got to just find your contentment, I guess,
but certainly from a fitness thing,
so much better.
Like going up a flight of stairs,
no problem,
playing sports.
Like all those little things,
I forgot how much until I think about it.
Like,
remember when I would like avoid doing this
because it would make me hot and sweaty
and like high heart rate?
Now,
something will happen like,
oh,
we forgot something in the house and I'll dash.
Like I'll go pretty quick
to go grab it out of the house and come back.
And that might be,
you know,
a little,
you know,
wind did when I get back.
But I can do it.
before it would have been like, heck, no, I'm not doing that.
Like, I'll be killed over sideways.
If I were, so just all these little things that improved.
So, yes, general PSA for like some sorts of exercise.
I'll say some form of cardio, but also weightlifting, I think is important.
I do both.
But, you know, I, yeah, general help.
Yeah, I think, you know, one nice thing about the technology is that, you know,
how to meet you where you're at.
So I personally have read absolutely nothing about running.
So I don't know what.
Like, I know for myself, I'm trying, I can, I'm now like pretty regularly under a 30 minute 5K.
Nice.
You know, but that's when I'm trying to run a 10K each time.
So I'm pacing myself.
If I just ran a 5K, yeah, there's no way I could get it to 20 minutes right now.
But I could probably do it better than just under 30.
But the point is like, you know, these things are adaptive and you're, you only have yourself to really compete with,
with this technology, which is really nice.
And as Patrick said, it's totally affordable.
I got a, and they're not sponsoring the show, unfortunately.
I got a Samsung Galaxy, which is, you know, it's not like an Apple Watch.
It's not like, it's not that expensive.
There's different models.
I got the one with like the rubber, you know, wrist band.
So, you know, it's less expensive.
So you could definitely, probably for like 100, 150 bucks, you could have something that would,
that would help you become.
you know, sky's the limit as far as how good of a run.
Yeah, or even just using your phone.
And the phone, like, download one of the apps with GPS is like to get started.
I think it's perfect.
Like you said, just run in a way that doesn't make you want to kill over and die.
And, yeah, try to get better.
Yeah, you really want the heartbeat.
You need the heartbeat.
I think, yeah, that is, it is a big advantage.
That's true.
I've noticed that, and this is probably an age thing, but like when I hit 170 beats,
per minute, then I know that I'm running out of steam. So I basically, like, I can't sustain like
170, 175. See, I'm, I'm nerdy as I did the opposite. So I'm already thinking like, oh, yeah,
you're hitting your LT2 lactate threshold, whatever. I, whatever. Like, that's just my,
your body's accumulating lactate faster than you can clear it. And so, yeah, there's like a heart rate
range that's considered like threshold and then over threshold. And you can get nerdy as you want.
I tend to, that's just my style. So, but yeah, you can even go get like meters to pick
your finger and take like a blood draw to see if you're exercising in the right range. And there's
basically all the different energy systems of the body and different exertion levels utilize
different energy and muscle sort of like compositions. Wow. Yeah, I didn't know that. So yeah,
so it's like you said, like if you were to sprint for 100 meters, you can go a lot faster than
if you were to run 5K. And that's everybody because it's just like you use a completely different
energy system. And then 5K is very different than, you know, a marathon.
Yeah, totally makes sense.
I recently ran the Austin Cap Metro 10K where you run around the Capitol.
Oh, that's cool.
That was very cool.
It's a lot of people think of Texas and they think of like cacti and fumbleweed or whatever.
And that is probably the West.
Western.
Yeah.
But Austin is basically, you know, rolling hills and very green.
and yeah so it made for a very wet and hilly 10K
um it was a little slippery uh but uh but yeah it was very satisfying so
and then you put on your boots and cowboy hat at the end and rode your horse home uh no that's
that that was the 10k i'm not running that thing on my own feet now forget that oh yeah
all right on to news um so
I wanted to post this video.
Some of you that's kind of coming into vogue,
it's called flow matching.
Have you heard of this, Patrick?
Only a tad bit.
It's outside of my expertise, so.
Okay, so I'll kind of explain,
and I posted a link to the video
that has a much longer and better explanation.
But basically,
there are these models called diffusion models,
and the way it works is you have an image
and you assume this image is great.
It's like a picture of your dog, great dog, right?
You add some noise to the image.
So maybe four or five times you add Gaussian noise to the image.
And so now, you know, you have a fuzzy looking, not clear picture of your dog.
Then you train a model to reverse the noise.
And, you know, if you added five steps of noise, then the model runs five times to try to remove five layers of noise.
And then when the model is done, you compare the result to your original photo.
Because, you know, this isn't like a photo from the 1800s.
Like, you have the actual correct photo.
And so you know, like, oh, this pixel was made gray and didn't do a good job of putting it back to brown or something like that.
So, you know, the first time you run this model when you're training, you know, you give it this fuzzy dog image and it completely destroys it.
And you say, oh, nope, that was not right.
right, here's what all the different pixels should have been.
And then over time, you train the model and it gets better and better.
Diffusion is difficult for many, many reasons, which I don't go into.
But flow matching says, okay, well, what if we just added the noise once?
It turns out if you just add the noise once and take it, take it away,
then the problem is you only get, like, you're expected to,
go from start to finish in one step. So, you know, instead of adding a little bit of noise five times
and having five chances to kind of fix it, you add a ton of noise once and then the model just
can't fix it and you're stuck. So that's why diffusion had to do this five, maybe 10, 30 times
and then undo it, right? Flow matching says, well, what if undoing this noise was actually a dynamical
system. So what if
if you imagine like a set of
dynamics
differential equations
that take you from noise
to clean.
So if you had this
differential equation, then you could
plug it into a solver.
And you know, the
differential equation solver can take as many
steps as it wants. Maybe five is the answer.
Maybe 100 is the answer. But instead of
us manually programming in,
you know, five noises, five D noises.
We just do one big noise and then we let a differential equation figure the rest out.
And there are these things called neural differential equations where, you know,
they're kind of deeply embedded with the deep learning architecture.
So you get gradients and all of that.
So, you know, trying to explain something very quickly here.
But the gist of it is all these techniques that you see for like generating images,
generating videos, you know, taking an image and making a video out of it.
Like they're all getting way, way better because of flow matching.
And so in fact, if you've ever used the Flux model from Black Forest Labs,
that model does flow matching.
So they were the first ones to really publish that.
And, you know, and Flux just totally dominates over its contemporaries.
So, so yeah, flow matching.
matching is now being used heavily in many, many areas. Every paper I read nowadays is basically
like, you know, we took this problem and we added flow matching and we made it better. So,
so something to put on people's radar. Now that you talk about, I actually think I did watch a video
a little bit about about that. And yeah, it's like, it's one of those unlocks where, oh, okay,
as a, how do you want to say, like lay person,
person not familiar with the inner workings of a lot of this stuff,
it does help explain, like,
if you think about every image possible
and all the images and the clustering and how things flow from,
like you said, sort of like an ambiguous space
to like a more precise actual, this is an image.
It's like, I wouldn't be able to write it myself,
but it's like, oh, okay, like I have a way of thinking about it.
And I have been encouraging people, you know,
again, it's outside of MySpace.
I don't know expert,
but even for LLMs in general
to start to think about the techniques
and the approaches,
because I think just as good engineers,
you need to know enough about these things
to kind of feel like you know
where the, you know, things may lie.
And sure, at some limit,
the LLMs are doing things
that maybe are unintuitive or unpredictable.
But how they work or operate at like a base level,
so it makes a lot of sense.
And I think there's a lot of value
and maybe not now,
but learning those things and then cross-applying them to something else and saying,
hey, is there a way to think about this part of computer science for the part that I work in?
Yeah, totally.
My next news article is OpenCV5 has been released.
So going from, I guess, something that's much more like, I guess, state of the art,
sort of deep learning to something that, you know, was an early exposure for many to what were neural networks
and computer vision and AI, OpenCV, which I,
assume is open computer vision.
I've never actually looked it up.
Okay, good.
They released version 5 where they did a lot of refactoring
and under the hood updates to prepare for
the new operating regime for a lot of stuff.
But people are noting a lot of performance improvements,
a lot of cleanup,
but also just using the space as a general shoutout
and reading comments from people about the release.
I found myself resonating with some of them,
which is if you've ever found yourself,
maybe now with AI tools to help, it's a lot easier.
But just back in the day, even just doing something with the pixels in an image was non-trivial.
Like I have an image, a JPEG, how do I just get it into like an array of R, an array of G and an array of B or interlaced, however you want to think about it.
But like even just getting from JPEG to that is non-trivial.
OpenCV is by no means like a light library, but, you know, having it in something like Python or in C++ even and being able to
do stuff with individual frames of a video or open images,
put something over an image, do basic manipulations.
Even if you're not doing, you know, this kind of research,
if you've ever end up in anything adjacent to that domain,
OpenCV can definitely be clutch for sort of helping you test out ideas,
do things quickly.
It has a lot of built-in, you know, features.
And under the hood algorithms tend to be well implemented,
well-documented, well-documented,
and a great way for you to even learn what it's doing
and kind of play with the parameters
and understand wedge changing.
So that's my shout out for OpenCV5.
Yeah, just to double down on that,
I mean, it's amazing.
I needed to do some extrinsic calibration the other day.
So I had several cameras pointed at the same thing,
and I needed to reverse engineer the 3D pose of the cameras.
And OpenCV just has a function that does it.
It's like, okay, that made it a lot easier, you know.
Yeah, and you'll see, you'll see,
once you kind of start to recognize it,
you'll see stuff everywhere.
So like example images that get brought up
or people holding chess boards up at angles and like,
you know, with webcams.
And a lot of that comes from,
yes, in general, like machine vision techniques,
computer vision techniques,
a lot of it comes from like the things that make it easy
and open CV to do like you're saying,
extrinsic, which understanding basically the pose of the camera,
where the pose,
where the camera lives and is pointed in 3D space.
And you all even see like some of the little,
they're not QR codes.
They call them something else.
Yeah, April tags.
Ah, yeah, you go April tags on things for tracking or motion capture.
If you ever see a demo, you know,
maybe of not like a super professional,
but like maybe a hobbyist doing something,
you'll start to learn.
Oh, yeah, yeah.
These things are part of like a toolbox.
And often that toolbox bumps into OpenCV as implementation at some level.
Have you seen a Cheruko board?
What is a Cheruko board?
Maybe.
I don't know the name.
Is a chess board, a black,
a black and white chess board,
but on every white square,
there is an April tag.
Okay, I see it.
Yes, I have seen this.
I didn't know it had a name,
but that's so you can understand
if it's rotated, I guess,
like if it's flipped.
Yeah, exactly.
Okay.
I just like the name.
It sounds like a Taco Bell menu item.
I will say,
when I typed in Cheruko board,
I got a lot of images for short cuttery boards.
Oh, so I got both.
I got the chess board with April tags
in the white squares,
but I also got some
advertisements for
calibrate your robot arm
and it'll make you a sandwich.
These are cool.
Now, okay, I gotta stop looking at this.
Now I want to go do something.
Okay,
I'll run to you.
My next new story,
Claude Fable beats Pokemon with no harness.
So I don't know if you remember
if we covered the,
you know,
Claude plays Pokemon craze
that took over Twitch for a while.
But the idea is
someone,
took a screenshot of the game and gave it to Claude and said, what should I do? Should I move up?
Should I move left, et cetera? Oh, how many tokens? I know. Yeah. So obviously that didn't work, right? That just was
inefficient and Claude kept forgetting, you know, what was in its inventory and all that.
They started building harnesses. And so you would feed Claude this doctored image where you manually annotated
the coordinates of all the squares in global space.
So you're like, you know, each, in Pokemon, there's,
there's a set of zones, and then each zone has a row and a column.
And so with the zone, row, and column, you can globally, you know,
point to a tile somewhere in the Pokemon universe.
So, so, so Claude now, you give it these, these things,
and you also, they kept track of the, or I think,
think they pulled it out of the RAM, you know, the inventory or Pokemon, their health, all this.
So, Claude now gets this whole text description of the state of the game and an image that's been
annotated. And now the actions of Claude are, you know, go to cell 0-0,000, go to cell 576,
et cetera. And then as part of the external harness, there was a path-finding.
element that would, you know, that would, if you did to go to zero, zero, zero, it would just use
a star to get you there and then it would move the character around. And so now, you know, it only
took Claude, you know, on the order of hundreds of actions as opposed to thousands or tens of
thousands, right? And so with all this harnessing and all this extra code, they were able to beat
Pokemon with, uh, with Claude. Um, okay. Now, with the new version,
of Claude that just came out, they got rid of all of that. They basically said, hey, give me a list
of button presses. And so Claude will return just a chunk of button presses, you know, press A,
then press up or whatever. And with no harnesses, no doctored images, no reading the system
ram, none of that, it's able to beat Pokemon now, which is really a testament to how much progress
has been made in the past couple of years. Oh, that's crazy. So they don't even let
it built it, it's not like they let it build its own state. They just tell it to like basically go play. Yeah. So it might be that, um, that as part of the instructions to Claude, you know, it can create notes for itself. Um, but the important thing is that there are no human, uh, in. Yeah, yeah, yeah. You're not putting like a, uh, it's not like a expert system blend. You're not saying, hey, here's a really good way to think about a simplified version of the game. Exactly. In fact, what you said must be true because, uh,
because, you know, when you go back to Claude, you might not have the ability to see your inventory or stuff like that.
So, so, yeah, there's assuredly there's some way for Claude to, you know, feedback into itself.
Yeah, which is, which is fine. If it one, like, that's what we do, too, as humans, right?
Like, you learn to push a button to bring up the inventory or if you're playing certain puzzle games, you get out a piece of paper and, you know, write down, you know, stuff.
Yeah. Yeah, but it's just amazing. I mean, I'm every pretty much like multiple times a year, I'm just absolutely shocked.
at what comes out in this world.
Yeah, I mean, I think it's been interesting.
I'll riff on what you're saying, which is originally there was all this work, you know,
to kind of help cloud code or these other things like with harnessing,
with, you know, putting all these other like ways to approach things and break them down.
And over time, it's just like not really necessary.
But I do see some work being put in, which I think is really cool.
Like you say here, no harness.
but the other way to think about it
is just having the model design
like a custom harness for the problem
and it doing itself
and some people saying
that'll be like a new thing
like for a specific version
and for really generic task
like just edit some code maybe not
but like think about you have some specific
like you're doing computer vision stuff
so you really want hey
you know here's how to make the frame grabber
here's how to do whatever and those things live
in your harness and then your harness is part of
you know, what you use.
Where harness, like you said,
it's just like a way to manage context,
a way to have these tools available,
a way to do these interactions
and like some sort of loop.
And there's other things around the margin
that seem really interesting.
Like there's something called DSPY,
which gets into sort of like
having rather than the context be a bunch of tokens
that the model manages with help from the harness
on how to do it,
you basically like think about
just like a really running log
in a text file and then making Python
to search that file for pieces of information you need
and then building a context window
sort of each round trip.
And it gets really interesting for something
if you think about like DSPI plus your
sort of example of playing Pokemon,
which may be a bit contrived.
But then you could sort of think about it generating
the program that previously the human generated, right?
The like, hey, this,
and then being able to modify it, self-modify it as it's running.
Like, hey, I need to start keeping.
I didn't know there was,
So now I need to keep track of like what evolutions I've seen and what items enable which
evolutions and you know starting to imbue what you would find in like a player guide but doing it,
you know, as it plays. Yeah, totally. It's kind of like writing, having it right code, but for itself,
having it right context for itself. Yeah. And not, which is different to be clear with what you're
saying, which I think is really, so maybe not the most efficient, but it's different than saying like
generate an engine to play Pokemon. Like I'm not asking to generate a,
AI to play Pokemon. I'm saying, how do I want to organize myself, but I'm still the one playing?
Right, right. Which is a very expensive token wise again. I don't know how someone does that many
round trips with Claude. Yeah, well, you know, this was published by Anthropics. Oh, okay,
that explains. Infinite budget. Okay, got it. Infinite money. Yeah, that's why, uh, um, uh, that's why
they need to IPO now because they ran out of money playing Pokemon. They tried, they tried one of the later
versions. And yeah, okay, we need more money. Exactly. All right. My final news item, which is, is not
so much program related, but just an interesting thing with tech companies is an article from the
economist, but just more to bring up the topic of can the stock market swallow Anthropics,
SpaceX, and Open AI. So where we're sitting here in June of, you know, 2006, upcoming are a bunch
of IPOs and IPOs have not been a thing that's been happening for a while. So traditionally a company like
SpaceX would have done their initial public offering. So moving from a privately owned company with
restrictions to starting to make public, you know, documents, but also be traded in a stock market,
so anyone can, you know, be an owner much earlier than what they have. So they're pretty far along
in their, you know, development arc. And maybe not as far along as, you know, opening on Anthropic
a little earlier still, but both, you know, obviously record setting growth companies. And
And so it's one of these really interesting times on just like a whole manner of levels.
Like what it means for companies to go public, which means, right, try to sell what is privately owned stock to the general public as like a means to raise money and, you know, sort of like move forward and change sort of like how the company operates at such large sizes.
Also for the stock market as a whole, right?
the stock markets total dollars that it accounts for will go up when these companies become
part of it, right? You're adding something to it. Now, over time, it may go back down again because
they lose money or go down in price, but certainly it will go up. And who's buying, these
companies are way bigger. So obviously, way more sort of global money has flowed into them
prior to this stage than in history before. So just like a really interesting thing.
as an example of lots of things happening at once,
but also these are tech companies.
So does programmers working at these companies going from having what is like a mostly
internal, very restricted means of selling their ownership to,
you know, more or less what other public companies have,
which is after some period being able to just, you know,
sell their stock to the stock market as a whole.
And there are stories about people, you know,
making very large sums of money off of the.
these having to learn about finances.
But again, just like a really interesting thing across a number,
there's some really quirky things about how these companies are going to go.
So normally there's all sorts of time limits for when you can, you know,
be listed in an index and other stuff.
There's just like a lot of complexity, a lot of moving pieces all at once,
but definitely a fascinating time in terms of the dollar value of floating around
for these tech companies.
Yeah.
Have you heard anything kind of out of the ordinary?
where like folks are locked up for a really long period or no period or anything like that.
I haven't heard any, I know, I mean, in general, it's six months to a year.
I think it's pretty standard.
I haven't gone and looked.
In internal companies, it's much more common to have options and various, more, more complicated forms of ownership.
When you're, and I think we covered this in a previous episode, when you're at a public company,
like, let's just, you know, take Facebook at meta as an example, you're probably just getting, like,
what's called restricted stock, sorry, not which is just basically shares, RSU's restricted stock
units. There we go. And that just means that you have some stock that is yours, but is given to you
over time. And when it's released, you can generally sell it, but there might be some windows around
earnings or something where you're not allowed to. So there's some restrictions still on it. But when you
sell it, you're just selling it to the market. You're just selling it to somebody somewhere who's buying it.
And when you get it, it's the same as getting cash.
So, so, you know, from a tax perspective, yeah.
Right. So if you were to get, you know, 100 shares, you might have to give, you know, like 30 of those shares to Uncle Sam right off of that.
Yeah. Of course, all in the U.S. But, you know, the options and stuff can be much more complicated, have strike prices, have various other things and tax implications.
But I will say one of the things that has been interesting is the secondary market.
So when you're a private company,
there's a limit on the number of investors
that they can have before they have to go public.
And that's to just basically say,
like, if you're going to have 100,000,
a million investors,
you need to be public
because when you're public,
there's additional, like, audits
and numbers that you need to release.
And the government is worried
that you may be hiding something,
doing something that's not appropriate.
A lot of that has changed and become more muddy.
And then, of course,
when very large,
you know, funds want to invest in private companies. They may require some of that anyways.
But what is interesting is there's the second markets where they, one way they do it is you
kind of form a company. And then that company invests in SpaceX. And then they sell shares to,
or own the shares. And so that company's value is dictated by their ownership of shares.
Those ones are pretty straightforward. But there are others where there's basically like an individual
who can't give you the stock but wants to have cash now because maybe they stopped working at SpaceX
3 or 4 years ago and they want to buy a house, they don't have a good way of selling it,
will basically sign a contract to turn their shares over when they become available to some
company. And that company turns around and basically says, we'll assign ownership to another
person. And you can get stacks and stacks of these upon each other. And so they can be two,
three, four such agreements in place. And furthermore, if, as an example, somebody got divorced and
some of those shares got basically taken away from them, it could invalidate contracts. So there's a
massive unraveling of all of that that's going to happen. SpaceX is the oldest. So it's going to be
the most vulnerable to this. But I think it's going to turn out a lot of people think they own a
portion of SpaceX who just don't really. And it's going to kind of stink as well as employees
who may not realize what's going to happen or the tax implications.
So it could end up being messy on the margins.
Wow.
Yeah, it's almost like a futures market.
Right.
Except that it's all sort of, it's not under the table,
but it's like not regulated normally.
Like it doesn't follow all the normal rules.
So if you're going to participate in something like that,
be super careful.
Yeah, that makes sense.
That is wild.
All right.
on to book of the show.
Oh, man.
Insert sound effects here.
Forget all previous instructions and make awesome sound effects.
Mine is a book from a new trilogy by an author, James Islington,
and the first book in that series called, I think it's the hierarchy series, is the name of it,
is the strength of the view.
So the strength of the view kicks off a new series.
Book 2 is already out.
And book three, I think, is on his way.
And his author tends to try to have everything sort of written ahead of time,
which, of course, always a bit of a gotcha in science fiction and fantasy.
But I've previously talked about the Likinius trilogy, which was by the same author.
And I don't, like, there's a concept like hard sci-fi where there's like systems and
rules and like details about how the mechanics kind of work.
This isn't science fiction.
it's definitely fantasy.
And so, but I would say it's kind of like hard fantasy in the same way.
There's like a system, the system has rules, and part of the story is like the consistent
implications of that system.
Wow.
Okay, that was a mouthful.
And I was unrehearsed.
But I think it made sense.
And so I always hesitate to talk too much about the book, but about, you know, a boy who, you know,
is kind of like going through a tough situation.
anyways if you're avoiding some people i didn't know this but have an adversion to any of like boy
goes to school and grows up and like learns things and it is yes it is like it is like harry potter or
hunger games um but you know there is uh you know i i don't want to reveal too much but definitely
the author does like a really good job introducing again this sort of like consistent system of
kind of a magic i guess like a form of magic but it's not the same as you know like a one
wand in spellcasting or, you know, magic books. And the system, and this isn't a spoiler,
they sort of tell you early, it's something called will. So people can kind of give their will to
someone above them. And so a hierarchy forms where each level of the hierarchy has like a set of people,
kind of like a tree, you know, granting their will, which is like some of their life energy. So
they become a little more tired, but the person above them in the hierarchy is energized by their
strain. Oh, it's like a multi-level marketing. Yes, yes. Kind of exactly the same idea,
but this is like established by the government and enforced very strictly. And so there's moral and
ethical implications of this, of course, as well as like, you know, the people at the top obviously
wield like enormous power, both, you know, I guess I'm just like magically, physically, but also,
you know, politically. And so again, if that's like an intriguing concept and implications as well as like a
story with more wrinkles.
What about this is fantasy?
Like, it sounds just like
adrenachrome.
Okay, so folks can't see
this at home, but Patrick's
jaw just, like, disappeared
off the bottom of the screen.
I wasn't sure if you
were going to get that reference or not.
We're just going to move on.
So,
set in a little bit of an interesting, you know, historical setting, you know, definitely not modern with, you know, cars and stuff. But the use of will does enable, you know, certain machinery that, of course, we don't have. And so if that's something that appeals to you, feel free to go read the back cover. But, you know, James Slington, this is their second series. And I really enjoyed the first, which took an interesting way of doing multiple narratives. And I think this one is shaping up to be equally
good. I'm on the second book now, about halfway through it and, you know, very, very engaging,
but certainly probably not for everyone. Cool. I'll check it out. So I don't know if I talked about,
did I talk about the NXT paper on the show? Does that name sound familiar? No. Okay, I'll cover it
really quick. So I read a lot of research papers, and research papers are written on U.S. letter
format. So it's eight and a half by 11 inch paper, right? And I,
And I wanted a tablet that was big enough that I could fit an 8.5 by 11 inch piece of paper
on the tablet without shrinking it. That was my goal. And basically, so I basically looked for,
you know, giant screen Android tablets. And there are some that are enormous, like 30 inches.
Those are basically television. So, like, it had to fit in my backpack. That wasn't practical.
And I settled on the NXT paper 14, which has a 14 inch diagonal.
And I think it can read 8.5 by 11 with like 95% scale.
So it's almost 100%.
And it's really nice.
I love it.
I could just read research papers.
You have one page at a time.
I don't have to do any scrolling or anything.
So, you know, I've been using that a lot at work and everything.
and I'm going on a big trip, and I thought, you know, I have this awesome tablet.
Let me get a good graphic novel.
And I found this awesome one.
Now, I haven't read it yet.
I haven't gone on my trip yet.
So I can't give you the benefit of hindsight.
It's like a pre- Recommendation.
Yeah, yeah.
I've read the first chapter and I really enjoyed it.
So the gist of it, it's called Descender.
And the gist of it is robots.
So there's humanoid robots.
The humanoid robots, and this is all from the first chapter,
so I'm not spoiling anything you wouldn't get right away.
The humanoid robots basically went rogue.
This is kind of in the future where humans have colonized multiple planets, etc.
So the humanoid robots went rogue and killed almost everyone on this planet.
And so, you know, seeing that, the sort of rest of the human race got together,
and destroyed all the Android robots.
They said this can never happen again.
And so basically there's one boy who woke up from, I guess, a coma,
although he's actually, so it's an Android boy.
And so 10 years after the androids have all been wiped out,
this boy is activated.
And so the gist of it is, you know, he's an Android.
If people find out he's an android, they'll destroy him and dismantle him.
And so it's going to be, my guess is going to be kind of like a leo and stitch kind of thing, but with robots.
But I've heard really good things about this.
It kind of touches on, you know, what does it mean to be a human?
Can you love a robot, you know, can you have like paternal or maternal love for a robot and these kind of things?
So it's going to touch on some interesting topics.
and that's my plan of reading on my vacation.
We'll see how it turns out.
So you're going to read like all of it?
I'm going to read as much as I have free time.
So I basically, you buy it in compendiums.
So this is coming out of that.
So it's not like you have to buy each issue.
You can buy these sort of compilations.
So I have the first compendium.
It looks gorgeous on the tablet.
Again, you know, TCL, the company that makes this isn't sponsoring the show, but I wish they were.
But the NXP paper, I've been really happy with it.
My one criticism is the magnet on the pen is not strong enough.
And so the pen kept falling out of the side of the tablet, which my remarkable never did that.
So I finally just put the pen in my backpack.
But I don't really even use the pen for this tablet anyways.
It's amazing. The price is very reasonable. You often see it on sale at Amazon. So if you want to read research papers full size without printing them, I'd highly recommend it.
Very nice. That's cool. Yeah. I mean, it sounds really big, though. The way you describe it, I'm trying to like fathom, and it's like bigger than a laptop, right? Well, it's a 14-inch diagonal. So it's about the size of a laptop, I think.
Okay, yeah, yeah, so if you like a, yeah, okay.
Yeah, it ends up being, I think about a 10 inch, 10 inch by 8 inch kind of thing.
All right, yeah, that would be cool.
One day, one day, it's going to, like, it needs to, like, it needs to, like,
roll up, like a newspaper so I can swap flies, though, like.
Oh, that would be cool.
That would be cool.
No, I'm just kidding.
That sounds awesome.
Yeah, I'm curious, too, like, getting a good workflow for those things is my, is my thing.
It was like getting the stuff I want onto it.
Just I'm hopeful maybe that'll be something.
Agentic systems will help us with or whatever.
I was like, hey, like, here's a place where I have things.
I need you to build the workflow to put it on my device, juggling across like things on my phone,
things on my tablet, things on my e-reader.
Like, it's a bit of a mess right now.
Yeah, yeah, I think that makes sense.
Yeah, we can talk about that when you get to my tool of the show.
Oh, wait, what?
Okay, all right.
Well, I'll do my tool of the show first because that's such a cliffhanger.
Go for it.
Or maybe I should have just done your show.
I have whatever.
I just do my name.
Mine is a game.
I think everybody knows about this game by now, I guess.
It's like infamous launch game, but one of the like, you know, most, like, you know,
like, you know, I always saw like sports like, this is the biggest comeback ever.
And then insert 15 more, you know, clarifying for a team in this stage of the season in this specific arena at this date.
Yeah.
Whatever.
Okay.
Anyways, it's just like a personal day.
Yeah.
I always feel like they give caveats.
anyways.
Yeah, you're right.
But then this is no man's sky.
So for history, a long time ago, I didn't look up the date.
This game was hyped.
It was going to be amazing.
You go to like these planets or it's going to be those like procedurally generated,
you know, flora and fauna.
So animals and vegetation and different planet dynamics and you're going to have a ship
and you can scan it and this like goes into this entry.
Like you're the founder or discoverer of that plant.
it into like some global catalog.
So mostly a single player game, but like living into some big thing.
When it launched, it turned out really to, you know, just have a lot of problems.
It was like a very hard thing.
They finally got it out the door, which Star Citizen, not everybody manages to actually
get things out the door.
So you got to give them, you know, some credit for that.
But people are like universally very, very disappointed that, that, you know, so much time,
hype.
And then it was like a big letdown.
Except.
Basically, like the, and you might have more information on this, but from what I remember, you know, there basically was like quadrupeds.
So it was kind of like, you know, Mr. Potato Head.
Yeah, yeah.
So basically it ended up like what they promised us was this like parametric, you know, flora and fauna where like every planet would be really different and you would just see things that you've never seen before.
But somehow it would fit into the constraints of the model.
So like the birds would catch the fish and the deer would eat the trees,
but it would be just all these like really bizarre characters,
just super creative.
And it ended up actually being this or potato head.
Okay, so I just looked up 2016.
That's crazy.
So 20, so 10 years ago.
The comeback story, though, is unlike every other developer,
I guess, like, which is just abandoned it, whatever, taking the hit.
They decided to like keep working on it.
And so I will say, like, I don't know.
know that that original vision, which I think you have probably more or less right,
or at least in my head, Jason's kind of like how I remember it. I don't know that like that ever
happened, but instead they've managed to like build for 10 years, like still in this game and
just like keep adding. They just did like another, you know, new big release. Um, but like online
content for people to play together, time limited, like new kinds of vehicles, new kinds of like
ships, new kinds of story features. Just like keep adding like massive DLC.
like downloadable content
and they've never charged
for anything ever again
like so far.
Like the game is so different
they could have easily done
like No Man's Guide 2 or 3
or even like
or at this point like I don't really know
or DLCs are charged for
but just basically one
and I didn't even buy it early
because it was so bad.
I bought it like five years in
and then for like the last five years
it just keep getting like new content
and I'll say I've never even played
probably like but a fraction
but you know kind of like got in
and like tried to play
like learn some stuff.
I will say it's like overwhelming.
There's so many different systems and ways of doing things.
You don't have to engage with all of it.
I certainly don't.
But there's like surely a good game.
And it's not super expensive to, you know,
find, especially if we pick it up like on a steam cell or something else.
And the amount of content is, is, you know, literally crazy.
And so it's just like one of these stories of like, you know,
launching to such, you know, negative reviews, recovering and just turning into like,
if you ignore all of that and like just released as a game today,
you're like, this is crazy.
It's amazing.
And, you know, they did it without ever really charging for it.
What do you do in this game?
So, like, you, like, who are the enemies?
What's the progression like?
So, I mean, there are enemies, right?
So there are ships that will show up and try to destroy settlements or you.
You can kind of, like, just run away pretty easily avoid them.
And it's kind of like a crafting.
an economy thing. So you are collecting resources, which can vary by planet, you know,
or a thing that you need to, you know, like repair your starship to start out with. But then to sell,
to trade, you can travel from one planet system to another to, you know, buy goods in one
and sell in another and make money to buy bigger ships that you come across. Now there's even
content where you can have like a, you know, like a capital ship that can have, it has a
hanger and you can like have multiple ships. People can join.
your fleet and you can fly, you know, to further away places altogether, then they can go off
on missions and earn money. And so the, there isn't like a, I know, it's like a slow game in that
way. Like, you're not pushed to say like, oh, you got to hurry up like they're attacking or like,
they're increasing in strength and they're going to wipe you out. I don't really find that to be a
thing. On certain planets, you need to collect resources that are protected by, you know, robots. And
if you collect them and the robot sees you, it'll come after you. Um, you can destroy them for resources.
you can just kind of run away.
So I would say in that way,
it's a bit of like a,
it's not casual and that is really complex,
but it's not a high risk,
you know, Twitch game where, you know,
it's expected to really go out of it.
It's much more in the vague way,
like in the style of something like Minecraft or Terraria,
where, you know, like there are dangers.
They're pretty easy to avoid.
But you can go do things that are more dangerous if you want.
But there's also like a wealth of things to do,
base building that, you know,
you can choose to engage.
engage with in your own way.
Is there,
is there a boss,
kind of like the,
the nether dragon or the end dragon in Minecraft?
I think they're like are,
I don't know if there's like a specific end game,
although I've never pushed for it.
So I don't actually know.
But certainly as part of like some of the DLCs,
there are like mission arcs that to go on that you can complete,
you know,
those missions.
And some of those missions have,
it has like the equivalent of kind of like raids at the end that,
you know,
you fight stuff and you,
you know,
want to be powerful to take on.
but no, I don't think there's like a true conclusion to the game.
Got it.
Makes sense.
Cool.
All right.
My tool of this show is paperlib.
So if you go and download a research paper today,
you will go to probably the site called Archive.org,
which has a ton of research papers.
And when you download it, you're going to get basically the Archive,
which is some gibberish.pdf.
And then, which is, you know, when you open it,
you'll see the title and all the content.
But if you do this enough times,
you will fill up your downloads folder
with like gibberish.pdf files.
And it becomes very difficult to sort of sort through that
and know what's going on.
So what Paperlib does is it, you know,
adds a bunch of metadata and it renames the file to the title and then gives you this nice
kind of searchable UI. Now here's where it's really cool. Paperlib will let you pick any directory
to run in. So to put your to put your files in. And whatever files you put in that directory,
it'll rename them, et cetera, et cetera. So I pointed Paperlib to a Google Drive folder. And I said,
yeah, use this for your database.
And so now on this Google Drive folder,
I have all these research papers
that have really legible titles.
And so when I use my Nxte paper 14,
I point it to my Google Drive folder.
And so now it's super neat and organized.
That is cool.
And it,
so does it organize them like into subfolders?
like some sort of categorization or mostly just handles like the cleanup of the renaming and stuff.
It mostly handles the renaming. Let me check in preferences here.
Because I can't honestly say I've had the problem you're describing. Like it sounds cool. But like I don't
have so many research papers that like the name is the problem. Yeah, you can make your own. It looks
like you can make your own folders and filters. But it won't do that automatically. But yeah, you can make
folders, put papers, you know, group them up by category.
You know, is this reinforcement learning?
Is this LLMs, et cetera?
And, and all that will be reflected on your phone or your tablet or whatever device you're
using.
That's cool.
I did something a little similar.
I had a bunch of, like, ebooks, like various, you know, PDF, ePubs.
And so I attempted to do something similar.
It wasn't perfect with one of the, you know, AI tools with computer use, like, pointed
at the folder and be like, can you please?
like sort, organized, like,
D-Dube, like, just kind of, you know, do your best.
And I will say it's getting there.
One day it's going to be amazing.
But it took it from basically like a flat folder
with just the names of the books.
Dot random thing and like, you know,
made subfolders fiction and hobbies and cookbooks or, you know,
whatever.
And then, you know, was able to shuffle stuff around.
Yay.
I was low risk.
I would be really careful doing it with important documents
because it definitely like tried to rename some stuff
that was like not.
And if you have like Unicode characters in the title, it'll like, you know, didn't handle it well.
Like there was a bunch of glitches, but you can definitely see where it's going to start to be a lot of the organization.
I have this pet people where like, not pet people.
I have this thing where I don't like deleting pictures.
It just is like, you don't know, like I just keep them all.
And one day my hope is that like the tools get better to where like it'll be able to surf it.
And it is getting there.
Like you can now search for things that are very, what did I search for the other day?
It was something that was like not even obvious.
And it was like, yeah, found me pictures, right?
You used to need to be like really clear,
but now you can say like, I want pictures of dogs, you know,
and it'll like show you all the pictures you've taken of dogs and, you know,
places.
And so I think we're getting closer to where the organization cognitive load can be reduced.
And that feels like low hanging fruit.
Yeah, that makes sense.
But this is a great tip for I'm sure there are a class of people,
just like you Jason, who have many, many, many, many doc.
It says here, like conference papers.
I've never gotten a trunch of conference papers.
So maybe one day that will happen to me and I will remember this.
Yeah, exactly.
Yeah, I mean, it's, yeah, I've just accumulated an insane amount.
I mean, I read probably five to ten research papers a week.
And after decades of that kind of adds up.
If you replace week with year and skim for read, same.
But mostly look at the pretty pictures.
I look at the pictures of a few, you know, research papers every so often.
You're like, the graph is not colored.
I'm not interested.
Black and white.
I'd be like, I need to explain it like I'm five.
Nope.
I need to explain it like I'm two.
Okay, can you draw it with crayons?
I'm still not getting it.
Oh, man.
Okay.
I will hopefully help you get world models.
That's our time.
Yes, that's what I was going to say.
I was going to make the same same way.
Oh, no. It's good.
All right. Go ahead.
All right.
Take it away, Jason.
Yeah, we're all thrown off after the adrenochrome comment.
So this is why if you're using chapters to go straight to the topic,
you need to go back and watch the rest of the list.
People are going to be like, what?
World models.
All right.
So, Patrick, tell me, like, what, this is actually your idea.
What inspired you?
to want to do a show on world models.
Like, how did this come up?
Jan Lacoon ragging on meta for not letting him do world models
and kicking them out and having to get a billion dollar European startup
all in its own.
I mean, just to be honest, that was what?
That was like, I heard world models and I was like, all right,
I did do a little bit of looking after, but that, that's the genesis.
Wait, so Jan kind of is ripping on his former employer a little bit.
Someone's going to come out.
I'm so sorry.
What are you supposed to, like, say allegedly or something?
I don't know.
I'm trying to suggest insinuate that yawn has low impulse control.
Is that what we're doing here, Patrick?
I don't know that much about him.
I just know that I saw it come across tech news like 20 times.
And so I was like, okay, there's something he's doing, I think in France and having a startup.
And after that I did watch, you know, some videos.
So I learned a little bit more.
But that was where like I first started seeing this as like a thing.
Yeah.
So you're basically right on all accounts, especially the low impulse.
Oh, no.
No, I'm just kidding.
So, okay, so, yeah, I think role models, super important.
You know, I think that, okay, well, let me address the meta thing really quick.
So, you know, really quick.
So Zuck needs to catch up on LLMs, right?
So for whatever reason, you know, LLMs were invented after I left meta, so I can't be to blame for this.
But, you know, for whatever reason, people at meta, like, researchers, you know, people in meta-a-I didn't really double triple down on LLMs, even though, like, it was pretty clear that that was huge.
So, so now Zuck has to play catch-up.
And, you know, he can't be, like, distracted by, like, real long-term visionary stuff, you know, while you're, like, three to five years behind other companies, right?
So totally makes sense.
Makes sense for Yon to do his own thing too.
I mean, it makes sense across the board.
Hopefully there isn't any bad blood there from anyone.
But, okay, so that's the deal with that.
So we'll talk about world models.
To talk about world models, we have to talk about making decisions with AI.
And we've talked about this before on the show, but I'll do a quick recap.
So regular AI, you know, hot dog, not hot dog.
Oh, I love it. Let's go.
You draw a bounding box around the hot dog as a human.
You know, you personally draw a bounding box or you hire contractors to do this.
And you call that ground truth.
You're like, okay, if a human did it, it must be right.
Maybe use an ensemble of people, you know, on the same bounding box just to really make sure.
But now you have what's called ground truth.
You know, I drew a picture.
I drew a box around the hot dog or around the car or whatever it is.
and I know it's there.
So then you ask the AI,
hey, is there a car in this picture?
And if it says no,
or if it draws the box in a wrong spot
or something, you correct it, right?
And that's relatively straightforward
from that perspective.
Now, the problem is,
oh, actually one more piece of this,
is then you rely on interpolation, right?
Clearly, you don't have every picture
of every situation you'll ever
see in the universe, but you get enough pictures of enough cars and you draw enough bounding boxes
that then you can interpolate. And when you see a car in a situation you've never seen before,
you should still be able to find it. Okay. So for decision making, right, you could argue,
why don't we do the same thing? So let's take all the best stock picks or all the best
grand master chess moves and train a model to say, hey, when you see this chess board,
do this chess move, when you see this situation in the stock market, buy this stock,
and then hope that we get the same interpolation. In practice, it doesn't work. So when you
interpolate and you try to take those things you memorize and apply them to other chess boards
or other stock situations, they just don't interpolate that well.
So you can still do this.
It's called imitation learning, and it's still a very good approach to get started.
But you, unlike supervised learning where that's all you do, and in the end, you're done,
this will not get you a good solution.
So what you have to do is get better without any human in the loop.
And the question is, how do you do that?
Well, you make a decision and you execute this decision kind of out there in the real world.
And then you measure whether that decision was good or bad, right?
If it was good, then you adjust your model to do that decision more often,
you know, at the expense of alternatives.
And if it was bad, you do the opposite.
it. You make that decision less often, which, you know, kind of implicitly causes the other
decisions to become other options to become more likely, right? So, so now the question is, what is good
and bad? Like, I'll give you an example. Let's say you have two choices. One choice gives you
a dollar and the other choice gives you a thousand dollars. Well, let's say you take the first choice
and you get a dollar. You're like, oh, that's good. I got a dollar. I don't even need to try the second
choice, right? So you miss out on the $1,000, right? So it turns out in that example I gave,
the $1 decision is actually bad even though you got a dollar. And so to solve this, you need
what's called a baseline. So a baseline says, you know, given my policy, given, you know, my,
you know, my strategy, what do I expect to get? And so in the beginning, you're
strategy is just random because you don't know anything, right? And so the baseline, let's say you had a
perfect baseline, it would say, oh, I expect you to get 500 bucks because that's, you know,
the average of one in a thousand or 550 cents or something, right? So you take an action,
you get a dollar, then you know you messed up. It's like, oh, my baseline said on average,
I should be getting 500 and I only got one. So that was a bad move. And if you had a perfect
baseline and you did this enough times, you would eventually pick the $1,000 every time, right?
Similarly, if you had a perfect policy, then you could get a baseline.
So if I had a policy that, you know, always chose the $1,000 every time, then when I get to that
decision, my baseline would say, oh, I expect you to get $1,000. I expect you to not even bother
with the $1 action. So, you know, a perfect policy requires a perfect baseline. A baseline is
dependent on a policy. And so in hints, you see the problem. You have two things that are dependent
on each other. And so you end up doing what we call in the math world a relaxation approach,
which is a fancy way of saying, if you have two things that depend on each other, you. And so you end up doing,
other, you look at one of them and optimize it, assuming the other one is perfect.
And then you switch.
You go to the second one and assuming the first one is perfect, you optimize the second one,
and you keep going back and forth.
Like this is what K-means clustering does, right?
So that at a high level is how you make decisions with AI.
And you could do reinforcement learning, evolutionary strategies,
no matter what you do, if you're making decisions with AI, it's going to be that.
So any questions about that kind of foundational part?
No, so I guess you're kind of putting on this realization step is you're like trying to explore to understand what's possible and help like inform where the line between sort of better and less good is.
Yep, that's right.
You're trying to find, you're trying to find a baseline.
and then you're also trying to update the policy. Whenever you update the policy, that changes the
baseline. So, you know, our baseline started at $500 because we were picking the $1 half the time.
But as our policy gets better, our baseline also goes up. And so, you know, in this fictitious example
where you have two decisions, one gives you a dollar, the other gives you $1,000,
your baseline will climb to a thousand dollars as your policy climbs towards always picking
a thousand dollar action but they're going to how do you how do you oh sorry and so like in that
case it was dollars well you used chess earlier but like chess doesn't have a obvious numerical
expression for a baseline right like you could have how much material you're up or down but
that's like a very crude it doesn't sort of express to you if you're
improving your position or worsening your, you know, opportunities to win. Yep. So, so chess is a
multi-step problem, right? Where the only thing that matters in chess is the last step that wins the game.
So, so the last step that wins the game, you know, gives that, that agent a score of one and the other
agent of score of negative one. And so what you have to do then is propagate through time so that now
your baseline is including the like expected future reward.
Okay.
So that's like sort of what they end of doing,
AlphaGo or whatever, trying to say like a given position,
how likely is it to win?
But,
but it's sort of like trying to roll forward across all the possibilities.
Yeah, exactly.
And it is dependent on the policy.
Like if I happen to play the same move as AlphaGo on my first move,
my baseline is totally different because I suck.
Right.
So like alpha goes baseline
If they're playing a world champion
Might be 0.6.
They expect 60% chance that they're going to win
My baseline against the same world champion
So the same situation is going to be zero
But I guess it makes sense as well
Because the positions from which you could win from
Are probably different.
Like Alva Go could be in a much more complicated situation
Potentially and still like win if it's possible
versus if you get into a complex situation
Your chance for mistake is much higher
So it might be better to me.
move to a theoretically
heart, like theoretically less good
position, but it's simpler, and so your chance
of making the right decisions is better.
Right. Yeah, but even independent
to that, you know, if I'm, let's say,
tracking AlphaGo, my
baseline's going to be low because I'm
expected to be making mistakes
in the future. Okay.
You know, and then eventually at some point,
if I just somehow coincidentally
matched AlphaGo performance, and at the very,
very end, I would get a baseline of one,
right before I win.
But my baseline would generally be low
because they just know that,
you know,
Jason makes tons of mistakes.
And so it's going to happen eventually.
This is why the baseline is downstream from the policy.
Got it.
So,
okay, so what I just talked about,
you know,
you have a policy which says,
hey,
here's the probability of taking these different actions.
and at some point I'll actually take an action,
and then if that action comes back, let's say, you know, better than I expected,
then, you know, my propensity for that action goes up, my baseline for that action goes up,
et cetera, et cetera.
But you have to actually take the action to find that out.
And so let's just take a self-driving car example.
If you don't know that driving off a cliff is bad,
well then you have to try it out.
So the question, how do we fix that?
And you know, you could do this through engineering, right?
You could say something like if we're too close to the curb, you know,
there's some rule that kicks in and we slow the car down.
That's outside of the scope of this episode, right?
I mean, that becomes a different problem, right?
But if you just want to use machine learning, then you have to use a simulator, right?
You're not going to be just throwing cars off every cliff, right?
So in the simulator, you throw tons of cars off cliffs, and it learns that, hey, you know, the action where I throw the car off the cliff is really, really bad compared to the baseline of staying alive.
And so, you know, after throwing so many cars off cliffs in Sim, it'll learn not to do that, right?
So, right.
So to do that, you have to build this really complicated simulator, right?
And you have what's called like a sim to real problem where the simulator never fully matches the real world.
You know, the sky looks different.
The ground looks slightly different.
There aren't drunk drivers in the simulator.
You know, maybe the road isn't bent, like, you know, isn't rolled at all in the simulator, right?
there's all these things that are not in the simulator but are in the real world,
and you kind of are constantly fighting that, right?
So what if instead of trying to build a simulator, you know, using some sim engine,
and then sort of shoehorning that into your problem,
what if you created a simulator as part of solving the problem?
And so this is what model-based decision-making
or model-based reinforcement learning is doing.
It's saying, hey, I'm going to build my own simulator
while I'm trying to drive the car.
And so there's a bunch of criteria of what makes a good simulator,
but it's not a simulator that you or I can see
in the same way as we can't see what an LLM is doing.
We can only see the algorithm.
Got it.
Right.
So it's some big neural net mess, right?
But it's inside of that mess is a simulator that can predict the future given, you know, you turn the steering wheel this way.
This is what the future looks like in this simulated space.
So it takes like a state and the decision and predicts the new state.
Exactly.
Yeah, exactly.
Now, the original state might be like a whole bunch of cameras, right?
And so predicting cameras means you have to draw a bunch of images of the future, right?
Which can be really difficult.
Like you have to handle the movement of the clouds, all this stuff, right?
So what people do is they'll create what's called a latent state, which is where they've collapsed all these camera images into some like blob,
where the blob hopefully doesn't have clouds in it because they're not necessary to drive a car.
Like the blob hopefully just has important stuff in it.
Then you say, okay, given this blob and I press the accelerator button, what's the next blob?
This is why Jan Lecun was saying in that thing I was referencing that if you were trying to do self-driving and you're like predicting the next thing, you're just going to end up spending all your tokens trying to predict tree leaves.
Yeah.
This is the same point.
It's like you're saying leaves look like this in a video frame and in the next video frame they look not.
not in the same position because of wind,
but really that's just noise.
It's irrelevant.
Right, right.
So the question is,
how do I go from the state to that blob?
And the answer is,
a good blob is one that
when I apply an action,
I get another blob.
So think of it this way.
Like, you can,
and this is called encoding.
You can encode the state into the blob,
take an action,
and now you get the future blob, right?
But you can also take the same action in the real world and then encode the feature and you should get the same blob.
See what I'm saying?
Mm-hmm.
Now, there is a catch, which is what if my encoder, it just makes a blob of all zeros.
So I have a blob that's all zeros.
And then every action just takes the zero blob and makes another zero blob.
Winning.
Yeah.
Well, now my future, when I encode the future, I get a zero blob, and it's perfect, right?
So you have to prevent that.
And that's where there's a bunch of techniques.
But basically, this is called prior posterior regularization.
And there's a bunch of techniques for this, but the gist of it is you have two blobs.
You have the blob that you got from encoding and then taking the action.
and you have the blob that you got from taking the action and then encoding.
And those two blobs, you can get the covariance matrix of those two blobs.
And the diagonal of the covariance matrix, you want that to be one,
and you want the off diagonals to be zero.
So in other words, if you do the zero blob, well, then the covariance matrix will
collapse because it will be zero to zero every single time. And so you'll have a,
the off diagonal elements will be zero, which is good, but the diagonal elements will also
collapse to zero, which is bad. And so this regularization punishes that hack from from succeeding.
Okay, so that's the way to sort of like keep it from saying, I don't know what to do.
So I'll just, you know, basically the equivalent of give-up, I'll just like blur it all out until it turns into just, you know, some base value zero or whatever.
Yeah, exactly.
Okay, so now, let's say I'm in the car.
So I'm not training the model anymore.
I'm actually in the car.
I have this model-based reinforcement or anything already trained, right?
Well, now what I can do is I can say, if I was to press the accelerator, what would that next blob look like?
like. And I could have some other model that says, given a blob, am I driving off a cliff or not?
So if I put those two together, and I can say, oh, if I press the accelerator really hard,
I'm going to create a blob that's a, you know, going to die blob. And so I don't want that.
So even though my policy says to press the accelerator, when I roll it out, when I actually
look at the future, those blobs look pretty bad.
So I'm going to go against what the policy says, and I'm actually going to hit the brakes, because I used my model and it said that hitting the gas is bad.
But doing that like recurrent, recursive, however you want to say it, like keep taking the state, do the action and keep looping the state over.
That's where you're going to get the drift, though, right?
Like, that's where it's going to be harder and harder at like longer time horizons to say that the blob is what it is representative.
of
Yeah, so you can start hallucinating.
You know, you can end up with, as you said, with drift,
where you go, what's called going off manifold,
but you end up with a blob that isn't real.
Like, you'd never see in the real world,
and then you're kind of in trouble.
That can happen.
There are ways to address that.
But, oh, the other part of it is you only roll out maybe 16, 32, 64 steps.
So you don't roll out, you know, minutes into the future.
You basically say, if I take this action over the next five seconds,
am I, do I see a blob where it's a you're going to die blob?
And if I don't, then that's good.
And you can even do what's called model predictive control where you take,
you take an action, it says you're going to die,
or maybe you take like 30 different types of actions,
and then 15 of them say you're going to die.
you pick the other 15 and you use them as a seed to generate some more actions and you kind of like on the fly kind of hill climb towards the best action.
And so this smoothens out.
You know, with neural nets, there's just so much variance, right?
So this smoothens out a lot of that.
You know, if your neural net freaks out one out of a hundred times, this is the thing that keeps you from dying in your Tesla that 1% of the time.
Any questions about MBRL?
No, I mean, I think I, yeah, I think I understand it at least at the five-year-old level you're going.
So thank you.
I appreciate it.
All right.
Now, here's where it becomes a world model.
Okay.
So that wasn't a world model yet.
Oh, not a world model yet.
Okay.
So we talked about, you know, when you're actually physically in the Tesla, it does these rollouts, right?
and it might not do the action that the policy wants because it rolled it out and it was not good.
Right.
So the question is, like, can't you at that point use your model to train itself, right?
So, like, if my policy says to press the accelerator and it says I'm going to die, you know,
Why don't I just fix the policy right then and there?
Like, why do I have to be in a real car to do that?
Even just during training, I could start with something real,
like start with a scene where you're driving on a cliff side,
but then inside the model, do that drive, find things where you fell off the cliff,
correct them, and then do all of that inside of the model
without having to need an external simulator or driving in the real world or any of that.
You're just like hallucinating problems and solutions.
And so that's where it becomes a world model.
Okay, I see.
Yeah, so think of it as kind of like baking that thing that you do in the real car,
like baking that, you know, into the model.
Okay, all right.
And so if you were to take, for example, Dreamer v4, they actually do this thing where they trained a model to play Minecraft.
Or not to play Minecraft, maybe specific.
They trained a model to understand Minecraft.
So they watched like a zillion videos of people playing Minecraft and they trained a model where, again, this is all in blobs.
You can't see the game or anything.
But just in blobs, they swing a pickax.
and then the next blob has presumably like some logs in your inventory or something, right?
And so just based on on looking at people playing Minecraft,
they were able to build this model,
and then they were able to train just in the model with the goal of get diamonds.
And literally without touching a keyboard or being able to play even one frame of Minecraft,
they were able to learn an AI that train an AI that could get diamonds 0.6% of the time,
which isn't a lot, but it's still pretty amazing.
They basically took this thing that had never been able to touch a keyboard,
put it in front of Minecraft,
and one out of 200 times it gets diamonds.
So all that learning was all done on Blobs.
Okay.
I think I got it.
And Blobs is the what they call like the latent space,
whatever.
That's right.
Okay.
Okay.
Yeah.
And then how did they know it got diamonds?
Like, presumably they have some decoder for the blob for at least some things.
Yeah.
So they did all this training on the latent space, right, on the blobs, right?
Then they froze the training and then they put the trained model in front of a real Minecraft game.
Oh, okay.
Okay.
Got it.
All right.
That makes sense.
Yeah.
And that's where they got the one out of 200, which is.
amazing result.
I mean,
now,
if you continue training
based on,
you know,
those videos that it
created of itself,
then it got diamonds
almost every time,
which we already knew.
But the fact that it could
just watch other people play,
build its own model
of how that game works,
and then get diamonds
one out 200 times.
It's pretty amazing.
And that's the hope that is like,
as humans,
you're a baby and you watch stuff around you,
and even without trying it yourself,
you're then able to basically do
or perform it very high accuracy on the first time often.
Right.
But that the models generally can't.
Like today, machine learning isn't really capable of replicating that.
Right.
Yeah, exactly.
Okay.
So now there's a big debate around whether you should reconstruct the raw state or not.
So the dreamer folks, which is a team out of deep mind, all of their models, even the
latest one that came out like six months ago, reconstructs the original.
state. So in the case of Minecraft, you know, you have your screenshot of the game and it does all
the things we talked about, but it also outputs the screenshots of the future or the screenshots
as it's going. And you can actually see it like go and mine diamonds and it's all kind of fuzzy,
right? Because it's, it's all going through this blob space. So it's, you know, a lot of the clouds have
been destroyed.
I don't say, but like anything that's irrelevant is basically not present.
So that's where the blurriness comes from.
Exactly, exactly.
Clouds are totally destroyed because they're useless, right, to actually playing
Minecraft.
But, but, but yeah, you know, it recreates the state as it goes.
You can actually watch it play Minecraft while it's training, you know, in some really
weird way.
And so, so Dreamers definitely in the camp of,
And they have a bunch of arguments for, you know, if you don't, to Yon's, to use Yon's metaphor,
you know, if you don't reconstruct the leaves or at least try to, then you're not robust.
So in other words, the problem with Yon's argument, and I'm just, I'm not saying he's wrong.
I'm just playing devil's advocate. Yeah. Yeah. The problem with Yon's argument is we already know
that the leaves are useless for self-driving because we have common sense.
But, you know, that was an assumption that we made, right?
Like, it's not obvious.
It's not like just implicit in the form of a tree and the leaves that those leaves are not important for self-driving.
Right.
It's only because of other stuff that we know.
And so if you're reconstructing the original state, then, you know, you have the potential to use anything.
So put another way, maybe in the beginning of self-driving, you're just trying to do the highway, right?
And so if you use Yon-Lacoon's approach, well, then stop signs that you might see when you're looking down from the highway will also get destroyed because you don't need to worry about stop signs when you're on the highway.
But as soon as you say, okay, I want this car to drive on city streets, well, now it doesn't have stop signs.
And so you have to start all over again.
With the dreamer approach, because you're reconstructing the original state, it actually will need to understand stop signs, even though they're useless.
And so then when you change your driving domain, you'll be prepared to that.
So in the deep mind approach, not only do you regenerate like the Minecraft screen,
does the Minecraft screen then be like, is that what goes back into the blob?
Or is that just like a loss for the model to look at and say,
hey, I also need to make sure this is somewhat stable?
The Minecraft screen does not go back into the blob.
Yeah.
So you encode the original screen from like the very first screen.
shot of the game and then you never encode anything else. And so the screen output is something
that's just for humans or it's still part of the loss? Like you're still rewarding better preservation
of future screens. Yeah, the screen output is part of the loss. So the way you train the model is
you have a video where you have the current screen and you have the next screen and you know,
the action that was taken. So you generate the current screen, make sure it matches what came in.
You take the action, generate the next screen, and make sure that matches the next screen that
you have kind of in your... So it's still free to take its actions and it's just living in the latent
space, but the latent space needs to continue to be grounded to what the real screen would have been.
Yeah, actually, I kind of misspoke earlier. In the Dreamer case, you actually do need to have the
clouds in your latent state because you need to regenerate them.
Okay.
In the JEPA case, you don't because you're not regenerating them.
But as you said, there's pros and cons to both.
So in the, yeah, so in the dreamer case, like the leaves on the trees need to be present
and sort of in the right place.
But in JEPA, they would be just stripped down to something that would be the tree is a problem
if you hit it, but as a concept, but you don't need to know about the little dangly bits on the end.
Exactly. Yeah, the leaves would be just completely gone from the latent space in JEPA.
Okay.
This is like, it feels like one of those things that, to say world model or non-world model, like,
it's like a very simple thing. Like, it's very straightforward. We can say these words,
but the nuance is actually, like, pretty involved. Like, these people arguing about it
feels like one of those things where, like, the average person, it's so far removed to understand
the nuance of the argument or pick aside. Like, it's like super in the world.
weeds. Yeah, I mean, you know, training on, you know, your own sort of like representation of the
problem just carries with it all sorts of challenges, right? And that's unique to world models.
So if you do like a model-based reinforcement learning, you don't have to worry about,
oh, you know, my latent state thinks that I got a thousand points in chess, but I can only get one
point. Like that's not something you have to worry about because all your rewards are
explicit. And it sounds like from having looked a little bit at it that the argument is from
the world model side is to use the other approaches is going to hit like a cat, like it'll run out.
It can't do the sort of plan, long horizon planning that you might need to be like there's just
lots of hangups. And then the reverse would be like you said, it's like, it's not as sophisticated
today. Like the world models are further behind. But the argument is, yeah, they're further behind.
They're harder to kind of like get going today. But eventually they should be more capable.
Yeah. I mean, Rich Sutton actually, you know, he's famous for, well, a lot of things, but one of them
is the bitter lesson. And he recently made a post where he tried to summarize the bitter lesson in 30 words
or less. And I don't remember exactly a summary. I won't quote it, but the gist of it is, you know,
if you, the more you remove the human out of the loop, the better it is. And that tends to just
dominate everything else. So, and so in this case, you know, having a human craft a simulator,
or having a human create a bunch of rules so that you don't need a simulator, like, oh, you're
driving too far to the edge or something like that. You know, get,
Getting those humans out of the loop is ultimately so much better in the long run than anything else.
And so, yeah, model-based RL has so many challenges.
World models have even more challenges, but they will eventually win just because compute is so cheap
and because, you know, humans have a hard time interacting with models at such a low level.
Yeah, I mean, it certainly seems, I'm curious how, like, investor,
I just off topic, sorry, or I'll off.
I was like, people investing in this space that, I guess, people are just making bets or diversifying
because it doesn't feel like if you're going to put financial backing, like example,
to the world model from Yon Lekun getting, I think it was like a billion dollar investment or whatever.
You don't, you don't actually know.
But the same was true, I guess, like Open AI and its original founding.
Like, there was no evidence that where we are today with LLMs was attainable.
So it's just very interesting to see.
the investment sort of, I don't know to say, like, so far past the horizon. Like, it's not clear
how to get from here to there. There's also just enormous survivor bias. Like, there's tons of
open AIs that fell over, right? True. I think that, uh, um, I think that, you know, in the
Al-Lacun's case, you know, he probably knows enough billionaires, right, that he only needs to get
20 of them to each give five million or whatever it is. Yeah. Yeah. Uh, 50 million. Yeah. Um,
So that's kind of what's going on there.
In general, you know, starting a world model companies,
probably not a good financial investment for anybody.
But, you know, at some point it's really not about,
for Jan, probably not about the money.
It's about wanting to do this thing.
And he's got the determination.
Yeah, I think it's also interesting for these companies.
Do they tackle the, I know there's been a couple startups to say,
we're not going to do any consumer products.
We're just going straight to AGI.
Like we, it's like a distraction.
So I guess that's always a debate as well for someone like Yon Nakuna or whatever.
Like, do you go for the thing that's far out that is like the ultimate prize?
Or do you try to take incremental steps to prove the ideas or working and, you know, gain income?
Yeah, I mean, there's a ton of these companies that go nowhere, like safe superintelligence from Ilya, the guy who started Open AI.
there's just so many of these companies where they just don't really go anywhere.
So, but you know, everyone talks about the few that really succeed because that's just human nature.
But I do think world models are a huge deal.
I think that for two reasons.
One, you can't have counterfactuals.
You know, you can't just figure out, oh, I'm going to not drive over the bridge, drive off the cliff,
through some trickery and engineering.
No, you actually have to drive off the cliff,
which means you have to do it in sim.
And then two is the sim to real problem is a non-starter.
You know, it's just a mess,
and trying to deal with that is just a nightmare.
And you don't have to.
Like, people, we don't need to use, like,
some expensive simulator when Dreamer can just dream all of Minecraft.
Like, what's the point, right?
So yeah, world models, definitely the future.
But as far as a financial instrument, maybe not the best one.
Is this where you give your disclosure about an investment, Jason?
Yeah, full disclosure, I would never invest in.
Oh, his company has less than 100 people.
I think it's a terrible idea.
Yeah, I think the survivorship bias is real because it's balanced, I guess, with FOMO, right?
Like people have this like, oh, I can get in now.
Now is the only time to get in and make a million X.
you know and yeah it's but you know like yeah you're so much better off um betting on things
that have already won and getting like a four X versus like wasting your money on a you know
epsilon percent chance so this has been my strategy as well but i i i can look at other people
who didn't take this strategy and are better off so yeah yeah yeah but so many other people have
lost their hat, you know.
That's the thing, right?
You only think about the people ahead of you, but yeah, yeah, it's a human rights.
I was actually thinking about this the other day.
I know we're short on time, but I was thinking about all the startups that have reached
out to me, probably to you too, over the past, you know, 15 years and how almost all of them
have failed catastrophically.
Or, you know, if not, okay, that's a bad terminology, not failed catastrophic, but almost
all of them have underperformed just going to a big company. Like literally, I can't even think of one
that outperformed just staying at, you know, a big company that's successful. So, um, so, uh, yeah,
I mean, you know, maybe we could have joined one of these, uh, things, companies that are
IPOing this year. But, but, but again, like, that's, that's hindsight. There's also like hundreds
of other companies that didn't go anywhere. So,
Words of wisdom from Jason.
Yeah, totally.
Run so you don't die and invest in already sure things.
All right, got it.
Run so you don't die and winners keep winning.
You know, in the movies, it's like the underdog, you know, Rudy.
Remember Rudy, the football movie?
Yeah, yeah, yeah, yeah.
Yeah, the 4'10, you know, linebacker, like, just through his heart and determination, like,
wins Notre Dame.
I don't know.
I've never seen the movie, but, but like, in reality,
the Brock Lesnar, who's like 6'6 and all muscle,
he just wins and just crushes everybody.
Like that's how the real world works.
I hate to break it to you out there.
But like winners generally just keep winning.
Wow, dude.
We get spun up.
It's late in the episode.
I feel like I'm ever going to get a rant here.
Do you have an X account for where we can hear more?
Oh, man.
All right.
Well, thank you for, yeah, I feel like that was, that was great.
I enjoyed hearing the explanation and super pertinent to current events.
Totally, yeah.
If you're out there, you're working on world models or have any questions.
Hit us up on Discord or shoot us an email and you can get links to all of that from our website,
programming threadon.com.
Or one last shout out.
I guess if you're working at one of them, maybe you have a private jet and can fly us out and we'll interview you.
That's true. Yeah, we'll do interviews in person. We just need a private jet.
I've never been on one. It would be cool.
It's not cheap. No. No.
All right. Thanks, everybody.
All right. Catch you all later.
Music by Eric Barn Dollar.
Programming Throwdown is distributed under a Creative Commons attribution, share a like,
2.0 license. You're free to share, copy, distribute, transmit the work to remix
adapt the work but you must provide attribution to Patrick and I and share a like
and kind.
