The Standup with ThePrimeagen - Casey Destroys Optimization Myths
Episode Date: June 5, 2026Don’t let bad code get merged without reviewing (hopefully not by merge cop!). Checkout out Code Rabbit at https://trm.sh/coderabbit We dive deep into technical performance debates, specifically th...e nuances of floating-point math—is multiplying by a reciprocal actually faster than division on modern CPUs? We also break down the latest from Microsoft and NVIDIA, including the "RTX Spark" and the vision of "unmetered intelligence." Plus, don't miss "Trash Facts" about professional bowling and the latest "Snack Picks." [Sponsors] Linear: https://trm.sh/linear Sentry: https://trm.sh/sentry
Transcript
Discussion (0)
If you guys do not know, chat, chat, everybody in here, Casey, trash.
The startup is going through a bit of a change.
All right, there's been some investment interest.
And we are now raising a series A like any proper multi-billion dollar unicorn startup.
Cool.
And so deck of corn.
Deca-a-corn, really.
We're a dec-a-corn.
And so I don't know what that means yet, but someone told me it's like, that's the best kind of unicorn.
I think it means it has multiple horns.
I don't know what it means.
10 horns.
10 more. Okay. Okay, nice. So it's kind of like the beast in the Bible, the devil, right? Very cool.
He doesn't have 10, does he? Yeah, yeah. I'm sorry. Fact checkers? Fact checkers. And so with that, T.I mean, I thought T.J. would fact check that one.
I actually thought T.J. was going to fact check me, but he fact checked me incorrect. It didn't seem 10, but it could be. Could be.
Bunch of Christians and no one's read revelation. What the hell's going on here, guys?
Well, that is kind of what it answers, you know. Anyway.
All right, so he has 10 horns.
Called it.
Okay, primary reference I got it.
Okay, I was correct on that one.
Boom.
Referenced.
All right.
So with that in mind, we...
We're just the AI for our theology.
Just let it happen, everybody, okay?
My gosh.
All right.
This is not starting off the way I wanted it to.
So anyways, nor does any stand up.
All right.
So with all of that said, that means as part of our amazing content,
I have now organized us an amazing linear board,
which will help us walk through a more organized stand-up
because if there's one thing I know about in raising a series A
is they want to see agile.
Okay, I'm pretty sure that's correct still in today's day and age,
so we're getting really agileized.
No, it's actually going to be,
it's going to make us hopefully have a more organized
and amazing stand-up
so that the world's most attended stand-up,
which is our current stand-up right now.
Right now, just in case you're wondering,
you know, you may not know this, Casey.
there's there is uh there's 2,500 people on the standup right now.
This is the world's most attended standup.
It's like you couldn't even like it we previously no one knew if that many people could
even stand at the same time.
We've just shattered that.
We've shattered that.
Oh yeah, definitely.
What was the record before?
Like 100 people maybe?
Yeah.
Yeah.
Probably.
Now we're at 2,500.
It's world.
It's world shattering.
Yes.
It's a game changer.
Yeah.
I want to add though, we didn't get an invite or the topics until this morning like two
No, yeah.
So this stand that's going to be amazing because Prime's like, we're going to talk about this new, well, you know what, I won't spoiler it.
But let's put it this way.
It probably was a topic that would have been nice to research beforehand if we're going to talk about it.
That's all I'm going to say.
I don't want to spoil anything.
I'll get the gym.
Let's say it would have been nice.
Well, you know, that makes this authentic case.
You could have react to Trash's workout gym knowledge.
Okay.
Rather than facts, you get trash.
I am not the expert.
Trash.
Trash.
He's the expert.
I was like,
dude, if there was a trash fact every episode, that would be, that would be fire.
It's literally going on the linear board.
I literally have endless lore.
Another thing that I would like to ask at this point is if we are serious about this series A, which I think we should be, is it time for us to start putting together a pitch deck?
Oh, wow.
I still don't know what one of those are, but.
I think we need to do that.
I think you're...
We need some slides.
My startup, daily BJJ.AI, is going nuclear.
Congratulations on product market fit.
Sorry, I called you the Diffler.
I knew you weren't the Diffler.
I always thought it was the intern with the weird glasses.
You should have seen his code.
But anyways, if you want to get in the ground floor of this thing,
you take an angel investments.
So, you finally cracked the case merge, cop.
put the white space in the diff. You made the synthetic traffic. You made me approve that PR. You're
the diffler and I always knew it. You've merged your last PR Diffler. I already clicked
to merge on my PR 35 seconds ago. Huh? My merge is blocked. Curse you, Code Robbits! You saved the city
from a code review related crime. Thanks, Commissioner. Well,
Technically, it was Code Rabbit that saved the day.
Not only do they have advanced AI features that can detect security vulnerabilities,
like what the Diffler was trying to merge.
They also have ways to enforce styling, linting, and a variety of other tools,
and so you can stop wasting time reviewing code that humans don't need to review.
You can try it yourself at codrabbit.a.i.
So, MergeCop, I would say it's a little misleading to say you did it all by yourself.
Huh. Would you like some cake?
Oh, sure, thanks.
We need slides that say like, here's our total addressable market, which is every human on earth.
The market is like everybody.
People in like uncontacted tribes in like, you know, indigenous places will be using the stand-up somehow.
They'll be standing up.
Because Casey, you know who our market is, right?
People who are affected by AI.
That's everybody.
That's tan.
Which is everybody.
Earth is done.
Is the Tam.
There you go.
Tam Infinite.
For next week.
Aliens on another planet, like, they're going to be getting the signals of this stand-up
beamed into space.
They're going to, they're part of the market.
They're being advertised.
When they show up to contact Earth, they will buy your product if it is advertised on the stand-up,
ladies and gentlemen.
True.
What else do you need to know?
They'll be chosen for whatever's on the stand-up.
That's right.
Trash a snack sponsor, the aliens will be like, ooh, I wonder what that tastes like.
Great point.
Great point.
I'd love to see next week maybe we each bring a pitch deck for what we believe the stand-up pitch deck would.
Oh, let's do that.
I don't even know.
Dude, it's going to be so good.
Let's each have a pitch deck.
We'll each present ours next week then.
What would your pitch for the stand-up getting a sponsor or like for your own personal sponsor?
No.
I think it's a pitch for the show.
The stand-up is worth probably a trillion dollars long-term.
A trillion dollars long-term, I'd say.
to be conservative.
I don't want to sound too aggressive,
but a trillion...
The pitch deck for why it's worth
a trillion dollars long term.
Yes, I like it.
We are a deck of corn.
All right.
All right.
So with that mind,
I figured we'd do our first task here,
which is trash fact.
Trash, can you hit us
with the trash fact really quickly?
I gotta think about this now.
Well, you said you have an endless supply of them.
It's true.
It's true.
I tried to become a pro bowler,
and my average was like 203.
What?
What?
How is it every time something I never could have imagined?
Trash.
That is a good bag.
I bowled for like 10 years and I went super hard.
And I had to stop because I rock climbed a lot.
So I was like, you know, rock climbing and bowling.
To competing.
Yeah.
But I used to do like cash leagues after work and would go play for money.
It was fun.
You were a bowling hustler.
I wasn't a hustler.
Like I competed like in leagues or whatever.
Okay.
But I had a really.
good average so I thought I could really like make it make it pretty far.
So what would be considered like a professional average?
That's good.
What is that like a professional average?
I think like a professional average. I think like 190 plus maybe is pretty good.
Really?
Okay.
I think my highest score is like 279.
So I was like one strike away from a 300.
Oh.
You know?
True story.
I grew up in Massachusetts.
And so I like the kind of bowling that everyone talks about when they say bowling is like
giant.
giant ass ball with like the
three holes in it, right?
Yes. You stick your fingers in. Okay?
Yeah.
Never seen that.
We didn't have that. We
had candle pin bowling because we're near
Rhode Island. And so it's this little
tiny ass ball. It's like about that big.
Duck bowling. We got like duck pin.
It's not duck pin. Duck pin's a different kind of bowling.
It's basically 10 pin
well it's not even 10 pin
it's long, thin, fairly heavy
pins and a tiny ass ball.
You're at, like,
There's no way you have a 200 average unless like the best bowler in the world.
It's like really hard because the ball is super tiny.
And it's pretty hard to get a strike.
Like a good game is, you know, I don't know, three strikes or something like that.
Whereas if in real bowling, you know, or 10 pin bowling, you'd be like, oh, three strikes, okay, like you're probably not.
That was, you're an amateur or something, right?
So I never heard of this game.
I've literally never heard of some alternative bowling game.
Then dunk pin bowling is even more narrow.
I think that's like Vancouver and Rhode Island are like and like that's it.
I don't even know who does duck pin.
Not many people.
Where did bowling originate then?
Did it not make its way up there?
I don't know.
I don't know.
But like in my town, there was only one bowling alley there and there was, and it was
Candlepin and there weren't, I don't even know where the next bowling alley would have been.
Are you like a small town?
Very small.
Yeah.
I think at the time there were 2400 houses in this town.
That's why the pins are smaller.
Oh, that's small.
Okay, that's actually small.
You're from a real small place.
Okay, I didn't realize that.
Yeah, yeah, very, I mean, it's not, there are tinier towns, but yeah, 2400 houses, I think, in it, when I was growing up.
What was the population?
Well, when I was, oh, and when I lived there, we, I lived to see the first actual stoplight put in.
Like a thing that could control.
Before that, we had only, we just had this hanging light that would just blink at the, at the one main intersection.
Are you from radiator spray?
rings because they have
TJ,
radiator springs,
they just have the
bridge in the yellow light.
That's for you,
buddy.
I think it's longer.
Wow.
But anyway,
moving on.
So that was an amazing
trash fact,
and now I'm just like,
what don't you do trash?
Like, is there a thing
you don't do?
You do,
Jitsu, bowling, climbing.
I will,
I will share more
as the episodes go on.
But I actually,
like,
I have like a weird person
out,
as soon as I get interested in something, I take it as far as I can go, get bored, and I find the next thing.
Got it.
And I take as far as.
So I just have, like, accumulated many random skills.
This is great.
You know, the thing is trash is I've also done that, except for I've done it on much more useless talents such as yo-yowing.
I used to be a professional grade yo-yroar at one point.
Yeah.
And so one time I went to a conference where it was called the server challenge, and I forget IBM, I believe, where you had the plug-in 18 internet cables and plug-in 12 hard drives.
and I was the fastest person at a 3,000 person conference and got my wife a MacBook Air for winning it.
It turns out like, they're like, hey, there's a useless talent award.
And I'm like, oh, that's me.
I'm that guy.
You get, oh, I was plugging in six Ethernet cables at a time.
And they're just like, this is a revolutionary status going on.
So it was pretty great.
All right, but hey.
Oh, yeah, we'll continue.
We can just continue on piece trash.
I don't know if you know this, but it's also your snack picks.
Okay, you dominate the intros around here.
Oh.
what I'm eating right now is the pipcorn spicy nachos and I got my favorite red yum earth
bag with me this is my current duo I already went through four four of these bags this week
which is terrible I have no excuse so I've been working out so I work I work I work out in the morning
and the afternoon so to do that to try and make up for the yum earth should like to make up for like
my sins that I'm doing I feel like so bad and I make sure like I wrist my mouth out with water
and all that stuff I show up with water
I told a trick to
in there to just clean the breath.
Well, that's what my dentist said.
He said,
don't let it sit in your teeth.
Just rinse through the water,
and that's a good thing.
You talk to your dentist about it,
you're saying?
Well,
fun fact,
almost all the teeth in my mouth
are like basically fake.
I could never have guessed.
We talked about that already.
That was a different fun fact.
I'm all messed up, dude.
All right.
Yeah.
We should honestly,
that's fantastic.
TJ,
you usually jump in
with a little bit of
Mack. I don't really have anything to that. I'm drinking a little bit of cinnamon
cinnamon spice tea because my nose hurts really bad because yesterday I was playing basketball
at our friend's kids' birthday party. And then one of the kids just ran directly into my nose.
And I had a ginormous nose bleed for like the rest of the evening. The kid did not attempt
to steal the ball. They just like lowered their head and ran straight in.
into my nostrils.
Like a, like a sheep or whatever?
Yeah, it was.
What was the plan?
What was like, were they trying to do some weird form of a box out that we haven't seen?
Some new, some new revolutionary.
Well, it was me versus like probably seven or eight children.
And then I was like, I would, uh, pretty pretty little like between five and eight.
Most of them.
Yeah, really short, I guess.
Yeah, yeah.
Well, you, you get low to do good basketball moves.
Okay.
Okay.
Yeah.
That's a good point.
So I'm spinning around and I'm dribbling like a maniac and then one of the kids just like wasn't, it just was excited, I don't know, and then ran right into my face.
And I don't know.
I mean, this could be a valid tactic in basketball, right?
Like you just sacrifice.
You just, you know, you take the technical.
One for one.
And you just, yeah.
I'd get so bad.
They tried it in karate kid.
Yeah.
Yeah.
Yeah.
Yeah.
There you go.
And then you just, you know, it says you're not only the star player of the other team, you're the only
player of the other team.
Taking you out would be very valuable for their
win chances.
Win by default. That's a good move, honestly.
Yeah. So I didn't
have any snacks right now.
Sorry. Sorry.
Sorry to this point.
That's okay. All right, well, let's let's
move on. By the way, I've been using
this also for my stream. It's shocking
how much more organized you are if you list
everything out. It's
it's crazy. It's crazy.
So Casey. What'd you do as a software
engineer? You're talking about Casey?
I kept it all in my head because Jiro was the worst experience of my universe.
And so I was just like, I would rather forget what I am doing than ever use that.
Like, I'd rather have to come to a meeting and explain to my boss why I didn't do something than use that.
Okay.
That makes sense.
That's a valid.
That's a valid excuse.
It's a valid excuse.
And you know you're at Netflix.
You know exactly the experience.
All right.
So, Casey, I saw this tweet and I figured you could maybe give me like a little, I think I understand the purpose here.
and what the problem is of it.
But I thought maybe you could help me a little bit on this one.
This is my blocker, by the way, for the day.
Which is, I saw this where someone said, hey, don't use division.
Use reciprocal math instead.
Can you, can you maybe give us, like, us dumb people,
why this is either good or bad idea?
Oh, boy.
I have no idea what that means.
So right here, he's doing array index I divided by pi.
versus array index I multiplied by the reciprocal.
So one divided by pie.
So he does the division thing once and then hits it with a bunch of multiplications.
Yeah.
So there's a lot of things that you would want to talk about here.
And that's, I actually saw that tweet that you're talking about.
Yes.
Did you respond to delete your account?
No.
I mean, see, here's the problem, right?
So delete your account.
I love delete your account.
And I love what Ryan's doing there.
But what mostly what he's doing there is he's just like, and he even described this on one of his own streams, he's just saying like, look, basically like you guys are all the people who like come up and and beg for change on the street or something.
Like you're just totally making the world difficult to live in by like trying to get all.
Like just just hounding everyone who goes into a particular area to try and get them to give you something, right?
Which is your attention or whatever.
They're like these AI shell accounts or whatever.
People trying to get engagement.
They're trying to get follows, whatever it is, right?
And it just really degrades the experience of being in that place.
That's basically what he said on like one of his streams.
I'm probably phrasing it poorly.
I think he said the phrase,
if you go to a third world country like France,
was literally I think the phrase he used.
I love Ryan.
In that stream.
So I don't know.
Don't play me.
Don't come after me for this.
I'm just repeating what I believe it was.
But anyway, the point is just that that was sort of the point of delete your account.
And he's right.
Like there's just all of this kind of, you know, really bad behavior that's going on there.
And this tweet, I don't really like being harsh on people who are trying to talk about performance
because I'm glad that we somehow went, you know, over the past, you know, 10 years or so to people.
to people actually starting to think that it matters that you're thinking about how the CPU is going to do things.
So I don't want to like be mean to anyone who's at least trying to say like, hey, let's think about what this is doing.
The problem is that it requires some actual knowledge to like actually get these things right.
And a lot of times if you just ask the AI or if you just read something on a blog and you then repeat that, it's not accurate.
Right.
Because performance is a very specific thing.
It's kind of like repairing a car or something.
You can't just know like one thing about like, oh, yeah, you know, the carburetor or something, right?
It's like you have to actually know how the whole thing works.
Right? Carbureator.
That's what car is short for.
Exactly.
I'm using that as an example.
Like I'd just be like, oh, the spark plug or something.
Right.
You know, it's like I don't even really know much beyond like there's a thing that ignites the gas.
Right.
Okay, that's it.
I couldn't tell you like, is it broken?
Is it working properly?
Does it need to be tuned?
What's going on, right?
Park plugs are for electric cars, right?
Because you get it from the battery.
It's sparks.
I mean, that's, there you go.
That makes sense to me.
Must be.
I mean, clearly.
So, yeah, this particular tweet is not good, though, because it misleads people, I think,
on a lot of points.
I'm not sure if the person actually tested this.
They've got, like, seconds listed there, which is a little weird because I don't,
those numbers, I mean, there's a lot of things that could have happened there.
I don't know.
So I don't know the circumstances of it, but I just don't want people to be misled.
When you're looking at a loop like this, both of those loops that are in that tweet, because I remember this, of course, I can't read what's on prime screen because Riverside won't let me, but I remember it fairly well.
You're basically going through arrays, and it's basically a read-modify right, right?
It's loading a value out of an array. The value is a floating point number.
loading a floating point value out of an array,
it's modifying that floating point number
and writing it back to the array, right?
Correct.
And the operation that's being done on the array,
the only difference between the loops is
one of them is doing a divide,
so it's loading, dividing, and rewriting,
and the other one is loading, multiplying, and rewriting.
Correct.
Now, obviously, in math,
if you do a
divide of a number,
if you say I'm going to divide by X,
there's no difference between that
and saying I'm going to multiply
by one over X, right?
True.
That's true in math,
meaning in like in grade school,
you write it down.
If someone said, oh, you know,
a divided by,
what's sorry?
Okay, all right, all right.
T.J. was a math major,
so, okay, he's flexing that he knows.
things.
Yeah, math major, get out of here.
Yeah, Tej.
This is, this is like in, in like, junior high, right?
Earlier, probably.
Yeah, yeah, yeah, totally junior high, yep.
Yeah, totally.
Come on, guys.
Like, seriously, like, I am not good at math.
I am not good at math.
If I know this, you know this.
Come on.
Yes, Casey, this is true.
This is true and straightforward, yes.
This is just something people did in grade school, right?
And so the idea, this was certainly true in the old days.
for computing, right, is, look, if I have floating point, right, if I have the ability to use
floating point numbers. Intigures, obviously if we're using integers, then you have to do a lot of
much more crazy stuff if you wanted to divide instead, I mean, multiply instead of divide.
Go into that, because that's a whole another ballgame. But if, assuming you're in floating point,
then that means that you could literally just store the number 1 over X, whatever you're going
to divide by, and then multiply by it instead. So if you're going to do a by building,
bunch of multiplies, this is potentially an attractive option if dividing is slow, right? So if we
look at divide and we go, the dividing is too slow, we could just do it once and then multiply if
the multiply is faster. Now, it's worth noting that these are not necessarily the same operation,
because in floating point, right, an actual divide, this divided by that, right, may produce a different
answer because you don't know, like, when you do an inversion like that, you're only capturing part
of the answer. Like the answer might have gone on for quite some time and you're only getting
a smaller part of it. So you're not doing infinite precision math. And so you also have to be
aware of the fact that this will change the answer. So you have to, you have to know, right,
if that mattered to you or not when you're doing this. So if it's just some random thing you're
doing for some UI element or something like that, then it probably will never matter. If this is
scientific computing, it could be a very big difference. And so this is the first thing that I want to
point out is that this is not a thing you can just tell somebody always invert your numbers and
then multiply because that may be very bad advice in certain circumstances that's thing one that
wasn't mentioned in the tweet that you have to understand the butterfly effect yeah well it's
yeah basically like you know Jeff Goldblum with the water going down his hand effect
it might even just be the the the 1,000 mile butterfly like it's this huge butterfly because it could be
literally it could just literally the answer to
you get right here could be wrong enough that you care, right?
Depending on the circumstances and depending on what those input numbers were and what their
scales were and all this stuff, right?
Who does you can just I, there is a decent chance a few people in the audience,
don't understand why that happens on a computer.
Well, so to be fair, I'm the wrong one to explain it because I don't work on scientific computing,
right?
And people who work on scientific computing are much better at analyzing things like, you,
you know, what's called ULP, or basically like the last,
what happens to the final bits in floating point numbers and things like that.
But in general, when you're working with floating point on a computer,
what's actually happening there is that you're taking the size of the storage that you, you know,
made for those numbers.
So let's say that you're doing 32-bit floating point numbers.
In scientific computing, it might be 64-bit numbers, whatever it is.
If you just have 32 bits, you only have a certain amount of space to store.
store what you would prefer was an infinite series. Like, let's say we want to compute with pie.
Right? Everyone knows from grade school, pie is this infinitely long series. Like, there is no storage
that can store pie. It just keeps going. So we have to approximate that, right, by putting it
into some format. And all the numbers that we work with in floating point computing, they have this
kind of format. Just something simple. Can I interrupt you for a quick second? Please. Trash, when he says
pie, he's not talking about a snack. Just in.
case you're wondering. I don't even like pie.
Actually, I actually hit pie. Me neither.
Trash. Yes. Another reason
trash is the star of the standoff. It's no pie zone.
Yeah. It's fruit.
Yeah. What's your favorite pie adjacent
kind of thing? Like is a cake or brownies or like
what is it then? Love brownies.
Pie is not adjacent to cake and you'll take that back.
We can, we know, we just got to move on because that was very
offensive. Okay. That's it.
We'll put that next episode.
I'm putting it in the backlog.
Yeah. How adjacent is
how adjacent is cake to pie.
So anyway, if you look at what happens to these things, during all your computations, every computation that you do is going to have some infinitely precise answer. That's the answer you actually wanted and that you would have wanted to carry through all of your subsequent operations. But at every step of the way, the computer has to like round it basically to whatever will fit back into that 32-bit encoding. And that 32-bit encoding is basically, you know, it's a sign bit that says whether the thing's positive or negative.
It's some number of bits for the exponent, right?
The mantisa, I think, is the sign?
The mantis is the next thing.
Oh, the next thing. That's right. That's right.
The exponent basically says what the scale is.
Like, how big roughly is this number?
And then the mantissa are the bits that actually say, like, what the number is.
So the reason it's called floating point is because it's broken up into these pieces.
Floating point, the point is the decimal point, if you will.
the floating part is that we store an exponent,
so we say where the decimal point is,
and then the mantissa is just the actual digits.
So we're basically floating the point to where it needs to go using the exploit, right?
And the reason we do this is because we're trying to compute with extremely large
and extremely small and everything in between,
and so we can't just take 32 bits,
which has a very small range if we just use an integer, right?
If we fixed point, if we just said the decimal point is here,
it would kind of suck.
That's what floating points for.
It's why we use it for everything.
That isn't just nice, you know, typical whole number kinds of stuff that we're doing, right?
You know, usually when we use just ins or things like that, just integer numbers.
The common one that people laugh at, right, is what 0.2 plus 0.1 in JavaScript and you get the 0.9,999.
Right.
So it's like this is why that happens, though.
So just I feel like there's probably a decent amount of people who are like just ha ha, Java,
JavaScript stupid.
I mean, you're right,
but like that's not,
that's like a thing that's not
the JavaScript problem.
Exactly.
That works in C.
It works in Java.
It works in all of them.
Yeah.
JavaScript is weird
in that it kind of only has
double precision arithmetic, period.
Like, it's just like,
that's just what we do everywhere,
I think.
Unless we use bit-wise operations,
then it converts to a signed 32-bit number.
Oh, just magically at that time.
Yes, for that one moment,
you can actually do a, you can actually do a overflow at 32-bit.
only if you use bitwise stuff.
Again, this is always why
JavaScript compilers are so hard, right?
Because they have to go like, oh, I don't,
you don't want to do double prison lift for like your loops
and like counter and things like that.
So it has to analyze and go,
okay, I can turn this into an integer for this code, right?
It's got to do that work.
Normal languages, it's already knows, right?
It's just been told, like, this integer
so it doesn't have to think about that, right?
But anyway, so that's the floating point math thing.
And that's sort of what this snippet was trying to say,
is just invert the floating point number so you only pay for the divide once,
then multiply by it because that'll give you the same answer.
Again, the tweet didn't say, no, it's not quite the same answer,
and it might matter, depending on what you're doing,
and depending on what the inputs are.
But anyway, the other problem I have with this is it misrepresents the performance as well,
for a number of reasons.
So the first reason it misrepresents the performance
is because it has some weird statements about how long a division,
is. So it says something like
20 to 40 cycles for
a divide or something like that,
I think? Yes, 20 to 40.
And I just have no idea
where that's coming from, right? Like, I have no idea
where they got that number from.
Like, I really wish it had said where it got that number
from, right? Came from his heart, all right?
It came from his heart, all right? It came
from his heart. Yeah. Yeah.
You could feel it. You could feel those cycles.
The CPU whisperer.
Yeah. So a floating point
divide,
you know, if you look,
So if you look at a modern chip, like let's say a Zen 4 or a Zen 5 chip, right?
Something that you might be running today that might be in a computer you'd buy today.
Typically, their floating point divide units are extraordinary, very fast.
I want to say that they can complete these things in, I don't know, five, six cycles, something like five cycles, maybe.
I'm not sure what it is.
We should probably look it up.
But 40 is nowhere to be found.
I mean, that's, it's not quite off by 10, but it's something like 10.
Like, it's very, very wrong.
I think it's from Chad Chippity, because if I'm not mistaken, I actually asked this question
not too long ago, and I got the answer of like 8 to 30 or something like that,
cycles for a divide depending on the CPU and architecture.
I mean, the part that Chad GBT got right is the defending on the architecture part,
because it does.
Okay, okay.
But like, let's, here, I'll just look it up on my machine right now.
Like, we'll just see what roughly it is.
So if I go in and I say, I want to do a, you know, 32-bit,
like we're going to do a div PS here or a div-s-s, let's say.
And we look at what's going to happen on a Zen 5 chip.
The total latency, right, is less than 10 cycles for that operation.
Latency, right?
So that means that if you...
What about for multiply then?
Sorry?
What about for multiply?
Multiply will be four.
I want to say, right?
So three.
How many FPS, though?
Zenn 5.
We just need to know how many FPS is 10 cycles.
Right.
If it were a bouncing ball.
Casey, how fast that ball?
So three on that, and I think four is four is on my CPU, which is Zen 4.
I think it should be 4, right?
Is that correct?
Zen 4.
Also 3.
So, all right, this is better than I would have expected.
I'm used to 4.
These get three.
Right.
Now, so,
in general, that is the number that you would use if you were trying to analyze how long it takes you to get the answer back, right?
These loops don't care how long it takes to get the answer back because nothing is waiting for that, right?
On the iteration of the loop, as you go through, right, you're going to do a load, the floating point up either divide or multiply.
and then the store, right?
When you do that, the next iteration of the loop doesn't have to wait for that to complete.
So those are just in the scheduling queue going.
They can be dispatched.
The next iteration of the loop is already decoded and probably also in the scheduler,
given this as a hot loop, presumably we're just running about, you know.
It's good looking.
It's good looking for sure.
It's just cooking.
So these are all going to pile up in the scheduler.
schedule is going to dispatch them every cycle as it can. And so what we care about is not how long
it takes to get the answer because no one cares. What we care about is how long will it take us to
issue the next one? So if we issue them on this cycle, how long does it take us to issue the next one,
right? And so that's what we typically call, that's a throughput number. It's what we typically
call that, right? And if you look at the throughput numbers for multiply, right, typically you can issue
two floating point multiplies per cycle,
which means it takes half a cycle effectively to multiply.
Half a cycle.
Okay?
That is the number you should be listing or thinking about when you're looking at this loop.
For divides, the throughput, again, is remarkably good on modern CPUs,
so it's like in the two to three range.
So is multiply faster than divide?
sure, right? Half a cycle versus three cycles when we're analyzing this loop, let's say, or something along those lines.
But then you have to ask the question, do you care? This is a load, a op, and a store.
Unless you're entirely out of the L1 cache, which is going to service, you know, at about, you know, like, let's say, one, you know, it's probably like two per cycle, three per cycle it could get you.
answers back from the L1 cache, right?
If you're talking about the L2 cache,
then, I mean,
I don't know, the number of clocks
it's probably going to take 14 cycle latency,
something like this. If we're
looking at how long it takes you to actually
get the memory back, it's
not clear to me that unless
the CPU was predicting extremely
well what you were going to do and always
had things, like you'd have to go look at the actual
bandwidth to see if you were actually going to get
enough back.
such that you could operate on them at that speed.
Maybe you can.
It's not in my brain right now as to whether you would actually be able to do that
at a speed that would matter at the like, let's say three cycles,
throughput number that you're going to get on the divide.
So it really just, it was very misleading the whole thing.
And as I, you know, you heard me say all that stuff.
I would want to go look at this loop and check it first.
And I'd want to think about it a little bit before I would,
post anything like that and it was just that clearly wasn't done because none of those numbers
make sense that are placed there. Is it true that multiply is faster than divide? Yes. Does it even
matter in this case? No. I don't think it probably does, but maybe. And furthermore,
usually in most cases, right, when you're looking at actual loops, more is going on. And so it's
also going to mislead people. So point number three, it's also going to mislead people in this final way,
which is to think that they always have to do this.
This is the problem with these kind of like out of context performance things.
People read it that think, oh, I just need to invert all my numbers.
And then my code goes fast.
And then they go around inverting all of their numbers and turning them into multiplies.
And the problem with that is, in a lot of cases, unless you were, unless you had serial dependency chains that really could have been unblocked by this change, you're just wasting your time.
The divide would have been just as fast.
It would have overlapped with other things you were doing in the loop.
the thing was waiting on an uncasted read or something,
so the whole loop was waiting 80 cycles anyway,
so none of this mattered, right?
So the problem, again, is that, like, performance is a thing.
Performance is a process.
It's an understanding.
You have to understand all those things that I just said.
They're not actually that hard to understand.
You just need to learn them.
You need to spend a few weeks learning this stuff.
And then you can think about,
do I need to invert this number,
do I not need to invert this at that point?
What you don't want to do is try to package it into a tweet like that
when you haven't really done the work of understanding why it matters or anything like that
because other people will take away the wrong lessons.
So that was a very long-winded way of trying to explain everything I didn't like about that tweet.
Again, no offense, sorry to the person who posted it.
It's just I don't think that kind of thing is helpful.
It's sort of like saying always call memset or never call memset
or these other weird performance things that people pass around,
which out of context don't help anybody
because you're just going to make incorrect decisions
if you think that the key to performance is always do X or never do X.
That's not how it works.
Spark plugs.
Great job, Casey.
Thank you.
I did my best.
I figured you'd be able to help me on that one because I had the rough idea that I would introduce error,
but I had no idea to the extent that divide and multiply may not make any sort of actual real difference in your program.
Plus, all the crappy things I do anyways, probably destroy my program's performance long before we get to this loop.
Yeah, Prime, we've got other issues.
before we get to whether we're multiplying or dividing, I think.
I just really want to take my Lua program seriously, TJ, okay?
It's usually, the one thing I can put in someone's head that's generally not wrong is,
don't confuse integer multiply with floating point multiply.
Floating point multiplies are much faster than integer multiplies typically.
And so some of the time when you see people saying, oh my God, like divide is so bad,
sometimes they're thinking about integer divide
and there's a lot of reasons
there are a lot of reasons
why integer divide is bad
a lot of reasons
I'm gonna put that on the blocker list for next week
you can we can talk about that
because that's actually a
that seems very interesting
okay
is it on the pause
the pregnant pause
I was like
sorry I was I was just putting it on
I thought Casey would finish anythings
now hold on can I just jump in for second
can jump in here
I'm gonna say that this instruction thing is done
Trash, I had a, TJ said you had a blocker
that you need to discuss. I don't know what it is.
It's an important blocker, trash, that we needed to get out to the world.
What is this blocker that you had to talk about?
Oh, dude, I'm locked out of my GoDaddy account that I set up like a decade ago
and I got a hit with like a renewal, whatever.
And I tried to get in.
Apparently it's using like my college email, which got deactivated God knows when.
So then I hit up, I hit up like their chat bot, not chat about, but there was like apparently
supposedly a real person.
Yeah, for real.
Sure.
And he just kept asking me, show me a screenshot of the air.
And I was like, dude, I'm locked out of my account.
I just need to change my email.
He's like, oh, show me a screenshot of the account.
And I was just, I got to the point why I just got so fed up because I was trying to do it for like an hour.
I went on Twitter and then I just like tagged them.
Now I'm like making some progress here because I want to get a refund from my like domain renewal.
Because before I was like a program, I did photography.
so it was a photography domain.
Of course you did.
Of course he did.
Dave that for next week's fun fact, bro.
You just spoiled it, dude.
Oh, fine, whatever.
But anyways, I'm just trying to get rid of that domain and get my money back.
It was like $200 or something, which is insane.
It's Pokemon cards.
I can't believe you could have something renew a decade later successfully.
Like, how did you keep hold of some sort of long-running credit card?
Like, what do you got on there?
I don't know.
I just got a random email.
I was like, this domain's going to renew.
And I was like, huh?
I haven't seen that.
expecting like
Domain and trash to be like
well before I was a photographer
I was in FinTech and
so I had this
I started a credit card company
Before I was an archaeologist
I used to do deep sea diving
The time where you hold your breath
I got some other things
I got some other things up my sleeve
But yeah that's my blocker
Hopefully go daddy if you're watching
I want my refund
and I need my email changed
ASAP please
Nice. I like it.
All right, well, we've now finally hit the part where I can introduce the main topic, everybody.
Today's stand-up, welcome to the stand-up.
We're going to be talking about NVIDIA and Microsoft, talking about what they're releasing.
I got some hot takes on Microsoft.
I didn't tell Casey long enough beforehand, so he probably has zero hot takes on what's happening.
But guess what?
We're still going to talk about it in today's stand-up.
One thing I'm truly disappointed by with today's stand-up introduction is that TJ has an in-progress task called Interrupt Prime's intro, and he did not, in fact, interrupt my intro.
I am now done.
TJ, what have you done?
See, in the business, that's what we like to call PowerPlay.
And I just made you do the whole intro waiting for me to interrupt the intro.
And I didn't even have to do anything.
Living in a nice.
in progress and I did nothing.
Anyways, that actually sounds like your daily life doing nothing.
Okay, so here we go.
Now that we got out of it, got them.
All right, so I think the first thing we have to talk about is this Microsoft Surface.
Has everybody seen the Microsoft Surface?
Yeah.
Oh, you mean in the 15 minutes between when you sent out the invitation and us recording this?
Did we see the, is that what you're talking about?
Yes, yes, yeah, yeah, yeah.
That's specific one.
that's the one that has the
DGX in there, right? Or the
RTFs in there, right? The RTX Spark
and so they're saying it has one petaflop, which
by the way, tossing out
one petaflop, I don't know what that
actually means. Is that a real word?
It's FDFLop, yes. I'd petaflop too if it looked good enough, but here's the
deal. I think it's for how badly
the product's going to fail.
It's going to be a petaflop.
It's a petaflop.
It's FP4 Tensorops, I'm pretty sure.
Okay, you just said a bunch of words at me.
Like, I should understand that.
I just think it just means it's, what is that?
That's one above Terra, right?
So that's like 12, E to the 12 amount of floating point operations per second?
Yeah, so, okay.
What did you say?
I said FP4, so 4-bit floating point, right?
FP4
TensorFlow ops
so meaning these have to go through
the tensor core
like largest
largest dimension mat-mol
okay
I'm probably just speaking gibberish
at this point
so we had this big long discussion
coincidentally about floating point
just moments ago
and we were talking about that whole
like exponent mantissa whatever thing
so in the quest to quote
higher and higher numbers
GPU vendors
have been making progressively smaller and smaller floating point types.
So we have 32-bit floating point.
They made 16-bit floating point.
That's been around for a while, actually, different kinds of 60-1.
What do they call a 16-bit floating point?
Because I know they call them doubles, floats for 32 back in the day and doubles for 64.
What's a 16?
They have various names because there's different encoding.
So there's like B-Float 16 and I don't know.
I don't use this stuff, so I have no idea what the, what the, there were, there were 16-bit floating point numbers for graphics, which were used for things like, oh, you know, we can, this frame buffer that we're storing results in.
We'd like to be able to store lighting results in something better than 8-bit color, right?
So we're going to go up to 16-bit color, where we store that as a 16-bit floating point number, because we don't want to go all the way up to 32-bit floats because that's a lot more memory bandwidth and blah, blah, blah, blah, blah, right?
You can imagine all the things that go downstream of this.
Stuff like that.
So 16-bit floating point is nothing new,
although the particular type of floating point you use can change.
Like, how many bits you dedicate to each thing, right?
Because remember we talked about exponent, Mantissa?
There's like, how many bits you dedicate to scale,
and how many bits you dedicate to the actual digits is up to you in something like this,
because you can move that tradeoff around depending on whether you wanted more range
or whether you wanted more digits in your number, right?
So, yeah, so we went down to 16.
Then there was eight, so eight-bit floating point numbers.
Again, very small amount of information, but, you know, still doing it.
And now they're actually at four.
So they actually do four-bit floating point, which I don't even, since I don't do AI.
I think they have one and a half bit.
Isn't that like the AI super quantized models, one-and-a-half-bit numbers representing?
I don't know.
I don't look at AI stuff.
So I'm not sure.
All I can tell you is I believe this is an FP4 number.
That's all I know.
And that could be wrong because, like I said,
I only had a very brief amount of time to even see what the heck this announcement is
we're supposed to be talking about.
So everyone out there take literally everything I'm about to say with a pretty big grain of salt.
I mean, always a good idea anyway, anytime anyone's talking these days,
take it with a big grain of salt.
But an even bigger grain than you ordinarily would have shaken out of the salt shaker into your mouth.
So the way that they quote these numbers now,
I believe, because they're marketing numbers,
is they take the smallest possible thing they can do on the chip.
So if your chip does FP4, you're quoting FP4 numbers, right?
Correct.
And you may ask why.
The reason is because all GPUs, they are wide ops.
They're like SIMD ops, single instruction, multiple data.
The lane width is typically, depending on how you want to consider it,
either 32 or 128, depending on how you look at it.
in terms of the size, the bites, that you're going to operate on at a particular time.
So when you talk about how many operations you can do,
they're not talking about how many like instruction operations,
like I told this thing to multiply.
They're saying, when I tell it to multiply,
it's going to take this giant bite pattern with a ton of stuff packed into it
and do this big multiply, right?
So they take that and they say, okay, how many of these can we pack in there?
Take the smallest thing, which is SP4, right?
Take the most efficient thing we have, which is when we multiply these two things together,
what instruction can do the most of these multiplies in the shortest amount of time, right?
And then that is the number that they quote you.
And just for good measure, because typically what you're doing is F-MADs,
Floating point multiply add, right?
It's a multiply in an ad in one step.
They, just for good measure, multiply the number by two because they're like, well, hey,
you're doing two operations.
So technically, it's twice as much, right?
So you have to do all of this to figure out the actual instruction rate, right?
But yeah, so that's what's happening.
Sometimes they quote dense.
Sometimes they quote sparse.
I don't know whether this is a dense or a sparse number.
I didn't have time to go look at the math to see.
you can tell pretty easily by just looking at the kudacore count of the thing in the clock rate.
So it's not hard to figure it out.
But and the way that that works is these kinds of cores have this ability to effectively pack data sort of, pack and unpack data as part of the operation.
So that if you have something where you know there's a lot of zeros in it, like it's a spark, like it doesn't have not all the numbers are always filled in.
they can exploit that sparseness and get higher throughput by just the chip actually seeing the sparseness
and only issuing the multiplies it actually needs.
So it can basically fit like more, you know, it can pack the zeros out.
So one of those big 128 byte lanes can hold more than 128 values because the zeros have been eliminated or something, right?
There's all these kinds of things they do in there.
And I don't study this.
So I don't study the specifics of that.
And I don't know which number this would be, whether it's a dense or sparse.
But that's everything.
When they quote these numbers, all those shenanigans are on the table.
Okay, this makes a lot more sense because that number, you know, in my head, I'm thinking, okay, this is like, you know, what's called architecture-wide floating point operation.
So they're doing, like, double, like a kind of flop of doubles.
I was like, this seems like that's really fat.
Like in my head, that's like supercomputer fast.
So I must be misunderstanding what we're talking about at this point.
Because then I went on the internets, you know, as I come and do from time to time.
And I started asking, like, I went to the AIs, of course, asked about the AIs.
Like, like, with the new Spark thing, like, how fast are these models actually being able to be ran?
Because it looks like it can be ran at, like, super fast, right?
You could have, like, local models running at infinity speed.
And a lot of the forums slash what the AI had to say was saying, you're going to be running at somewhere between, like, 8 to 16 tokens a second, which on a 32B model.
so they're not the speed and all that I don't I don't know actually what these speeds mean because that just seems like the ultra fast number I've ever seen in my lifetime but really it may not be nearly as fast as as at least my brain would assume that means well again I mean so if you look at the actual chip like like what's in this thing right it's basically like the same number of kudacores as like a 5070 right nice
It can run Fortnite really good, so I assume it was a pretty good, you know, graphics unit.
Well, I mean, maybe the other problem, though, is it's, so the memory is going to be different, right?
Because this thing is LPDDR.
So low power, the stuff that's in laptops.
And so if you compare it to something like a 5070, a 5070 is using GDDR, which is the like the high clock rate,
high power consumption, high clock rate.
The gains.
Right.
And so the tradeoff,
the actual thing that you have to think about here is
a 50-70
has like three times the memory bandwidth or something of this thing.
So this thing is actually a third the speed
of getting you the actual memory.
So I don't know how good it would be at Fortnite.
Maybe fine. I'm not sure.
But it's definitely worse than a 50-70 at Fortnite.
So you're kind of doing a trade.
You're trading the fact that this thing
will have much worse memory bandwidth.
I mean, a third of the memory bandwidth is not good, right?
Which I think is roughly what it would get.
Again, take it with a grain of salt because we had all of zero seconds for me to look at this stuff.
But the trade is that because this is using LPDDR, right, it can have a ton of memory.
Can have 128 gigabytes of memory.
Whereas an RTF20 is like, I don't know, it was like 12 gigabytes of memory.
It's like a tenth of the memory.
So you're making a trade, and I assume that this is a trade that AI people are willing to make, because although it'll be slower, so evaluating any particular part of your model, in theory, would probably be slower on this chip because it's got one third of the memory bandwidth for pulling it in.
So it can only go at one third the speed of pulling it in if it's memory bound, which, again, I don't study this stuff.
My understanding is that AI workloads often can get memory bound, depending on.
and what you're doing.
But on the flip side, you can fit a much larger model in there.
Correct.
You can fit 10 times larger model.
So I'm assuming that's a trade that AI people would be happy with.
Otherwise, why did they make this thing, right?
Can I ask a dumb question really quickly?
I probably can't answer if it has to do with AI.
No, it has to do with memory.
And specifically why you said like one-tenth of memory,
why does the low power enable more memory versus the high-power counterpart?
Is it purely just a power draw problem, or is there something I don't understand?
I really couldn't tell you.
So I would assume that it is probably more to do with cost than anything else and or market segmentation,
meaning I don't think there's anything in particular that's – well, you know there isn't
because you could go buy a 50-90, right?
And what does that have?
24 gigabytes?
I assume it's 24.
I don't remember what they have, but they've got –
way more, and that's on a consumer level card.
You're not talking about buying like a blackwell for the data center.
It's 32, right?
So obviously you could have stuck more GDDR on there.
It's just the cost is higher.
And the, you know, so I think they're trying to get
128 gigabytes into something that doesn't cost you 10 grand.
If you wanted to go pay data center prices, you can get that.
I mean, you can get high bandwidth memory for that matter.
HBM modules on their, like, actual data center parts,
they can do it, right?
So it's really, I think it's just a cost issue in terms of the manufacturing of these things.
I imagine GDDR is just a lot more expensive than LPDDR.
That'd be my guess, but I don't know.
Probably a little bit for laptop form factor, too.
I mean, it's just like you'd rather not have necessary, like the, I don't know how much power drop,
but I imagine going faster means more power, more heat, bad battery life, etc.
Well,
especially with GDGR.
I don't know if you've seen his pictures,
but he has a battery life
that he could run for days on a laptop.
Well,
I didn't show,
I didn't post that picture,
but I can grab it if you want,
probably good picture.
So,
Casey,
I did,
I did verify the RTX Spark
has 300 gigabytes a second
for the unified memory system,
whereas the 5070 has
670 or something like that.
Yeah.
It has an over double
of whatever it is.
And then the 5090 has like 1.
300 sounds a bit high.
300 sounds high.
for the spark? Where are you getting that from?
This is the summary from
an Instagram Verge post
and then Tom's hardware as well
as reporting apparently that exact
number.
Interesting. Right here. So Tom's hardware
is saying it's 300 gigabytes a second right
here. Because it looks like 273
on the
Oh. Okay, so
they may have done a little bit of rounding.
It's like 273 versus
672. So, yeah.
I mean, yeah, maybe
maybe saying a
third is being a little bit harsh, but I wouldn't say double.
So it's more than double.
Let's take a look at what it actually is if people want to get that specific about it.
2.5. So split in the middle.
Split the baby.
So 2.5.
And again, that's the trade, right? You're just trading like, okay, we got, you know,
we're going to have less speed at pulling this stuff in from memory, but you're going to get a lot more memory and presumably
that's what you want. Otherwise, again, I don't know why they would build this, right?
Presumably, that's what people want because if it wasn't, I then what's, I don't know what the
pitch for this thing is, right? So I think I understand more of what this pitch is going to be in kind of
the general direction Microsoft's taking is that they kind of, he said a very interesting thing,
which I just wanted to bring up in here. Who's he? Who's he? Well, I'm about to show you. His name is
Satya. And he said the following. He said, our goal is to deliver unmetered intelligence to every home and
every desk with Windows.
And then they do this RTCS spark.
And part of RTF Spark is that you can run a bunch of local models yourself.
So you don't have to necessarily rely on the cloud models.
You can run models yourself as the general kind of tenor or goal, it seems like,
for where Windows is going.
And there's a very specific phrase in there, which is unmetered intelligence.
Sam Alvin for the last two years has been saying that he envisions open AI to be a
metered intelligence.
It should be no different than your electricity, all this.
So this kind of seems like, hey, they're kind of going against this general idea.
And instead, they're moving a lot of it down into the individual's laptop and then building harnesses.
So my personal guess is that what Windows is wanting to do with these larger memories, but a bit slower,
is that you can run a bunch of these smaller models to do mundane tasks, simple tasks,
and then have a cloud subscription to run a model up in the cloud that doesn't need nearly as much power
because they're just telling little models what to do.
and you're running like four models in parallel to go and do these basic operations.
And so thus that's why they're fine with these, say, lower bandwidth memory,
these four to tens tokens per second kind of production because it's not really,
they don't really have to care about that because they have multiple going and they have
the big powerful one up in the cloud.
And they're just trying to win the basic consumer by reducing cost and reducing data center.
So that's what I think is generally happening here.
And that's why they've chosen this general architecture.
texture.
I mean, I got two things that it'd be fun to chat about for it.
I think the first one is, is it not a little bit weird, like that it feels like a direct
shot at Open AI that like Microsoft has a really big stake in and has spent a lot of money
on?
It feels like, like Sam's thing's been the, it's like it's too cheap to meter or like it should
be treated like utility or we should like have UBI and tokens.
They kind of gave up on that train, but that was, I recall talking about it on the POTS.
I'm interested to hear sort of what people are thinking for like, is there some trouble in Paradise over there?
What's going on?
Well, we do know from the court leaks that Satya was a major reason why Sam, you know, held on to a CEO ship of OpenAI.
And so there is something, it is very interesting that he's using very, because we all know that Sata does not tweet a tweet that has not been looking.
at by human resources, by legal, by, yeah, especially not this one.
Like, they've really, this is a very intentional tweet, which makes it seem very intentionally
against whatever that vision is of the world, which to me, I think, I think part of this
is also in response to like the data center hate is that generally speaking, people seem to be
very anti-Data center in our current, in our current world, whether it's founded for good
reasons or bad reasons people generally are. And so having the ability to have agents run on
your computer as opposed to up in the cloud means less data centers. Thus, it feels more tenable,
I guess. And so my two guesses is that there is trouble in paradise when it comes to
this, because Windows has also created their own frontier model. And so there's like 900
things going on to where it seems like they've went from competing, like competing against
the other models using Open AI to saying now we're competing against Open AI with our own
models and we're like directly going against it and their contracts i think run out 2032 so maybe there's
something about that trash thoughts no i was going to say i'm actually i've been thinking about doing
local a stuff we have like a local ai group that i'm part of and they all have like these like dGS sparks
are just way too expensive um so i feel like this was very good time wait wait trash turn around
what what yeah too expensive it's you said it's too expensive dude djc spars is like
5,000 or like almost $5,000.
It's crazy.
Do you know how many Pokemon
Pekarza could buy?
It's a fraction of this, but
No, but
like to be serious here for a second,
I think it's pretty neat.
The one thing I don't like is like there's no Linux support.
Now I'm a Linux knob now.
Thanks to the last.
I was like, no Linux support.
I'm not even taking a part of this,
part of this show right now.
But apparently it might.
be coming. But I think the interesting thing is I'm curious how, because I already know a lot of
these companies are underwater with as far as revenue goes. So if I have this computer, I don't
need to subscribe to chat GPT anymore. I can do my mundane task. I feel like most consumers aren't
doing anything crazy besides actual enthusiasts and developers. So your average consumer doesn't need
that like crazy power. So they can run a Q132 model on their machine and get away with it.
And then if they're really, I don't know much about what Microsoft's releasing with their seven frontier models.
I'm assuming they're shoving that down their employee's throids right now and it probably sucks.
They did cancel code.
Yeah, exactly.
I'm assuming that's definitely like leading up to this moment.
But anyways, I'm just curious how it's going to work out for Anthropic and ChatGBT, BT, because I know like the big thing was like, we need more compute.
We don't have enough compute.
But now everyone has it on their machine now.
That's no longer going to be an issue.
And also, you're going to lose all those subscribers.
and that's the next issue.
So I'm just curious how chat Chb-T is going to, like, respond to this.
And if they're going to think, like, they're potentially in trouble here, I don't know.
Yeah, the other trash which you mentioned, which I was kind of interested in, is I'm wondering how much of this is like a, you know how when someone says no to you?
So you just reframe it like it was always your idea and that you're really happy that you can't do option A and now you have to do option B.
like I do wonder if some of it is like, hey, we thought we'd be able to like build all these data centers and everyone would be like stoked about it and we'd get to like do all that stuff. And then that doesn't seem like it's going to be possible. They can't get the electricity. They can't get anything else. And so are they doing like a pivot to try and be the first like, you know, first past the finish line on these sort of like home AI setup things because they can't get the data center stuff approved? Right. You know what I'm saying? Like in terms of like, well, we need to get.
AI to everybody, obviously.
But we can't do the data center thing.
You know, is I don't know.
I'm sure that's a part of it, but I feel like they were probably working on this well
before like the hate on data centers potentially.
Well, some of these have been worked, but like the packaging is very different.
Like packaging it inside of a surface laptop.
Like they've been building all these components, but putting it together and one is interesting.
Yeah.
I'm curious how it's all going to shake out.
I don't know that.
I don't think they mentioned prices.
Did they prime?
I don't, I didn't see the prices.
I haven't watched the, I watched the keynote, and the keynote didn't seem to mention
prices or at least that I remember.
I couldn't find prices anywhere.
I'm assuming it's just going to be kind of insane.
High and 28 gigs of RAM, even if it is low power along with a 5070.
Come on, come on.
Well, I mean, I can't, is like, what would even be the explanation for how it would be
cheaper than a DGX?
the thing that trash doesn't want to buy
because it costs too many Pokemon cards.
Like, you take
some of the networking out of that thing,
but now you got to pay for a screen and a keyboard
and chassis and all this stuff
that you didn't have pay for before, right?
And a battery.
So how does it get cheaper?
Like, it seems like it will just cost the same
as the thing you can already buy, right?
I mean, I don't know, I don't know why this would be cheaper.
I'm not sure about it.
Is that there's a huge amount of people
that are just Windows and,
shall we say they don't they're not really into the computer side of the computing they don't
really want because if you've ever played with local models at least last time i did it last year
there's a lot more work to it than simply you know plug it in hey i'm up and running and i think
that's what microsoft is selling is this really kind of like hey oh you you're you're just in
excel you're just in these things well here's are integrated it runs locally so it's fast it you don't
even have to pay money you just have to hey just buy a microsoft 365 and pay the once a year fee for that
and you're good to go.
It runs all locally.
It's all yours, right?
And so to me,
they're making a play to purchase people in
by saying it's free effectively,
but do you have to buy all their services anyway?
So they're like still winning in the end.
It's just giving people the access.
Well,
I'm just pointing out for this particular product,
I'm not really seeing how this comes in
at a price point below, you know,
even $3,000 possibly higher.
So it doesn't seem like,
like this is not something very many people
are going to buy.
didn't think. It's geared towards programmers predominantly. Like that's what at least their presentation
really made it seem like a programmer thing. The person that did the examples talked about how they got
all the like effectively Unix tools like grep is now in Windows. All these things are now in Windows.
They have like 53 tools or something like that. And so it did seem like it's geared towards
the working professional as a working professional computer. Right. And that's why I was saying like so it,
so this isn't really something that you're going to sell to, you know, somebody who just runs Windows and
doesn't know anything, right?
Like, it's not really targeted at them.
It's only targeted at people who already know what they're doing and are using the command
line and stuff, right?
I think.
Casey, you're not thinking about this.
You're not thinking about this in the right way.
Okay.
So here, we were talking about Tam earlier, obviously the Tam for this podcast and the
company that it represents everything that's ever lived.
Yeah.
That coincidentally, everything that's ever lived, they can be programmers now, Casey, because
of AI.
Right.
haven't even thought of that probably.
True.
So everybody at the company is a programmer now in 2026.
So you just have to convince the CEO of the company that since everyone's a programmer,
they need a programmer computer.
And therefore, they need to get this computer for every single person at the company so that they can run it.
Hey, Casey, do you want a $500 million token bill?
Or would you rather just pay, let's just say, a million dollars for a few laptops?
and then it's all run locally.
What do you think, Casey?
Well,
Hey, I'm just saying.
Which one you want, Casey?
Or Prime, a thing that looks like a griddle that you can get.
Did you see that thing?
It looks like a hot plate that you could cook on, which I think is kind of fun.
Like, can you make an egg on this thing by putting a like a piece of like Teflon on top of it?
I want to know.
I did.
It actually did remind me very, I bet I bet you I even have one in here.
An Xbox 1x.
the top to the Xbox 1X
they just lifted the design of that
and just put it as the like the grating or whatever
the grit on top of the griddling
here's what I want
and I'm dead serious
I want someone to take one of these things
the griddle one
I want them to put a solid metal
piece in place of the waffle pattern
and then I want them to create a thing
called Bacon Mark.
And Bacon Mark is basically an AI throughput benchmark
where you put a piece of bacon onto it
and then the benchmark result is just a picture of the bacon
after the workload has run.
Okay, okay.
And how do you determine if it's good result
at the bacon still raw?
Literally.
The raw or the bacon, the better the result?
Yes.
Yes.
Right?
I mean, maybe or, but yes, like if the,
the if the thing computes the same
if it's computing the same answer
I guess then the more
like intact the more virginal the bacon
the more efficient this AI
should be a fried egg
unfortunately Casey this we did miss
the boat on this one this would have been
perfect for Apple
to release because then we could have called it let
Tim Cook okay that
would have been the perfect too late
too late he's gone it's past
it's over
yeah yeah
unfortunately they did let
Tim Cook and he got kicked out of the kitchen
He couldn't handle the eat. His cooking days are over.
I do think this is really interesting though.
I am very curious where this is going
because there is something that's very
unique about this moment that Windows did have
a three and a half hour keynote and this was obviously like the kind of the
centerfold of this. I'm not sure if that's the appropriate term
now that I think about it because I do remember that song.
My baby is the centerfold.
Do you really want that time?
Do you really want to unfold that?
I don't think I can't do that.
I can't,
I literally can't put the centerfold back in the magazine at this point.
So we're just going to keep on going.
And someone mentioned that they never mentioned any of their pro,
any other programming apparatus other than AI except for VS code once.
And so like they are fully going towards the programmer,
but they're going towards them exclusively through.
model usage.
Like they don't talk about
typescript.net.
Like dot net, they've always talked about dot net.
They're not talking about dot net anymore.
They're not talking about C sharp.
They're talking about
models.
AI.
I'm curious how those models are going to be.
I said their flagship one is like equivalent
to 4.8 or something.
Honestly,
he misses you a lot and gets distracted.
Goes on X
and tells me about things that are happening.
Okay, I do have to stay.
So, Casey, 4.
Opus 4.8 came out, so I gave it a test run. And it was hilarious because the very first message it sent back to me, I haven't been using Claude Models in quite a while. The very first message goes back. It goes, excellent insight. That's straight to the point. And I'm like, ah, bro. Bro, put the code in the bag.
Bro, stop spending my tokens. I'm telling me how smart I am. I already know that. Okay.
Please. Got to get a little glaze in there.
Got a little glaze on that code donut.
It was crazy.
Immediately.
It was like so obvious.
And I was like,
this is why people are falling in love with Anthropic.
It tells them how smart they are every time they send something to it.
Oh, man.
Yeah.
As far as models go,
as like which one's better?
Honestly, at this point in my life,
I can't really like,
besides for like some kind of weird implementation details,
like some will create many tests,
some will create one test.
Like there's like kind of these weird like kind of variations in that.
As far as performance goes, I'm having a harder and harder time disambiguating which one is better, right?
Like I can't.
If you just show me a function, a snippet of code, it'd be hard to understand which one's better.
So I'm sure the Microsoft one does fine, right?
Like I'm sure you wouldn't even, you would probably not be able to understand at a at a more micro level.
The difference.
I could also imagine if I used my powers of imagination to really.
imagine in an imaginie way.
Yeah. Yeah.
That if your goal is to program
on Windows, which at this
point, I don't know.
That's the imagination part, Casey.
That's the imagination part.
Microsoft obviously has a huge history
of source code control check-ins
to Windows codebases
that nobody else has, because they're all
internal. They weren't done publicly.
And so their AI is presumably trained
on all of the entire development of the
NT code base and all the development of
Windows 95 code base and all of the development of Microsoft Office and all that stuff,
which nobody has as a training dataset, except people who maybe like hack sorted out of Microsoft
or were paid to smuggle it or whatever happens in AI circles now.
I'm sure it's all very illegal and that they're all fine with it because they've convinced
themselves of something.
But I would imagine that Microsoft is probably the only people who can train on that unless
Open AI got the rights to train on that as part of one of their sleazy backroom deals
that they do or whatever.
However,
I would just like to bring this back to
another point which is,
again,
I think it's worth underscoring
that this is a non-announcement.
Like, the only thing that's actually,
as far as hardware is concerned,
this is just an announcement
that Microsoft is like getting Windows to run
on this thing.
Right?
Like, that's it.
Because there's no, it's just a, like,
I guess you can,
say that the laptop part is a form factor announcement.
It's like someone's going to put this thing you could have bought from Nvidia since last year, right?
We're putting it into a laptop for you if you prefer a laptop.
Like that's the hardware announcement.
There isn't another thing right there.
It's not a new chip.
It doesn't have new capabilities.
That's it.
So really, the announcement is more of a software announcement.
It's like Microsoft got Windows on Arm running on this thing.
That's really the announcement, right?
I mean, and laptop form factor.
So just want to make sure that that's got, like, there is no hardware.
There's no chip thing here, right?
I mean, are we all?
Yeah, I think the chip thing is just like how small and integrated it is, right, with unified memory as opposed to having its own dedicated V RAM, right?
Isn't that like the...
No, that's the same, that's the thing Trash was talking about that he didn't want to buy because it was too expensive.
It's the same chip, same memory.
I'm an idiot.
Sorry.
So we took this thing that you could have bought since last year
And we're putting it into a different thing now
That's it for the hardware side of things as far as I can tell
I didn't see any announcement of like a new actual
Capability or a more modernized chip right
So it's really more of a software announcement like this
The actual new thing here in terms of what's going to change for you
If you instead of buying one of these Nvidia
little desktop boxes
we'll call it a Mac
Mini compete or whatever. Instead of
buying one of those,
the only difference really is that now you'll be able
to run Windows. Right?
You'll be like before you couldn't
run Windows on those. I don't think Windows on Arm worked
on that little Nvidia
box. But it's the same box
just in a different form factor now.
It's a nice little skinny laptop.
And Casey, if I'm wrong,
this was, there was
some sort of like Jonathan
Blow was tweeting like that what if it was like a completely different style of computer right that was
like in terms of the things of like yeah and it definitely isn't that like right there's the mysterious
mysterious tweets and all of these secret little signs and stuff and you're like oh wow they're
going to release like a GPU computer or something insane no and then it was like yeah whatever
john blow wanted he definitely didn't get it because literally the chip because literally the chip is
the same chip that you could have bought last October, right?
So there was no, like, they didn't announce any new, like, architectural, interesting things where
John would have been like, oh, that's a step towards the thing I want it.
Like, no, there was no step at all.
Well, because they were doing the mysterious, like, oh, Nvidia sent the date and Microsoft sent the date.
And you're like, whoa.
No.
Is this going to be like the first time it was a location in Taiwan, if I'm not mistaken.
It was the location of the, it was the, it was the, the, the, the, it was the, the, the,
GPS coordinates of where they were going to give the keynote, right?
Right.
Yeah, yeah.
So you thought it was going to be like, whoa, I bet their stock's going to double after this.
Like, it's going to change the game.
Let me look at the stock, actually.
Oh, yeah, trash.
Let us know.
Yeah.
So that's the longest short is if you want this thing early, just go buy an Nvidia,
one of those little Nvidia Spark boxes, whatever they called it, the RTX Spark or something.
I don't remember what it was called.
Yeah. Just go, you can literally buy this ship right now if you want it.
By the way, you can see it.
in the stock price.
It actually did directly affect the stock price.
And then after everybody saw it,
it directly unaffected the stock price.
People are like,
literally,
I mean,
look at this.
You're talking about somewhere in the teens,
415,
we'll just call it 415 for average right there,
all the way up to 460.
So a 10, 15% boost all the way back down to now 427.
So,
oh, well.
Oh, well.
I mean,
So to be fair, though, I am going to at least try to glaze it a little bit here.
Yeah.
I think that trying to prevent, like, first off, 128 gigabytes of unified memory and then saying
this 120B plus, this is very misleading as far as what you can run locally.
But excluding all of that, the idea that you should be able to run your stuff locally,
and then that, like, that's a good idea, I think is a really good message to be said,
because, you know, the whole Dario's of the world talking about how everything's dangerous
and they should own everything you're doing.
There are so many companies that have really, like,
competitive data about a activity, right?
They have users who pay them money,
who send them specific data.
They have, like, all this good stuff,
and they have to effectively ship everything off to open AI
if they want to use any sort of these models
to be able to do any sort of kind of cool things
that they say can't afford a team of data scientist
to figure out how to do.
And they don't want to share this data with anybody.
I'm thinking of, like, mining factories.
I'm thinking about just a lot of just more kind of brick-and-mort
tech kind of combo places.
And the ability to run things locally, I think, becomes really appealing to them.
And so this idea of moving towards it, I think is really good.
I mean, I personally think this is a good direction overall for people, because you stop
having to be tied to large companies.
And I think that that is overall a good thing.
So I would agree with that.
And like, it would be nice if Microsoft decided to go that direction for some reason.
They're not.
They don't, they don't.
They don't seem like.
like it because, you know, Sacha Nadella is famously the person who came in and basically like, I mean, I'm not going to say killed Windows, but he was definitely the person who came in and said, look, cloud is, cloud is what we do. That's where we're going to make all our money is subscription, soft as a service. We don't really want Windows to be good because the better, the better things are locally, the less we can charge you remotely. That's Satch and Adela. So if he is, what? Well, yeah, but that, yeah, but that's kind of separate from this, this thing. But, but.
So if he's actually now changing tack on that and he's going to turn Microsoft back to an actual personal computer company, that would be fantastic.
I don't believe it.
But if that actually happened, that would be a great outcome.
Yeah.
So I'm trying to like be at least Microsoft positive because I'm always pretty negative on Microsoft.
I think they've somehow just fumbled the bag so hard for the last decade.
I cannot believe that I think Steve Balmer is a better CEO.
than Saki, like, I'm saying it right now.
That's my personal opinion.
Halo 3, and he was ruling during that time.
Absolutely love the man.
And so it's just,
this has to be a good,
and that's what I'm hoping for, is that people realize
the need for local,
and perhaps that will push Microsoft.
But Microsoft, they did it with VS code,
which is they built a bunch of cool open source stuff,
and then they just immediately started EEEing people.
I assume this thing is going to be the exact same thing,
like, oh, all these cool local models.
but if you really want to actually use it with Windows,
you have to pay for all these services to be able to build your soul
or your memory thing and your management and all this kind of stuff, right?
And so it's just going to be, they're going to re-EEEU again.
They're just in the embrace season of local AI right now.
And someone's asking what EEE is, embrace, extend, extinguish, right?
You become the good guys.
Everybody comes to you.
You extend it and make it really, really functional,
but then it becomes also like privatized
to where it's very hard for
anybody to compete with you, but you're also now
extracting value. And then you extinguished
by diving your prices and effectively
killing everybody else out while you can then reap
all the benefits of being a single
monopoly winner. Classic Microsoft,
by the way. So I'm trying
to give them a little bonus. Okay. Hey,
I hope they're good. All right. Now that
the second donut has been glazed.
We were going to talk about
Nvidia as well, like the Spark. Is there
anything interesting in that or is it really just purely a
this is just a it's just the same thing because i was reading through it and there's nothing
that really feels uh interesting about it it just i mean they really do love d lss s and i know that
that just just just ruffles some serious feathers out there if you mentioned d ls s s s s s
that's all i know yeah i mean this this thing is not uh like i said it's not a new chip
if you want to know how it performs with gdd or no it's still no no in fact it's still
LPDDR in the other box. So you can literally just go see how the thing performs. And I guess we don't know what kinds of concessions might have to be made for like the laptop version. So maybe the laptop version performs worse for some reason because of some thing that has to happen there because of the power. That part I don't know. But in general, like if you want to know how this thing works, well, you can just go look at the benchmarks for the thing that we already have. And I believe the benchmarks for that thing put it somewhere around to 3080.
so like if you want to know how your graphics would run on this thing it's like a 3080
roughly um which is fine which is fine uh performance i think like again take all this with a grain
of salt but i think that's roughly where the spark is uh and again that's just because as we said
before you know games care a lot about the memory bandwidth on a GPU um and so you know
having not that much memory uh you know on a GPU is actually fine if the bandwidth is really
high. That's why we don't, typically gamers don't
have 128 gigabyte GPUs,
right? That's not a common kind of skew.
And that's because, you know, games
are used to fitting in 8 gigabytes,
12 gigabytes, 16 gigabytes now.
So they just don't
need it as much, but they do need that memory of.
Everybody to have 128 gigs
so that games can start taking up that
much space. That's all that will happen. That's
exactly right. All that will happen is
now the games will just take 128
gigabytes. They're all going to ship
JavaScript. It's going to be so good.
Sorry, sorry for ruining everyone's day.
You guys thought it was bad downloading 100 gigs for Call of Duty.
They're going to have it be several terabytes, and it's going to be 128 gigs of RAM required.
It's so good, though?
And I believe, I mean, I think the, again, didn't have time to actually look at this myself.
But vague understanding is, you know, so it's media tech, right?
I mean, presumably you guys have heard of Media Tech, right?
they make cell phone CPUs.
TV CPUs too.
And TV, CB, they make
CPUs that are in a lot of
1.8 gigahertz quad cores, right?
Something like that.
Well, it depends entirely
on which one of their
11 million.
But anyways, yeah, sorry, keep going.
They're the, like, they make this,
Nvidia doesn't make the CPU part of this thing.
It's two, it's actually two, like,
kind of things on a package
where the, the Nvidia side,
is the part with
the GPU stuff and then the CPU side
is made by Media Tech
and it's got
just the standard
the CPUs that they make right
so it's it's a I think it's 10 cores
of the X
of the X line and 10 cores of the
A line so that they have like
there's like 10 performance cores and 10
like more efficiency
kind of cores that's
that's what it is so it's
not a mystery like you can just go look like
like like a
I said, they made it sound like this was this big grand unveiling of stuff. It's all stuff.
We already know what it is. You can already look up the benchmarks. It's not as good, for example,
as a Qualcomm, like the Qualcomm Windows on Arm CPUs. Those CPUs are like better for performance,
I believe, than these are. So on the CPU side of things, if you think of one of those Qualcomm
notebooks, that's probably like best case CPU performance for, at least for like single-threaded,
like the things that you would normally be looking at for CPU there, maybe
maybe if you're doing like a big multi-threaded
Lur Club maybe it would change.
But so it's, again, it's a really kind of a nothing burger
from the standpoint of like hardware announcement.
There just wasn't much to look at here,
which is good because Prime only gave us about 15 minutes to look it up.
Casey.
What was what?
They said something about a thousand games.
Like what does that mean?
You know what I'm talking about?
I'm looking at it.
Over a 1,000 RTS enhanced and accelerated games and apps are available.
they probably just mean that they tested them and they run because remember they all have to run through a compat layer because nobody ships arm compiled binaries so they have to run through a compatibility layer that makes the x64 code run on arm right um so they probably just mean that that's or like it's got the lSS support or something or it's got some of the other like AI capability enhancement things maybe I don't know but yeah I
I don't think there's anything special about that
because again, like, you know,
that's just what you would expect
from a Windows PC.
Like, what you expect from Windows PC is all the games.
All the games run.
Exactly.
So it's not much of, again, not much of announcement,
but I don't know what they mean there.
I didn't see that part of the announcement.
Okay, well, I'm watching a little video of it going,
and this one looks like you're playing on a PlayStation 1,
and this one looks fantastic,
so I don't really know what's exactly happening.
Oh, is that that new, um,
I don't know what's going to do for DLSS5.
or whatever where they just,
where they just replace the graphics with different graphics that are like better.
Looks like Minecraft on the left.
Yes.
Yeah.
So this thing is just like slowly zooming in.
Yeah,
I'm not,
I'm not really sure what this comparison is because I don't think I've played a game that's
looks like this.
I'm actually replaying Dark Souls 1 right now and it doesn't look this bad.
So I don't really know what's going on here.
I don't really get this example,
but whatever,
I guess I'm not a,
I'm not in the no on whatever that is.
You're a boomer.
I'm a boomer.
Okay,
All right.
Well, I think that's about it, you know?
Appreciate you guys.
Great work, everyone.
I think we really,
I think we really knocked it out of the park this time.
I think we did.
Honestly, the floating point operation thing
was super interesting,
talking about the local agents.
I think is actually not only interesting,
it's actually important.
I think it's really good.
I'm hoping that more people realize
that they don't need to be beholden
to these big companies.
That would be great if the industry
go in that way,
because that's much better than the current thing.
Yeah.
I'm against current thing.
I am against the current thing and for the new current thing.
Me?
I support current thing.
Oh, dang.
I actually hate you is what that means, TJ.
What is the current thing?
I don't even know.
Trash, that's actually,
smashes red button.
Yeah, red button trash.
That is red button behavior right there, trash.
Are we talking about the police thing again?
No, I don't know what the boys thing is, but we're not talking about it was our grand first blue button.
I thought that's what we were talking about.
Yeah, that has nothing to do with police, but I really appreciate you.
What the mini-fig situation?
Dude, the mini-fig situation is insane.
Should we talk about that next week?
What's mini-fig?
Fender guitars?
Oh, my gosh.
Can we do a slideshow as the blocker for next week and show Casey the Lego situation?
Brick by, break.
Break by, break.
What is this?
Okay, how about this one?
Trash and Casey, if you see anything about Lego news,
great point.
Just navigate away.
I will prepare a roughly 10 slide presentation.
Okay.
About it.
By the way, TJ, there's been some...
There's a lot of pictures required.
There's a lot of pictures, so I'm going to have to figure out how to, like, make it so that
the blocker isn't the entire episode.
I think it could be the entire episode.
It might actually have to be, honestly.
Just do it.
Casey, you will like it.
It will be shot.
Just do it, Prime.
Just do it.
Lego Pokemon out there.
Yeah, I know I'm very interesting.
Hey, dude, I got a lot of Lego over here.
But it's a very interesting story.
But, TJ, there's a lot of lawyers and other lawyer tubes coming out saying that the inverse party is completely incorrect.
And that minifigs is the correct one.
And so there's a lot of drama going on.
And I've been trying to keep up with all of it.
So I'll be glad to do a whole, like.
I have no idea what you're talking about.
Dude, you just.
What's a mini fig?
Oh my gosh
Again trash
Do not even Google it
If you Google minifig
You'll be led into the entire situation
So you've got to stay away from it
Is that your wife's name?
No,
Alexa
Your wife comes in
Please could you play
Yes sir
Play Despacito
And she starts playing it on the end
Imagine someone just walks into my frame
And just does whatever
That'd be insane
That would be good.
All right.
Okay, we'll do that next week.
The receipt app is working.
Pitch deck, I don't think, honestly, TJ, I don't think we can pitch deck and do mini-figs.
The mini-figs situation is next time.
Mini-figs is next time.
That's what we're doing.
That's great.
Dude, I've been holding my peeve for like 45 minutes.
All right.
All right, that's it.
Bye, everybody.
That's the end of the episode.
Bye, everybody.
Put up the day.
Vibe coding.
Errors on my screen.
Terminal coffee
And him
Living the dream
