Limitless: An AI Podcast - THIS WEEK IN AI: NVIDIA Dominates Google | America vs Open Source | Tesla Starlink V5
Episode Date: July 24, 2026🌌 LIMITLESS HQ ⬇️EMAIL US: lkwhit14@gmail.comNEWSLETTER: https://limitlessft.substack.com/FOLLOW ON X: https://x.com/LimitlessFTSPOTIFY: http...s://open.spotify.com/show/5oV29YUL8AzzwXkxEXlRMQAPPLE: https://podcasts.apple.com/us/podcast/limitless-podcast/id1813210890RSS FEED: https://limitlessft.substack.com/------In another jam-packed week in AI, we discuss NVIDIA’s Vera Rubin chips, Google’s latest disappointments, and the competition between US frontier labs and Chinese open-source models. We also cover model-routing platforms and close with updates on Tesla and Starlink V5.------TIMESTAMPS0:00 AI Costs Crashing Down0:36 Vera Rubin Lands2:39 NVIDIA’s Market Edge5:27 Google’s Disappointing Week8:12 Google’s Chip Race12:34 Open Source Under Fire16:58 Model Routers Rise19:28 Tesla’s Internet Network------RESOURCESJosh: https://x.com/JoshKaleEjaaz: https://x.com/cryptopunk7213------Not financial or tax advice. See our investment disclosures here:https://www.bankless.com/disclosuresJosh works with Anthropic as a contractor. All views expressed are his own and do not represent Anthropic, its leadership, or its affiliates. Nothing in this episode is investment advice.
Transcript
Discussion (0)
Invidia just announced the key to one of AI's biggest problems, cost.
Companies spend between tens to hundreds of millions of dollars every year and are running out of money.
Google, for the first time yesterday reported a negative cash flow on their quarterly earnings.
They're officially losing more money than they are making.
That's because they're spending so much on AI infrastructure.
Nvidia's new Vera Rubin officially got stood up yesterday and it saves you 10x on the tokens that you spend,
which means you can have a 10x better model for the same cost.
There's all of this and Google's new Gemini 3.6 models,
which is very, very underwhelming on today's roundup.
Yeah, so we have to start with the big news of the day,
which is that Vera Rubin is actually shipping.
And for those who are not familiar,
Jensen came on stage at Nvidia GTC and announced these a few months ago.
The news today is that they're finally live and they're actually operational,
and we started to get an idea of what these things look like.
And I want to preface this section with the idea that we just covered
an episode yesterday about how GPT6 or whatever the new internal open AI model was, they broke out
of internal air gap containment and hacked into a public-facing website into their production
database to steal secrets. That was on Blackwell chips. The new chips are Verra Rubin. No models have
been trained on Verar Rubin chips, but the math behind how much more powerful they are is so unbelievably
impressive. It's like, oh my God, this feels like the end game. It's like once these Verirubin chips come
online at scale, what on earth are these models going to look like if we already have Fable and
GPT6 class models, it's going to be pretty wild for some numbers. 10x, the efficiency, which is
crazy. So for every single megawatt you put into a GPU, it will give you 10 times the amount of tokens.
This is a huge unlock for a lot of models. The second thing is in terms of density of transistors,
there's $336 billion.
That is a 62% increase over the Blackwell GB300.
And this is on TSM's 3-nanometer technology,
which is basically the cutting edge.
There's 22 terabytes per second of memory
that is going through this whole thing.
And basically, they ran out of room for a single piece of silicon,
and they glued these two maxed out dies together and call it one.
So that's kind of where we are.
It's like this chip is going to be 10x more performant per watt,
and it's going to have a just unbelievable baseline relative to the GB300.
A lot of people are saying about four times the baseline.
So imagine what we get when these models are trained on not only four times the baseline,
but also efficiency improvements in terms of algorithms.
Like the next generation of software built on these things is going to be a monster.
And I want to translate what this means for the wider market.
And for Nvidia stock, which has pretty much just been flat for the last like five months,
I think this is Nvidia's star moment for this year.
Now, the reason is it's really costly to train and inference models these days.
And so any way that a company and enterprise that is spending 10 to hundreds of millions of dollars
every year can save money is a big deal.
Now, a big way that they can save money is infrastructure.
Right now, people are paying around 50 to 60% profit margin on every dollar spent on an AI token
to Nvidia.
They are like monopolies in this way.
So, Nvidia giving them a 10x efficiency increase means that you can create or run your Fable5 model,
your local open source model, Kimi K3, at one-tenth of the cost purely because of infrastructure.
Now, the way that this GPU works, you can't just run one Vero Rubin on its own.
You need to run 72 of them in one server rack, and they're serviced by 32 of Nvidia's CPUs.
Now, the reason why they have so many CPUs is because when you're running these AI models these days,
you're not really running one instance of an AI model.
You're running many.
It's called AI agents and they need to access different tools.
That's what the CPUs are for.
So it's this collective thing with the software and hardware integration,
the way it's set up, that makes Nvidia's GPUs so, so effective.
And so if I had to make a call here, what is their most direct competitor, Josh?
It has to be Google.
Google has tried to go for the Nvidia throne with their own custom-made GPUs called
TPUs, their tensor processing units,
and they had their quarterly earnings results yesterday,
and they made a significant chunk of money on TbUs,
but they also announced that they're going to be releasing a new chip,
and I just don't think that they can catch up to Nvidia.
So Nvidia is running away with it,
and this is a huge advancement for AI models in general.
Nvidia's got a hell of a lead,
and one of the headlines that I do want to finish the segment on
is the amount that they're actually able to produce of this,
because we have a lot of people that are interested in the investment angle.
We have an interesting investment angle here,
and the number is 1,000,
a thousand of these racks per day.
And each, as like you mentioned, each rack contains a lot of these chips inside of it.
And that means that they are projected to make $630 billion per quarter.
Assuming they can get this up at scale.
Throw that into your earnings per share forecast.
So yeah, whatever forecasting you're doing for her in video, like perhaps think higher.
That is a tremendous amount of money for a single skew.
Like this is one of the many products lines.
It's important to note also that Verra Rubin has CPUs.
And the CPUs have been shipping since.
last month. So there's a huge business being built on this new Varirium platform. I think everyone
trying to make like what, $20 billion this year, just from CPUs alone? A gazillion at this point.
I don't know. Nvidia, just like there is no end in sight to the amount of GPUs that they're
going to be able to sell to people. And this number I found pretty large. Now, you just mentioned
Google. And I feel like we do have to talk about Google because they just had their earnings report
as well as announcing a lot of things. And to preface the section, we were possibly the number one fanboys on
the internet of Google and Gemini and Deep Mind. And we were talking about them almost every day.
I don't think we found an episode on them in the last, I don't know, two, three months. It's been
kind of rough. So are things getting better? Okay, so Google had a big week of announcements.
Unfortunately, 100% of those announcements kind of sucked. So the first major announcement is they
released a series or rather their new flagship AI model. It's called Gemini 3.6 Flash. This is an
iteration on Gemini 3.5, and unfortunately, it's not that very good. Now, they advertise it as
not being frontier, but being cheap, affordable, and cost efficient. Now, if you're saying those
words, it better be worth the cost that I'm paying for it for the intelligence that I receive.
Fortunately, it's dumber than ChatGBTGPT 5.6, Luna, which is their worst model of their
frontier labs that they release. And it gets worse, Josh. It's worse than GLM-5, as well. It's worse.
5.2, not even 5.5.3, GLM 5.2 and Kimi K3, the open source models from China that have a fraction of the budget that Google, remember, Google is spending $200 billion this year. These Chinese AI labs are spending 100th of that, and they're able to pull off a better model and they're open sourcing it for everyone to run at a much more cost-effective way. So the question I have for Google is, one earth are you doing? These models are just terrible. Sorry.
Yeah, it's disappointing to see because, you know, Google was kind of crushing it.
They were winning across the board and they were working on these world models.
Like Google was the world model company and they had all of these individual pillars.
And now when we look at this chart, I mean, they're like, they're not even in the middle.
They're at the back of the attacks.
I think that's probably most embarrassing.
It's like, dude, you're getting beaten by meta.
Come on now.
And this seems to be the trend.
And I get upset with myself a little bit for like they fooled me once and then they fooled me again.
where Google has had this period of time where like, hey, they literally invented the transformer
architecture that every single AI language model runs on, but they weren't able to turn into a product.
Then the CEO, the co-founder, or the old co-founder, I guess, Sergey Brin, he comes back into the office,
he whips the company into shape, he gets them right back at the frontier.
And then we're kind of going to this fall off again, where perhaps like culturally at Google,
there is this deeply rooted inability to actually convert these ideas to compelling products
and to do so at the scale required to compete with frontier AI labs.
And I think that's the moment of time that we're at now
where Google very much feels like they are in crisis mode,
given the things that they have released so far.
And relative to the things in the market,
it's like, it's tough to find a good thing to say about what's gone down this week.
There is some promise and some hope in the sense that while I was reading
through everything from Google this week,
they have this new chip architecture that they're trying to builds,
and that seems like it could possibly give them an edge.
I mean, when we think about the GPU CPU architecture world, there is like general GPUs with
Nvidia, but then there's accelerators like the TPUs and like the custom A6 and like the etched
chips that have LLM architecture baked right in. So what is the deal with this new Google chip that
seems to be competing directly with that? Yeah. So it's called or code named Frozen V2 and it's
basically an iteration of their existing chip architecture that they're built, tensor processing units
or TPUs. Now, if you look at their quarterly earnings, which are about to in a second,
they made quite a decent chunk of money selling their TPUs to external customers. Now, this is a
big jump from the previous quarter because Google primarily built and used TPUs to train their
own Gemini models. What they've done, the change between Q1 and Q2, is they've started selling
it to other vendors, such as Anthropic and other AI labs, to train their models to create a new
revenue business line for them. Now, Frozen Chip V2 is meant to be six to eight times.
more efficient than TPUs.
And if you remember, the original pitch of TPUs
was more efficient training and inference
for your AI model.
So they seem to have revived a new chip architecture
and they want to build it.
Now, the bad news is this thing isn't coming online
until 2028.
Vera Rubin that we just discussed
came online yesterday.
And it has 10 times more efficiency
than the last predecessor of MVIDIA,
which the TPUs couldn't keep up with.
So my question to you is,
even if they achieve six to eight times more
efficiency in two and a half years time, who the hell cares? Because
Nvidia will have a better chip architecture by then that is massively more efficient. So
I don't see a way that like Google kind of like capitalizes or keeps up or catches up rather
with Nvidia on the chip side of things. Now just briefly on the on the model side of things,
I can't emphasize enough how bad it is that Google hasn't delivered a frontier model in the last
what is it, four months, right? They have all the data in the world.
world. Gmail, search, browser, Android, G-sweet, they can use that data to create an amazing
model. Sucks. Coding models where they should have diverted all their compute to absolutely failed.
What were they doing instead? Oh, let's just sell all our compute to other vendors.
Okay, but what if they weren't selling their compute to other vendors? What were they doing?
They were distributing it amongst teams. Teams literally had to fight internally to get compute.
DeepMind had to fight for the compute. Now, if you compare that to Anthropic, you compare that to OpenAI,
It is a completely different philosophy.
It's like, give the researchers the compute because they're going to build a better model,
and that model eventually down the line will bring every other business unit up.
And Google fumbled that back.
A lot of critics are going to respond to this and say, well, Google had an amazing revenue quarter.
Yeah, for now.
Like, what does that look like in like five months from now?
And there's a reason why Google stock is down 5% as we're recording this.
Because even though they had an amazing record core, no one cares.
Yeah, I think that's the idea is a lot of earnings now,
although they are making a lot of money, they are up, oh, I think revenue is up 24% year over year.
Yes, 24%.
They're still growing fairly quickly, but no one really cares about growth.
That is expected.
There is this AI golden era where if you are serving any sort of tokens, you are making money on them with a fairly large margin.
The thing people are interested in is that looking forward.
Like, what does the forward look?
What is their chance of capturing this new wave of AI token economics?
And their probabilities continue to go down.
And we see this in the stock price.
Like, look at this chart.
This is so ugly.
It's so bad.
really tough to see. And I think it's because, like, you read the numbers now. You're like,
okay, great, Google search totally isn't dead. It's actually doing much better than it ever was
before. I think, what, it's like up 17% on the year or something like that, which is really
great. But when you look forward, like, what is the odds that Google is at the frontier
12 months from now? I think that number is lower. And it's lower than it has been before. And that
is discouraging to investors. So I'm wishing Google all the best. I really am. I freaking
love this company. And I want them to win so bad. It's just, it's difficult to,
get excited when there's no tools that you actually want to use because everything else
that's on the market is so far superior. So moving on, a big debate this week, Josh, we'll cover this
in a few episodes this week, is the China versus USA, open source versus close source AI debate.
Now, for those of you who are unfamiliar, definitely go check out some of our previous episodes,
but basically China released Kimi K3 and soon to be a series of open source models that are
90% capable of the frontier models that the US have made.
So we're talking about Fable 5, GPT 5.6.
But at a fraction of the cost, and you can host it on your own server, which means that,
you know, it could be even cheaper.
And so the question all these companies are asking, similarly to, you know, the new Vera Rubin
update is, why should I pay all this money to access Fable or GPT5.6 when I can just spend a lot
less.
And yeah, maybe the Chinese models are a little slower, but I get to own my data and I need to,
and I basically save a ton of money.
So why would I do this?
And so the US has responded and said,
hmm, these Chinese labs, I think,
have distilled a lot of proprietary information
from US American labs.
And so we are now considering banning open source models.
Now, I just want to repeat that for a second.
The capitalist freedom country
is considering potentially banning open source models.
And meanwhile, the communist country is like,
hey, open source is a strategy going forward.
We had a comment from Liang Wen Feng from Deep Seek.
he said, we're going to remain open source for as long as that we can because that is what we believe in.
So it's just weird dichotomy. And if US does end up banning open source, it's going to mess up
a lot of young startups in the US who are currently relying on open source models to build
their products because they can't afford the expensive American frontier labs. Yeah, this doesn't make
any sense to me at all. I don't think, like I'm not even sure this is going to be possible.
It seems to me like this is posturing for probably setting up some sort of legislative plan
or some sort of form of action, it's not actually going to happen.
Because when you think about what it means to ban these Chinese models, this open source code,
I mean, ignore the idea that they're Chinese.
Just think of it as open source code.
This is just binary numbers in a page.
This is a near impossibility in this world to ban when everyone is actively seeking it.
I mean, you think about internal secrets at frontier AI labs and how quickly those leak out.
If this is out in the public, there's no way that people aren't going to clone it and run it locally.
and who's going to stop them from doing that?
It's like this is not really a practicality.
And when you think about, I mean, the country as a whole,
just from an ethical standpoint,
open source code very much is reminiscent of open and free speech.
And just like banning that seems like a slippery slope that is very problematic.
So I am hopeful and feeling fairly certain that this is posturing for something larger,
just kind of signaling an intention versus actually taking action on it.
because, I mean, I would hope that they are, you know, reasonable enough to recognize that
banning Chinese open source models is just never going to happen. I mean, if the code is online,
we are going to get it. We are going to use it. We are going to run it locally. The fact that it runs
locally is even more against it because it's easier to just take it on a computer and run an air
gap and no one can ever discover it. So I think it's, it seems like we'll see. We'll keep monitoring
this one and see like what the larger idea behind this is. But I don't think it's actually
abandoned because that just goes against so many ethical things that we've set up, so much of
like just practical the way this works. It's just, it doesn't make sense. I think what's happening
at the core of all of this is there's a big fear in the US that China is catching up. And
originally it was like China's three years behind, both in model development and chip architecture.
And you need both of those things. You need compute, but China has a lot of compute in order to
build frontier models. Recently, they've caught up a lot quicker than we expected. The models are getting
better. They're owning their own chip architecture in Xi Jinping's speech last week. There was also
a co-announcement from Huawei, which basically said, hey, we have this new chiprack server and
every single Chinese frontier lab is going to use it. And then followed by that, we have
Chippew, one of their biggest AI labs, announced that they've just completed the construction of a gigantic
data center that is only going to run Chinese-made chips. No Nvidia chips inside. Now, if you
rewind like three months ago, Invidia, Jensen really wanted to sell China chips because they will then
rely on American frontier chip architecture, and therefore we could kind of like regulate and
keep an eye on these things. That is no more. So I can understand why the Trump cabinet is getting
a bit worried, and I'm interested to see kind of what process they take going forwards. But
moving on, Josh, there is a trend that is developing this week, a weird one that not a lot of
people expected. Over the last three days, three different companies, three major companies,
announced that they're creating what's known as a model routing platform. Now, if you want to know
what that is, when you send a prompt, typically, you're using a chatchbtee subscription. So it goes to
chatGPT or using a Cloud subscription. It goes to Claude. A model routing platform takes your prompt
and decides which parts of your prompt to send to different models for different types of function.
And the reason why they do this is, one, to give you a better answer output than sending it to a
single bottle, and two, to save you a heck ton of money. That's what Cursa did for coding, and it's what
led to them getting acquired for $60 billion. And Cursor themselves is coming out with a new
platform called Cursor Router, which is a more generalized platform for any kind of LLM prompt.
The other two companies that also announced that they're building a similar product is
meta and RAP. Ramp is in the finance game. And META, as you know, looked at Open Router,
looked at Ramp and was like, I need to do this internally because I'm spending tens of billions
of dollars every single year to build a better A.m. Why don't I just create a cheaper,
effective routing platform? It's interesting to see. Yeah, the interesting thing. I mean,
cursor's claiming Frontier quality at 60% lower cost.
And if that is true, like, that seems like a pretty compelling argument for a lot of
labs to come and use this or for a lot of users to come in and use this.
Now, I find it interesting that we're seeing this from cursor, from ramp, from meta,
and we're not seeing this from OpenAI, from Anthropic, from Google,
because they very clearly have the suite of models that would imply that they have the ability
to do this.
I mean, we have, like, the sole lunataura from OpenAI and GPT, where it's, like, very clear
that they would benefit from a router.
So you have to imagine they're kind of working on this.
And I wonder if this is an instance,
kind of like the early internet era where like your feature or your company becomes a feature.
And if cursor router is not able to kind of sustain a moat of an audience,
what's stopping chat GPT and OpenAI from launching their own and doing this own router with
their own tokens? And as they get more efficient, they're able to route it more effectively.
And they have more data than any other company because they have all of the queries that
they can just train off of and imply this. So it's exciting to see this now.
And I think it's incredibly important for companies who are trying to save money to,
to use this type of technology.
But man, it's going to be tough to compete, I think.
It's like, well, you're kind of going like,
you're going up against the big guys here.
And like, maybe this works for now.
But I mean, you're one feature away from people just not really caring anymore.
I think so.
I think so.
Okay, I think to round up the docket today, Josh, something's going on with Tesla.
Oh, this is so sick.
Wait, yes.
Can we talk about Tesla first.
Yeah, what's going on?
Tell me what's going on.
Oh, my God.
Okay.
So first, Starlink V5.
If you've ever used Starlink before, this is.
is like a big satellite dish, sits on your roof,
can get you connection anywhere in the world.
If you've ever traveled to a country that isn't very well industrialized,
you've seen these on roofs everywhere.
It's how everyone gets their internet.
This new version, Starlink V5, is about 40% smaller.
It is significantly more efficient in terms of how much you can get per watt.
And more importantly, it is being integrated into every single cyber cab being produced
ever, which is unbelievable because now there is this super cluster of computers
that can run anywhere in the world that will be connected to.
to the internet. So if there's ever a problem, if the car ever needs help from anyone,
if it ever needs to share data back to the mainframe, it's able to do that using this
inter-rayed Starlink antenna. And we're entering this world between this and the Starlink mobile
satellites that are coming with Starlink V-3 on Starship, where you will never go anywhere on
planet earth and not have connection to the internet. And that to me seems like so cool and really
powerful. And when you think about how much compute is going on these cybercabs, training cluster,
I don't know, just saying.
Yeah.
So my biggest bull case for Tesla has got nothing to do with cars and honestly nothing to do with robots.
It's to do with this decentralized network that he's building that he can lean on to train models, to inference prompts and just to kind of like power this new AI economy.
That sounds like a lot of jargon.
It's because this is new.
This is net new and we haven't ever seen anything like this.
But Elon, who is the master of hyperscaling any kind of infrastructure target, he sets his eyes on, is the guy to do this.
think Tesla, it's the largest network of robots. If you want to argue, like, what a robot is
and where the robots are roaming around today in the world, Tesla is a prime example of that.
And so, I'm excited to see this happen. I'm so, so pumped. Sorry, I know this is a layman thing
to say, but to just have internet wherever I go. I'm tired of getting on a, on like a, what,
a delta flight and having crappy internet. Just allow me to stream, allow me to watch videos,
allow me to call my friends. That would be amazing. So I'm excited to see this come out.
But I believe that is the end of the docket for this week.
That's it. Yeah, if you're listening to this, you are all officially caught up. I've heard rumors that there's some potentially new features and models coming out. You'll hear about all of that next week, I think on Tuesday. But aside from that, if you're listening to us on YouTube, please subscribe. Please leave us a comment. Give us a rating. Turn on notifications. It helps us out massively. And I think that's it, Josh. Is there anything else?
That's everything. Thank you so much for joining us for another amazing week of AI Frontier Technology.
Next week, buckle up. The rumor mills be going crazy. I think by them we're going to have a lot of news,
so I'm excited to talk about it here on the show. Don't miss it. We'll see you guys next week. Have a great weekend.
See you guys.
