Limitless: An AI Podcast - Are Cheaper AI Models Better than Claude and ChatGPT?
Episode Date: July 16, 2026We're discussing new AI model releases from xAI, Meta, OpenAI, and Anthropic, and the shift toward cheaper, more efficient models.Focusing on Grok 4.5 and Meta’s MuseSpark 1.1, there's also... a broader move toward model routing and enterprise use.------🌌 LIMITLESS HQ ⬇️EMAIL US: info@limitless.fmNEWSLETTER: https://limitlessft.substack.com/FOLLOW ON X: https://x.com/LimitlessFTSPOTIFY: https://open.spotify.com/show/5oV29YUL8AzzwXkxEXlRMQAPPLE: https://podcasts.apple.com/us/podcast/limitless-podcast/id1813210890RSS FEED: https://limitlessft.substack.com/------TIMESTAMPS0:00 AI Price Competition2:48 Grok 4.5 Breakdown7:50 Meta Enters The Race15:21 Agents Change Everything17:08 Comparing Model Prices21:24 Cheaper Tokens, More Usage23:30 Profitability Still Matters24:34 A Multi-Model Future------RESOURCESJosh: https://x.com/JoshKaleEjaaz: https://x.com/cryptopunk7213------Not financial or tax advice. See our investment disclosures here:https://www.bankless.com/disclosuresJosh works with Anthropic as a contractor. All views expressed are his own and do not represent Anthropic, its leadership, or its affiliates. Nothing in this episode is investment advice.
Transcript
Discussion (0)
Just last week, in the span of about 48 hours, three of the most powerful AI labs on the planet
all shipped brand new models. Elon shipped Grok 4.5, Open AI took 5.6 global, and meta, for the very
first time in history, put a price tag on its own frontier model. And here's where it gets interesting.
For the last five years or so, the deal in AI was that every time a model shipped, it got smarter
and cheaper at the same time. But that kind of died this year with Anthropics Fable 5 and these new
Frontier launches like GPT 5.6. So today, we're asking the question that probably everyone should.
When Frontier Intelligence costs less than a cup of coffee, when they cost just a few pennies per
millions of tokens, what does that look like in terms of your costs and how much you use these
models? I mean, this is a totally different paradigm now. We have a very clear separation
between Fable and 5.6 Sol and GROC 4.5 and Meta's new model, MewSpark 1.1. And it seems
like these models are kind of diverging in a way that's really interesting, particularly centered around
price. For the last couple of years, the ultimate validation of whether your AI model is good or not
is how intelligent it is, how smart it is, and no one really cared about cost. And then Fable 5 kind of
came on the scene, and I think it was like $20 or $30 input and like $80 output. And companies
that were spending tens to hundreds of millions of dollars, something to think, is this right? Does
this make sense? Do I need the smartest model to do every single task?
And so we hear the likes of XAI, Elon's company.
We look at the likes of meta.
They release these models and we compare it to Fable Fibble Fibor, like they're not actually
that smart.
But that's the whole purpose of the models that they're releasing.
They're not trying to be as smart as Fable.
In fact, they're betting on the opposite.
They're betting that the cheapest model per unit intelligence is the model that will
ultimately fill the middle ground.
That will be the ultimate model that is embedded in every single enterprise and used by
every single user because the truth is you don't need the most intelligent model.
to do your task. Maybe if it is 90 to 95% of the intelligence of the smartest model ever,
but it costs one-tenth of the price. It's a no-brainer that you're using these models at scale.
And like you said, Josh, over the last week, there have been two particular companies,
Meta, who we've known and spoken about a lot on this show, who have spent upwards of, I think,
$35 billion to try and build the world's best model, came out with their new MetaMew Spark 1.1,
and then you had SpaceX AI who recently IPOed,
and there's a lot of pressure writing on them,
release their new GROC 4.5 models.
Now, are they as good as FABEL? No,
but are they cheap enough to use at scale
when you're using or spinning up agents
or when you're trying to figure out that long, complex task
that's going to take tens of hours
and you don't want to burn very expensive FABEL tokens,
these models might be the ones to choose,
and I think it's probably worth covering a bunch of them.
Yeah, there's a new meta almost in town
where it's like a new model doesn't necessarily mean
higher intelligence. It could just mean higher efficiency. And I think that's where we're going to start
with with the GROC 4.5 release, because this is kind of like an opus class claim, but at a third of the
price, which is a pretty big deal. So this new GROC 4.5 model, it's built on XAIs or I guess SpaceX AIs,
their version 9 foundation model, which is about 1.5 trillion parameters. And for reference,
the version we've been using all year, if you've used GROC at all this year, that is the version
8 small, which is about 500 billion parameters. So we're looking at about a three times multiple in
parameter count, which generally speaking is three times better, but probably a little bit more.
It seems like this is a serious increase relative to what we've been using.
And Elon has called this Opus Class model.
But instead of just being this highly expensive, very slow model, it's much more quick and
it's much more efficient.
When it comes to how many tokens you're able to generate for that same dollar, and this is
very clearly the route that you can see SpaceX AI has been trying to go for a long time.
They're very hardcore engineers.
They love the engineering challenge.
and what they're trying to do now is figure out how you can kind of sculpt these GPUs that they're training on to get as efficient as possible.
We spoke a few weeks ago, he does about the etched guys, like the startup who is building their own training architecture chip stack, and they're basically building their own servers.
And within that, they're pretty big to make one specific thing work, which is the transformer.
And we talked about how GPUs are not very efficient.
They don't actually use, like, sometimes up to 60% of the GPU isn't used.
what GROC and the SpaceX AI team are doing
is they are taking that code base
and they're really getting down to the bare metal
to figure out how to squeeze the most juice out of it.
And that's what this model is.
That's what 4.5 is.
It has $2 in, $6 out per million tokens generated.
And it seems like it's incredibly efficient.
When comparing it to other models like Opus,
it appears as if it's up to four point times fewer tokens needed
in order to reach the same task.
So that's like an adjusted multiple of what.
It says 17 times less than Opus.
That's like a really big deal for a model as it relates to cost, at least.
The major unlock that you have with GROC 4.5 is it's a model meant for building.
So if you're a hobbyist out there, if you're a software engineer and you want to tinker
with some of these models and try and build something, but you know it's costing inexpensive.
You don't want to use the Fable 5 API.
This is a really, really good model to use because it's so efficient.
So some of the stats you are referencing right there, Josh, basically it uses four times fewer
tokens than Claude Opus 4.8. What that results in is a 17x less price or cost to do the same
task. And the way that they've been able to achieve this is arguably the funnest or coolest part.
They spin up a bunch of different agents to solve different problems of the tasks that you've
asked it to at the same time in parallel. And this is a growing trend you'll see with the other
models that we're going to talk about on this show today, where agentic coding or agentic reasoning,
or using agents to solve your problem
has been a major unlock
for these cheaper models
to be cheap in the first place.
The other thing I want to mention
about GROC 4.5 in particular
because it's very unique to them
and SpaceX AI is
they're just acquired a company called Curse
for a low price of $60 billion.
And the major unlock that cursor gave them
is they have all this data
around how users use coding models.
Now, it's very specific.
It's not what they're coding,
it is how they use the models.
And this data was very important in training GROC 4.5
to become smarter at routing people's requests.
So let's say you ask it to build an app
that can change your wardrobe into something way better.
And you send it like pictures of your camera,
or whatever that might be.
We talked about this as an example on yesterday's episode.
It'll be able to know which model to use
at what time for how long
and calculate the costs preemptively to make sure that it's not burning your wallet.
It's a really, really smart model.
And Elon himself has said, listen, this isn't going to be the smartest model.
We're working on better models, but it is a really good daily driver.
It's a really good workhorse.
And if you're anyone at a company that doesn't want to burn, you know, $30 in and $80
out on Fabo 5 or whatever the cost is, something that's like 10x more, you should use this model
for 80 to 90% of the work that you're trying to do and then use Fable to kind of orchestrate
the plan or the design or whatever that might be. The final point I'm going to make in SpaceX AI's
favor, because you might be thinking, okay, fine, whatever, but like when's the next model going to come?
They took, like, whatever, nine months to release this upgrade. Elon is cooking up three more
models that are in order of magnitude larger than the model that we're talking about today.
So right now, in under a month, we're expecting to see GROC 5, and GROC 5.5 is also being cooked
at the same time as well. These are like five to 10 trillion parameter models. So it's
feasible to say that we'll have a Mito's class model from SpaceX AI in a couple of months' time,
which is also.
Yeah, this is the part that's most exciting to me is like, it's very obvious that SpaceX AI
is the best at the engineering part of generating tokens.
They have the whole vertical integration now.
They have the ability to build out these data centers.
They have the data centers running.
And now they have the harness through cursor.
And they also have the data set that they've used through cursor in order to kind of achieve
this fully integrated stack.
And what's funny now is if they are continuing to release models that are bigger and
bigger. If they ever do run up against a compute wall, remember, they have that deal with
Anthropic and with Google, I believe, and they're going to have to figure out who's going to get
the short end of that stick. Because it doesn't seem like they're going to have GPUs for everyone.
But like you mentioned, the exciting thing here is that like this is version one of, it feels like
SpaceX AI 2.0. This is their second try at getting to the frontier. The first time they may have
gotten there for a couple of days, but it wasn't very long lived. Now we're getting new models every single
month and the goal around August is a two trillion parameter model. And then the goal for GROC
5, which is coming hopefully not too long after, is six to 10 trillion parameter models. Like this is
going to put them right up at the frontier. And if they're able to serve these tokens at a fraction of
the cost that say GPT6 is going to be or mythos six or whatever comes next, that's going to be a pretty
serious competitor in the AI space because they're going to be right at that frontier with a very
low-cost model that I think a lot of people are going to find a lot of use for. Now, that is the
GROC update. There is a second update that is just as noteworthy and probably even more noteworthy,
actually, because this is the first time that meta is charging for tokens. Meta, famously,
they have been the open source kings. They have always wanted to publish the models open source
to move the needle forward to kind of make this open source developer community thrive.
unfortunately, if you are a participant of that, your time has ended because now meta is closed source
and they are releasing these closed source models. They are charging via the API. It's not a lot of money,
but it was big enough news for Mark Zuckerberg to come back on X after a, what was it,
a three or four year hiatus and actually announce Muse Spark 1.1, which he describes as a strong agentic
and coding model at a very low price. It's available through our new meta model API and in the
meta AI. So, EJES, the question I have for you, because, you know, I haven't,
been the most excited about meta recently. Their offerings have left a lot to be desired as it
relates to AI. Is this a serious model? Like, is this worth actually paying for relative to all
the other models that exist today? Short answer is yes. And this comes from a professional meta hater.
Which is shocking. Like, this is a novel breakthrough. This is an exciting announcement for meta.
Well, I'm just happy to see something competitive enough for a lot of people to use in the AI market.
And I'll tell you why they would use this model.
Number one, the reason right at the top is it is the cheapest model per unit cost of intelligence.
So what do I mean by that?
You can have cheap models, but they're pretty crappy and you won't end up actually using it for serious work.
This is a model that you'll end up using for serious work, and it costs 25% of the price of the frontier model.
Now, can it do everything a frontier model can do, like what Fable 5 does?
No.
And Mark Zuckerberg openly admits that.
but it is an absolute workhorse.
You can throw it at a task and it can work for hours,
or you could throw it at multiple problems at once,
and it can figure it out.
But there's a few other advantages that this model in particular has.
So aside from the price, which is, by the way,
$1.25 in per million tokens and $4.25 out.
Pretty cheap.
If you want a comparison as to how crazy cheap,
that is, that is cheaper than GLM's model,
GLM 5.2, which is the leading open source Chinese model right now,
and they're known for being the cheapest model.
So the fact that Mark Zuckerberg, there's some blissful irony there, actually,
with the open source thing and the fact that, you know,
he was competing with China and they beat him,
to come back and offer the cheapest model is pretty amazing.
But the second thing is, like GROC 4.5,
it uses agents very intelligently.
So it has this thing called a master agent in this model.
And the master agent reads your prompt,
and it thinks very diligently about a plan
to answer and execute your prompt.
Now, that sounds very vague, right?
You're like, aren't the other models doing it?
No.
When you look at like Fable 5,
when you look at GPD 5.6,
it processes its entire model weights,
which is Gargantuan, by the way,
and that ends up being very costly.
Meta found a sneaky little way to circumvent this,
which is have a master agent,
have it plan, and then have it delegate
to a bunch of different agents.
The third thing that is very impressive about this model
is that it's an Omni model.
So it can take in video,
it can take in images,
It can take in text all at once and understand how to use that intelligently.
So if you give it a video, you can extract that video.
It knows that you want to list it or use it for a post that you have on Facebook Marketplace,
and it'll be able to do that in one shop.
And then the fourth and final thing that it's very good at is computer use.
So this thing can take over your computer, take over your account or whatever that might be,
and just know what you wanted to do.
Now, version one is very fine-tuned to Metas products, Instagram, WhatsApp, Facebook itself.
But the idea is you can pretty much use this for any computer use work going forward.
And I'm looking forward to version 1.2 and 1.3.
They pulled this off, by the way, within a month of releasing Mew Spark 1.
And they also released Muse image and video in the same week,
which it tells me that the cycle of release, similar to SpaceX AI,
is getting much quicker.
And if you ask me why that is, it's because both Elon and Zuck have one thing
that the two frontier labs opening and Anthropic don't have.
A crap ton of compute.
They have so much compute.
They are the most aggressive hyperscalators
and arguably the most successful hyperscalers.
They have amassed the most GPUs.
And if you still believe that compute scaling laws matter,
Zuck and Elon are not out of the race.
In fact, this proves that they're back into it.
One of the things that I found noeworthy of this
is that this API presumably runs on meta's own silicon.
It's the chips that they have been making and producing,
and likely running. In fact, there's a report that they're spending $250 billion,
including chips over the next year in order to build out something like 14 gigawatts of
data center capability. And that is mostly going to be running their new MTIA 400 chips,
which are basically the in-house meta-silicon, which is 400% faster than the previous generation
and uses 51% more HBM. So for the HBM folks, when we talk about all our investing videos,
these chips are going to be using a lot more of it.
And meta is now charging a quarter of the rival's prices,
which is a really kind of interesting and noteworthy thing.
And that combined with the switch from open source to closed source,
it kind of implies that.
Like, Mark Zuckerberg very clearly thinks the model is now the product
instead of the moat around the model.
So I think traditional meta would have been,
no, no, no, we're going to release the model.
We're going to build a moat around it,
and that is going to be the product we monetize.
This new shift implies like, no, actually the model is the product,
and we're going to integrate it into all of our services.
But in order to do that, we're going to do this huge, tremendous data center build out,
and they're going to build their own proprietary chips.
And it seems like both of those things are actually going well.
And I have to ask, it's like, okay, what happens if they actually do it?
Like, what happens if there is 14 gigawatts of compute running next year?
Like, that seems pretty considerable, particularly running on their own silicon,
which we know has tremendous competitive advantages when it comes to training your own models.
Well, here's the bet that like we're making,
we've made on this show multiple times on previous episodes,
is that the future of AI isn't you tapping a bunch of buttons and approvals every single second.
It is you write a prompt, you send the prompt,
and then the AI just kind of knows what to do.
It autonomously works.
If you believe in that world, then you believe in a world of AI agents.
And if you believe in a world of AI agents,
these agents are going to be making tens to hundreds of thousands of tool calls per month.
You don't want to be there clicking a proof the entire time, and those tool calls are pretty expensive.
So you want to go the cheapest route when you're using an agent to get the work done.
That's basically the entire thesis for why cheaper models are more effective.
And therefore, the labs who have the most compute and can use it most effectively, you mentioned, you know, their custom chips.
I believe Elon and SpaceX Air are also working on their own chips.
Open Air is doing their own with Halapeno.
This is a growing trend.
It makes sense that the people that have the most compute
and the best chip architecture will end up winning.
And that's basically the bet that Meta and Elon are going after.
I have to say, like, if we ground ourselves a second, right,
and look at the counter thesis,
I do think Zuck is heavily subsidizing this model.
I don't think there's any chance in hell
that it actually costs $125 and $425.
output. I think he's subsidizing this massively. And I think he needs to prove a point to his
investors or shareholders that it is worth the AI cap expense that he's probably going to
announce at the end of Q2. So that's my bet. I think it's strategic. I think it's the right move.
But I don't think he's unlocked some kind of major architecture redesign just yet.
Regardless, these are two really solid models between GROC 4.5 and NUMU Spark 1.1. And to kind of
place them in a spot relative to others, we can go down the price list of,
other models and kind of compare what they're like. So at the top of this list is Claude Fable 5,
and that's $10 in, $50 out per million tokens. GPT 5.6 comes next at just a little bit,
close to half, at $5 in, $30 out. Then Opus is $5 in, $25 out. And then it kind of goes down the
line until we get to Grock 4.5, which is $2 in, $6 out. Then we have Muse Spark,
125 and 425.
And at the very bottom, believe it or not, is Open AI with a dollar in and six dollars out for
GPT 5.6 Luna.
So the middle section and the upper section is kind of where the war is.
It's like we have those two models at the top.
We have Claude and we have GPT 5.6 soul.
Those are very much competing on the intelligence curve.
But then below that where we have this like kind of cluster of you can think Gemini and
the smaller GPT models.
and GROC and meta 1.1, that's where we are seeing this like price competition. And it's funny to see
this kind of divergence and strategies. And I think that's a good way of looking at the frontier when
you evaluate these new models is, okay, is this a frontier intelligence model or is this a frontier
price model? And those two things now are very different because they're going to be used for a very
different set of use cases. And I think one of the more interesting applications that we're going to be
following on the show is how people route through these models to do different tasks. Like you said,
some agentic tasks don't require necessarily the highest frontier intelligence,
maybe you can get away with using a Muse Spark 1.1 for a lot of that,
and perhaps using a Fable 5 for orchestration of those agents, things like that.
So we're going to see what I suspect as a new meta start to come into play of this
orchestration at a high level and using different models for more particular things,
more specific things.
Josh, have you heard of something called the Silicon Token Expenditure Index?
it's basically this index which tracks the spending of companies or enterprises.
It is, I think the Bloomberg ticker is SDLMTK, but the point of this index is it captures how much money is being spent by the top Fortune 500 on AI specifically.
So if you're listening to us talking about cheap models and, you know, talking about this thesis of cheaper models will be used more and you don't believe it, well, you have.
to look at the customers who are actually ingesting this AI and the movements and actions that
they're taking. And if I show you the index, you'll notice a particular trend over the last
couple of weeks and months, which is, it's down 20%, which basically means people are spending
much less on tokens or per token, and they're also spending much less overall. What that indicates
is they are looking for cheaper alternatives. And we're seeing this from the headlines that we've
seen from Uber, from META themselves, from Microsoft, who are now adopting Chinese open source
models into their product, into their co-pilot product. We are seeing this shift of enterprises
realizing that it's not about using the smartest model for every single task. It's about
finding that middle ground. It's about finding the right model or maybe the right types of models
to use for the middle ground of tasks. And that brings me to another point, which is,
I don't think it's necessarily just going to be one cheap model that dominates everything.
I think they're going to realize after they've shifted to a cheaper model that they could use
certain models for specific tasks. And that takes us down the path of routing, which is what
cursor has infamously figured out and what SpaceX figured out and acquired them for $60 billion.
So there's these really interesting conversations around this trend, but it's being validated by this
index that companies are going to be looking for cheaper models. And the irony is,
is what is the paradox, Josh, that we spoke about early on in the AI thesis where like the
cheaper something gets, the more you end up spending on it? What is the name of that paradox?
Jevins paradox. Jevins paradox. That's it. So the cheaper something gets, I expect to see way more
tokens being spent because the output will actually be worth it instead of spending a very expensive
prompt and getting, you know, maybe a mediocre answer from it. Yeah, and this is in line with
expectations. Goldman, they published this prediction.
that token consumption would multiply 24 times between 2026 and 2030.
That's like a tremendous amount.
That's 120 quadrillion tokens per month, which is unbelievable.
So even if the price of these tokens does go down some extent, a multiplier of that much is still increased spend, which I think is important to note.
It's like, I don't think we're going to see a fable class model be 1.24th the price very shortly,
but they're expecting that much increased demand.
And every single kind of,
every single prevailing wind
is pushing towards longer turn agentic coding sessions
where I feel like rarely do people actually now,
or people who are using it for productive uses at least,
are using it just as a chat box.
They're using it as an agent to do longer and longer and longer term tasks.
And what we notice with these cheaper models in particular
is that they actually do oftentimes consume more tokens
to get the same output at a higher quality.
I know we were mentioning this with 5.6,
the other day, where some of their lower models, they actually, they cost less, but they do use
a considerable amount more tokens to get to the answer. So we're in this weird crossroads, where we have
these, like, forcing functions that are pushing people to want to generate more tokens. Companies
want these tokens to be highly intelligent because they don't want subpar work if they're paying
for it. But then they have this like enterprise belt tightening that's kind of happening where
like Amazon killed. It's like, I think they had their internal token leaderboard and they totally
removed it. Coinbase and Walmart are setting up usage caps. A lot of companies are
kind of like crunching down on usage of AI internally because they can't really figure out
how to justify the value that's been given to the company. So there's a lot of these weird,
it very much feels like right at crossroads now and there's this new paradigm shift that is
happening as it relates to cost in particular. And it's going to be really interesting to see how
this plays out. There is one metric that I will be tracking to see whether these cheap models
are actually good enough. And it's simple. Are these companies,
making money. When do they turn profitable? With meta, I'm almost certain that they're subsidizing
their cost of their model. With SpaceX AI, I don't know maybe they're doing something similar,
but until these guys post a positive revenue earnings on their quarterly earnings or whatever that
might be, I'm not going to be convinced that cheap models are the way forward. I think it will
direction you play out, but that's going to be a big metric to track. Now, if you compare that to
the likes of Anthropic. They're not a public company yet. They're rumored to IPO maybe later this
year, but there are rumors that they have turned profitable in Q2. And they will, if true or if
confirmed, they'll be the first AI lab to do so. And the major way that they've been able to do
that is not only creating amazing models that are used by every single enterprise in the world,
but because it's expensive, because they pay for the cost of that you're getting. And, you know, at the
end of the day, if you're making a lot of money and people continue using it, maybe that is the
right way to go through. So that, that would be the only counterpoint to this. The final thing I'll say
on this topic, I think, is what I was alluding to earlier, which is I'm now convinced that, number one,
it's not just going to be open AI and anthropic ruling the entire world. I think it's going to be
multiple labs. And I think that's ultimately a very good future, right? Maybe five, six, maybe 10 labs,
right? And I think there are going to be so many different models to use for so many different
purposes. I mean, look at Metamuse Spark 1.1. It topped the health benchmark. Did you know that?
Like a random social media company's model topped the health benchmark. GROC 4.5, really good at
computer use and spinning up agents. Luna from Open Air also really good at computer use.
So all these different models will be good for various different tasks. And if you don't want to
spend time thinking about which subscription to get or what model to use or when. I think
routing companies, so what cursor basically built, which is you can type in a prompt and they
decide which parts of the prompt gets used by which kind of models will be the killer platform
or product. I don't know of any company that is building this that hasn't been acquired,
aka cursor. Maybe open router. We've had the founder, Alex Atala, on our show before. Maybe they
pivot into some kind of a routing product, but I'm convinced now that it's going to be a
layer that sits on top of these models potentially. Yeah, I think the way I'm looking at it is like,
okay, now there's two markets. Like this used to be a singular commodity, a singular race to the top.
Now there's these two things. And like cost per intelligence is the right lens for looking at maybe 80%
of the work that's become like kind of routine. But it is the opposite and exactly the wrong
lens for like the 20% where capability and reliability and verification matter. It's like if you
are doing really complicated work where mistakes cost you a tremendous amount of money in time.
If you are doing frontier math or problem solving or you're building complicated things
as a business, there's almost no limit to the amount that you'll spend in order to accomplish
your goals. Because I'm sure many of these companies have a tremendous amount of ideas that
they want to implement to the market and the constraint is their workforce, that knowledge-based
that could actually deploy that. There is no limit to the amount that they will spend on these
Fable models, on these GPT 5.6 and GPD 6.0 models. But for the rest of the world,
Maybe that's not the case. They are price sensitive. They don't need the frontier intelligence for everything. And that's where a lot of these models are going to really play a serious role. And I think that is cool that like we no longer have intelligence as the singular commodity. It is now split between intelligence and price. Those two things both matter. And I think that's kind of where we are in the market right now. It's like we are at this crossroads. The commodity has diverged into two. And we have some pretty serious players in the pricing game who have a very serious trajectory of actually continuing to be serious threats.
SpaceX AI is building a huge amount of data centers.
They have a very clear trajectory to a 20 trillion parameter model.
Meta is going to put 15 gigawatts on the ground, powered on,
and they're going to do that vertically integrated with their own silicon.
And that's a pretty serious threat to the pricing war,
because you have to imagine they're going to get some amazing efficiency from that.
So, yeah, I think that's the update for the state of low-cost models.
It's been very interesting week for that.
As we're recording this episode, I'm realizing that we need to make another episode on these
trends that we've spoken about during this talk.
Primary ones being, okay, if the most expensive model lab or the most intelligent model lab
isn't necessarily going to win over the next couple of years, and it's going to be this
middle ground, this cheap middle ground, which companies are there?
Where are the bottlenecks?
How are people using these things?
What will agents look like when they're working on long horizon tasks?
And most importantly, which companies, which customers are going to be using this thing?
I think, I mean, you let me know, as you as the audience that are listening to this right now, would that be an interesting episode to sort of unpack and dig into?
All of that said, I think that cheaper models are here to stay, and I think they're here to stay in a very big way.
And I think they're going to come out, thankfully, from some frontier Western labs.
We're shifted away from building the most expensive and intelligent model that is being kind of blocked to anyone and everyone to a $1.25 in, $4.25 out model from meta, which,
surprisingly works and is intuitive.
I'm also excited to see what people do with these cheap models.
I don't think they're necessarily going to build the same things
that people using Fable 5 are going to do,
simply because they're not prohibited by cost anymore.
And I think that kind of unlocks a new line of creativity
when you're building these apps.
You're like, oh, you know what?
Maybe I will try this crazy thing because it's certainly going to cost me $10.
Heck, I'll just do it.
So if you're listening to this and you're inspired by some of the things
that we're spoken about, go and try some of these models.
super cheap to use, and tell us what you end up building with it.
Tell us if there's anything creative that you haven't seen being built by Fable 5.
We would love to see it.
Demo it, heck, maybe even show it on an episode in the future.
But if you enjoyed this, please, please,
share this with your friends.
If you're not subscribed, please subscribe.
If you haven't left us a comment, we always want to hear from you.
We love hearing from you guys.
And if you're listening to us on Spotify or Apple Music,
please, please give us a rating.
It helps us out pretty massively.
Josh, any final thoughts?
No, that's it. Thank you guys so much for watching. On the sponsor front, we're chatting with some people. This has been going well. So for the people who have reached out, thank you. We are working our way through those. We are slowly working towards becoming a self-sufficient entity, which has been amazing. So thank you so much for the support on that. Again, if you know anyone, please refer them, send them our way, either on X, email and description, whatever it may be. But that is another episode. Next coming up is the roundup, which is, I think, our favorite episode of the week. We just throw everything that we haven't been able to talk about into one and do a full recap on everything that has happened this week, which is a lot.
We got to talk about the data center bands, dude.
Like, New York is my city.
Why are we banning the day?
Oh, you can have a lot of conversation about that one.
So stay tuned.
Stay tuned for that grieving session.
But anyways, that is the episode today.
Thank you all so much for watching, as always.
And we will see you in the next one.
