Limitless: An AI Podcast - Are Cheaper AI Models Better than Claude and ChatGPT?

Episode Date: July 16, 2026

We're discussing new AI model releases from xAI, Meta, OpenAI, and Anthropic, and the shift toward cheaper, more efficient models.Focusing on Grok 4.5 and Meta’s MuseSpark 1.1, there's also... a broader move toward model routing and enterprise use.------🌌 LIMITLESS HQ ⬇️EMAIL US:           info@limitless.fmNEWSLETTER:    https://limitlessft.substack.com/FOLLOW ON X:   https://x.com/LimitlessFTSPOTIFY:             https://open.spotify.com/show/5oV29YUL8AzzwXkxEXlRMQAPPLE:                 https://podcasts.apple.com/us/podcast/limitless-podcast/id1813210890RSS FEED:           https://limitlessft.substack.com/------TIMESTAMPS0:00 AI Price Competition2:48 Grok 4.5 Breakdown7:50 Meta Enters The Race15:21 Agents Change Everything17:08 Comparing Model Prices21:24 Cheaper Tokens, More Usage23:30 Profitability Still Matters24:34 A Multi-Model Future------RESOURCESJosh: https://x.com/JoshKaleEjaaz: https://x.com/cryptopunk7213------Not financial or tax advice. See our investment disclosures here:https://www.bankless.com/disclosures⁠Josh works with Anthropic as a contractor. All views expressed are his own and do not represent Anthropic, its leadership, or its affiliates. Nothing in this episode is investment advice.

Transcript
Discussion (0)
Starting point is 00:00:00 Just last week, in the span of about 48 hours, three of the most powerful AI labs on the planet all shipped brand new models. Elon shipped Grok 4.5, Open AI took 5.6 global, and meta, for the very first time in history, put a price tag on its own frontier model. And here's where it gets interesting. For the last five years or so, the deal in AI was that every time a model shipped, it got smarter and cheaper at the same time. But that kind of died this year with Anthropics Fable 5 and these new Frontier launches like GPT 5.6. So today, we're asking the question that probably everyone should. When Frontier Intelligence costs less than a cup of coffee, when they cost just a few pennies per millions of tokens, what does that look like in terms of your costs and how much you use these
Starting point is 00:00:41 models? I mean, this is a totally different paradigm now. We have a very clear separation between Fable and 5.6 Sol and GROC 4.5 and Meta's new model, MewSpark 1.1. And it seems like these models are kind of diverging in a way that's really interesting, particularly centered around price. For the last couple of years, the ultimate validation of whether your AI model is good or not is how intelligent it is, how smart it is, and no one really cared about cost. And then Fable 5 kind of came on the scene, and I think it was like $20 or $30 input and like $80 output. And companies that were spending tens to hundreds of millions of dollars, something to think, is this right? Does this make sense? Do I need the smartest model to do every single task?
Starting point is 00:01:25 And so we hear the likes of XAI, Elon's company. We look at the likes of meta. They release these models and we compare it to Fable Fibble Fibor, like they're not actually that smart. But that's the whole purpose of the models that they're releasing. They're not trying to be as smart as Fable. In fact, they're betting on the opposite. They're betting that the cheapest model per unit intelligence is the model that will
Starting point is 00:01:47 ultimately fill the middle ground. That will be the ultimate model that is embedded in every single enterprise and used by every single user because the truth is you don't need the most intelligent model. to do your task. Maybe if it is 90 to 95% of the intelligence of the smartest model ever, but it costs one-tenth of the price. It's a no-brainer that you're using these models at scale. And like you said, Josh, over the last week, there have been two particular companies, Meta, who we've known and spoken about a lot on this show, who have spent upwards of, I think, $35 billion to try and build the world's best model, came out with their new MetaMew Spark 1.1,
Starting point is 00:02:23 and then you had SpaceX AI who recently IPOed, and there's a lot of pressure writing on them, release their new GROC 4.5 models. Now, are they as good as FABEL? No, but are they cheap enough to use at scale when you're using or spinning up agents or when you're trying to figure out that long, complex task that's going to take tens of hours
Starting point is 00:02:41 and you don't want to burn very expensive FABEL tokens, these models might be the ones to choose, and I think it's probably worth covering a bunch of them. Yeah, there's a new meta almost in town where it's like a new model doesn't necessarily mean higher intelligence. It could just mean higher efficiency. And I think that's where we're going to start with with the GROC 4.5 release, because this is kind of like an opus class claim, but at a third of the price, which is a pretty big deal. So this new GROC 4.5 model, it's built on XAIs or I guess SpaceX AIs,
Starting point is 00:03:09 their version 9 foundation model, which is about 1.5 trillion parameters. And for reference, the version we've been using all year, if you've used GROC at all this year, that is the version 8 small, which is about 500 billion parameters. So we're looking at about a three times multiple in parameter count, which generally speaking is three times better, but probably a little bit more. It seems like this is a serious increase relative to what we've been using. And Elon has called this Opus Class model. But instead of just being this highly expensive, very slow model, it's much more quick and it's much more efficient.
Starting point is 00:03:40 When it comes to how many tokens you're able to generate for that same dollar, and this is very clearly the route that you can see SpaceX AI has been trying to go for a long time. They're very hardcore engineers. They love the engineering challenge. and what they're trying to do now is figure out how you can kind of sculpt these GPUs that they're training on to get as efficient as possible. We spoke a few weeks ago, he does about the etched guys, like the startup who is building their own training architecture chip stack, and they're basically building their own servers. And within that, they're pretty big to make one specific thing work, which is the transformer. And we talked about how GPUs are not very efficient.
Starting point is 00:04:16 They don't actually use, like, sometimes up to 60% of the GPU isn't used. what GROC and the SpaceX AI team are doing is they are taking that code base and they're really getting down to the bare metal to figure out how to squeeze the most juice out of it. And that's what this model is. That's what 4.5 is. It has $2 in, $6 out per million tokens generated.
Starting point is 00:04:36 And it seems like it's incredibly efficient. When comparing it to other models like Opus, it appears as if it's up to four point times fewer tokens needed in order to reach the same task. So that's like an adjusted multiple of what. It says 17 times less than Opus. That's like a really big deal for a model as it relates to cost, at least. The major unlock that you have with GROC 4.5 is it's a model meant for building.
Starting point is 00:05:00 So if you're a hobbyist out there, if you're a software engineer and you want to tinker with some of these models and try and build something, but you know it's costing inexpensive. You don't want to use the Fable 5 API. This is a really, really good model to use because it's so efficient. So some of the stats you are referencing right there, Josh, basically it uses four times fewer tokens than Claude Opus 4.8. What that results in is a 17x less price or cost to do the same task. And the way that they've been able to achieve this is arguably the funnest or coolest part. They spin up a bunch of different agents to solve different problems of the tasks that you've
Starting point is 00:05:37 asked it to at the same time in parallel. And this is a growing trend you'll see with the other models that we're going to talk about on this show today, where agentic coding or agentic reasoning, or using agents to solve your problem has been a major unlock for these cheaper models to be cheap in the first place. The other thing I want to mention about GROC 4.5 in particular
Starting point is 00:05:57 because it's very unique to them and SpaceX AI is they're just acquired a company called Curse for a low price of $60 billion. And the major unlock that cursor gave them is they have all this data around how users use coding models. Now, it's very specific.
Starting point is 00:06:14 It's not what they're coding, it is how they use the models. And this data was very important in training GROC 4.5 to become smarter at routing people's requests. So let's say you ask it to build an app that can change your wardrobe into something way better. And you send it like pictures of your camera, or whatever that might be.
Starting point is 00:06:35 We talked about this as an example on yesterday's episode. It'll be able to know which model to use at what time for how long and calculate the costs preemptively to make sure that it's not burning your wallet. It's a really, really smart model. And Elon himself has said, listen, this isn't going to be the smartest model. We're working on better models, but it is a really good daily driver. It's a really good workhorse.
Starting point is 00:06:59 And if you're anyone at a company that doesn't want to burn, you know, $30 in and $80 out on Fabo 5 or whatever the cost is, something that's like 10x more, you should use this model for 80 to 90% of the work that you're trying to do and then use Fable to kind of orchestrate the plan or the design or whatever that might be. The final point I'm going to make in SpaceX AI's favor, because you might be thinking, okay, fine, whatever, but like when's the next model going to come? They took, like, whatever, nine months to release this upgrade. Elon is cooking up three more models that are in order of magnitude larger than the model that we're talking about today. So right now, in under a month, we're expecting to see GROC 5, and GROC 5.5 is also being cooked
Starting point is 00:07:39 at the same time as well. These are like five to 10 trillion parameter models. So it's feasible to say that we'll have a Mito's class model from SpaceX AI in a couple of months' time, which is also. Yeah, this is the part that's most exciting to me is like, it's very obvious that SpaceX AI is the best at the engineering part of generating tokens. They have the whole vertical integration now. They have the ability to build out these data centers. They have the data centers running.
Starting point is 00:08:02 And now they have the harness through cursor. And they also have the data set that they've used through cursor in order to kind of achieve this fully integrated stack. And what's funny now is if they are continuing to release models that are bigger and bigger. If they ever do run up against a compute wall, remember, they have that deal with Anthropic and with Google, I believe, and they're going to have to figure out who's going to get the short end of that stick. Because it doesn't seem like they're going to have GPUs for everyone. But like you mentioned, the exciting thing here is that like this is version one of, it feels like
Starting point is 00:08:31 SpaceX AI 2.0. This is their second try at getting to the frontier. The first time they may have gotten there for a couple of days, but it wasn't very long lived. Now we're getting new models every single month and the goal around August is a two trillion parameter model. And then the goal for GROC 5, which is coming hopefully not too long after, is six to 10 trillion parameter models. Like this is going to put them right up at the frontier. And if they're able to serve these tokens at a fraction of the cost that say GPT6 is going to be or mythos six or whatever comes next, that's going to be a pretty serious competitor in the AI space because they're going to be right at that frontier with a very low-cost model that I think a lot of people are going to find a lot of use for. Now, that is the
Starting point is 00:09:13 GROC update. There is a second update that is just as noteworthy and probably even more noteworthy, actually, because this is the first time that meta is charging for tokens. Meta, famously, they have been the open source kings. They have always wanted to publish the models open source to move the needle forward to kind of make this open source developer community thrive. unfortunately, if you are a participant of that, your time has ended because now meta is closed source and they are releasing these closed source models. They are charging via the API. It's not a lot of money, but it was big enough news for Mark Zuckerberg to come back on X after a, what was it, a three or four year hiatus and actually announce Muse Spark 1.1, which he describes as a strong agentic
Starting point is 00:09:54 and coding model at a very low price. It's available through our new meta model API and in the meta AI. So, EJES, the question I have for you, because, you know, I haven't, been the most excited about meta recently. Their offerings have left a lot to be desired as it relates to AI. Is this a serious model? Like, is this worth actually paying for relative to all the other models that exist today? Short answer is yes. And this comes from a professional meta hater. Which is shocking. Like, this is a novel breakthrough. This is an exciting announcement for meta. Well, I'm just happy to see something competitive enough for a lot of people to use in the AI market. And I'll tell you why they would use this model.
Starting point is 00:10:33 Number one, the reason right at the top is it is the cheapest model per unit cost of intelligence. So what do I mean by that? You can have cheap models, but they're pretty crappy and you won't end up actually using it for serious work. This is a model that you'll end up using for serious work, and it costs 25% of the price of the frontier model. Now, can it do everything a frontier model can do, like what Fable 5 does? No. And Mark Zuckerberg openly admits that. but it is an absolute workhorse.
Starting point is 00:11:05 You can throw it at a task and it can work for hours, or you could throw it at multiple problems at once, and it can figure it out. But there's a few other advantages that this model in particular has. So aside from the price, which is, by the way, $1.25 in per million tokens and $4.25 out. Pretty cheap. If you want a comparison as to how crazy cheap,
Starting point is 00:11:25 that is, that is cheaper than GLM's model, GLM 5.2, which is the leading open source Chinese model right now, and they're known for being the cheapest model. So the fact that Mark Zuckerberg, there's some blissful irony there, actually, with the open source thing and the fact that, you know, he was competing with China and they beat him, to come back and offer the cheapest model is pretty amazing. But the second thing is, like GROC 4.5,
Starting point is 00:11:47 it uses agents very intelligently. So it has this thing called a master agent in this model. And the master agent reads your prompt, and it thinks very diligently about a plan to answer and execute your prompt. Now, that sounds very vague, right? You're like, aren't the other models doing it? No.
Starting point is 00:12:03 When you look at like Fable 5, when you look at GPD 5.6, it processes its entire model weights, which is Gargantuan, by the way, and that ends up being very costly. Meta found a sneaky little way to circumvent this, which is have a master agent, have it plan, and then have it delegate
Starting point is 00:12:19 to a bunch of different agents. The third thing that is very impressive about this model is that it's an Omni model. So it can take in video, it can take in images, It can take in text all at once and understand how to use that intelligently. So if you give it a video, you can extract that video. It knows that you want to list it or use it for a post that you have on Facebook Marketplace,
Starting point is 00:12:41 and it'll be able to do that in one shop. And then the fourth and final thing that it's very good at is computer use. So this thing can take over your computer, take over your account or whatever that might be, and just know what you wanted to do. Now, version one is very fine-tuned to Metas products, Instagram, WhatsApp, Facebook itself. But the idea is you can pretty much use this for any computer use work going forward. And I'm looking forward to version 1.2 and 1.3. They pulled this off, by the way, within a month of releasing Mew Spark 1.
Starting point is 00:13:11 And they also released Muse image and video in the same week, which it tells me that the cycle of release, similar to SpaceX AI, is getting much quicker. And if you ask me why that is, it's because both Elon and Zuck have one thing that the two frontier labs opening and Anthropic don't have. A crap ton of compute. They have so much compute. They are the most aggressive hyperscalators
Starting point is 00:13:37 and arguably the most successful hyperscalers. They have amassed the most GPUs. And if you still believe that compute scaling laws matter, Zuck and Elon are not out of the race. In fact, this proves that they're back into it. One of the things that I found noeworthy of this is that this API presumably runs on meta's own silicon. It's the chips that they have been making and producing,
Starting point is 00:13:55 and likely running. In fact, there's a report that they're spending $250 billion, including chips over the next year in order to build out something like 14 gigawatts of data center capability. And that is mostly going to be running their new MTIA 400 chips, which are basically the in-house meta-silicon, which is 400% faster than the previous generation and uses 51% more HBM. So for the HBM folks, when we talk about all our investing videos, these chips are going to be using a lot more of it. And meta is now charging a quarter of the rival's prices, which is a really kind of interesting and noteworthy thing.
Starting point is 00:14:33 And that combined with the switch from open source to closed source, it kind of implies that. Like, Mark Zuckerberg very clearly thinks the model is now the product instead of the moat around the model. So I think traditional meta would have been, no, no, no, we're going to release the model. We're going to build a moat around it, and that is going to be the product we monetize.
Starting point is 00:14:51 This new shift implies like, no, actually the model is the product, and we're going to integrate it into all of our services. But in order to do that, we're going to do this huge, tremendous data center build out, and they're going to build their own proprietary chips. And it seems like both of those things are actually going well. And I have to ask, it's like, okay, what happens if they actually do it? Like, what happens if there is 14 gigawatts of compute running next year? Like, that seems pretty considerable, particularly running on their own silicon,
Starting point is 00:15:17 which we know has tremendous competitive advantages when it comes to training your own models. Well, here's the bet that like we're making, we've made on this show multiple times on previous episodes, is that the future of AI isn't you tapping a bunch of buttons and approvals every single second. It is you write a prompt, you send the prompt, and then the AI just kind of knows what to do. It autonomously works. If you believe in that world, then you believe in a world of AI agents.
Starting point is 00:15:46 And if you believe in a world of AI agents, these agents are going to be making tens to hundreds of thousands of tool calls per month. You don't want to be there clicking a proof the entire time, and those tool calls are pretty expensive. So you want to go the cheapest route when you're using an agent to get the work done. That's basically the entire thesis for why cheaper models are more effective. And therefore, the labs who have the most compute and can use it most effectively, you mentioned, you know, their custom chips. I believe Elon and SpaceX Air are also working on their own chips. Open Air is doing their own with Halapeno.
Starting point is 00:16:20 This is a growing trend. It makes sense that the people that have the most compute and the best chip architecture will end up winning. And that's basically the bet that Meta and Elon are going after. I have to say, like, if we ground ourselves a second, right, and look at the counter thesis, I do think Zuck is heavily subsidizing this model. I don't think there's any chance in hell
Starting point is 00:16:42 that it actually costs $125 and $425. output. I think he's subsidizing this massively. And I think he needs to prove a point to his investors or shareholders that it is worth the AI cap expense that he's probably going to announce at the end of Q2. So that's my bet. I think it's strategic. I think it's the right move. But I don't think he's unlocked some kind of major architecture redesign just yet. Regardless, these are two really solid models between GROC 4.5 and NUMU Spark 1.1. And to kind of place them in a spot relative to others, we can go down the price list of, other models and kind of compare what they're like. So at the top of this list is Claude Fable 5,
Starting point is 00:17:20 and that's $10 in, $50 out per million tokens. GPT 5.6 comes next at just a little bit, close to half, at $5 in, $30 out. Then Opus is $5 in, $25 out. And then it kind of goes down the line until we get to Grock 4.5, which is $2 in, $6 out. Then we have Muse Spark, 125 and 425. And at the very bottom, believe it or not, is Open AI with a dollar in and six dollars out for GPT 5.6 Luna. So the middle section and the upper section is kind of where the war is. It's like we have those two models at the top.
Starting point is 00:17:58 We have Claude and we have GPT 5.6 soul. Those are very much competing on the intelligence curve. But then below that where we have this like kind of cluster of you can think Gemini and the smaller GPT models. and GROC and meta 1.1, that's where we are seeing this like price competition. And it's funny to see this kind of divergence and strategies. And I think that's a good way of looking at the frontier when you evaluate these new models is, okay, is this a frontier intelligence model or is this a frontier price model? And those two things now are very different because they're going to be used for a very
Starting point is 00:18:29 different set of use cases. And I think one of the more interesting applications that we're going to be following on the show is how people route through these models to do different tasks. Like you said, some agentic tasks don't require necessarily the highest frontier intelligence, maybe you can get away with using a Muse Spark 1.1 for a lot of that, and perhaps using a Fable 5 for orchestration of those agents, things like that. So we're going to see what I suspect as a new meta start to come into play of this orchestration at a high level and using different models for more particular things, more specific things.
Starting point is 00:19:00 Josh, have you heard of something called the Silicon Token Expenditure Index? it's basically this index which tracks the spending of companies or enterprises. It is, I think the Bloomberg ticker is SDLMTK, but the point of this index is it captures how much money is being spent by the top Fortune 500 on AI specifically. So if you're listening to us talking about cheap models and, you know, talking about this thesis of cheaper models will be used more and you don't believe it, well, you have. to look at the customers who are actually ingesting this AI and the movements and actions that they're taking. And if I show you the index, you'll notice a particular trend over the last couple of weeks and months, which is, it's down 20%, which basically means people are spending much less on tokens or per token, and they're also spending much less overall. What that indicates
Starting point is 00:19:58 is they are looking for cheaper alternatives. And we're seeing this from the headlines that we've seen from Uber, from META themselves, from Microsoft, who are now adopting Chinese open source models into their product, into their co-pilot product. We are seeing this shift of enterprises realizing that it's not about using the smartest model for every single task. It's about finding that middle ground. It's about finding the right model or maybe the right types of models to use for the middle ground of tasks. And that brings me to another point, which is, I don't think it's necessarily just going to be one cheap model that dominates everything. I think they're going to realize after they've shifted to a cheaper model that they could use
Starting point is 00:20:39 certain models for specific tasks. And that takes us down the path of routing, which is what cursor has infamously figured out and what SpaceX figured out and acquired them for $60 billion. So there's these really interesting conversations around this trend, but it's being validated by this index that companies are going to be looking for cheaper models. And the irony is, is what is the paradox, Josh, that we spoke about early on in the AI thesis where like the cheaper something gets, the more you end up spending on it? What is the name of that paradox? Jevins paradox. Jevins paradox. That's it. So the cheaper something gets, I expect to see way more tokens being spent because the output will actually be worth it instead of spending a very expensive
Starting point is 00:21:21 prompt and getting, you know, maybe a mediocre answer from it. Yeah, and this is in line with expectations. Goldman, they published this prediction. that token consumption would multiply 24 times between 2026 and 2030. That's like a tremendous amount. That's 120 quadrillion tokens per month, which is unbelievable. So even if the price of these tokens does go down some extent, a multiplier of that much is still increased spend, which I think is important to note. It's like, I don't think we're going to see a fable class model be 1.24th the price very shortly, but they're expecting that much increased demand.
Starting point is 00:21:59 And every single kind of, every single prevailing wind is pushing towards longer turn agentic coding sessions where I feel like rarely do people actually now, or people who are using it for productive uses at least, are using it just as a chat box. They're using it as an agent to do longer and longer and longer term tasks. And what we notice with these cheaper models in particular
Starting point is 00:22:18 is that they actually do oftentimes consume more tokens to get the same output at a higher quality. I know we were mentioning this with 5.6, the other day, where some of their lower models, they actually, they cost less, but they do use a considerable amount more tokens to get to the answer. So we're in this weird crossroads, where we have these, like, forcing functions that are pushing people to want to generate more tokens. Companies want these tokens to be highly intelligent because they don't want subpar work if they're paying for it. But then they have this like enterprise belt tightening that's kind of happening where
Starting point is 00:22:48 like Amazon killed. It's like, I think they had their internal token leaderboard and they totally removed it. Coinbase and Walmart are setting up usage caps. A lot of companies are kind of like crunching down on usage of AI internally because they can't really figure out how to justify the value that's been given to the company. So there's a lot of these weird, it very much feels like right at crossroads now and there's this new paradigm shift that is happening as it relates to cost in particular. And it's going to be really interesting to see how this plays out. There is one metric that I will be tracking to see whether these cheap models are actually good enough. And it's simple. Are these companies,
Starting point is 00:23:26 making money. When do they turn profitable? With meta, I'm almost certain that they're subsidizing their cost of their model. With SpaceX AI, I don't know maybe they're doing something similar, but until these guys post a positive revenue earnings on their quarterly earnings or whatever that might be, I'm not going to be convinced that cheap models are the way forward. I think it will direction you play out, but that's going to be a big metric to track. Now, if you compare that to the likes of Anthropic. They're not a public company yet. They're rumored to IPO maybe later this year, but there are rumors that they have turned profitable in Q2. And they will, if true or if confirmed, they'll be the first AI lab to do so. And the major way that they've been able to do
Starting point is 00:24:12 that is not only creating amazing models that are used by every single enterprise in the world, but because it's expensive, because they pay for the cost of that you're getting. And, you know, at the end of the day, if you're making a lot of money and people continue using it, maybe that is the right way to go through. So that, that would be the only counterpoint to this. The final thing I'll say on this topic, I think, is what I was alluding to earlier, which is I'm now convinced that, number one, it's not just going to be open AI and anthropic ruling the entire world. I think it's going to be multiple labs. And I think that's ultimately a very good future, right? Maybe five, six, maybe 10 labs, right? And I think there are going to be so many different models to use for so many different
Starting point is 00:24:55 purposes. I mean, look at Metamuse Spark 1.1. It topped the health benchmark. Did you know that? Like a random social media company's model topped the health benchmark. GROC 4.5, really good at computer use and spinning up agents. Luna from Open Air also really good at computer use. So all these different models will be good for various different tasks. And if you don't want to spend time thinking about which subscription to get or what model to use or when. I think routing companies, so what cursor basically built, which is you can type in a prompt and they decide which parts of the prompt gets used by which kind of models will be the killer platform or product. I don't know of any company that is building this that hasn't been acquired,
Starting point is 00:25:37 aka cursor. Maybe open router. We've had the founder, Alex Atala, on our show before. Maybe they pivot into some kind of a routing product, but I'm convinced now that it's going to be a layer that sits on top of these models potentially. Yeah, I think the way I'm looking at it is like, okay, now there's two markets. Like this used to be a singular commodity, a singular race to the top. Now there's these two things. And like cost per intelligence is the right lens for looking at maybe 80% of the work that's become like kind of routine. But it is the opposite and exactly the wrong lens for like the 20% where capability and reliability and verification matter. It's like if you are doing really complicated work where mistakes cost you a tremendous amount of money in time.
Starting point is 00:26:17 If you are doing frontier math or problem solving or you're building complicated things as a business, there's almost no limit to the amount that you'll spend in order to accomplish your goals. Because I'm sure many of these companies have a tremendous amount of ideas that they want to implement to the market and the constraint is their workforce, that knowledge-based that could actually deploy that. There is no limit to the amount that they will spend on these Fable models, on these GPT 5.6 and GPD 6.0 models. But for the rest of the world, Maybe that's not the case. They are price sensitive. They don't need the frontier intelligence for everything. And that's where a lot of these models are going to really play a serious role. And I think that is cool that like we no longer have intelligence as the singular commodity. It is now split between intelligence and price. Those two things both matter. And I think that's kind of where we are in the market right now. It's like we are at this crossroads. The commodity has diverged into two. And we have some pretty serious players in the pricing game who have a very serious trajectory of actually continuing to be serious threats. SpaceX AI is building a huge amount of data centers.
Starting point is 00:27:16 They have a very clear trajectory to a 20 trillion parameter model. Meta is going to put 15 gigawatts on the ground, powered on, and they're going to do that vertically integrated with their own silicon. And that's a pretty serious threat to the pricing war, because you have to imagine they're going to get some amazing efficiency from that. So, yeah, I think that's the update for the state of low-cost models. It's been very interesting week for that. As we're recording this episode, I'm realizing that we need to make another episode on these
Starting point is 00:27:43 trends that we've spoken about during this talk. Primary ones being, okay, if the most expensive model lab or the most intelligent model lab isn't necessarily going to win over the next couple of years, and it's going to be this middle ground, this cheap middle ground, which companies are there? Where are the bottlenecks? How are people using these things? What will agents look like when they're working on long horizon tasks? And most importantly, which companies, which customers are going to be using this thing?
Starting point is 00:28:11 I think, I mean, you let me know, as you as the audience that are listening to this right now, would that be an interesting episode to sort of unpack and dig into? All of that said, I think that cheaper models are here to stay, and I think they're here to stay in a very big way. And I think they're going to come out, thankfully, from some frontier Western labs. We're shifted away from building the most expensive and intelligent model that is being kind of blocked to anyone and everyone to a $1.25 in, $4.25 out model from meta, which, surprisingly works and is intuitive. I'm also excited to see what people do with these cheap models. I don't think they're necessarily going to build the same things that people using Fable 5 are going to do,
Starting point is 00:28:53 simply because they're not prohibited by cost anymore. And I think that kind of unlocks a new line of creativity when you're building these apps. You're like, oh, you know what? Maybe I will try this crazy thing because it's certainly going to cost me $10. Heck, I'll just do it. So if you're listening to this and you're inspired by some of the things that we're spoken about, go and try some of these models.
Starting point is 00:29:11 super cheap to use, and tell us what you end up building with it. Tell us if there's anything creative that you haven't seen being built by Fable 5. We would love to see it. Demo it, heck, maybe even show it on an episode in the future. But if you enjoyed this, please, please, share this with your friends. If you're not subscribed, please subscribe. If you haven't left us a comment, we always want to hear from you.
Starting point is 00:29:31 We love hearing from you guys. And if you're listening to us on Spotify or Apple Music, please, please give us a rating. It helps us out pretty massively. Josh, any final thoughts? No, that's it. Thank you guys so much for watching. On the sponsor front, we're chatting with some people. This has been going well. So for the people who have reached out, thank you. We are working our way through those. We are slowly working towards becoming a self-sufficient entity, which has been amazing. So thank you so much for the support on that. Again, if you know anyone, please refer them, send them our way, either on X, email and description, whatever it may be. But that is another episode. Next coming up is the roundup, which is, I think, our favorite episode of the week. We just throw everything that we haven't been able to talk about into one and do a full recap on everything that has happened this week, which is a lot. We got to talk about the data center bands, dude. Like, New York is my city.
Starting point is 00:30:14 Why are we banning the day? Oh, you can have a lot of conversation about that one. So stay tuned. Stay tuned for that grieving session. But anyways, that is the episode today. Thank you all so much for watching, as always. And we will see you in the next one.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.