Limitless: An AI Podcast - Ox Alpha Revealed: China's GLM-5.3-Flash Gave Away Free AI
Episode Date: August 27, 2026We discuss the World Humanoid Olympics in China, where humanoid robots competed in human-style events and set new performance marks. We also cover autonomy, reward-based training, and the com...parison between China and the U.S. in robotics and AI.------🔒 Check Out Our Sponsor: LEDGER AGENT STACK 🔒https://developers.ledger.com/docs/ai-tools/overview/?utm_source=Audio&utm_medium=Podcasts&utm_campaign=Limitless------🌌 LIMITLESS HQ ⬇️NEWSLETTER: https://limitlessft.substack.com/FOLLOW ON X: https://x.com/LimitlessFTSPOTIFY: https://open.spotify.com/show/5oV29YUL8AzzwXkxEXlRMQAPPLE: https://podcasts.apple.com/us/podcast/limitless-podcast/id1813210890RSS FEED: https://limitlessft.substack.com/------TIMESTAMPS0:00 Mystery Model Appears1:30 Stealth Launch On OpenRouter4:31 Chasing The Model’s Origin6:09 Benchmarking Ox Alpha6:59 Flash Model Theory10:00 Blind Taste Test Buzz12:11 Why Continual Learning Matters14:58 Data, Subsidies, And Strategy18:47 China’s Copycat Advantage20:21 Mystery Solved For Now21:11 More Models Incoming------RESOURCESJosh: https://x.com/JoshKaleEjaaz: https://x.com/cryptopunk7213------Not financial or tax advice. See our investment disclosures here:https://www.bankless.com/disclosuresJosh works with Anthropic as a contractor. All views expressed are his own and do not represent Anthropic, its leadership, or its affiliates. Nothing in this episode is investment advice.
Transcript
Discussion (0)
There's a new top dog in town, and the question on everyone's minds is who is this?
There's a secret model codename zero X alpha that's been live since August 20th, the last six days,
and it is offering the seemingly unbelievable things to the public.
They're offering 100 trillion tokens of free usage to anyone who wants to go and get it.
They're offering a million token context window, and the benchmarks of this thing are pretty unbelievable.
There's also this weird thing going on in the background where seemingly every single day,
the outputs of the model are getting better.
There were tests on day one that were far inferior to the test that were currently on day six.
As of this morning, we think we have the answer to who this person is.
But before we talk about that, we have to discuss what is zero X alpha?
You guys, this has been the mystery of the week.
It seems like frontier level intelligence.
It's offering all of these unbelievable things, like 100 trillion tokens for free.
Do you know how much that would cost if you were to use like Fable or GBT's 5.6 salt?
I think it's like 15 to like 20 million bucks a day.
It's a lot of money a day.
Okay, so just with that figure alone,
who on Earth is giving away
$20 to $30 million on inference per day
in this economy, right?
It has to be a big frontier lab.
The second clue is, Josh, it's not zero X alpha.
It's ox alpha, which seems kind of weird, right?
Like, who's talking about the animal ox?
Until you realize that it's probably part of like
the Chinese kind of folklore and zodiac side of things,
which, again, might be a little hint as to where we go.
But anyway, let's rewind six days ago very quickly.
OpenRouter, which were covered on the show before, it's a platform that kind of like hosts a bunch of different models, allows you to pick and choose, revealed a stealth model.
Now, a stealth launch is where they launch the model.
You go to their website.
You can access and use the model for free, but they don't tell you the name of the model.
Now, they've done this with previous launches through Google Gemini's model.
They've done it through Anthropic.
They've done it through Open AI as well.
And they haven't had a stealth model launch in a while.
But now this is the first one, and it was in high demand.
So as you mentioned earlier, 100 trillion tokens, which means that you can effectively go there
and do all your sorts of coding works, very heavy AI usage tasks.
And you can kind of use that in a day, again, $20 to $30 million worth of inference costs.
It has a massive 1 million context window, which is very competitive with the top models that
we see from Fable 5 as well as GPT 5.6 Seoul.
And this is the most important part, it's a use.
model, which means that you can use not just text, but images, models, it's really good at rendering
all these different types of medium. Now, the final point, which I think is the most exciting one
that you referred to earlier, Josh, is this concept of continual learning. So people, as they
started to use this model, realized a very curious phenomenon. They realized that as they were using
this model, let's say on day one, day two and day three, the model got exponentially better at
doing the very same task that they fed it in day one. And so people started to be suspicious
that this is a model that doesn't just kind of take your input and spit back out an answer,
which is what a lot of models do right now, but it can continually self-learn and improve 24-7
to give you a better response. So the question on everyone's mind is, who created this model at all?
And when you look at the kind of usage or when you look at the kind of uptake on Open Router
itself, although no one knew, it became the most viral launch on Open Routier.
It became the new number one model and dethroned Deepseek's Flash, which launched about a couple weeks ago.
And Deep Seek has been kind of like the name brand on Open Routo for the longest time ever.
So just a really exciting model to see.
Yeah, it was not only the most popular model on Open Router immediately upon launch, but by a factor of two.
And it more than doubled Deep Seek's use.
Like this is hugely popular and very interesting because of how well it performs.
The benchmarks from it, the Deep SWE subset benchmark, it had an 80%.
pass rate, which for those aren't familiar, Fable 5 was at 65% and GLM 5.3 was at 62%.
Hold on, hold on.
So you're saying that it's crushing the METOS level models and the GPT 5.6 models at coding.
In one specific particular subset of coding benchmarks.
I definitely don't want to say it's better than the frontier models because upon further
investigation, we will come to find out it's not.
And it's pretty far from it.
But it's an incredible model for what it is.
And the reason I was comparing it to GLM 5.3 is because there's this thing when you're producing
models called the tokenizers, how tokens get generated. And it was just so happened to be identical
to GLM. So people were like, huh, this is interesting. Maybe we can go further down the rabbit
hole and see if there's more here. So you can test oftentimes whether a model is Chinese or not
based on asking it a series of questions that are not really allowed to be answered in China,
one of which is asking whether Taiwan is part of China.
And what's funny is if you're using like an American frontier lab,
they'll very clearly give you the answer to this question.
But Chinese models do not.
And what we realize is that the output of the answer to this question
was pretty much identical to the GLM model.
And so like, okay, it's got a similar tokenizer.
It's giving similar outputs, but we still don't really know.
Up until this morning in which we discovered that GLM is actually the new version of
Jipu. It is, or sorry, Ox Alpha is the new version of JPU. And it seems like this is going to be a new
GLM model that we have. And to that, when I heard it, it actually landed me even more suspicious
because over the course of this last week on X and all of its social media, everyone has been
kind of vague posting about who this is, why they're doing it, why they're offering so many tokens
and so much compute. And then the Google team started chiming in, which I thought was really bizarre.
A lot of developers on the Google DeepMind team who work on Gemini, they started vague posting about this model.
And as of this morning, Bloomberg released a post, and it seems like it is everything but confirmed.
The model actually is from Z.A.I. And Ox Alpha is now known. And I think that was like a really interesting way of coming to this conclusion.
And now that we know, I just have more questions. How on earth are they serving all this inference to all these people for free?
Like the cost must be tremendous. I mean, let's see how this model shapes up against all the other
models. So we have a chart over here which shows kind of like the average tokens output per task
on the DeepSuite benchmark that you mentioned. DeepSuite, for those of you don't know, is a good
test of a model's capability to actually code really well. And code is kind of like the premium
benchmark that a lot of companies and AI labs kind of like benchmark their model assets against.
And so Ox Alpha is actually quite high up there. It's up there in terms of reduced token
output, but still maintaining a very high intelligence score. Now, as you mentioned earlier on,
that's not actually quite what was revealed when people started playing around with this model.
It ended up achieving not really 80% that you mentioned earlier, but around 63%. And that's under,
like, real usage. So there's a suggestion that the Chinese lab, Cheebu, kind of like maybe
benchbacks this. But to just talk about the model and why it's so impressive, just in its own
merit, there's a few things to look at. Number one, the rumor has it that this model isn't a foundation
model. This is a flash model. So what that means is it's a smaller derivative of a much larger
model that Jipu is training or has already trained. So what does this mean? Let's take Mithos or let's
take GPD 5.6. These are rumored to be around 6 to 10 trillion parameter models. Now, when you
look at Jipu's flash model, if that's true, you're only looking at around like two to three trillion
parameters in size, which, you know, is a decent size, but it's still very small compared to these bigger
models. And then the fact that this model is not only likely going to be cheaper in its usage,
let's say maybe 30 to 50 cents per usage, so they gave away like 15 to 30 million dollars per day
of 100 trillion tokens, it is able to work much faster and they're going to open weight the entire
model. So, I mean, to answer your question, I don't really have an answer, Josh, as to how they've
been able to kind of create this, aside from the usual, or maybe they distilled a Western model,
or maybe they had a breakthrough in continual learning that we suggested.
earlier. But the point is, this is a very serious competitor, and we saw a kind of exchange between
Elon Musk and the founder of Jipu. Where about a month or two ago, Elon Musk was responding
into a thread as to when he believed that Chinese or open source models in general will become
mythos level, you know, have that cybersecurity risk. And he said, probably Q1 of next year of
27 and the founder G-Tung of Cheapu replied says it won't take that long and he goes on to say it
doesn't show it in this thread right here that it'll probably take two to three months and he was true to
this word we now have an ox alpha mythos level model that is from jipu that is flash that is smaller
that is cheaper that anyone and everyone can use and it's going to be open-weighted a few weeks from now
and that's probably how it was served it's just a low-cost model to serve and it's really funny watching the
reaction to this and kind of, I don't want to say overreacting, but getting very excited and
enthusiastic about this, because a lot of the posts about this were like, oh, it does have
continual learning and it is getting better every day and suspicions as to like, why are there
more tokens being available every single day? And we have this post from Jeffrey Emanuel,
who's actually, I guess, on the limitless show. And he just said, just thinking through this
free zero X alpha model or OX alpha model, how could it possibly make sense to give out this many
tokens for free? Maybe there's been a continuous learning breakthrough and the best way to pour more
gas on the fire is to source as many coding-related requests as possible. I don't know if that's right.
I think the kind of where I'm landing here is that this is a flash model. This is some sort of
distilled model that is meant to serve inference very cheaply. And what we're seeing,
instead of continual learning, it's just these new checkpoints that they're kind of releasing
every day. And we're able to track the checkpoints as they come into play. And they're just kind
of generally iterating over time. I'm not sure it's self-improving. We are going to find out more
on that soon. The interesting phenomenon that I have really enjoyed seeing over the last week is
the blind taste test that's been going on. This is known kind of like as the Pepsi challenge,
where you put like Pepsi and Coke against each other in a blind taste test and people don't know
what they're getting. And one of the things I found really interesting is how excited people got
about this model because of the mystery and the lore. And it really makes you question the value
of true frontier level intelligence when people were doing seemingly incredible things
with this open source free flash model.
And when you're comparing the two next to each other,
for some tasks, it's very clear that the frontier models are stronger.
But for many other tasks, that's not apparently obvious.
And if you are just sending a prompt into a text box
and it is this Ox Alpha model or it is like a mid-tier chat chattchipt or a cloud model,
the outputs are mostly the same.
And people didn't realize and people don't really care where the tokens come from.
So I find this to be an interesting phenomenon because rarely do we get
a true stealth model that is good.
Oftentimes it's connected to a brand and there's a very strong brand affinity.
So when you remove the brand affinity from the equation, what happens?
You kind of get this phenomenon here where it turns out people just want pretty smart tokens
and they don't actually care where it comes from.
It's interesting to see the continual improvement or whatever's happening here where there's
an example of a rocket ship taking off and it very clearly looks like day over day,
its abilities to render 3D graphics and understand physics and create this real world
emulation are significantly better.
So I'm really excited to learn more as to what's going on behind the scenes.
Like are these pre-prepared checkpoints?
Are they actually taking all the data they're collecting from these 100 trillion tokens
and using it to continuous to apply like RL on top of the model?
Like, I don't know.
These are things I'm very excited to find out over the next couple of days.
Yeah, I think in the future, no one's really going to care which model they use.
They just want to make sure that the output is what they expected it.
They want to have the most effective output for the cheapest possible cost that they can pay
for that output. And we're seeing, like, a lot of companies move to some of these Chinese
open source models, or just other models, cheaper models from other smaller American labs,
just so that they are able to save on costs and create a more fine-tuned model for this specific
bit of work. Now, I want to spend a bit of time to talk about if this model has achieved
continual learning, why that's important and why that puts China, if this is Jee Poo's actual
model, which they confirmed, in a very advantageous position, because it kind of shifts the
Game of Thrones around a bit.
Well, right now, you look at Anthropic, you look at Open AI, they have the world's leading
models, there have internal models that they haven't even released yet that are supposedly
even more intelligent than the ones that they have publicly available.
And they have all the compute in the world, and they're spending so much money trying to train
the best model.
So if you were to reason within yourself, maybe no one else can have the capability to catch
up in China, although they've been able to put out very good open source models, hasn't
been able to surpass the frontier.
one way that you can kind of cheat that dynamic is if you have a continual learning model.
Now, the best way to think about what this means is think of a really intelligent 15-year-old
or 18-year-old. When you teach it how to ride a bike on day one and then let's say you teach
it how to make pancakes day 50s, 50 days later, it'll still remember or know how to learn how
to ride a bike or how to ride a bike. Current models right now doesn't work like that. They lose
side of their memory, it's something called catastrophic memory loss or something like that,
and they're unable to kind of remember some of the smart intelligence stuff that they learned
earlier on. Continual learning solves for that. Now, you can imagine if a model is actually able to
learn, like a real human is, like a really smart kid, you could reason to say that that model
will be more valuable to you, especially if you could run it locally at home on your local
device on all your types of private data. So I can see two types of individuals looking at this kind of a model,
and being really incentivized to use it over a clod or over an OpenAI chat GPT type model.
And that one is enterprises who don't want to give all their data away to major American labs.
They would want to use an open weights model that can continue to learn.
And the second is any kind of respective retail user who doesn't want to give away any of their personal records to Anthropic and Open AI.
Let's say it's medical records or finance information.
Let's say they want to have an AI agent that handles their finances.
You know, you may have different reasoning behind whether you want to use a.
certain brand or a certain lab, but if you could run it privately at home, maybe you might want to
use something like this, but continual learning, just assume you have this model that learns 24-7,
even when you're asleep, it could probably catch up with the frontier models.
And I think it's really interesting to see.
I don't know if this is that.
I'm not convinced that it is, but if China has made a breakthrough, that's going to put them
in a very advantageous position.
Yeah, I also don't think it's that.
I find it hard to believe that this is what it looks like and that they've reached that
point so far.
I mean, that'd be an amazing breakthrough.
In the case they do, you're going to need to be prepared.
you're going to need to be safe. And Ledger, our sponsor, is here to keep you safe because if you are
building with AI agents and you're giving prompts to free open source models that are probably owned by
the Chinese, chances are you're going to want some safety and some protection. And the ledger agent
stack protects this. It gets this job done. It offers a series of tools that allow you to work
with agents in a three-step process where the agent proposes, the human approves, and then the ledger
signer enforces. So if you're building with security and you want to securely use these agents,
Ledger is for you. It's available on Claudecote, on Codex, on cursor, pretty much anywhere in which
you engage with LLMs, and is open source and available today. So you can find a link in the description
down in our show. Thank you, Ledger for sponsoring this episode. And then I guess we kind of have
to talk about that. Like, is it really free case? Because like, is free ever really free?
And I guess there's two possible answers to this question. One is yes. It's just being heavily
subsidized, perhaps by the Chinese government as a vampire attack to pull some demand from the
frontier webs into these Chinese models. The other,
is perhaps they're just using the data.
One of the big problems that we have been seeing
people complain about is data retention
and zero data retention policies
because people don't want to feel like
they are having their data harvested
and used to train other models.
And when you think about giving away these free models,
China, for a fact, is collecting all of the inputs
from the context window.
What's been happening is a lot of people are using this
on perhaps proprietary code bases.
People have been testing this model
in their own private projects.
They've been using it for their own personal work.
and a lot of data gets derived out of that. I mean, it's a million token context window in this
one chat box and you just paste all in and you let it go to town. So I'm sure the amounts of
data that's collected from offering all these tokens just through collecting all the inputs
has been tremendous. And I'm sure that gives them a unique advantage in a way that some frontier
labs might not because oftentimes they are criticized for their retention policies and people really
want to make sure that there's that zero data retention. China doesn't face the scrutiny,
in particular doing it in stealth kind of shields that for at least the first week while people
have used it. So they were just kind of freely and liberally using this thing, probably aware that they
were having their prompts used for future training, but not really caring too much. So I thought
that was also interesting. The other thing is that like the status of Jeepo as a company still isn't
very strong. They had, I think, like $100 million in revenue last year and they had a net loss of like
seven times that or something like that. It's currently priced at 750 times revenue. The company
itself clearly can't sustain this type of funding and this free giveaway forever. And I'm curious
to know what the type of business model eventually, if any, there is going to be in order to keep
this thing alive. Like, is it just going to be a government subsidy until they figured out?
The business model is its state back, dude. Yeah. All these Chinese labs are backed by the state.
So that's how they get their infinite amounts of capital to be able to kind of fund all these things.
And this is a strategy that China has been pursuing for three plus years now. They want to
catch up with the frontier Western models. They distill attack a bunch of these model labs so that they
can train similarly capable models. And then they flood the market with cheap intelligence,
which means that Anthropic can't charge the highest price that they can for the high quality
model that they have. Open AI can't do the same. So this drives model costs down. This is what
kind of meta and Mark Zuckerberg is also trying to do, but failing to do as effectively as the Chinese
labs, they're just very good at doing it. The second thing I'll say is, at least according to Open Router,
from their official statement.
They said that this stealth model provider,
which we now know to be GPU,
is retaining prompts and responses,
but explicitly state that they're not using it for training.
Now, this isn't verifiable in any kind of way.
Like, okay, China, you've never done anything nefarious before, I'm sure.
You're really going to give this way totally for free.
Like, no way.
Coming from the people that stole all the, I can go and generate a full episode of SpongeBob
on those models.
Exactly.
Exactly.
Well, this is another unique thing, right?
culturally in China, and this applies for every vector of tech that they're built in the past,
electric cars, mobile phones, etc. They don't really see copyright as a thing or IP infringement
as a thing. They see, okay, if you have a good product, I can copy that. And Chinese in particular
are very, very good at copying effectively. There was this great interview I watched with,
who was it? It was Travis Kalanick on David Semra's podcast. I don't know if you called that, right?
Absolutely legend, right? Twice. And he talks about his
story where, you know, how he brought Uber over to China. And he made it a success there,
but he said it was very tough. But there was also some really unique cultural things that he
needed to understand. And he realized that in America or in the West, at least, the innovation
level is super high. Coming up with net new zero to one ideas, super high. In China, less so,
but they are masters of copying. And he originally looked down at copying as like a really
kind of like menial thing. But he said that the way that they do it is master's.
masterful and artful in itself. And this, I feel like, is what, like, some of these Chinese
labs are really good at doing. It doesn't matter if they're earning $100 million. Their
IPOs are popping off anyway. They're never going to compare to the Western IPO market at all.
But the point is they want to flood the market and kind of like keep America at a cost level
where maybe they're spending too much money than they actually need to until they're able to figure
out some kind of a breakthrough, which might be continual money. Who knows? Maybe. We'll see. It's an
exciting story. And we're going to find out more. So this will be covered and followed up
in the Roundup tomorrow.
Stay tuned for that.
There's also a reference.
It's kind of cool each is the main character of the story is OpenRouter.
And we just covered an entire OpenRouter episode just last week.
So everyone who is unfamiliar with OpenRouter and why it is so valuable,
please go and check out that episode.
And that's mostly the state of the mystery as of now.
Ox Alpha has been solved, but there is still a lot of behind the scenes that is left to uncover.
We're hoping to get that information shortly.
When we do, we will share it with you right here on the
round up, hopefully this week, so stay tuned for that one. This has been kind of a crazy week.
I mean, it's China's always throwing, like, they're throwing wrenches in everyone's plans.
Yeah. First it was Deep Seek. Now they're doing this like kind of really fun stuff model.
Did you see Quinn as well? Yeah, we got a new Quinn model. We're going to talk about that tomorrow too.
Dude, they're just like spaying out these models. I know. And you know the craziest part about all of this
is I'm seeing rumors from trusted sources that a bunch of American frontier labs are all
about to like drop a few models. So like I'm just saying here like, you know, we're getting like
the best from the Chinese. We're getting the best from the West and there's just too many models to
use right now. But if you're listening to this and you're curious about this Ox Alpha model,
or if you don't believe anything that we're saying, if you don't think it's good enough,
go out and try it right now. You can use it. It's all free. On open router. Go check it out.
You don't, there is no promo code. Just sign up, make an account and you can kind of like
inference it and use it yourself. And I'm curious like, send us, exactly. Like use the free
tokens, please use the Chinese state funded money budget to kind of do your AI task.
Just don't give it your proprietary information.
Do not definitely don't do that because they're definitely using it to train the next model.
But like, let us know in the comments whether you actually like the model, whether it actually
is good at the task that you give it.
And if it magically improves the next day, because it's continual learning.
We don't know.
Like, let us know in the comments.
DM us.
We're available on all socials.
Yeah, anything else, Josh?
Yeah, no, just again, share this with your friends if they enjoyed it.
If you have something to say, please leave a comment.
We try to read all of them.
If you enjoy it on your favorite podcast player, you can leave a five-star review there.
If you're listening to this, you should be watching.
It's very fun visually.
You can watch this on YouTube, on Spotify.
And yeah, as always, thank you so much for watching.
Feel free to go get caught up on the other episodes.
And tomorrow, we got a banger roundup.
Like, it's just going to, there's so much stuff to talk about.
It is our favorite episode of the week.
We've been in like a lull lately where there has.
I mean, it's been exciting, but it hasn't been like, oh my God, this is crazy.
there's new frontiers forming.
And hopefully this is the first of many instances
in which we're going to rev back up that engine
and get a whole bunch of new models.
So I'm very excited to stay tuned for all of that
where we'll be covering it.
And yeah, we'll see you guys in the next episode.
