Limitless: An AI Podcast - Ox Alpha Revealed: China's GLM-5.3-Flash Gave Away Free AI

Episode Date: August 27, 2026

We discuss the World Humanoid Olympics in China, where humanoid robots competed in human-style events and set new performance marks. We also cover autonomy, reward-based training, and the com...parison between China and the U.S. in robotics and AI.------🔒 Check Out Our Sponsor: LEDGER AGENT STACK 🔒https://developers.ledger.com/docs/ai-tools/overview/?utm_source=Audio&utm_medium=Podcasts&utm_campaign=Limitless------🌌 LIMITLESS HQ ⬇️NEWSLETTER:    https://limitlessft.substack.com/FOLLOW ON X:   https://x.com/LimitlessFTSPOTIFY:             https://open.spotify.com/show/5oV29YUL8AzzwXkxEXlRMQAPPLE:                 https://podcasts.apple.com/us/podcast/limitless-podcast/id1813210890RSS FEED:           https://limitlessft.substack.com/------TIMESTAMPS0:00 Mystery Model Appears1:30 Stealth Launch On OpenRouter4:31 Chasing The Model’s Origin6:09 Benchmarking Ox Alpha6:59 Flash Model Theory10:00 Blind Taste Test Buzz12:11 Why Continual Learning Matters14:58 Data, Subsidies, And Strategy18:47 China’s Copycat Advantage20:21 Mystery Solved For Now21:11 More Models Incoming------RESOURCESJosh: https://x.com/JoshKaleEjaaz: https://x.com/cryptopunk7213------Not financial or tax advice. See our investment disclosures here:https://www.bankless.com/disclosures⁠Josh works with Anthropic as a contractor. All views expressed are his own and do not represent Anthropic, its leadership, or its affiliates. Nothing in this episode is investment advice.

Transcript
Discussion (0)
Starting point is 00:00:00 There's a new top dog in town, and the question on everyone's minds is who is this? There's a secret model codename zero X alpha that's been live since August 20th, the last six days, and it is offering the seemingly unbelievable things to the public. They're offering 100 trillion tokens of free usage to anyone who wants to go and get it. They're offering a million token context window, and the benchmarks of this thing are pretty unbelievable. There's also this weird thing going on in the background where seemingly every single day, the outputs of the model are getting better. There were tests on day one that were far inferior to the test that were currently on day six.
Starting point is 00:00:33 As of this morning, we think we have the answer to who this person is. But before we talk about that, we have to discuss what is zero X alpha? You guys, this has been the mystery of the week. It seems like frontier level intelligence. It's offering all of these unbelievable things, like 100 trillion tokens for free. Do you know how much that would cost if you were to use like Fable or GBT's 5.6 salt? I think it's like 15 to like 20 million bucks a day. It's a lot of money a day.
Starting point is 00:00:59 Okay, so just with that figure alone, who on Earth is giving away $20 to $30 million on inference per day in this economy, right? It has to be a big frontier lab. The second clue is, Josh, it's not zero X alpha. It's ox alpha, which seems kind of weird, right? Like, who's talking about the animal ox?
Starting point is 00:01:19 Until you realize that it's probably part of like the Chinese kind of folklore and zodiac side of things, which, again, might be a little hint as to where we go. But anyway, let's rewind six days ago very quickly. OpenRouter, which were covered on the show before, it's a platform that kind of like hosts a bunch of different models, allows you to pick and choose, revealed a stealth model. Now, a stealth launch is where they launch the model. You go to their website. You can access and use the model for free, but they don't tell you the name of the model.
Starting point is 00:01:47 Now, they've done this with previous launches through Google Gemini's model. They've done it through Anthropic. They've done it through Open AI as well. And they haven't had a stealth model launch in a while. But now this is the first one, and it was in high demand. So as you mentioned earlier, 100 trillion tokens, which means that you can effectively go there and do all your sorts of coding works, very heavy AI usage tasks. And you can kind of use that in a day, again, $20 to $30 million worth of inference costs.
Starting point is 00:02:14 It has a massive 1 million context window, which is very competitive with the top models that we see from Fable 5 as well as GPT 5.6 Seoul. And this is the most important part, it's a use. model, which means that you can use not just text, but images, models, it's really good at rendering all these different types of medium. Now, the final point, which I think is the most exciting one that you referred to earlier, Josh, is this concept of continual learning. So people, as they started to use this model, realized a very curious phenomenon. They realized that as they were using this model, let's say on day one, day two and day three, the model got exponentially better at
Starting point is 00:02:53 doing the very same task that they fed it in day one. And so people started to be suspicious that this is a model that doesn't just kind of take your input and spit back out an answer, which is what a lot of models do right now, but it can continually self-learn and improve 24-7 to give you a better response. So the question on everyone's mind is, who created this model at all? And when you look at the kind of usage or when you look at the kind of uptake on Open Router itself, although no one knew, it became the most viral launch on Open Routier. It became the new number one model and dethroned Deepseek's Flash, which launched about a couple weeks ago. And Deep Seek has been kind of like the name brand on Open Routo for the longest time ever.
Starting point is 00:03:32 So just a really exciting model to see. Yeah, it was not only the most popular model on Open Router immediately upon launch, but by a factor of two. And it more than doubled Deep Seek's use. Like this is hugely popular and very interesting because of how well it performs. The benchmarks from it, the Deep SWE subset benchmark, it had an 80%. pass rate, which for those aren't familiar, Fable 5 was at 65% and GLM 5.3 was at 62%. Hold on, hold on. So you're saying that it's crushing the METOS level models and the GPT 5.6 models at coding.
Starting point is 00:04:07 In one specific particular subset of coding benchmarks. I definitely don't want to say it's better than the frontier models because upon further investigation, we will come to find out it's not. And it's pretty far from it. But it's an incredible model for what it is. And the reason I was comparing it to GLM 5.3 is because there's this thing when you're producing models called the tokenizers, how tokens get generated. And it was just so happened to be identical to GLM. So people were like, huh, this is interesting. Maybe we can go further down the rabbit
Starting point is 00:04:34 hole and see if there's more here. So you can test oftentimes whether a model is Chinese or not based on asking it a series of questions that are not really allowed to be answered in China, one of which is asking whether Taiwan is part of China. And what's funny is if you're using like an American frontier lab, they'll very clearly give you the answer to this question. But Chinese models do not. And what we realize is that the output of the answer to this question was pretty much identical to the GLM model.
Starting point is 00:05:05 And so like, okay, it's got a similar tokenizer. It's giving similar outputs, but we still don't really know. Up until this morning in which we discovered that GLM is actually the new version of Jipu. It is, or sorry, Ox Alpha is the new version of JPU. And it seems like this is going to be a new GLM model that we have. And to that, when I heard it, it actually landed me even more suspicious because over the course of this last week on X and all of its social media, everyone has been kind of vague posting about who this is, why they're doing it, why they're offering so many tokens and so much compute. And then the Google team started chiming in, which I thought was really bizarre.
Starting point is 00:05:43 A lot of developers on the Google DeepMind team who work on Gemini, they started vague posting about this model. And as of this morning, Bloomberg released a post, and it seems like it is everything but confirmed. The model actually is from Z.A.I. And Ox Alpha is now known. And I think that was like a really interesting way of coming to this conclusion. And now that we know, I just have more questions. How on earth are they serving all this inference to all these people for free? Like the cost must be tremendous. I mean, let's see how this model shapes up against all the other models. So we have a chart over here which shows kind of like the average tokens output per task on the DeepSuite benchmark that you mentioned. DeepSuite, for those of you don't know, is a good test of a model's capability to actually code really well. And code is kind of like the premium
Starting point is 00:06:28 benchmark that a lot of companies and AI labs kind of like benchmark their model assets against. And so Ox Alpha is actually quite high up there. It's up there in terms of reduced token output, but still maintaining a very high intelligence score. Now, as you mentioned earlier on, that's not actually quite what was revealed when people started playing around with this model. It ended up achieving not really 80% that you mentioned earlier, but around 63%. And that's under, like, real usage. So there's a suggestion that the Chinese lab, Cheebu, kind of like maybe benchbacks this. But to just talk about the model and why it's so impressive, just in its own merit, there's a few things to look at. Number one, the rumor has it that this model isn't a foundation
Starting point is 00:07:10 model. This is a flash model. So what that means is it's a smaller derivative of a much larger model that Jipu is training or has already trained. So what does this mean? Let's take Mithos or let's take GPD 5.6. These are rumored to be around 6 to 10 trillion parameter models. Now, when you look at Jipu's flash model, if that's true, you're only looking at around like two to three trillion parameters in size, which, you know, is a decent size, but it's still very small compared to these bigger models. And then the fact that this model is not only likely going to be cheaper in its usage, let's say maybe 30 to 50 cents per usage, so they gave away like 15 to 30 million dollars per day of 100 trillion tokens, it is able to work much faster and they're going to open weight the entire
Starting point is 00:07:56 model. So, I mean, to answer your question, I don't really have an answer, Josh, as to how they've been able to kind of create this, aside from the usual, or maybe they distilled a Western model, or maybe they had a breakthrough in continual learning that we suggested. earlier. But the point is, this is a very serious competitor, and we saw a kind of exchange between Elon Musk and the founder of Jipu. Where about a month or two ago, Elon Musk was responding into a thread as to when he believed that Chinese or open source models in general will become mythos level, you know, have that cybersecurity risk. And he said, probably Q1 of next year of 27 and the founder G-Tung of Cheapu replied says it won't take that long and he goes on to say it
Starting point is 00:08:42 doesn't show it in this thread right here that it'll probably take two to three months and he was true to this word we now have an ox alpha mythos level model that is from jipu that is flash that is smaller that is cheaper that anyone and everyone can use and it's going to be open-weighted a few weeks from now and that's probably how it was served it's just a low-cost model to serve and it's really funny watching the reaction to this and kind of, I don't want to say overreacting, but getting very excited and enthusiastic about this, because a lot of the posts about this were like, oh, it does have continual learning and it is getting better every day and suspicions as to like, why are there more tokens being available every single day? And we have this post from Jeffrey Emanuel,
Starting point is 00:09:21 who's actually, I guess, on the limitless show. And he just said, just thinking through this free zero X alpha model or OX alpha model, how could it possibly make sense to give out this many tokens for free? Maybe there's been a continuous learning breakthrough and the best way to pour more gas on the fire is to source as many coding-related requests as possible. I don't know if that's right. I think the kind of where I'm landing here is that this is a flash model. This is some sort of distilled model that is meant to serve inference very cheaply. And what we're seeing, instead of continual learning, it's just these new checkpoints that they're kind of releasing every day. And we're able to track the checkpoints as they come into play. And they're just kind
Starting point is 00:09:54 of generally iterating over time. I'm not sure it's self-improving. We are going to find out more on that soon. The interesting phenomenon that I have really enjoyed seeing over the last week is the blind taste test that's been going on. This is known kind of like as the Pepsi challenge, where you put like Pepsi and Coke against each other in a blind taste test and people don't know what they're getting. And one of the things I found really interesting is how excited people got about this model because of the mystery and the lore. And it really makes you question the value of true frontier level intelligence when people were doing seemingly incredible things with this open source free flash model.
Starting point is 00:10:32 And when you're comparing the two next to each other, for some tasks, it's very clear that the frontier models are stronger. But for many other tasks, that's not apparently obvious. And if you are just sending a prompt into a text box and it is this Ox Alpha model or it is like a mid-tier chat chattchipt or a cloud model, the outputs are mostly the same. And people didn't realize and people don't really care where the tokens come from. So I find this to be an interesting phenomenon because rarely do we get
Starting point is 00:10:59 a true stealth model that is good. Oftentimes it's connected to a brand and there's a very strong brand affinity. So when you remove the brand affinity from the equation, what happens? You kind of get this phenomenon here where it turns out people just want pretty smart tokens and they don't actually care where it comes from. It's interesting to see the continual improvement or whatever's happening here where there's an example of a rocket ship taking off and it very clearly looks like day over day, its abilities to render 3D graphics and understand physics and create this real world
Starting point is 00:11:26 emulation are significantly better. So I'm really excited to learn more as to what's going on behind the scenes. Like are these pre-prepared checkpoints? Are they actually taking all the data they're collecting from these 100 trillion tokens and using it to continuous to apply like RL on top of the model? Like, I don't know. These are things I'm very excited to find out over the next couple of days. Yeah, I think in the future, no one's really going to care which model they use.
Starting point is 00:11:50 They just want to make sure that the output is what they expected it. They want to have the most effective output for the cheapest possible cost that they can pay for that output. And we're seeing, like, a lot of companies move to some of these Chinese open source models, or just other models, cheaper models from other smaller American labs, just so that they are able to save on costs and create a more fine-tuned model for this specific bit of work. Now, I want to spend a bit of time to talk about if this model has achieved continual learning, why that's important and why that puts China, if this is Jee Poo's actual model, which they confirmed, in a very advantageous position, because it kind of shifts the
Starting point is 00:12:28 Game of Thrones around a bit. Well, right now, you look at Anthropic, you look at Open AI, they have the world's leading models, there have internal models that they haven't even released yet that are supposedly even more intelligent than the ones that they have publicly available. And they have all the compute in the world, and they're spending so much money trying to train the best model. So if you were to reason within yourself, maybe no one else can have the capability to catch up in China, although they've been able to put out very good open source models, hasn't
Starting point is 00:12:54 been able to surpass the frontier. one way that you can kind of cheat that dynamic is if you have a continual learning model. Now, the best way to think about what this means is think of a really intelligent 15-year-old or 18-year-old. When you teach it how to ride a bike on day one and then let's say you teach it how to make pancakes day 50s, 50 days later, it'll still remember or know how to learn how to ride a bike or how to ride a bike. Current models right now doesn't work like that. They lose side of their memory, it's something called catastrophic memory loss or something like that, and they're unable to kind of remember some of the smart intelligence stuff that they learned
Starting point is 00:13:35 earlier on. Continual learning solves for that. Now, you can imagine if a model is actually able to learn, like a real human is, like a really smart kid, you could reason to say that that model will be more valuable to you, especially if you could run it locally at home on your local device on all your types of private data. So I can see two types of individuals looking at this kind of a model, and being really incentivized to use it over a clod or over an OpenAI chat GPT type model. And that one is enterprises who don't want to give all their data away to major American labs. They would want to use an open weights model that can continue to learn. And the second is any kind of respective retail user who doesn't want to give away any of their personal records to Anthropic and Open AI.
Starting point is 00:14:18 Let's say it's medical records or finance information. Let's say they want to have an AI agent that handles their finances. You know, you may have different reasoning behind whether you want to use a. certain brand or a certain lab, but if you could run it privately at home, maybe you might want to use something like this, but continual learning, just assume you have this model that learns 24-7, even when you're asleep, it could probably catch up with the frontier models. And I think it's really interesting to see. I don't know if this is that.
Starting point is 00:14:41 I'm not convinced that it is, but if China has made a breakthrough, that's going to put them in a very advantageous position. Yeah, I also don't think it's that. I find it hard to believe that this is what it looks like and that they've reached that point so far. I mean, that'd be an amazing breakthrough. In the case they do, you're going to need to be prepared. you're going to need to be safe. And Ledger, our sponsor, is here to keep you safe because if you are
Starting point is 00:15:01 building with AI agents and you're giving prompts to free open source models that are probably owned by the Chinese, chances are you're going to want some safety and some protection. And the ledger agent stack protects this. It gets this job done. It offers a series of tools that allow you to work with agents in a three-step process where the agent proposes, the human approves, and then the ledger signer enforces. So if you're building with security and you want to securely use these agents, Ledger is for you. It's available on Claudecote, on Codex, on cursor, pretty much anywhere in which you engage with LLMs, and is open source and available today. So you can find a link in the description down in our show. Thank you, Ledger for sponsoring this episode. And then I guess we kind of have
Starting point is 00:15:39 to talk about that. Like, is it really free case? Because like, is free ever really free? And I guess there's two possible answers to this question. One is yes. It's just being heavily subsidized, perhaps by the Chinese government as a vampire attack to pull some demand from the frontier webs into these Chinese models. The other, is perhaps they're just using the data. One of the big problems that we have been seeing people complain about is data retention and zero data retention policies
Starting point is 00:16:03 because people don't want to feel like they are having their data harvested and used to train other models. And when you think about giving away these free models, China, for a fact, is collecting all of the inputs from the context window. What's been happening is a lot of people are using this on perhaps proprietary code bases.
Starting point is 00:16:20 People have been testing this model in their own private projects. They've been using it for their own personal work. and a lot of data gets derived out of that. I mean, it's a million token context window in this one chat box and you just paste all in and you let it go to town. So I'm sure the amounts of data that's collected from offering all these tokens just through collecting all the inputs has been tremendous. And I'm sure that gives them a unique advantage in a way that some frontier labs might not because oftentimes they are criticized for their retention policies and people really
Starting point is 00:16:49 want to make sure that there's that zero data retention. China doesn't face the scrutiny, in particular doing it in stealth kind of shields that for at least the first week while people have used it. So they were just kind of freely and liberally using this thing, probably aware that they were having their prompts used for future training, but not really caring too much. So I thought that was also interesting. The other thing is that like the status of Jeepo as a company still isn't very strong. They had, I think, like $100 million in revenue last year and they had a net loss of like seven times that or something like that. It's currently priced at 750 times revenue. The company itself clearly can't sustain this type of funding and this free giveaway forever. And I'm curious
Starting point is 00:17:29 to know what the type of business model eventually, if any, there is going to be in order to keep this thing alive. Like, is it just going to be a government subsidy until they figured out? The business model is its state back, dude. Yeah. All these Chinese labs are backed by the state. So that's how they get their infinite amounts of capital to be able to kind of fund all these things. And this is a strategy that China has been pursuing for three plus years now. They want to catch up with the frontier Western models. They distill attack a bunch of these model labs so that they can train similarly capable models. And then they flood the market with cheap intelligence, which means that Anthropic can't charge the highest price that they can for the high quality
Starting point is 00:18:06 model that they have. Open AI can't do the same. So this drives model costs down. This is what kind of meta and Mark Zuckerberg is also trying to do, but failing to do as effectively as the Chinese labs, they're just very good at doing it. The second thing I'll say is, at least according to Open Router, from their official statement. They said that this stealth model provider, which we now know to be GPU, is retaining prompts and responses, but explicitly state that they're not using it for training.
Starting point is 00:18:32 Now, this isn't verifiable in any kind of way. Like, okay, China, you've never done anything nefarious before, I'm sure. You're really going to give this way totally for free. Like, no way. Coming from the people that stole all the, I can go and generate a full episode of SpongeBob on those models. Exactly. Exactly.
Starting point is 00:18:48 Well, this is another unique thing, right? culturally in China, and this applies for every vector of tech that they're built in the past, electric cars, mobile phones, etc. They don't really see copyright as a thing or IP infringement as a thing. They see, okay, if you have a good product, I can copy that. And Chinese in particular are very, very good at copying effectively. There was this great interview I watched with, who was it? It was Travis Kalanick on David Semra's podcast. I don't know if you called that, right? Absolutely legend, right? Twice. And he talks about his story where, you know, how he brought Uber over to China. And he made it a success there,
Starting point is 00:19:25 but he said it was very tough. But there was also some really unique cultural things that he needed to understand. And he realized that in America or in the West, at least, the innovation level is super high. Coming up with net new zero to one ideas, super high. In China, less so, but they are masters of copying. And he originally looked down at copying as like a really kind of like menial thing. But he said that the way that they do it is master's. masterful and artful in itself. And this, I feel like, is what, like, some of these Chinese labs are really good at doing. It doesn't matter if they're earning $100 million. Their IPOs are popping off anyway. They're never going to compare to the Western IPO market at all.
Starting point is 00:20:01 But the point is they want to flood the market and kind of like keep America at a cost level where maybe they're spending too much money than they actually need to until they're able to figure out some kind of a breakthrough, which might be continual money. Who knows? Maybe. We'll see. It's an exciting story. And we're going to find out more. So this will be covered and followed up in the Roundup tomorrow. Stay tuned for that. There's also a reference. It's kind of cool each is the main character of the story is OpenRouter.
Starting point is 00:20:27 And we just covered an entire OpenRouter episode just last week. So everyone who is unfamiliar with OpenRouter and why it is so valuable, please go and check out that episode. And that's mostly the state of the mystery as of now. Ox Alpha has been solved, but there is still a lot of behind the scenes that is left to uncover. We're hoping to get that information shortly. When we do, we will share it with you right here on the round up, hopefully this week, so stay tuned for that one. This has been kind of a crazy week.
Starting point is 00:20:54 I mean, it's China's always throwing, like, they're throwing wrenches in everyone's plans. Yeah. First it was Deep Seek. Now they're doing this like kind of really fun stuff model. Did you see Quinn as well? Yeah, we got a new Quinn model. We're going to talk about that tomorrow too. Dude, they're just like spaying out these models. I know. And you know the craziest part about all of this is I'm seeing rumors from trusted sources that a bunch of American frontier labs are all about to like drop a few models. So like I'm just saying here like, you know, we're getting like the best from the Chinese. We're getting the best from the West and there's just too many models to use right now. But if you're listening to this and you're curious about this Ox Alpha model,
Starting point is 00:21:32 or if you don't believe anything that we're saying, if you don't think it's good enough, go out and try it right now. You can use it. It's all free. On open router. Go check it out. You don't, there is no promo code. Just sign up, make an account and you can kind of like inference it and use it yourself. And I'm curious like, send us, exactly. Like use the free tokens, please use the Chinese state funded money budget to kind of do your AI task. Just don't give it your proprietary information. Do not definitely don't do that because they're definitely using it to train the next model. But like, let us know in the comments whether you actually like the model, whether it actually
Starting point is 00:22:03 is good at the task that you give it. And if it magically improves the next day, because it's continual learning. We don't know. Like, let us know in the comments. DM us. We're available on all socials. Yeah, anything else, Josh? Yeah, no, just again, share this with your friends if they enjoyed it.
Starting point is 00:22:18 If you have something to say, please leave a comment. We try to read all of them. If you enjoy it on your favorite podcast player, you can leave a five-star review there. If you're listening to this, you should be watching. It's very fun visually. You can watch this on YouTube, on Spotify. And yeah, as always, thank you so much for watching. Feel free to go get caught up on the other episodes.
Starting point is 00:22:36 And tomorrow, we got a banger roundup. Like, it's just going to, there's so much stuff to talk about. It is our favorite episode of the week. We've been in like a lull lately where there has. I mean, it's been exciting, but it hasn't been like, oh my God, this is crazy. there's new frontiers forming. And hopefully this is the first of many instances in which we're going to rev back up that engine
Starting point is 00:22:54 and get a whole bunch of new models. So I'm very excited to stay tuned for all of that where we'll be covering it. And yeah, we'll see you guys in the next episode.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.