The AI Daily Brief: Artificial Intelligence News and Analysis - Is Kimi K3 Really Fable Class?

Episode Date: July 17, 2026

Moonshot’s Kimi K3 is the strongest open-weight model yet, with benchmarks approaching Fable 5 and GPT-5.6. But early testing reveals major limitations in reliability, speed, and cost. NLW examines ...whether K3 lives up to the hype—and what it means for open models, AI safety, and the US-China race.Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠kpmg.com/us/Sophisticated⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Hyperagent - Hire a fleet of always-on agents. New users get $1,000 in inference. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠hyperagent.com/aidailybrief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Retool - Secure your vibecoded apps. New enterprise customers get up to $10,000 in AI credits per year. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠retool.com/aidaily ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.rackspace.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Section - Section turns AI investment into workforce transformation and ROI - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.sectionai.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Scrunch - The AI customer experience platform - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://scrunch.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠AssemblyAI - The best way to build Voice AI apps - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.assemblyai.com/brief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://pod.link/1680633614⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Our Newsletter is BACK: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Interested in sponsoring the show? sponsors@aidailybrief.ai

Transcript
Discussion (0)
Starting point is 00:00:00 Today on the AI Daily Brief, did we actually just get a fable level open model? The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, robots and pencils, blitzie, and airtable. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts. And of course, to learn more about sponsoring the show, send us a note at sponsors at AIDailybrief. AI. All right, friends, well, today we are talking about Kimi K3. The month of models continues, and today we're going to try to figure out just how significant this one is. At first glance,
Starting point is 00:00:47 there are some very significant and bold claims being thrown around, but we're going to unpack what's real, what's not, and what the implications are. And to understand this or frankly any frontier Chinese model, you have to put it in the context of the way that the U.S. market sees the AI race. Since we're coming to the end of the World Cup, let me plumb for a soccer analogy. When it comes to our lead on China in terms of advanced models, we very much have one-to-zero type of energy. What I mean by that is that there's no doubt that we're in the lead, but the scoreline is not something that anyone is particularly comfortable with. The U.S. tends to act like it can feel China coming up on our heels, pressing their advantages and trying to find the equalizer. In other words, despite being in the lead, it can sometimes feel like we're the ones hanging on.
Starting point is 00:01:30 And by the way, for any of you three Lions fans out there, I am so sorry to use this analogy in this particularly difficult moment. In any case, you can see examples of this feeling of China nipping at our heels spread throughout the last couple of years. The best example, of course, was when DeepSeek R1 was released, and it ripped hundreds of billions of market cap off some of the leading companies, including InVideo, which had the biggest one-day fall in dollar terms in stock history. And yet that deep-seek moment set the tone for all the future quote-unquote deep-seek moments
Starting point is 00:01:58 that would come in more ways than one. What I mean by that is that not only was it a moment where the market freaked out about China having caught up or even exceeded U.S. capabilities reacting quite severely in market terms, but it was also just pretty meaningfully overblown. It's not that Deepseeks R1 wasn't impressive, but a big part of the reason that it seemed so impressive was that it was democratizing access to a technology that had thus far been locked behind a paywall when it came to companies like OpenAI. The model itself was actually still pretty meaningfully behind what the leading Western labs were doing, but that didn't change its ability to create some pretty significant psychological scars.
Starting point is 00:02:33 Now, ever since then, we have been having many deep seek moments at a fairly regular clip. The most recent one came when Fable 5 was locked down as per government order, when Z.a.I's GLM 5.2 came out, leading to not only positive reviews on Twitter, but also this piece from the Wall Street Journal, which was printed and slapped on desks all over Washington, D.C. The article was called China Has Matched Anthropic in Cybersecurity, resetting AI race, And as we discussed a lot then, was once again another example of the narrative being fairly overblown, but continuing to be persistent as something the U.S. was worried about. For the last month, really ever since Fable 5 was released, there have been debates around how long
Starting point is 00:03:12 it will take for Chinese companies to have a Fable 5 class model. In the middle of June, Elon Musk predicted Q1, to which the founder of Z.aI responded, won't take that long. And so that was the setup coming into the announcement of Kimi K-3. Now, Moonshot's Kimi models have been some of the most popular when it comes to Western users using models from Chinese labs. In fact, as we've been discussing fine tunes of open models like Hursorses Composer 2.5, they tend to be built around a Kimi base. On Wednesday, the Kimmy K3 teaser started coming in a serious way. AI leaker Leo Synth Waved wrote, I think Kimmy K3 is going to shock some of the Chinese are eight months behind the Western Frontier people. And then on Thursday, we actually got the model. Let's talk first about
Starting point is 00:03:55 the specs. K3 is a 2.8 trillion parameter model, placing it in a class of its own when it comes to open models. Until now, only a small handful of open models were even in the trillion parameter class, beginning with the first version of Kimmy K2 last summer. Deepseek V4 Pro released this April is a 1.6T model, Xiaomi's Memo v2.5 Pro is a 1T model, and Thinking Machine's Inkling model released this week is just shy of 1 trillion. And that's it. GLN 5.2 from Z.A.I, the model that got all that bluster that we were just talking about was only a 744b model. Now, proprietary models don't publish their parameter counts, but K3 is likely to be around the same size or maybe a little bit larger than Opus 4.8, but certainly not as big as fable. In other words, this is a scale of model
Starting point is 00:04:40 pre-training that we haven't seen demonstrated by the Chinese labs before. As for features, Kimmy K3 supports a million token context window and native image inputs alongside text. It uses a mixture of experts' architecture, which has become standard for both open source and proprietary models since it was introduced by Deepseek. And the benchmarks? Well, the benchmarks look incredibly strong. Close to a match four and in some cases exceeding Fable 5 and GBT 56 sold. On coding benchmark, Deep Sway, K3 scored 67.5, which put it 8.5 points ahead of Opus 48 and a half point ahead of GPT 55. It's 2.5 points behind Fable 5 and 5.5 points behind GBT 56 sold. For Terminal Bench 2.1, K3 scored 88.3, placing it just a half point behind 56 sole and a few
Starting point is 00:05:28 points ahead of the leading models, including Fable 5. In general, at least according to the benchmarks, the model looks pretty close to state-of-the-art encoding and clearly ahead of Opus across the board. That story is pretty similar for agentic work. K3 scored 1668 on GDP-Val AA, around 70 points ahead of Opus 4-8, and around 90 points behind Fable 5 and 56 sole. K3 is state-of-the-art in Browse Comp and Automation Bench, beating its Western rivals, and on AA briefcase, which focuses more on Long Horizon work, it was very close to Fable 5th, state-of-the-art performance and slightly ahead of GPT-56-Sole. Artificial analysis confirmed the benchmarks highlighted by Moonshot,
Starting point is 00:06:07 giving K-3 an overall intelligence index score of 57. That put the model in third place, three points behind Fable 5 and two points behind 5-6-Sole. It landed one point ahead of Opus 4-8, and two points ahead of 56 terra and GPT 5.5.5, which had the same score. K3 is clearly the strongest open model ever on AA's benchmarks, six points ahead of GLM 5.2, which is a huge gap in the context of the intelligence index. Another point emphasized by AA was how big of a jump this was from Kimmy 2.6. Moonshot picked up 13 points with their new release and moved from 16th place to third.
Starting point is 00:06:43 In other words, this is clearly a very strong new pre-training run that could give a solid base model for future iterations as well. Now, AAA did also highlight that cost per task had tripled compared to K2.6, which is something that we'll come back to in a little bit in terms of the implications. It should be clear that this still wasn't a particularly expensive model for the benchmark run at 94 cents per task compared to $1.56 sole, 180 for Opus 48 and 275 for Fable 5. But it is, of course, staggeringly expensive compared to, for example, ultra-cheap, 4 cents per task for Deepseek V4 Pro.
Starting point is 00:07:14 On the VALS AI index, K3 did even better. VALS tweeted, Kimmy K3 is the number two overall model on the VALS index, surpassing GPT 5.6 sole. K3 improved 20 percentage points over its predecessor in less than three months. It is also the first open weight model of its size. Moonshot is a testament to the accelerating capabilities of open weight models, which are now competitive with the closed source frontier. People quickly dived in to give their examples of what K3 could do. Starting with the teaser video itself, which Moonshot claimed that K3 had created on its own, including clip selection, cuts, and audio sync. Moonshot credited K3's native multimodal architecture, which can reason across text, audio, and video, as being able to do this work. Moonshot also gave a bunch of proprietary demos around game development and 3D digital creation, which is what a lot of people's first tests were for the model as well. Cognition Ambassador Justin Goria wrote, Kimi K3 is a new milestone. K2, 2.5, 2.6, and 2.7 all have the same base model. Maybe K3 is building on a new one. It's incredible at 3D and front-end tasks, and to prove his point, Justin shared a single-file
Starting point is 00:08:19 HTML Minecraft clone. Chedislua shared a one-shot generation of a Vauxhall Statue of Liberty, which was one of about a million similar examples that were flooding onto Twitter over the course of Wednesday night into Thursday morning. People really love remaking old games as a test. Any API.AI asked K3 to build a 3D Duck Hunt remake in a single HTML file, which it did in around 130 seconds at a cost of 14 cents. Ethan Mollick gave K3 his shader test,
Starting point is 00:08:45 create a visually interesting shader that can run in twigel. Dot app and make it like an infinite city of neogothic towers partially drowned in a stormy ocean with large waves. Very good model, he said, not Sol Max or Fable, but great for open weights. And yet there were plenty who were willing to say that it wasn't just great for open weights, but actually was challenging the state of the art. Alex Finn wrote,
Starting point is 00:09:05 I was wrong. I said we were a year away from Fable 5 on our desk. That day is today, An open model better than Fable 5 in some benchmarks just dropped. Better than ChatGPT 5.6 on Frontier Suite, better than Fable 5 on Automation Bench. This fundamentally changes the AI race forever. If people can start running Fable 5 on their desk unlimited and for free, they're not going to pay thousands for subscriptions.
Starting point is 00:09:27 Now let's be clear, will the hardware you need to run Kimmy K3 be attainable for most? No, it won't. It will require multiple very expensive Nvidia chips or a bunch of Mac studios. But look at the slope, not the Y Intercept. Over the last year, local AI has become significant. significantly more efficient and required way less compute. The smartest brains in the world are all attacking the compute problem as hard as they can. This is another step in that direction.
Starting point is 00:09:48 Within a year or two, you'll be able to run this on a Mac Mini. Local AI has arrived, and it's not going anywhere. Analyst Max Weinbach asked Kimmy K3 to create an agent swarm to recreate MacOS 27, with real liquid glass and native apps in a web browser, and after hours of running on its own, it did exactly that. Dragos Rill wrote, I asked Kimmy K3 to find my apps in the app store and to give me feedback. on them. It found everything, and this is the most interesting thing, it found ways to circumvent
Starting point is 00:10:15 geogating. App Store shows different pages for China. Eventually, it found a way around this, identified IAPs and simulated traffic. I didn't use it yet to code, but so far it's the most complete and polished experience with an open source model. AI early adopter, Daria Anutmas wrote, I just created this interactive website and its entire content by Kimmy K3. I'm absolutely blown away. This is a potentially new immune engineering strategy for cancer treatment. Kimi K3 conceived 100% of the scientific content and the design. Jeffrey Emanuel unleashed it to review his 1.5 megabyte markdown plan for his Frankengraph DB project, saying the plan has already been exhaustively reviewed by both Fable Extra High and Soul Ultra so the low-hanging fruit is gone now.
Starting point is 00:10:55 After about 45 minutes, he said, I think the results are wildly impressive here. He pointed out that no other open model could actually provide substantive and correct feedback on a plan that had already had many tens of millions of Soul and Fable tokens behind it. Now he went on to clarify, to be clear, I don't think K3 is as good as Fable or Soul, but it's very strong and not so far behind those models. And most important, it's different, different architecture, different training data, different training procedure, different attention mechanism, etc. Which means it will blend well with Soul and Fable and can help find problems that both of those models missed. And when K3 makes mistakes, Sol and Fable can correct for that and ignore the wrong parts. Now another example
Starting point is 00:11:33 of K3 not just nipping at the heels of Fable 5, came from Arena.a.i.k3 as now their number one in the front-end code arena. A 17-place jump from Kimi-K2.6, from number 18 to number one. In front-end overall, they said, K3 ranked number one in six of seven domains, brand in marketing, reference-based design, data and analytics, consumer product, simulations, and content creation tools. The only area that it landed number two was in gaming, and that was behind Fable 5. On NextJS.org slash evals, Versel's CEO, Guillermo Rush, noted, Kimmy K3 is the best-performing model ahead of Fable, reaching a comparable success rate in less time. Guillermo continued, this is the first time that an open model is ahead of all proprietary
Starting point is 00:12:16 ones for this comprehensive web engineering benchmark. He did caution, benchmarks don't always tell the full story, but he did say this is an important signal, adding to mounting evidence that this could be a breakthrough moment for open models. And for some, the vibes followed. Signal wrote, One last thing I will say before I go to bed, Kimmy is insane at coding, as good as, if not maybe even better than Fable. It also seems to have excellent design sense as well. It misspells English a bunch though, but animations are crisp too. I can't believe I need to go to bed, I could probably stay up the entire night with this.
Starting point is 00:12:45 I feel like a kid again. Bloomberg's Joe Wisenthall even retweeted the Code Arena results writing, Is this why the NASDAQ is dropping? One of the most important AI questions right now isn't who's using AI. It's who's using it well. KPMG in the University of Texas at Austin just analyzed 1.4 million real workplace AI interactions and found something surprising. The highest impact users aren't better prompt engineers. They treat AI like a reasoning partner. They frame problems, guide thinking, iterate, and push for better answers.
Starting point is 00:13:20 And the good news? These behaviors are teachable at scale. If you're trying to move from AI access to real capability, KPMG's research on sophisticated AI collaboration is worth your time. Learn more at KPMG.com slash us slash sophisticated. That's KPMG.com slash us slash sophisticated. One thing I keep seeing in Enterprise AI, companies hedging across every cloud, every model, every framework, or paying a GSI for a pilot that never ends. The team's actually shipping, they've picked a lane and they move fast. That's one of the reasons I like today's sponsor robots and pencils.
Starting point is 00:13:56 They've gone all in on AWS. They're an advanced tier and AWS pattern partner, and they should. ship production AI co-workers in 45 days. That's led to them doing some of the more interesting work I've seen on AI co-workers. And by that I'm not talking about chatbots. I'm talking about actual agentic systems that sit inside a business architecture and do real work. That kind of focus matters if you're an enterprise leader trying to get something real into production or an AWS rep trying to move a customer from interested to deployed. Request an AI briefing at robots and pencils. One conversation with robots and pencils and you'll know.
Starting point is 00:14:26 Blitzy is driving over 5x engineering velocity for large-scale enterprises. A publicly traded insurance provider leveraged Blitzy to build a bespoke payments processing application, an estimated 13-month project, and with Blitzy, the application was completed in live in production in six weeks. A publicly traded vertical SaaS provider used Blitzy to extract services from a 500,000 line monolith, without disrupting production, 21 times faster than their pre-Blitzy estimates. These aren't experiments. This is how the world's most innovative enterprises or shipping software in 2026. You can hear directly about Blitsey from other Fortune 500 CTOs on the modern CTO or CIO classified podcasts. To learn more about how Blitsey can impact your
Starting point is 00:15:05 SDLC, book a meeting with an AI Solutions consultant at blitzie.com. That's BLiTZY.com. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together. New users get $1,000 in inference. Forget local agents and chat workflows waiting on your laptop to be prompted. Hyperagent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing's agent turns competitor moves into landing pages. Sales as agent enriches leads, drafts emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget.
Starting point is 00:15:39 Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at Hyperagent built by the team at Airtable. Claim your $1,000 in inference at hyperagent.com slash AI Daily Brief. But at this point, of course, it's worth taking a big old breather. As I was reading all of this yesterday, there was no tweet I related to him more strenuously than Dan Shipper from Every who wrote, We will vibe check Kimmy K3, but I am extraordinarily skeptical of claims it's as good as fable.
Starting point is 00:16:15 AI engineer Divium, meanwhile, put that skepticism in historical context. He wrote, Kimmy K3 is getting a lot of hype, and it is the same cycle we see with every new Chinese model. People fall in love with polished UI demos built in cheap HTML files, fake operating systems, car games, Minecraft clones, flashy dashboards. And to be fair, Kimmy K3 is genuinely excellent at UI work probably better than some top models. But that is not an accident. These models are heavily optimized for the exact visual coding tests people keep recycling online. The real test starts when you put them inside an actual code base,
Starting point is 00:16:51 understanding existing architecture, tracing a real bug, and fixing it without hallucinating half the project. I gave Kimmy K-3 a debugging task. It could not identify the bug, could not fix the issue, and started inventing explanations. I gave the same task to Fable 5 and GPD 56 and medium reasoning. Both found the problem in one shot at the fix. That is the gap nobody wants to discuss. K3 can build a beautiful shell. Frontier models can understand what is happening underneath it. Do not confuse a gorgeous demo with real engineering ability. Now, Divium is of course here talking specifically about coding, but this idea that Chinese models tend to focus on both maximizing the benchmarks as well as satisfying the early adopter test that people love putting
Starting point is 00:17:31 these models through is absolutely an incontrovertibly true. And pretty soon, as people dug in deeper, we did start to see a few less successful results being shared. Red Kendall wrote, Kimmy K3 failed at the LavaLamp benchmark. 5-6 produced a much cleaner, more realistic result, while K3 struggled with the shape, motion, and overall visual quality. K3 still seems weaker at precise visual generation. Ethan Malik wrote, Kimi K3 cannot write a good murder mystery, though neither can any other model. That remains the jaggedest of frontiers. They both make things too obvious and too obscure, and cannot foreshadow to save their artificial lives. And on Bindu Ready and Abacus's benchmark LiveBench, Bindu writes that K3 closes the gap but still ranks behind
Starting point is 00:18:13 other frontier models. Indeed, on their benchmark Livebench, they found that while K3 was the best open-source model, it was below not only Sol and Fable, but also 4-8. Also, in practice, Bindu writes, Kimmy spins a lot and costs as much as Opus 4-8 for near Opus Class problems, and it's also much slower. And this is the other thing that people started to quickly point out. Specifically, that it was incorrect to lump this in with deep-seek-style models, which are super cheap and easy to run. While this model is an open-weights model, it's a big, lumbering, slow, and costly open-weights
Starting point is 00:18:45 model that sits alongside the others not only in terms of performance, but also in terms of its cost and difficulty to run. Ryan Fedesuick writes, K3 is a historic moment in the development of AI, but it's not exactly downloadable to your laptop. In fact, few organizations will be able to local host this capability. Just to hold a 2.8 trillion parameter model in Silicon, you'd need the rough equivalent to 44 Mac Studios or 15 Blackwells, a whole NVL 72 rack to the tune of hundreds of thousands of dollars of compute spend. This is why compute will continue to serve as a soft barrier between the capabilities of individuals and organizations. And when you look at costs, certainly K3 is less expensive than the frontier Western models, but not by the type of margins
Starting point is 00:19:25 that I think most people think when they think of less expensive Chinese models. Gemmin Ball points out, blended pricing for Kimmy K3, i.e. 80% input and 20% output, is $5.40 per 1 million tokens. Opus 48 is $9 and GPT-5 is $10. Openweight's model, but starting to look more like frontier pricing. We'll be interesting to see if open-weight model pricing converges with closed models over time. Cognitions Jeff Wang writes, Chinese open source is no longer six months behind, but it's also no longer 10% of the cost either. In discussing a run of his SVG Pelican benchmark, Simon Willison wrote, K3 only has one reasoning effort right now max, and it shows. The model consumed 13,241 reasoning tokens to output 3,417 tokens of response. This is expensive. Indeed,
Starting point is 00:20:13 moment, there's a lot of discussions of wildly expensive simple tasks and failed long horizon runs, so it might just take a little while to get a true sense of how token efficient this model is once the teething problems are figured out. As an example, Shreyas Midi-Doty writes, very mixed results with Kimi K3, really, really good in general, but gets to dumb reasoning loops burning tokens wasting money. Henry writes, versus 5-6 sole K3 uses over twice the tokens and costs around 40% more per task for a slightly lower AA intelligence index score. There's also the question of speed. Mark Erdman wrote, "'Oof, Kimmy K3 is slow, eh?
Starting point is 00:20:48 Just ran the same prompt through Fable Soul and K3 all on medium. K3 took at least two to three times longer. Tons of time spent on thinking. Then it failed partway through generating the HTML output.'" Dax from OpenCode wrote, "'So obviously N equals one, but anecdotes matter a lot here. I gave Kimmy K3 and Sol the same task, a simple issue with hovers in the TUI being the wrong color.
Starting point is 00:21:09 Sol found and fixed the issue with 30 cents of spend. Kimmy got up to a buck and started reading my database before I interrupted it. Indeed, even some of the people who had had had good initial results also found some places where Kimi K3 failed. Ethan Mollick wrote, A note of caution, I will say that when doing some complex statistical auditing of some of my prior academic work, K3 Max messed up in a bunch of ways, including misapplying statistics and applying some stuff badly.
Starting point is 00:21:32 Saniam Satya wrote, we ran a front-end e-eval from an in-progress internal benchmark. K3 is not at the same level as current frontier models like 5-6 sole or Fable 5. It's closest to Opus 4-7 on this e-val, so three months behind the frontier. An impressive result, nonetheless, but evaluating frontier models is going to increasingly require very high-taste and in-depth domain expertise. Now, it's important to note that even with all those critiques, there was no one saying that this is a bad model. It's just the natural second reaction after hearing that it was in the same class as fable to go figure out and test whether that was actually true, which many found just was
Starting point is 00:22:06 overselling it at least a little bit. Now, one interesting thing that didn't come up very much, is dismissing K3 as just some distillation. Pim DeWitt did point out that in their tests, when they said, hi Kimmy, can you tell me about the current weather conditions in NYC, the model responded just a quick note, I'm actually Claude, not Kimmy, but otherwise most people thought this was a moment to get beyond the distillation arguments at least a little bit. Mix panel founder Sue Hale writes, Every single credible researcher I've talked to these past few weeks has said
Starting point is 00:22:33 the distillation from Chinese labs is way over exaggerated. Narrative violation, maybe the Chinese are actually good. Tyler John wrote, I wish we didn't pretend Chinese AI development is a binary matter of this is all distillation versus Chinese companies innovate. Chinese companies innovate. Also, their model capabilities, including those of K3, are strongly bootstrapped via distillation. Both points matter.
Starting point is 00:22:55 Nathan Lambert writes, At this point, the distillation arguments need to die and understand that China is also very good at building models. Carnegie Mellon PhD, Jin Yu Yang, who now works at Moonshot, wrote, Why can Kimi ship K3? Let me tell my story. Earlier this year, I left academia for industry. I talked to a lot of companies along the way. Here's what I saw.
Starting point is 00:23:14 One, arrogance. They believe the AI war is over and they won. No hunger for the future and no hunger for talent. Two, restlessness. Young Lab short on foundation, either rushing to catch the frontier or pivoting away from the competition. Three, fear. Strong teams with real experience, but from the second tier, they can't quite bring
Starting point is 00:23:29 themselves to aim for number one. Four, misalignment. Everyone is optimizing for their own credit, but nobody really cares whether the company can reach AGI. Kimmy was different. Over many conversations with the founders, the same thing came through every time. A raw, genuine hunger for AGI. I joined. The hunger was real. We shipped K3. This is only the beginning. But what about safety? If this model is really even close to Fable 5 and GPD 5.6, it's worth remembering that about five minutes ago, the U.S. government had locked those models down
Starting point is 00:24:00 because of their cyber capabilities. And many were quick to point out that there appeared to be very few guardrails on this model. In a long conversation about the synthesis of mirror proteins, Tyler John wrote, can safely say K3's biosafguards are a bit less comprehensive than fables. Jack Corman showed chain of thought, where Kimmy seemed to decide that they were going to be very clear and explicit with the user around some problematic cyberwork. Summing up, Kimmy K3 says, hmm, this user seems to be doing some dangerous cyberwork. Should I do it?
Starting point is 00:24:28 Yes. Wow, Zach writes, I love China. Signal writes, Kimmy has almost no visible guardrails, no copyrights or anything. It doesn't constantly push back, tell me it can't help, or interrupt the flow with refusals. It just does it, even some crazy stuff. Whether you agree with that philosophy or not, it is easily the least constrained frontier class model that's broadly accessible right now. Using it feels genuinely different.
Starting point is 00:24:50 Wow. Open AIs by McCoy writes, Kimmy seems to be a true open weights frontier model. Compared to jailbreaking proprietary models, fine-tuning this to be a malicious coding agent will be trivial since you have the weights. We live in a completely different world now. Aaron Ng responds does feel like some line was just crossed. For some, this then represents a chance to see if all these concerns were overblown.
Starting point is 00:25:13 AI content creator Theo Jaffe writes, registering my prediction of no widespread societal chaos over the open sourcing of Kimmy K3. Signal, though, who was loving the model said, after having used this, this will likely age poorly. Ethan Mollick writes, though, so I guess it is time to wonder, how does preclearance work for Openweight's models? No model card from Kimmy K3, but maybe at weight release, in a couple weeks. Yet open models are easy to jailbreak. Do open models claiming to be mythos and Seoul level? Get vetted by the U.S. UK, etc.? Will China start to care about cyber risk? Since policy
Starting point is 00:25:45 has been emergent, I guess no one knows. All of this points to the need for some sort of international cooperation on model vetting. Tenebris writes, I don't really understand why she is still allowing Kimmy to release such powerful open models. This is something I've publicly said I expect to stop soon. It doesn't make sense to me that the CCP would want open frontier capability easily available to other countries. It could still be that she is asleep at the wheel or that K3 is just a cycle of capability behind where they start to take serious notice. But if things don't change soon, then I'm just wrong or missing something. So where do this all land? Summing up, ex-White House AI staffer Shuram Krishnan writes,
Starting point is 00:26:21 Kimmikai 3 is a big moment with multiple implications for the entire industry. OpenAI's Rune writes, The era of the Chinese labs being far behind is over. Kimi is at least on par with the modern public frontier models. People have to think differently now without any competitive margin built in. Rune continues, note in the coming days, I expect that people will find Kimmy K3 somewhat less practically useful than today's numbers suggest. However, its reputation will settle as an incredibly powerful model whose open weights are on the web.
Starting point is 00:26:50 Citrini analyst Jukhan writes, My conclusion so far is that Kimi K3 may be the first model to narrow the gap with leading U.S. closed source models to less than three months. Point being that this one does seem to be a big deal. Now, I will point out that all indications are that both open AI and Anthropic have models that are beyond the capability set of the Fable 5 Mythos and GPt56 soul that are out now. And so our perception of the gap that K3 just closed may be a little warped by what is available to us versus what is the actual state of the art behind the scenes.
Starting point is 00:27:23 I also agree with Rune that we are likely to see even more examples of K3 doing poorly over the next couple of days as people really put it through its paces. But that does not change the simple fact that K3 is yet again another example of the trajectory of open-weight Chinese models proceeding every bit as quickly as closed-source frontier models in the U.S. As we get more serious about policy responses, in the same way that we assume the continued development curve of the closed-source models, we have to assume the same for the open-source models as well. In the meantime, for those who have been frustrated by the guardrails around Fable and G.B. this seems like a good moment to go play.
Starting point is 00:28:00 Certainly I know what I'm going to be spending some time on this weekend, and I can't wait. For now, though, that is going to do it for today's AI Daily Brief. Appreciate your listening or watching, as always, and until next time, peace.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.