Tech Brew Ride Home - It's All China AI All The Time

Episode Date: July 20, 2026

Alibaba launched Qwen3.8 Max as Moonshot paused Kimi K3 signups amid demand. The Trump administration weighed a slow squeeze on Chinese AI, Hugging Face used China's GLM-5.2 after US guardrails blocke...d its breach forensics, and Google built a Gemini chip. Alibaba launches a 2.4T parameter Qwen3.8 Max preview that it says rivals frontier AI models and is second only to Fable 5, plans to make it "open-weight soon" (Bloomberg) The Trump administration has reportedly explored sanctions, security warnings, and executive-order requirements since 2025 to build a slow, durable squeeze on Chinese AI models instead of pursuing an outright ban (The Decoder) Ben Thompson argues the market reaction to Kimi and other Chinese models is overblown, since it's compute scarcity — not a real Chinese cost advantage — that's keeping frontier-model prices high (Stratechery) Hugging Face says it used the open-weight GLM-5.2 hosted on its own compute for breach forensics, after US frontier model safety guardrails blocked the requests (The Stack) Sources: Google is developing a specialized server chip, informally dubbed "Frozen v2", that integrates its Gemini AI model blueprint into the silicon, for 2028 (The Information) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices

Transcript
Discussion (0)
Starting point is 00:00:03 Welcome to the TechBoo ride home for Monday, July 20th, 20th, 2026. I'm Brian McCullough today. Alibaba launched Quinn 3.5 max as Moonshot, pause Kimmy K3, signups amid demand. The Trump administration weighed a slow squeeze on Chinese AI. Hugging Face used China's GLM 5.2 after U.S. Guard Rails blocked its breach forensics, and Google is building a chip with Gemini baked in. Here's what you miss today in the world of tech. Every day, shareholders meet to discuss important matters about the companies you invest in. Now you can make your voice heard too.
Starting point is 00:00:41 Vanguard Investor Choice makes it easy to set your proxy voting preference for your eligible Vanguard Index funds. Whether you hold a Vanguard fund directly or through another brokerage firm, all it takes is a few clicks to select your proxy voting preference and be heard on important shareholder topics like executive pay and director elections. Visit vanguard.com slash investor choice to learn more. It's your shares. It's your voice.
Starting point is 00:01:05 It's easy. Vanguard investors own shares of their index funds and those funds own shares of the companies they invest in, Vanguard Marketing Corporation Distributor. There was another one over the weekend. Alibaba launched a $2.4 trillion parameter model called Quinn 3.8 Max into preview that it says rival's frontier AI models and is second only to Fable 5 on various benchmarks. Meanwhile, Moonshot AI was forced to pause new subscriptions saying that over the past two days,
Starting point is 00:01:37 Kimi K2 demand nearly past the limit of its capacity. and they're also reworking their planteers accordingly. But back to the new Quinn model first, quoting Bloomberg. The Sunday release came only days after startup Moonshot AI unveiled a powerful new offering that's royal markets and triggered concern in the U.S. about China closing the gap on global leaders like Anthropical Open AI. Quinn 3.5 Max has 2.4 trillion parameters joining Moonshots Kimmy K-3 as the heavyweight in the open weight class.
Starting point is 00:02:09 with 2.8 trillion parameters, K3 rival's top offerings, and Alibaba is setting similarly high expectations. Developers can now access Quen 3.8 max through Alibaba's coding platforms, including Coder. Alibaba plans to make the model open wait soon, expanding access beyond the preview release. Interest in these made-in-China artificial intelligence systems and models is so high that Moonshot was forced to pause taking on new subscriptions late on Sunday to manage overwhelming demand. While optimism around Alibaba is growing other contenders in China's hotly contested AI race have suffered a drop in the wake of the new Kimi release. Rival Jipu declined nearly 30% on Friday and added a further 14% to the losses on Monday after being one of the star debut stocks in Hong Kong for much of this year. Alibaba, China's e-commerce leader and one of its biggest investors in AI recently scored another victory after Beijing approved Apple Intelligence,
Starting point is 00:03:02 the software suite for iPhones, iPads, and other Apple gear, which will use Alibaba technology. in that country, end quote. So again, Chinese AI and open source AI is having another moment, but quoting the decoder. The Trump administration is considering measures against Chinese AI models that could amount to an effective ban. According to Axios, the Department of Commerce, the NSA, and the White House have explored several options since 2025. These include placing Chinese AI labs on a sanctions list, issuing security warnings, and using an executive order to impose security requirements and liability on U.S. companies that host Chinese models. The Commerce Department reportedly drafted rules as early as the summer of last year to protect
Starting point is 00:03:42 domestic supply chains from Chinese open source models. Advisors who favored a lighter regulatory approach initially blocked these efforts, but the release of China's Kimi K-3 model and personnel changes in the White House have helped supporters of tighter restrictions regain influence. A direct ban wouldn't even be necessary. A source close to the government told Axios that what's actually happening is slower and more durable, pointing to procurement rules, sanctions threats, and public pressure campaigns against U.S. companies that use Chinese models. Rather than ban Chinese models outright,
Starting point is 00:04:14 the administration could focus on potential backdoors and security flaws. Another source said that closely matches the FUD strategy, short for fear, uncertainty, and doubt that open AI strategist Dean Ball recently predicted. Soft guidelines and public warnings could deter companies without imposing binding rules. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models. This will just drive startups to sketch your providers. There's a happy middle ground here, Ball wrote. Commercial and economic interests may also be driving the push. U.S. companies are increasingly using Chinese open source models because they're
Starting point is 00:04:54 cheaper and nearly as capable. Restrictions would protect the market dominance of Google, open AI, and anthropic. The AI sector is also driving much of the U.S. stock markets' gains under President Trump. If Chinese models threatened the business of major U.S. providers, the fallout could hit markets hard, end quote. You know what? Let's turn to Ben Thompson to get a read on this new China AI moment, quoting Ben. Invidia's CEO Jensen Wong has described what Nvidia is building as token factories, and from Nvidia's perspective, that framing makes sense. Invidia GPUs are model agnostic. They generate tokens and do so in the fastest and most efficient way possible. That leads to measurements like tokens per second, time to first token, tokens per watt, token costs, etc. And
Starting point is 00:05:38 Wong argues that these metrics will be the basis for decision-making. This is a framing that definitely made sense during the first paradigm of AI, the chat GPT era where tokens were delivered straight to the end user. The second paradigm of AI, however, the reasoning era, confounds this measurement. Reasoning entails an explosion in chain of thought tokens and different models need different amounts of reasoning tokens to arrive at the right answer. Kimmy, for example, reportedly uses significantly more tokens than Sol, rendering its price advantage moot. Agents introduce a similar dynamics. Some models are more efficient than others in terms of the number of tokens they need to execute agentic workflows. What this means is that tokens are not a commodity. The defining
Starting point is 00:06:19 characteristic of a commodity is that it is fungible. A gallon of oil is a gallon of oil. A ton of copper is a ton of copper. A bushel of wheat is a bushel of wheat. A token from one model, however, is not the same as a token from another model. What is fungible is what is constructed from tokens, which is to say intelligence. In other words, if both Kimmy and Saul generated the right answer, then that answer is fungible. The difference in tokens generated to get to that right answer is a contributor to a difference in the costs of goods sold. The reason this matters is that we are rapidly approaching a state in which intelligence for many economically beneficial tasks is in fact a commodity. Anyone building a basic crud app, for example, can likely do so using models for multiple
Starting point is 00:07:03 providers. And in a commodity market, the route to profitability is not through charging higher prices. Again, you can or you will soon be able to make the exact same app using multiple models, but rather through having a superior cost structure. Right now, none of the above analysis applies because demand exceeds supply for frontier models and supply is limited by a lack of compute. This compute shortage doesn't just mean that a compute supplier like Nvidia makes very large margins, but also that Nvidia's customers like SpaceX AI can turn around and resell compute at high margins as well to a company like Anthropic. Anthropic, meanwhile, can pay the markup because they can sell tokens with a higher markup still. It's not just excess demand that gives Anthropic great margins, however.
Starting point is 00:07:46 Anthropic and OpenAI likely have among the lowest costs per unit of frontier quality intelligence, thanks to model capability, serving scale, and token efficiency. They are serving models at a particular capability level for months before their competitors and are simultaneously applying the best models to optimize those costs. It's also worth noting that the market is not yet treating intelligence like a commodity. Demand is for Anthropic and Open AI specifically, and much less for models that aren't as good, thus SpaceX AI and Meta selling capacity to Anthropic. One way to think about the push for optimizing costs is that it is a function of defining
Starting point is 00:08:20 jobs to be done by intelligence levels such as intelligence buyers can create a market where intelligence is commoditized. In the long run, however, whoever is on the frontier is the best place to dominate non-frontier markets as well, which are just the frontier minus and months, i.e. months in which the frontier model makers have been optimizing their cost of serving. All of this is to say that I think the reaction to Kimi and Chinese models generally is pretty overblown, at least from an economic perspective. Right now there is a price umbrella that is downstream of the lack of compute. I highly doubt that Chinese models are cheaper to serve on a marginal cost basis. They just seem cheaper because Anthropic and Open AI are so supply constrained that they are charging far more
Starting point is 00:09:02 than they would if there were sufficient supply to meet the demand for intelligence. I think the frontier labs are anchored in a world where training costs dominated their financial modeling. As long as training costs consume more GPUs than inference, it was critical to maximize inference revenue to help fund the next training run, which meant charging very high prices for inference. Going forward, however, I expect the inference market to grow much faster than training costs, and that includes the assumption that training costs will continue to skyrocket, which means they really can make it up in volume. It wasn't clear this would be the case as recently as eight months ago, but the agent paradigm unlock is so massive that Frontier Labs should have more confidence that they cannot just
Starting point is 00:09:41 survive, but thrive with lower prices once they have sufficient compute. Second, Intelligent isn't in fact a perfect commodity, in part because applied intelligence makes itself smarter. Specifically, whoever is running inference is also collecting data, and that data goes into making the next iteration of the model better. This is, on one hand, all the more reason for the frontier labs to lower prices and increase usage, as more compute comes online. On the other hand, this is also why companies like Microsoft are increasingly obsessed with helping companies run their own models. This is much more viable if Chinese models are a viable, alternative, end quote.
Starting point is 00:10:26 Hear that? It's your money calling. It wants a promotion. Elevate your savings with the Scotia high-interest savings account. Always earn high regular interest rates that grow the more you save and invest. Conditions apply. Visit scotiabank.com slash h-I-SA to learn more. Scotia Bank. You're richer than you think. When critical company knowledge isn't documented, there's a major ripple effect. work becomes inconsistent. Tools don't get adopted, and knowledge walks out the door when someone
Starting point is 00:11:01 leaves. Thankfully, our sponsor, Scribe, was built to fix that. Their workflow AI platform is trusted by nearly half of the Fortune 500 to capture workflows in real time. Here's how it works. You turn on the Scribe browser extension or desktop app, do a process as you normally would, and Scribe will build a guide as you go. It automatically redacts sensitive information and even suggest improvements to your existing workflows. To see what Scribe could look like for your org, head to Scribe.h. How slash ride home and mention ride home for your first month of Scribe capture free on select plans. That's S-C-R-I-B-E. Dot How slash ride home. So get this. Hugging Face says an agentic AI system hacked into its data pipeline, accessing several
Starting point is 00:11:48 internal clusters and credentials. But here's the rub. Hugging Face's own AI-based triage is what caught the breach. Hugging Face says it used the open-weight G-I-R-E-R-E. LM 5.2 model hosted on its own compute for breach forensics, after U.S. Frontier Model Safety Guardrails, read Anthropic, blocked requests by it to do similar forensics, quoting the stack. The platform security team were initially stymied in their incident response by unnamed U.S. LLM Frontier Model guardrails, quote, which cannot distinguish an incident responder from an attacker, they said. So HuggingFaces Defenders turned instead to the open-source GLM 5.2 model from China's Z-AI Lab, running it on their own infrastructure to analyze the more than 17,000 logs or
Starting point is 00:12:34 footprints that the attackers left behind. That's a striking public admission for the New York headquartered Hugging Face, which lets users collaborate on models, datasets, and applications, and which this summer hit the $100 million ARR mark. In an incident report, the company recommended that defenders have a capable model you can run on your own infrastructure, are italics, vetted and ready, before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment. The attacker, the unknown attacker, abused two code execution paths in our dataset processing, a remote code dataset loader, and a template injection in a dataset configuration to run code on a processing worker, or compute
Starting point is 00:13:16 instance said Hugging Face in a detail-th-thin July 16th incident report. They then escalated to node-level access, harvested cloud and cluster credentials and moved laterally into several internal clusters over a weekend using what the firm said was a swarm of short-lived sandboxes with self-migrating C2 staged on public services. The company's lesson for defenders in their incident write-up on July 16 stood out to both Infosec practitioners and tech investors. When we started the log analysis, we first used frontier models behind commercial APIs. This did not work. The analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the provider's safety guard rails.
Starting point is 00:14:00 HuggingFace said it then ran the forensic analysis instead on China-developed GLM 5.2 An Openweight model on their own infrastructure. HuggingFace added this had a second benefit. No attacker data and none of the credentials it referenced left our environment. HuggingFace's incident report was published the same day that Chinese AI startup moonshots Kimmy K3 model rocked global markets. The 2.8 trillion parameter model is the largest open-weight AI model to date. Blind developer testing by Arena, a platform created by researchers at UC Berkeley for its front-end code evaluation test, put Kimmy K-3 ahead of Anthropics Fable 5 and OpenAIs GPD 5.6 last week.
Starting point is 00:14:39 Chinese frontier models are also notably cheaper than their U.S. counterparts, as data from artificial analysis shows below. Hugging Face did not say which commercial frontier models it had first tried to use for its IR, nor was it clear what model the attackers used, end quote. Sources tell the information that Google is developing a specialized server chip, informally dubbed Frozen V2, that integrates its Gemini AI model blueprint right into the actual silicon, quoting the information. Google intends the new chip informally dubbed Frozen V2 to help it address a major shortage in AI computing capacity that has fueled internal tensions and compelled Google Cloud
Starting point is 00:15:23 to turn down deals with outside customers. Google employees working on the chip have projected that it could be 10 to 6 times more efficient than the newest version of Google's existing line of homegrown AI chips when it launches based on the number of tokens, a basic unit of AI consumption. It can serve per unit of power, according to the people. The frozen name comes from the idea of permanently etching parts of the model into the silicon. Engineers are still deciding what the new chip's main features will be and how its different parts will work together. Google plans to deploy the chips as soon as 2028, the people said.
Starting point is 00:15:58 The Frozen Project is intended to create a new branch of homegrown chips that's different from Google's tensor processing units rather than to replace the TPUs. Like Nvidia's graphics processing units, the most widely used AI chips today, TPUs are made to work with many different AI models. As a result, those chips must make a lot of time-consuming decisions as they interact with whichever model they're running.
Starting point is 00:16:20 Whereas Frozen V2 would have some of the decisions for Google's Gemini models built in, reducing the number of steps the chip must make and take and the amount of data it has to move around. This could make the frozen chip faster at responding to queries, potentially enabling new AI applications for Google, one of the people said, but it also means the new chip represents a bet that Google will stick with its current approach to building its Gemini models. The chip would be usable for later generations of Gemini AI models only if they are based
Starting point is 00:16:47 on the same underlying model architecture in use when it was designed. Google would be able to make changes to the same. the chip, including updating the model to take advantage of new model weights, essentially the settings that determine how models respond to questions, one of the people said. A range of companies, including smaller startups such as Sambanova and D-Matrix, and giants like OpenAI and Microsoft, are developing similar chips for AI inference, the computing work to run AI models. The goal is to handle inference more efficiently than Nvidia's GPUs to alleviate shortages of computing power that have driven up costs. Invidia itself struck a $20 billion deal in December to license technology
Starting point is 00:17:24 from Grock, a startup that designs inference chips, end quote. Nothing more for you today. Talk to you tomorrow.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.