Tech Brew Ride Home - Hey, Want Some Models?

Episode Date: September 2, 2026

Anthropic shipped Claude Fable and Mythos 5.1 with a 75% cache-price cut, Google countered with Gemini 3.8 Flash, OpenAI teased Astra's public release, a judge spared Google's ad exchange from a sale,... and the FBI probed a 153M-license leak. Anthropic releases Claude Fable 5.1, which is generally available, and Mythos 5.1, for trusted partners and with cybersecurity and life sciences safeguards (Anthropic) Fable 5.1 cuts cache-read pricing 75% to $0.25/1M tokens, lowering effective costs ~25% for typical workloads and up to ~45% for agentic ones, and debuts Enterprise Frontier Safeguards (VentureBeat) Google launches Gemini 3.8 Flash, three weeks after 3.7 Flash launch, for an introductory price of $0.75/1M input and $3.75/1M output tokens until December 31 (9to5Google) OpenAI says it plans to publicly release a version of Astra "soon", but will give access to "its most advanced cyber capabilities" only to testers and partners (Wired) Sources: Astra's "recurrent depth" technique obscures chain-of-thought reasoning, sparking concerns inside OpenAI and across the industry, though OpenAI limited its use so Astra's thinking stays legible and monitorable (The Information) A US federal judge rules that Google does not have to sell off its ad exchange and instead must make its ad tech tools work with those operated by rivals (Bloomberg) The FBI is investigating Nexus, a new ID theft service on the dark web claiming to sell digital scans of 153M+ drivers licenses from people in the US and Canada (Krebs on Security) The leaked licenses, including SecDef Pete Hegseth's, trace back to ID-verification provider IDScan, used by Hertz and Planet13; the FBI's New Orleans office has opened an investigation (Tom's Hardware) Acer unveils the Swift Blade 14, an ultralight laptop weighing 1.76lbs with Intel Wildcat Lake processors, and the Swift Air 16, set to launch in December (The Verge) Subscribe to the ad-free feed.

Transcript
Discussion (0)
Starting point is 00:00:04 Welcome to the Tech Brew Ride Home for Wednesday, September 2nd, 2026. I'm Brian McCullough. Today, Anthropic shipped Claude Fable 5.1 and Methos 5.1 with a 755% cash price cut. Google countered Gemini 3.8 Flash. OpenA.T.'s Astras public release, a judge spared Google's ad exchange from a sale, and the FBI probed a 153 million driver's license leak. Here's what you missed today in the world of tech. If you are anything like me, you're tracking every metric your wearables can give you. the hardest one to track is the same one impacting almost all of your metrics, how the temperature
Starting point is 00:00:43 of your bed affects your sleep. That's where the pod comes in. The pod by eight sleep is a smart mattress cover that goes over your existing mattress and actively cools or heats each side of the bed independently from 55 degrees Fahrenheit to 110 degrees Fahrenheit. It tracks your sleep, heart rate, HRV, and respiratory rate through the night without sleeping and a wearable. What I love about the 8-sleep is the simplest thing. The ability to be cool when it's time to go to sleep and warm when it's time to wake up, and it's entirely under my control. My wife can be whatever temperature she wants on her side of the bed, too. Use code ride home at 8Sleep.com slash ride home for up to $350 off the pod five. That's ride home at 8Sleep.com slash ride home.
Starting point is 00:01:35 Well, say hello to Claude Fable 5.1 and Mythos 5.1. Quoting Venture Beat, the two names refer to the same underlying model. Fable 5.1 is the generally available version with Anthropics' production safeguards in place. Methos 5.1 is available through restricted access programs for vetted cybersecurity and life sciences organizations that need capabilities normally constrained by those safeguards. For enterprise buyers, however, the release is about more than another round of benchmark gains. Anthropic is simultaneously changing the economics of running persistent agents, reducing the cost of cached context by 75% and introducing a new security architecture called Enterprise Frontier Safeguards, or EFS, designed to let organizations retain
Starting point is 00:02:22 monitoring data inside infrastructure they control. Taken together, Fable 5.1 looks less like a conventional model refresh than an attempt to solve three increasingly intertwined enterprise problems. How to make agents capable enough to finish difficult work, economical enough to leave running for hours and governable enough to give access to sensitive systems. Anthropic is positioning Fable 5.1 primarily around sustained problem solving. Investment firm Millennium told Anthropic that Fable 5.1 traced an extremely rare software crash to a bug inside an external vendor library after the problem had resisted explanation for four to five years.
Starting point is 00:03:00 Corporate expense management provider Ramp described an unattended 38-hour machine learning run in which the model re-evaluated a previous result, launched six experiments, and returned with findings and proposed next steps. Browser base said Fable 5.1 completed 82% of tasks on its hardest browser agent benchmark versus 74% for Opus 5 and 57% for Fable 5. Fable 5 retains Fable 5's headline API rates of $10 per 1 million input tokens and $50 per million output tokens. That makes it considerably more expensive on uncash tokens than other models in Anthropics. lineup. Opus 5, for example, costs $5 per million input tokens and 25 per million output tokens, while Sonnet 5 costs two and $10 respectively. The important change is cashed input. Anthropic has cut a Fable 5.1 cash hit to just 25 cents on input down from a dollar for Fable 5. That's also
Starting point is 00:03:56 just 2.5% of Fable's normal input token price of $10 rather than the 10% multiplier used by most other Claude models. That produces an unused. usual pricing profile. Fable 5.1's ordinary input and output are twice as expensive as Opus Fives, yet its cashed input is half the cost of Opus V's cash reads. Its cash read price is only 25% above Sonnet Fives, despite Fable's base input being five times higher. That matters for agents because they repeatedly revisit the same code base, system instructions, tool definitions, documents, and accumulated conversation history. Anthropics says the lower cash price reduces Fable 5.1's effective costs by around 25% for typical workloads and as much as roughly
Starting point is 00:04:41 45% for highly agentic workloads in which cached context accounts for a larger share of usage. This is a more useful enterprise framing than simply comparing per token list prices. Model selection for an agentic workflow increasingly depends on cost per successfully completed task, including retries, context replay, tool calls, and the number of tokens a model consumes before reaching a usable result. The cash price reduction also may be an effort to help Wu increasingly price-conscious enterprises. A Financial Times report found that more than two months after launch, Fable 5 accounted for only about 11% of anthropic model spending among roughly 70,000 companies represented in Ramps's transaction data, while the cheaper Opus 5 and Opus 4.8 actually
Starting point is 00:05:27 gained share. Fable 5.1, nevertheless, remains expensive relative to much of the broader market. Open AIs current promotional API pricing for GPT 5.6 sole is $4 per million input tokens, 40 cents for cashed input, and $20 per million output tokens through at least November 21st. Google's Gemini 3.7 Flash currently lists at $75 per million input and $3.75 per million output through the end of 2026. Fable, therefore, needs to justify its premium through higher task completion, lower token consumption, or the ability to replace more expensive human or multi-stage workloads, not simply through raw API price, end quote.
Starting point is 00:06:10 Only one new model today. No no, Monfrere. Google has launched Gemini 3.8 Flash just three weeks after 3.7 Flash for an introductory price of 75 cents per 1 million input and 375 per 1 million output tokens until December 31st, quoting 9 to 5 Google. Gemini 3.8 Flash delivers substantial gains over its
Starting point is 00:06:37 predecessor and various benchmarks with Google also noting how it is often approaching the performance of higher cost frontier models. Similar to the previous model, Gemini 3.8 flashes at knowledge cutoff date is March 2026 for some domains, while in other users may experience the model's knowledge is limited to January 2025. Google is once again offering an introductory price of 75 cents per million input and 375 per million output tokens until December 31st, same as 3.7, Gemini 3.8 is already live in the Gemini app for Google AI Pro and ultra-subscribers, AI mode, and Gemini and Google Sheets. It's also available for developers in Google Antigravity, AI Studio, and the Gemini API. Google Today also announced Gemini 3.8 Flash Cyber,
Starting point is 00:07:23 replacing 3.5, for trusted users via a new Fair Wind program with Frontier-Level performance in autonomous vulnerability discovery, end quote. Meanwhile, OpenAI wants you to know that it is planning to publicly release a version of its Astra model soon. That's the model that we think did the hacking dirt recently. But we'll give access to its most advanced cyber capabilities only to testers and partners. But what seems to be worrying some people is the technique behind this new model. Quoting the information. Open AI says its forthcoming AI model Astra marks a step up in capabilities such as coding and operating applications on a computer.
Starting point is 00:08:06 But an innovative technique that improve the model's performance also means that the model, and others like it, will reveal less of their thinking, making them harder to monitor for signs of bad behavior, according to a person with knowledge of Astra's development. While the limitation isn't necessarily a significant issue with Astra, the technique has triggered concerns inside OpenAI and across the industry about whether AI developers that adopt and supercharge it will struggle to guard against the kind of rogue AI that recently hacked OpenAI's own systems and those of other companies such as Hugging Face. The new technique OpenAI is using known as recurrent depth or looped transformer allows an AI model to improve its
Starting point is 00:08:45 answers by processing the same text multiple times. Unlike commercially available state-of-the-art models which show in writing how they are thinking about a task before completing it, the new technique works in a way that obscures some or all of the AI's reasoning otherwise known as its chain of thought. That means the steps that the model takes to accomplish a task can't easily be read understood by humans. Open AI has limited its use of the recurrent depth technique with Astros, so the model still produces a legible chain of thought, and the company's researchers can still sufficiently monitor its reasoning, said the person with knowledge of its development. Open AI said in a blog post Tuesday that it will launch Asthma with additional chain of thought
Starting point is 00:09:23 monitoring to rapidly detect and contain potential misbehavior. Tuesday night, OpenAI chief scientist Yakub Pachaki said in a post on X that although monitoring of models chains of thought was fragile and unfortunately heading in a negative direction. He wanted to discourage an industry-wide race toward developing models that don't produce the kind of legible reasoning that AI researchers need to understand the model's behavior. He didn't comment about the techniques Open AI is using, but said the complexity or depth of its leading models, including ASTRA, are within a factor or two of GPT4, which Open AI released in 23. The comment suggests that the structure of Open AI's current models themselves don't pose a big
Starting point is 00:10:02 problem, though other factors are making chain of thought monitoring harder. Strengthening such monitoring is a core goal of our current research program, he said. Researchers at OpenAI and elsewhere worry that some AI developers may not impose the same kinds of limits Open AI did if they adopt the same technique for their own models, and that unfettered use of the technique could potentially lead to runaway AI, whose actions can be hard to oversee. For instance, the UK AI Security Institute, the British government's main point of contact with the AI industry, wrote in a May report that opaque reasoning risks, quote, severely undermining current monitoring approaches concerns them. Open AI is preparing to release Astra, at one point the company considered labeling at GPT6,
Starting point is 00:10:44 after CEO Sam Altman went on a tour of podcasts and meetings in Washington with officials to showcase and describe the model's strengths. He hasn't publicly discussed the loop technology, that underpinned part of Astra's training, and which is also used when Astra answers questions, end quote. Fall is almost here, which means less time in the sun and way more time in your car. Whether you're headed to a big meeting, picking up the kids from school, or starting your cross-country road trip, you'll want to stay connected on your drive. With AT&T connected car, your eligible vehicle can become a Wi-Fi hotspot so you and your
Starting point is 00:11:23 passengers can happily stream, browse, and even email from the road. Got a gamer in the backseat, help keep them connected and in the game with AT&T connected car. Being on the road more doesn't mean you have to put your whole life on pause. Stay connected no matter where you're going. See if your car is eligible at ATT.com slash tech brew. That's ATT.com slash tech brew. Requires eligible vehicle, service and coverage not available everywhere. Restrictions apply.
Starting point is 00:11:53 Is your multi-entity management creating more confusion than clarity you need the Intuit ERP, Intuit Enterprise Suite. It's the AI-Native ERP solution that's powerful, painless, and proven. Learn more at Intuit.com slash ERP. A U.S. federal judge has ruled that Google does not have to sell off its ad exchange and instead must make its ad tech tools work with those operated by rivals. Quoting Bloomberg, U.S. District Judge Leone Brachima's decision allows Google to avoid a forced sale after an April 2025 ruling that the company illegally monopolies to advertising technology markets. The Justice Department had asked Brinkima to force Google to sell its exchange and make
Starting point is 00:12:41 public the auction logic that decides which advertisement will show on a website. Most of the parties proposed behavioral remedies as modified by this court be and are accepted, Brinkima said in her order, the full decision with redactions will be released later this month. Google's ad tech stack sits between publishers' selling display space and advertisers bidding on it, giving the company a powerful position in how prices are set and which ads appear. Regulators have long argued that controlling both the dominant publisher, ad server, and one of the largest ad exchanges allowed Google to advantage its own systems at multiple points in the process. The Justice Department could still appeal the ruling,
Starting point is 00:13:20 and Google continues to face scrutiny from state attorneys general, private plaintiffs, and European regulators over its advertising practices. The pace and scope of Google's implementation will likely determine whether the remedy produces meaningful competitive shifts in the display ad market, end quote. This one passes my what is potentially a big enough deal to weren't mentioning a hacking incident test, quoting Tom's hardware. More than 153 million U.S. and Canadian driver's licenses, as well as other identity documents, have reportedly become available for purchase on the dark web for a limited time. According to cybersecurity journalist Brian Krebs, the service was called nexus, and although it's no longer available at the time of writing, it claimed to have possessed
Starting point is 00:14:08 153 million driver's licenses, 10 million ID cards, 1.9 million travel documents, 1.3 million international driver's licenses, 579,000 medical cards, 429,000 common access cards, 91,000 residents cards, 77,000 employment authorization records, and 5 million other documents allegedly sourced from an ID authentication service based in Louisiana. The service was advertised on the Russian Cybercrime Forum exploit where whoever was promoting it posted the driver's licenses of Krebs as a free example, which caught the journalist's attention. He was also able to see a preview of U.S. Secretary of Defense Pete Hegsef's information on the database, a concerning breach of security for someone with such a sensitive position in the government. After further
Starting point is 00:14:55 investigation, they concluded that the service seemed to have possessed legitimate data, especially after searching for the data of several of his friends and family members with their consent. One thing that all the people he found in the database had in common was that they all used Hertz to rent a vehicle. Krebs also talked with security and privacy researcher Zach Edwards, who said that their information was also found on Nexus. Edwards said that they did not rent a car recently but used their ID at a Planet 13 marijuana dispensary. The timestamps found on the scanned images of the driver's licenses and other identity documents coincide with the time that the victims use their IDs at the said companies, confirming that they were sources of the leaks. However, Planet 13 and Hertz do not do their own authentication. Instead, they contract a service provider for the service.
Starting point is 00:15:44 Now, it turns out that both Planet 13 and Hertz use the company for identity, verification, and ID authentication ID scan. Based on the evidence gathered by crabs, it seems that the leak is centered around ID scan. He has already contacted the company about the issue and they said they were investigating the matter. At this point, I'm not able to share any additional information, but the updates you have provided have been welcome and helpful to our team's investigation, a marketing and operations leader at ID scan told the journalist. The FBI has also started looking into the leak with its New Orleans field office opening an official investigation into the breach.
Starting point is 00:16:20 The massive amount of data that was briefly available on the dark web is certainly concerning. A similar data breach hit Discord after its third-party service provider was hit and resulted in the exposure of 70,000 government IDs. Incidents like these have got privacy experts concerned with the push for online age verification requirements, which is why the Electronic Frontier Foundation is asking the California governor to veto such a law requiring IDs like this, end quote. Finally today, long-time listeners know I can't resist super thin and light laptops, quoting the verge. Acer is announcing a new laptop poised to rival the MacBook Air's thinness and lightness, and in the process, it copps another laptop's name. The Acer Swift Blade 14, not to be confused with the Razor Blade 14,
Starting point is 00:17:13 is a productivity machine weighing just 1.76 pounds, and measuring as little as 0.051 inches thin. That's around 1.7 millimeters thicker than Apple's Svelt, 13-inch MacBook Air, but Acer's blade is nearly a pound lighter. Its featherweight stature is partly due to Acer using carbon fiber in the laptop's lid and bottom plate. Inside the Swift blade will be a range of Intel Wildcat Lake processors from the Lully 5-core Core 3304 that I tested in the Chewy Uni Book and the 6-core Intel Core 7350. and regardless of which exact chip is inside, theacer will come with 12 gigabytes of RAM, so it's not likely that the acer will go toe to toe to toe with the MacBook Air's M5 chip in terms of raw performance, but it could make a compelling case for someone who prefers their laptop to be about as light as they come. The Swift Blade will come with a 1920-120-160-I-PS LCD screen with just 300 nits of brightness,
Starting point is 00:18:14 but there will be a couple of OLED options as well. The OLED versions make the chassis ever so slightly thicker, and the laptop will have three USBC 3.2 Gen 1 ports, one of which will be data only, and a 3.5 millimeter audio jack. Acer isn't sharing pricing for any configurations of the Swift Blade 14 until closer to its December launch, end quote. Nothing more for you today. Talk to you tomorrow.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.