Tech Brew Ride Home - AI Models Just Got Cheaper?

Episode Date: September 23, 2026

Anthropic launched Claude Opus 5.5, cheaper and safer than Opus 5, right as OpenAI dropped GPT-6 Sol and Luna with half the errors and lower prices. YouTube countered with Shorts Series, an AI creator... agent, video A/B-testing, and Custom Feeds. Anthropic launches Claude Opus 5.5, its first model since Dario Amodei's "pace the frontier" essay, and says it has enhanced safeguards to combat risky behavior (The Verge) OpenAI launches GPT-6 Sol and Luna, saying Sol makes about half as many mistakes as GPT-5.6 Sol and Luna matches GPT-5.6 Sol's performance at ~1% of the cost (ZDNet) YouTube rolls out tools for microdramas, an AI storytelling assistant for script analysis, video A/B-testing, and more, as it battles Netflix over top creators (WSJ) YouTube launches Shorts Series, letting creators organize their Shorts into TV-style seasons and episodes with custom thumbnails, so fans can watch a season as one continuous story (TechCrunch) YouTube unveils an AI agent that mines creators' back catalogs for trending video ideas and lets them A/B-test three full video versions against real audience segments (The Verge) At its Made on YouTube event, YouTube unveils Custom Feeds, an LLM-powered feature that lets users generate and save tailored homepage video feeds, for US users (Wired) Subscribe to the ad-free feed.

Transcript
Discussion (0)
Starting point is 00:00:04 Welcome to the Tech Brew Ride Home for Wednesday, September 23rd, 2026. I'm Brian McCullough today. Anthropic launched Claude Opus 5.5. Cheaper and Saper 5. Right as OpenAI dropped GPT6, Sol and Luna with half the errors and lower prices. And YouTube is trying to make both viewers and creators happy with custom feeds and AI everything creation tools. Here's what you miss today in the world of tech. I've got good news. IT pros. Ace of Uptime is back with new decks. You won't want to make. This online card game made by Eaton gives you the chance to battle IT villains and become an IT hero. And now there are three new decks made for Enterprise Data Center and lone IT pros. Each one is built for those unique worlds with their own challenges. Ace of Uptime has all the decision-making you already know, just faster, funnier, and without the ticket backlog. And keep in mind, no game plays out the same way twice. Try the new decks and see how your strategy holds up.
Starting point is 00:01:06 Head to Eaton.com slash ace to play. That's eaton.com slash ace. Well, yesterday afternoon was a big day of model launches. Anthropic led with Claude Opus 5.5, its first model since Dario Amadai's Pace the Frontier essay. And Anthropic says it has enhanced safeguards with this new model to combat risky behavior. Anthropic says Opus 5.5 matches Fable 5.1 on most tasks while costing about 40% less to run than Opus 5. Opus 5.5 costs $4 per 1 million input and $20 per 1 million output tokens. Anthropic also raised its five-hour usage limits for users by 20% on pro, max, and team plans,
Starting point is 00:01:53 and gave subscription users a rate limit reset like a get-out-of-jail-free card, quoting the verge. Anthropic says Opus 5.5 is the strongest performing model on the company's most comprehensive alignment test. During testing, it attempted to circumvent boundaries 85% less. than Opus 5 or Claude Mythos 5.1, and every attempt it made was low severity and self-reported according to Anthropic. It also comes with improvements to biased or motivated reasoning, which contributed to the recent AI hacks. Opus 5.5.5 costs 40% less to run than Opus 5.5, but matches the performance of Fable 5.1 on most work. It also comes with safeguards similar to the ones offered by Anthropics' more advanced Fable 5.1 model. This means Opus 5.5 will reroute certain cybersecurity-related
Starting point is 00:02:39 request to the less powerful opus 4.8, while biology-related requests flagged by its safeguards will go to Opus 5. Anthropic says Opus 5.5 was tested by outside partners, including Frontier Design and Metter before release. The company also plans to launch Claude Sonnet 5.5 and Haiku 5.5 in the coming weeks, end quote. And quoting the decoder. Anthropics says total operating costs, which account for both token prices and token usage, should be about 40% lower than Opus Fives. The model uses fewer tokens and generates output more than 30% faster. Five-hour usage limits for subscribers will increase by 20% with the model's lower cost.
Starting point is 00:03:21 Anthropic says those limits stretch 25% further overall. Users can also save a limit reset for when they need it most. Anthropic is using coding benchmarks to make its case on price and performance on frontier code. The company says Opus 5.5 beats Open AIs GPT6 Astra at about 20% of the cost per task. On Terminal Bench, it claims the same performance as Astra at 40% of the cost. On cursor bench, it says Opus 5.5 beats GPT 5.6 sold by 11 points at one third of the cost. The price cut is a response to pressure from OpenAI and especially Chinese AI models, which offer lower performance but cost a fraction as much. Opus 5.5 is also supposed to communicate more naturally than earlier models. Anthropics says it puts the most important information first, uses less jargon, and follows writing instructions more closely. Early testers described its writing as clearer and easier to understand, which Anthropic says makes it a better partner for long work sessions. Current Claude models have drawn plenty of criticism for their formulaic convoluted writing, sometimes called clawdish. Independent platform artificial analysis confirms Opus 511.5.5.5.5.5.5.coms
Starting point is 00:04:32 5.5's strong showing at max effort. The model scores 58 on the artificial analysis intelligence index, the highest ever, and several points above the previous leader. It leads six of ten evaluations, including Humanity's last exam at 61.4%. The previous best was Fable 5.1 at 59.1%. And Cicode at 66.9%, 63.1% was the previous best, also Fable 5.1. On Terminal Bench, 4.0, Opus 5.5.5 hits 50 59.6% tying GPT6 Astra and beating Opus 5 by 11 points. Artificial analysis says Opus 5.5 brings Anthropic to parity with GPT6 Astra on terminal bench 4.0 and automation bench AA, while extending its lead in agentic knowledge work on the private AA briefcase benchmark. Opus 5.5 hits an ILO of 1822, up 143 from Fable 5.1,
Starting point is 00:05:29 marking the first time an anthropic model has beaten GPD 5.6 sole on presentation quality. On GDP Val AA and OpenAI benchmark for real world white color work, Opus 5.5 outperforms Astra on every reasoning mode except low. Artificial analysis does flag an efficiency tradeoff, though. At max effort, Opus 5.5 burns about 119,000 output tokens per task, far more than Opus 5.5. 73,000 or Fable 5.1 at 78,000, or GPT6 Astra at only 27,000. Lower token prices keep its cost per task in line with Opus 5, but not cheaper. But four of five Opus 5.5 effort levels land on the Pareto frontier, matching or beating every other model above 50 on the index for cost per task, end
Starting point is 00:06:20 quote. And right on top of that, opening I launched GPT6. Sol and Luna, saying Soul makes about half as many mistakes as GPT 5.6 and that the new Sol and Luna match GPT 5.6 Sol's performance at around 1% of the cost. So again, the name of the game here is better and cheaper. Everybody claims, quoting ZDNet. Open AI says that GPT6 Sol makes about half as many mistakes as its predecessor. The company also says GPT6 Luna also improves substantially at higher effort levels, it matches GPD 5.6 sole at about a hundredth its cost. In other words, if GPT5.6 was making factual errors 20% of the time, GPT6 is making them only 10% of the time based on the same queries. GPD6 Luna, the low-cost engine, now performs as well as its
Starting point is 00:07:19 more capable tier in the previous release. Let's put these models into some perspective. Both OpenAI and Anthropic have flagship, workhorse, and cheap model offerings. Think of OpenAI's Astra workflows as super elite specialists. If they were doctors, they'd be the experts flown in to save the life of a head of state. They're good, but very expensive. Methos and Fable in the Claude world are at that same elite level. Now, think of Sol and Opus on the Claude side as senior staff, the top-tier doctors at major hospitals who do the bulk of the serious work. They're costly, but not at the level of cost where you're booking flights on private jets for them to intercede in an emergency. Using the same analogy, Luna and Haiku Fore Anthropic are more like the medical
Starting point is 00:08:02 residents. They're capable, but they don't have the deep expertise and understanding. They're more prone to mistakes, but they can handle routine work quite well. They're also very inexpensive, even compared to senior tier models. That leads to the most scarily impressive detail of this announcement, which Open AI isn't even showcasing, the rate of improvement. GPD 5.6, Sol and Luna were released in July less than three months ago. And less than three months, I managed to double the factual accuracy of the workhorse models, which is roughly equivalent to making every top doctor in a hospital as good as the super experts. On top of that, they also managed the equivalent of skilling up low-tier residents, so they
Starting point is 00:08:42 perform as well as the top doctors, again, in less than three months. But back to the rush towards affordability. Yes, cost matters, but how much it matters probably depends on your plan and usage pattern. The headline pitch in Opening Eyes release is, and this is the only bold-faced text in the entire opening paragraphs, reducing API prices for Sol and Luna by 50% compared with their GPD 5.6 promotional pricing. You can get a total headache trying to compare input API and output API pricing and promotional pricing for both input and output tokens. The company says its new models cost half as much as the older models. Since it's compared to the promotional prices for the older models,
Starting point is 00:09:22 those paying the regular price would save even more with the new GPT CPCS. releases. For API users, this is a very big deal. But for those of us on plans, whether that's the $20 a month chat GPT Plus plan or the $100 to $200 pro plans, these cost savings don't directly apply unless the cost is really 50% lower with GPT6. Does that mean plan usage will be consumed half as fast when running the newer models? Opening Eye says it's improved the AI's communication style for GPT6, Seoul and Luna. The company describes it as expect to see more clarity, less jargon, fewer odd turns a phrase, fewer low value details, and slightly shorter answers overall without losing substance. Continuing with the cost-saving theme of this
Starting point is 00:10:06 release, Open AI has improved prompt caching in GPT6. When you issue a prompt, the AI has to process it and understand what that means. If you ask the same prompt over and over, or ones that are very similar, or you ask the AI to review older prompts, requiring the AI to essentially recompile the prompt each time is expensive. So prompt caching is the practice of storing those process prompts and then retrieving them as appropriate. OpenAI says that using cached prompts can cost up to 90% less than evaluating an uncashed prompt. This is particularly beneficial for long-running agents who won't have to evaluate the same instructions over and over again. Regardless of the output results, the drumbeat of lowered costs runs through every point raised by OpenAI. This is as important
Starting point is 00:10:51 for the company as it is for its users. The lower the cost, the more cycles Open AI can mine from its ever-growing and ever-more costly data center footprint. If Open AI can get the same amount of work out of fewer machines or scale up output at a higher rate than scaling up data center resource utilization, that's ultimately a win for everyone, end quote. Is your multi-entity management creating more confusion than clarity you need the Intuit ERP, Intuit Enterprise Suite. It's the AI Native ERP solution that's powerful,
Starting point is 00:11:31 painless, and proven. Learn more at Intuit.com slash ERP. Teams that struggle to scale aren't falling behind because they're lacking headcount or budget. It's because critical knowledge isn't documented. That's what today's sponsor, Scribe, was built to fix. When specialists or experts leave your team, their knowledge disappears too. That's where Scribe comes in. Scribe is a specialized intelligence platform trusted by 90 percent of the Fortune 500. It captures workflows in real time and automatically generates the documentation for how your company works with no manual writing. It automatically redacts sensitive information like names and account numbers from screenshots. Plus, your team can launch real-time
Starting point is 00:12:13 on-screen guidance showing exactly where to click step-by-step inside the actual tool. And these captured workflows give your AI agents real context so they can operate on how work actually happens, not guesses. Learn more at Scribe.com. slash ride home and mention TechBrew Ride Home for your first month of Scribe capture free on select plans. That's S-C-R-I-B-E dot how slash ride home. YouTube is holding its big made for YouTube event and they made two big announces. Let's cover the announces for the creators first. Then we'll cover the announces for actual viewers.
Starting point is 00:12:53 YouTube wants creators to get into the microdrama business, quoting TechCrunch. As microdrama apps gain traction, YouTube is clearly betting that episodic short form storytelling will stick around and that the format belongs on shorts. At its made on YouTube event on Wednesday, the company announced Short Series, a new tool that lets creators organize their shorts into TV-style seasons and episodes. The feature brings a more traditional episodic structure to YouTube shorts, complete with custom thumbnails and playback. With shorts series, viewers can watch a selection of shorts as a continuous story rather than, simply scrolling from one standalone clip to the next. The series will live on a creator's YouTube channel.
Starting point is 00:13:37 The feature is rolling out to creators in the YouTube partner program across web, mobile, and connected TVs. The move comes as no surprise, given that YouTubers have already been experimenting with microdramas, and increasingly they're considering it a serious business venture, end quote. And quoting the verge. Part of the job of a content creator is to figure out how to get their work in front of the most people. Cracking or fighting, the algorithm has historic. been a frustration for creators, but YouTube is increasingly simply telling creators what they should do. At the annual creator focus made on YouTube event, the company announced updates to the suite of
Starting point is 00:14:12 AI-powered creator tools that unveiled in 2025. In its initial iteration, the dashboard of creator tools included things like A-B testing for video thumbnails and a chatbot interface to query how content was performing. This year's update adds AI tools for more complex tasks. Perhaps most significantly, YouTube is introducing an AI agent that will work in the background to optimize a creator's channel. For example, the tool will monitor a creator's back catalog of videos looking for content that has taken on a new life, is relevant to the news, or has started trending. The agent will offer suggestions for tweaking the thumbnails or titles for older videos to ride the new wave of relevance. The tool can also put together things like pitches for brands by coming through a creator's YouTube channel
Starting point is 00:14:56 and pulling out demographic and audience data to make a case to potential sponsors, says Amjad Hanif, vice president of creator products at YouTube. The new features iterate on previously existing AI-powered tools. The agent will now generate entire thumbnails and titles based on the content of an uploaded video, and creators can ask the tool for feedback on their own thumbnails. Viewers may also see different content as creators run tests via YouTube studio. A dynamic thumbnail tool allows creators to upload up to three different images for each video and YouTube system will assign the best one to different audience segments to try to boost
Starting point is 00:15:30 watch time. YouTube is taking it a step further to by allowing creators to test three different versions of videos, a clip with different intros and a hook, for example, or with different structure to see how each performs. The different versions of the same video will be fed to small audience segments to test which results in the highest watch time. Creators can go forward with the winner to make it the permanent video or YouTube will automatically do so after seven days. Hanif says the YouTube system ensures the three video versions are not dramatically different from one another, end quote. And say hello to custom feeds on YouTube, an LLM powered feature that lets users generate and save tailored homepage video feeds. Quoting Wired,
Starting point is 00:16:19 as part of its made on YouTube event Wednesday, Google released several new features, including a custom feeds button for adjusting what video. pop up on your front page. It's a rare chance for users to directly tweak their algorithmic feed on YouTube in ways that go beyond just pressing thumbs up on a video or blocking a creator. To try it out, users can click the Your Custom Feeds button below the search bar and type a request for the feed they'd like to see. YouTube uses large language models to process your request and whip up a freshly personalized home page. I want a rainy Shanghai vibe, cafe jazz, street food, and vlogs. Read one of my suggested YouTube tested custom feeds as an experimental feature earlier this year before deciding that it was a good fit.
Starting point is 00:17:03 The company plans to roll it out soon as an option for all U.S. users, whether viewing on mobile desktop or TV. Users are always telling us that they want more control, says Emily Moxley, YouTube's vice president of product management. They want more ability to steer their algorithm with custom feeds. We're giving them a natural language way that they can do that. They can tell us exactly what they want. So I told YouTube exactly what I wanted to see. Gouy Macaroni and Cheese recipe videos, Caleb Heron clip compilations,
Starting point is 00:17:34 and Crystal Singing Bowl videos under 15 minutes. And the feed it generated was exactly that. My perfect Zen blend for rotting on the couch after a long day. YouTube suggested potential additions to expand my prompt like uploaded this month and without talking or commentary. Not bad. The custom feeds button sits next to other topic. tabs below the search bar like music and podcasts.
Starting point is 00:17:59 Experimenting with this tool doesn't replace your standard homepage, which you can access anytime by clicking the all button. Users can save multiple custom feeds. This feature feels most helpful when blending multiple preferences. It's better than a standard YouTube search at responding dynamically to complex video requests, whether that's showing videos on multiple topics under a certain length or attempting to capture a specific vibe for your current mood. The product team at YouTube also tested a feature that lets users directly steer their overall recommendation algorithm, but ultimately decided against shipping it. We found users were nervous and afraid to do that, Moxley says. They were worried about messing it up.
Starting point is 00:18:36 The custom feeds button is designed to give users a low stakes place to experiment, end quote. Nothing more for you today. Talk to you tomorrow.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.