Tech Brew Ride Home - The Day Of All The Models

Episode Date: July 9, 2026

SpaceXAI debuted Grok 4.5 with Cursor, targeting Opus-level performance at lower cost. Meta launched Muse Spark 1.1 via API, OpenAI rolled out full-duplex GPT-Live voice models, PrismML ran the larges...t AI model on an iPhone, and Character.AI launched AI microdramas. SpaceXAI debuts Grok 4.5, its first model built in partnership with Cursor, designed to "handle difficult, long-running tasks" across finance, legal, and coding (Bloomberg) Meta releases Muse Spark 1.1, capable of more advanced coding and a "step-change" from the first generation, available to US developers via a public API preview (The Verge) OpenAI launches GPT-Live, new voice models powering ChatGPT Voice and built on a full-duplex architecture, meaning they can listen and speak at the same time (OpenAI) OpenAI launches GPT-Live, new voice models powering ChatGPT Voice and built on a full-duplex architecture, meaning they can listen and speak at the same time (VentureBeat) PrismML says it ran a 27B-parameter Qwen 3.6 model on an iPhone 17 Pro, bigger than any prior on-device model; sources: Apple held talks with PrismML about it (The Information) Character.AI launches three human-written, AI-generated microdramas, whose characters users can chat with, and aims to eventually let users make their own shows (TechCrunch) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices

Transcript
Discussion (0)
Starting point is 00:00:00 This spring, denim gets a softer, lighter update. Introducing Old Navy's drapey denim wide leg, a new fit that moves with you. It's everything you want denim to feel like for summer. Easy, breathable, and effortlessly cool. With a fit that creates natural movement and a wide leg that feels modern, not overwhelming. Plus, that signature, wait, for this price, moment. Old Navy's drapey denim wide leg. Welcome to the TechBoo ride home for Thursday, July 9th, 26.
Starting point is 00:00:34 I'm Brian McCullough today. SpaceX debuted Rock 4.5 with Cursor. Meta launched Muse Spark 1.1. OpenAI rolled out full duplex GPT live voice models and GPT 5.6 is still coming later today. PrismML ran the largest AI model on an iPhone and character AI launched AI microdramas. Here's what you miss today in the world of tech. Every day, shareholders meet to discuss important matters about the companies you invest in. Now you can make your voice heard too. Vanguard investor choice makes it easy to set your proxy voting preference for your eligible Vanguard index funds. Whether you hold a Vanguard fund directly or through another brokerage firm, all it takes is a few clicks to select your proxy voting preference and be heard on important shareholder topics like executive pay and director elections.
Starting point is 00:01:24 Visit vanguard.com slash investor choice to learn more. It's your shares. It's your voice. It's easy. Vanguard investors own shares of our index funds and those funds own shares of the companies they invest in. Vanguard Marketing Corporation distributor. We're in this weird place this week where OpenAI is going to release its new models probably today, and everyone seems to be rushing to front-run that. For example, SpaceX AI has debuted Grok 4.5, its first model built in partnership with
Starting point is 00:01:55 Cursor, designed to handle what it calls difficult, long-running tasks across finance, legal, and coding. Quoting Bloomberg, the software called Grok 4.5 marks the first joint AI model developed by the two companies and comes just weeks after SpaceX formally agreed to acquire Cursor in a deal that values the startup at $60 billion. The work with Cursor is part of a broader effort by Elon Musk's company to catch up in the AI race and attract more business customers. Musk said earlier this year that his AI startup, known as XAI, before it merged with SpaceX, had fallen behind on coding, prompting a wave of staffing changes to rebuild the venture. SpaceX AI, as the AI outfit is now known,
Starting point is 00:02:32 released its first coding agent in May to compete with Anthropics offerings, end quote. and quoting TechCrunch. In a blog post published Wednesday, SpaceX AI characterized its new release as a workhorse that can tackle all of the typical tasks that the AI industry has sought to automate, coding an app building, office and clerical work, research, writing, and other forms of routine knowledge work. Grock can supposedly do all this for less spend too, as SpaceX AI says, that its model has twice greater token efficiency than other leading models. If it carries through to real-world use cases, that efficiency would be a big advantage for SpaceX AI since the cost of tokens has been a growing concern for AI consumers. The company released benchmark metrics Wednesday
Starting point is 00:03:14 that appeared to show GROC's competitiveness with other top models from SpaceX AI competitors, although just short of best in class. In a post on his social media platform X, founder Elon Musk, compared the model to Opus, Anthropics' LLM designed for intensive and complex tasks. based on strong positive feedback from customers in our beta test program at SpaceX AI will make GROC 4.5 available to the public tomorrow. It is an Opus class model, but faster, more token efficient, and lower cost, wrote Musk in his post on X. Musk later added, our internal assessment is that GROC 4.5 is roughly comparable to Opus 4.7, but much faster. The combination of capability, faster speed, and lower cost is what makes it competitive. SpaceX AI says its new model
Starting point is 00:03:58 cost $2 per million input tokens and $6 per million output tokens. That's quite competitive if GROC's capabilities match SpaceX AI's rhetoric. Opus 4.7 by comparison costs $5 per million input tokens and $25 per million output tokens. Open AI has tiered costs for different model versions. Seoul, its most expensive, costs $5 for 1 million input tokens and $30 for 1 million output tokens, while its least expensive, Luna costs $1 per 1 million input and $6. per one million output tokens, end quote. Okay, so Zuck wants to match that, quoting the verge. After re-entering the AI race with its first in-house Muse Spark model in April,
Starting point is 00:04:46 meta is now opening up the doors to developers with a new model that can plug into AI coding software with the new meta model API. Meta says that Muse Spark 1.1 is a step change from the first generation with improvements based on feedback from developers. The company says it's capable of more advanced coding, including detection and fixing of complex bugs, better supports end-to-end agentic workflows across a range of apps, including multi-agent systems,
Starting point is 00:05:12 and has native multimodal perception across images, videos, and documents. The Muse Spark 1.1 launch follows this week's launch of Muse Image and Image Generation model that's proved controversial for its ability to incorporate other users' Instagram content into its generations. It's part of Meta's race to just, testify the billions it's spent on catching up in the AI competition and attempting to achieve parity with companies like OpenAI, Google, and Anthropic, especially after a slew of high-profile hires and a company restructuring last year.
Starting point is 00:05:40 The new 1.1 model is available now in thinking mode through the Meta-I app and Meta-I website. It will also be accessible through a new Meta-Model API available from today in public preview for U.S. developers. Meta is including $20 worth of free credits with every new Meta-Model API account. Muse Spark was initially only available directly through meta-a-I before eventually powering the chatbots inside Instagram and WhatsApp and the latest meta smart glasses, end quote. Muse Spark 1.1 reportedly costs $1.25 per 1 million input tokens and $4.25 per 1 million output tokens, so less than even Grock. Quoting Bloomberg, Mews Spark will be among the most affordable options on the market. Zuckerberg said in an interview ahead of the release, since this is not an open source model, this is, I think, the first time that
Starting point is 00:06:28 we're doing a real serious API, Zuckerberg said, referring to the application programming interface used to access meta's AI. And the pricing is going to be very aggressive and attractive, he said. The new model's stand-on improvement is in its agentic capabilities, the meta-chief executive officer said. Agents are the big theme of AI this year with the label applied to systems that can complete multi-step tasks on behalf of a user. Zuckerberg described Muse Spark 1.1 as having state-of-the-art or very close to it, agentic reasoning, and tool use. The model is also going to be. greatly improved when it comes to coding and meta employees are using it internally to build products and features for various apps, he added.
Starting point is 00:07:05 Meta will also introduce a new meta model API system, which will be used to collect fees from developers. Its API pricing is roughly 25% of the cost advertised by other top models from OpenAI and Anthropic. Developers will be able to use Meta's model for free, but only up to a point. They'll be required to pay for access after reaching a certain token threshold, Zuckerberg said. The pricing from some of the other labs is very extreme.
Starting point is 00:07:28 and has very high margins, Zuckerberg said, underscoring that his strategy is to get meta's technology in front of as many people as possible. We think that there's a real ability to be able to offer frontier or very high-level intelligence at a much more affordable cost, end quote. I wasn't able to get independent benchmarks to see how competitive this is with the frontier stuff, but meta's own charts of benchmarks claim it is very competitive indeed. But hey, why can't open AI pile on here too, even as they're are still going to launch GPD 5.6 imminently, quoting Venture Beat. OpenAI on Wednesday launched GPT Live, a pair of new voice models that fundamentally redesign how people talk to chat GPT,
Starting point is 00:08:14 replacing the company's existing advanced voice mode with an architecture that can listen and speak simultaneously, much like an actual human conversation. The two models, GPT Live One and GPT Live1, are rolling out globally starting today across iOS, Android, and chatGPT.com. GPT Live1 becomes the default voice model for pay GPT users on the Go Plus and Pro tiers, while GPT Live1 Mini serves free tier users. OpenAI also plans to bring the models to the API, and developers can sign up to be notified. The release marks the third generation of ChatGPT's voice technology in roughly two years,
Starting point is 00:08:49 and OpenAI is clear as bid yet to turn its chatbot into something that feels less like querying a search engine and more like talking to. to a colleague. The defining technical advance in GPT Live is what OpenAI calls a full duplex architecture in telecommunications. Full duplex means both parties on a phone call can talk and listen at the same time. Applied to AI, it means the model continuously processes your incoming audio even while it generates its own spoken response. No more waiting for a clean silence gap to figure out when you've finished a thought. Instead of processing a sequence of separate messages, GPT Live continuously processes input while generating output. OpenAI wrote in its research blog.
Starting point is 00:09:28 The model can therefore make interaction decisions many times a second, whether to speak, continue listening, pause, interrupt, or invoke a tool. In practice, that translates to a voice assistant that can insert conversational acknowledgments like, mm-hmm, yeah, got it, while you're still talking, pick up on a natural pause while jumping in prematurely and handle rapid interruptions without derailing the entire exchange. GPD Live introduces a second structural change, that may prove just as consequential for enterprise adoption, though, it decouples the voice interaction layer from the reasoning layer. When a user asks a straightforward question, GPT Live handles it directly.
Starting point is 00:10:03 But when the query demands a web search, deeper reasoning, or more complex agentic work, GPT Live delegates the task to a frontier model running in the background. At launch, GPD 5.5, the large language model OpenAI released in April, and continues talking with the user while the computation happens asynchronously. While it works, GPD Live can keep talking with you and maintain the flow of conversation. OpenAI explains, as we release new frontier models, will continuously update the model used by GPT Live. This delegation model is a meaningful architectural bet rather than building a single monolithic voice model that tries to be both conversationally
Starting point is 00:10:37 fluid and deeply intelligent. Open AI has split the problem in two, a voice native model optimized for real-time interaction and a separate reasoning engine that can be swapped out as the state-of-the-art improves. It is, in effect, a modular design, one that allows open AI to upgrade the intelligence of its voice assistant without retraining the voice model itself. The implications for enterprise and developer workflows are significant. A voice agent built on this architecture could maintain a natural conversation with a customer while simultaneously querying databases, searching the web, or performing multi-step reasoning, tasks that would have introduced several seconds of dead air under the old pipeline, end quote.
Starting point is 00:11:13 Ever spent hours of your workday tracking down information only to find that it wasn't documented at all? It's an age-old corporate conundrum. Critical knowledge isn't available or accessible to the folks who need it. That's the very problem our sponsor, Scribe, was built to fix. Scribe is a workflow AI platform that automatically turns your workflows into clear-cut documentation. All you do is turn on the Scribe browser extension or desktop app and go, through your processes as you would normally. Scribe builds a guide as you go, capturing every click, step, and screenshot automatically. So what would have taken hours gets done in under a
Starting point is 00:11:56 minute to see what Scribe could look like for your org. Head to Scribe.com, slash ride home. And mention ride home for your first month of Scribe capture free on select plans. That's S-C-R-I-B-E. dot how slash ride home. When your company deploys a customer-facing AI that misses its mark, who do you think those customers blame? Well, according to the 2026 Delight AI Index, 83% of customers surveyed blamed the brand, not the tech they were using. That's why it's important to have tech you can trust.
Starting point is 00:12:32 Delight.a.i is a customer-facing AI concierge that delivers hyper-personalized experiences on your behalf. With zero-touch improvement, it can continuously monitor its own performance, find failure patterns before they spread, write the fix, and ship it. It gets smarter with every conversation automatically. The longer it runs, the more your customers can trust it. To learn more, just head to delight.a.ai slash brew. That's delight.com. A.I. slash brew. Meanwhile, in a roundabout way, could this news lead to future moves by Apple in the AI space? Quoting the information. Apple is on a quest to shrink powerful AI models to run on iPhones,
Starting point is 00:13:11 which could cut down on cloud computing costs and enhance user privacy. But a small startup that emerged from stealth mode earlier this year says it recently got an AI model running on an iPhone bigger than any previous mobile model. The startup, PrismML, said it has shrunk down Quen 3.6, an open-source large-language model developed by Chinese internet giant Alibaba to run on an iPhone 17 Pro. The model has 27 billion parameters, which are roughly similar to the synapses in a brain, and can help determine the complexity of the data a model can process. In contrast, most models that run on mobile phones have only a few billion parameters active at a time.
Starting point is 00:13:49 The largest AI models, which can measure in the trillions of parameters, are still far too big to run on mobile devices. But the Model PRISMML has working on an iPhone is capable of tasks like complex chat, reasoning, fully autonomous agents and software coding, the startup said. The open source model will be available for download next week on Tuesday. The milestone, which hasn't been previously reported, reflects a broader push to get AI running on devices instead of expensive high-powered servers in data centers. Microsoft, Amazon, meta-platforms, and others are spending hundreds of billions of dollars racing to build those data centers to keep up with the demand they're anticipating for AI. Apple, though, has largely stayed on the sidelines of the costly data center race while also being a vocal proponent of making sure as many of the iPhone's AI functions as possible run on the devices rather than in the cloud. The company believes on-device AI will better allow it to deliver on its privacy and security promises to customers.
Starting point is 00:14:42 In an interview, Bebek Hasebe, CEO of PrismML, predicted that the vast majority of AI will eventually be processed on devices. Imagine a world maybe three years from now where 95% of the intelligence that you need is available to you locally on your phone, on your laptop, on your appliances, and it's really on the last maybe 5% of high-end stuff that you'll need to go to the cloud, Hasibi said. I think that's how people are seeing the way forward. Shrinking models to run on devices, he added, fundamentally changes the economics of AI. Prism ML uses a mathematical trick to shrink the Kuen 3.6 model to a fraction of its original size. Shrinking models typically results in worse performance, but the company claims its technique for miniaturizing AI model sizes doesn't hinder their performance. Prism ML has compressed the size of Kwen 3.6 to less than 4 gigabytes down from around 54.
Starting point is 00:15:32 PRISMML plans to continue shrinking larger AI models even at the scale of a trillion parameters which will bring it into the realm of cutting-edge models such as OpenAIs GPD and Anthropics Claude, said Hasebe. PrismML's approach may particularly appeal to Apple at the company's worldwide developers' conference in June. It announced its long-delayed Siri overhaul based on Google's Gemini models. The most advanced parts of Siri are still so big that they require Apple to tap into Nvidia chips running in Google Cloud.
Starting point is 00:15:59 Apple is currently on the hunt for acquisitions of companies that can help it run more AI on device. The information previously reported, Apple has held meetings with Prison ML about ways it could use its technology for people familiar with the talks said, end quote. And Character AI has launched three human-written AI-generated microdramas whose characters users can chat with and aims to eventually let users make their own shows. Quoting TechCrunch, microdramas are such a rage these days, that nearly every kind of company in the attention economy space, be they dedicated microdrama apps, social media giants, TikTok, and Instagram, or streaming services like Peacock, Amazon Prime, and India's Geo Hot Star, is building a product to tap the opportunity.
Starting point is 00:16:48 Character AI, which lets people chat with customized AI avatars, is also tapping this budding market by producing its own microdramas using AI characters. But there's an interesting twist that takes advantage of the company's core product. Users older than 18 can chat with the show's characters, ask them questions and even roleplay different storylines. The startup is launching three microdramas to start with a romance series dubbed last summer, a horror show titled The Nighttime Game and a Hunger Games-like survival microdrama called Eden Fall. Character AI says these dramas were created using AI production tools,
Starting point is 00:17:20 and in the long term, it aims to help users create their own characters and series. This is the latest in a slew of recent features from the startup following its shift toward entertainment-focused features last year. In April, it teased a tool called The Lorraine. book that users can employ to create world-building information that characters can reference and launched another feature called books that lets users insert themselves into select classic literature titles or roleplay as characters from them. The company said on Thursday that it is also testing a feature dubbed C.A.fm that will let users put together audio series and
Starting point is 00:17:52 another that lets you create fiction called C.A.I. Reads. The audio series feature is currently available to select users under its experimental C.A.A. program, which the company says professional writers are using to create serialized audio dramas, end quote. The most exciting times I've seen in tech in my lifetime has been when the chessboard is thrown up in the air and the pieces still haven't landed yet. Think of the early 80s. What do you use?
Starting point is 00:18:30 PC, Mac, other. What software do you use? Word, WordPerfect, Lotus Notes. What browser do you use? What search engine do you use? Are you going iOS? Are you going Android? Maybe Palm Pre. Times when people were just, you know, trying things out. Switching by the day, we're in a moment like that right now, I think.
Starting point is 00:18:50 I downloaded Codex on the plane ride home last night, and I'm more than happy to jump over there if GPT 5.6 proves to be better. And it's so funny how granularly these things can be in terms of being better or worse. Shifting back and forth between Opus and Fable is like switching, conversations between a drunk at a bar and a no-nonsense junior executive. If you use them occasionally, you can't tell the difference, but if you use them every day, it's weird how varied these things are at the moment. Talk to you tomorrow.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.