Tech Brew Ride Home - The Day Of All The Models
Episode Date: July 9, 2026SpaceXAI debuted Grok 4.5 with Cursor, targeting Opus-level performance at lower cost. Meta launched Muse Spark 1.1 via API, OpenAI rolled out full-duplex GPT-Live voice models, PrismML ran the larges...t AI model on an iPhone, and Character.AI launched AI microdramas. SpaceXAI debuts Grok 4.5, its first model built in partnership with Cursor, designed to "handle difficult, long-running tasks" across finance, legal, and coding (Bloomberg) Meta releases Muse Spark 1.1, capable of more advanced coding and a "step-change" from the first generation, available to US developers via a public API preview (The Verge) OpenAI launches GPT-Live, new voice models powering ChatGPT Voice and built on a full-duplex architecture, meaning they can listen and speak at the same time (OpenAI) OpenAI launches GPT-Live, new voice models powering ChatGPT Voice and built on a full-duplex architecture, meaning they can listen and speak at the same time (VentureBeat) PrismML says it ran a 27B-parameter Qwen 3.6 model on an iPhone 17 Pro, bigger than any prior on-device model; sources: Apple held talks with PrismML about it (The Information) Character.AI launches three human-written, AI-generated microdramas, whose characters users can chat with, and aims to eventually let users make their own shows (TechCrunch) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices
Transcript
Discussion (0)
This spring, denim gets a softer, lighter update.
Introducing Old Navy's drapey denim wide leg, a new fit that moves with you.
It's everything you want denim to feel like for summer.
Easy, breathable, and effortlessly cool.
With a fit that creates natural movement and a wide leg that feels modern, not overwhelming.
Plus, that signature, wait, for this price, moment.
Old Navy's drapey denim wide leg.
Welcome to the TechBoo ride home for Thursday, July 9th, 26.
I'm Brian McCullough today. SpaceX debuted Rock 4.5 with Cursor. Meta launched Muse Spark 1.1.
OpenAI rolled out full duplex GPT live voice models and GPT 5.6 is still coming later today.
PrismML ran the largest AI model on an iPhone and character AI launched AI microdramas.
Here's what you miss today in the world of tech.
Every day, shareholders meet to discuss important matters about the companies you invest in.
Now you can make your voice heard too.
Vanguard investor choice makes it easy to set your proxy voting preference for your eligible Vanguard index funds.
Whether you hold a Vanguard fund directly or through another brokerage firm, all it takes is a few clicks to select your proxy voting preference and be heard on important shareholder topics like executive pay and director elections.
Visit vanguard.com slash investor choice to learn more.
It's your shares. It's your voice.
It's easy.
Vanguard investors own shares of our index funds and those funds own shares of the companies they invest in.
Vanguard Marketing Corporation distributor.
We're in this weird place this week where OpenAI is going to release its new models probably
today, and everyone seems to be rushing to front-run that.
For example, SpaceX AI has debuted Grok 4.5, its first model built in partnership with
Cursor, designed to handle what it calls difficult, long-running tasks across finance, legal,
and coding.
Quoting Bloomberg, the software called Grok 4.5 marks the first joint AI model developed by the two
companies and comes just weeks after SpaceX formally agreed to acquire Cursor in a deal that
values the startup at $60 billion. The work with Cursor is part of a broader effort by Elon Musk's
company to catch up in the AI race and attract more business customers. Musk said earlier this year that
his AI startup, known as XAI, before it merged with SpaceX, had fallen behind on coding,
prompting a wave of staffing changes to rebuild the venture. SpaceX AI, as the AI outfit is now known,
released its first coding agent in May to compete with Anthropics offerings, end quote.
and quoting TechCrunch. In a blog post published Wednesday, SpaceX AI characterized its new
release as a workhorse that can tackle all of the typical tasks that the AI industry has sought
to automate, coding an app building, office and clerical work, research, writing, and other forms
of routine knowledge work. Grock can supposedly do all this for less spend too, as SpaceX AI says,
that its model has twice greater token efficiency than other leading models. If it carries through
to real-world use cases, that efficiency would be a big advantage for SpaceX AI since the cost of
tokens has been a growing concern for AI consumers. The company released benchmark metrics Wednesday
that appeared to show GROC's competitiveness with other top models from SpaceX AI competitors,
although just short of best in class. In a post on his social media platform X, founder Elon Musk,
compared the model to Opus, Anthropics' LLM designed for intensive and complex tasks.
based on strong positive feedback from customers in our beta test program at SpaceX AI will make
GROC 4.5 available to the public tomorrow. It is an Opus class model, but faster, more token
efficient, and lower cost, wrote Musk in his post on X. Musk later added, our internal assessment
is that GROC 4.5 is roughly comparable to Opus 4.7, but much faster. The combination of
capability, faster speed, and lower cost is what makes it competitive. SpaceX AI says its new model
cost $2 per million input tokens and $6 per million output tokens. That's quite competitive if
GROC's capabilities match SpaceX AI's rhetoric. Opus 4.7 by comparison costs $5 per million input
tokens and $25 per million output tokens. Open AI has tiered costs for different model versions.
Seoul, its most expensive, costs $5 for 1 million input tokens and $30 for 1 million output
tokens, while its least expensive, Luna costs $1 per 1 million input and $6.
per one million output tokens, end quote.
Okay, so Zuck wants to match that, quoting the verge.
After re-entering the AI race with its first in-house Muse Spark model in April,
meta is now opening up the doors to developers with a new model
that can plug into AI coding software with the new meta model API.
Meta says that Muse Spark 1.1 is a step change from the first generation
with improvements based on feedback from developers.
The company says it's capable of more advanced coding,
including detection and fixing of complex bugs,
better supports end-to-end agentic workflows across a range of apps,
including multi-agent systems,
and has native multimodal perception across images, videos, and documents.
The Muse Spark 1.1 launch follows this week's launch of Muse Image
and Image Generation model that's proved controversial
for its ability to incorporate other users' Instagram content into its generations.
It's part of Meta's race to just,
testify the billions it's spent on catching up in the AI competition and attempting to achieve
parity with companies like OpenAI, Google, and Anthropic, especially after a slew of
high-profile hires and a company restructuring last year.
The new 1.1 model is available now in thinking mode through the Meta-I app and Meta-I website.
It will also be accessible through a new Meta-Model API available from today in public preview
for U.S. developers.
Meta is including $20 worth of free credits with every new Meta-Model API account.
Muse Spark was initially only available directly through meta-a-I before eventually powering the chatbots inside Instagram and WhatsApp and the latest meta smart glasses, end quote.
Muse Spark 1.1 reportedly costs $1.25 per 1 million input tokens and $4.25 per 1 million output tokens, so less than even Grock.
Quoting Bloomberg, Mews Spark will be among the most affordable options on the market.
Zuckerberg said in an interview ahead of the release, since this is not an open source model, this is, I think, the first time that
we're doing a real serious API, Zuckerberg said, referring to the application programming interface
used to access meta's AI. And the pricing is going to be very aggressive and attractive, he said.
The new model's stand-on improvement is in its agentic capabilities, the meta-chief executive officer said.
Agents are the big theme of AI this year with the label applied to systems that can complete
multi-step tasks on behalf of a user. Zuckerberg described Muse Spark 1.1 as having state-of-the-art
or very close to it, agentic reasoning, and tool use. The model is also going to be.
greatly improved when it comes to coding and meta employees are using it internally to build
products and features for various apps, he added.
Meta will also introduce a new meta model API system, which will be used to collect fees
from developers.
Its API pricing is roughly 25% of the cost advertised by other top models from OpenAI
and Anthropic.
Developers will be able to use Meta's model for free, but only up to a point.
They'll be required to pay for access after reaching a certain token threshold, Zuckerberg
said.
The pricing from some of the other labs is very extreme.
and has very high margins, Zuckerberg said, underscoring that his strategy is to get meta's technology
in front of as many people as possible. We think that there's a real ability to be able to offer
frontier or very high-level intelligence at a much more affordable cost, end quote.
I wasn't able to get independent benchmarks to see how competitive this is with the frontier stuff,
but meta's own charts of benchmarks claim it is very competitive indeed.
But hey, why can't open AI pile on here too, even as they're
are still going to launch GPD 5.6 imminently, quoting Venture Beat. OpenAI on Wednesday launched
GPT Live, a pair of new voice models that fundamentally redesign how people talk to chat GPT,
replacing the company's existing advanced voice mode with an architecture that can listen and speak
simultaneously, much like an actual human conversation. The two models, GPT Live One and GPT Live1,
are rolling out globally starting today across iOS, Android, and chatGPT.com.
GPT Live1 becomes the default voice model for pay GPT users on the Go Plus and Pro tiers,
while GPT Live1 Mini serves free tier users.
OpenAI also plans to bring the models to the API,
and developers can sign up to be notified.
The release marks the third generation of ChatGPT's voice technology in roughly two years,
and OpenAI is clear as bid yet to turn its chatbot into something that feels less like
querying a search engine and more like talking to.
to a colleague. The defining technical advance in GPT Live is what OpenAI calls a full duplex architecture
in telecommunications. Full duplex means both parties on a phone call can talk and listen at the same
time. Applied to AI, it means the model continuously processes your incoming audio even while
it generates its own spoken response. No more waiting for a clean silence gap to figure out when
you've finished a thought. Instead of processing a sequence of separate messages, GPT Live
continuously processes input while generating output. OpenAI wrote in its research blog.
The model can therefore make interaction decisions many times a second, whether to speak,
continue listening, pause, interrupt, or invoke a tool. In practice, that translates to a
voice assistant that can insert conversational acknowledgments like, mm-hmm, yeah, got it,
while you're still talking, pick up on a natural pause while jumping in prematurely and handle rapid
interruptions without derailing the entire exchange. GPD Live introduces a second structural change,
that may prove just as consequential for enterprise adoption, though,
it decouples the voice interaction layer from the reasoning layer.
When a user asks a straightforward question, GPT Live handles it directly.
But when the query demands a web search, deeper reasoning,
or more complex agentic work, GPT Live delegates the task to a frontier model running in the background.
At launch, GPD 5.5, the large language model OpenAI released in April,
and continues talking with the user while the computation happens asynchronously.
While it works, GPD Live can keep talking with you
and maintain the flow of conversation. OpenAI explains, as we release new frontier models,
will continuously update the model used by GPT Live. This delegation model is a meaningful architectural
bet rather than building a single monolithic voice model that tries to be both conversationally
fluid and deeply intelligent. Open AI has split the problem in two, a voice native model
optimized for real-time interaction and a separate reasoning engine that can be swapped out as
the state-of-the-art improves. It is, in effect, a modular design, one that allows open AI to
upgrade the intelligence of its voice assistant without retraining the voice model itself.
The implications for enterprise and developer workflows are significant. A voice agent built on
this architecture could maintain a natural conversation with a customer while simultaneously
querying databases, searching the web, or performing multi-step reasoning, tasks that would have
introduced several seconds of dead air under the old pipeline, end quote.
Ever spent hours of your workday tracking down information only to find that it wasn't documented at all?
It's an age-old corporate conundrum.
Critical knowledge isn't available or accessible to the folks who need it.
That's the very problem our sponsor, Scribe, was built to fix.
Scribe is a workflow AI platform that automatically turns your workflows into clear-cut documentation.
All you do is turn on the Scribe browser extension or desktop app and go,
through your processes as you would normally. Scribe builds a guide as you go, capturing every
click, step, and screenshot automatically. So what would have taken hours gets done in under a
minute to see what Scribe could look like for your org. Head to Scribe.com, slash ride home. And
mention ride home for your first month of Scribe capture free on select plans. That's S-C-R-I-B-E.
dot how slash ride home.
When your company deploys a customer-facing AI that misses its mark, who do you think those
customers blame?
Well, according to the 2026 Delight AI Index, 83% of customers surveyed blamed the brand,
not the tech they were using.
That's why it's important to have tech you can trust.
Delight.a.i is a customer-facing AI concierge that delivers hyper-personalized experiences on
your behalf.
With zero-touch improvement, it can continuously monitor its own performance, find
failure patterns before they spread, write the fix, and ship it. It gets smarter with every
conversation automatically. The longer it runs, the more your customers can trust it. To learn more,
just head to delight.a.ai slash brew. That's delight.com. A.I. slash brew.
Meanwhile, in a roundabout way, could this news lead to future moves by Apple in the AI space?
Quoting the information. Apple is on a quest to shrink powerful AI models to run on iPhones,
which could cut down on cloud computing costs and enhance user privacy.
But a small startup that emerged from stealth mode earlier this year says it recently got an AI
model running on an iPhone bigger than any previous mobile model.
The startup, PrismML, said it has shrunk down Quen 3.6,
an open-source large-language model developed by Chinese internet giant Alibaba to run on an iPhone 17 Pro.
The model has 27 billion parameters, which are roughly similar to the synapses in a brain,
and can help determine the complexity of the data a model can process.
In contrast, most models that run on mobile phones have only a few billion parameters active at a time.
The largest AI models, which can measure in the trillions of parameters, are still far too big to run on mobile devices.
But the Model PRISMML has working on an iPhone is capable of tasks like complex chat,
reasoning, fully autonomous agents and software coding, the startup said.
The open source model will be available for download next week on Tuesday.
The milestone, which hasn't been previously reported, reflects a broader push to get AI running on devices instead of expensive high-powered servers in data centers.
Microsoft, Amazon, meta-platforms, and others are spending hundreds of billions of dollars racing to build those data centers to keep up with the demand they're anticipating for AI.
Apple, though, has largely stayed on the sidelines of the costly data center race while also being a vocal proponent of making sure as many of the iPhone's AI functions as possible run on the devices rather than in the cloud.
The company believes on-device AI will better allow it to deliver on its privacy and security promises to customers.
In an interview, Bebek Hasebe, CEO of PrismML, predicted that the vast majority of AI will eventually be processed on devices.
Imagine a world maybe three years from now where 95% of the intelligence that you need is available to you locally on your phone, on your laptop, on your appliances,
and it's really on the last maybe 5% of high-end stuff that you'll need to go to the cloud, Hasibi said.
I think that's how people are seeing the way forward. Shrinking models to run on devices, he added,
fundamentally changes the economics of AI. Prism ML uses a mathematical trick to shrink the Kuen 3.6 model
to a fraction of its original size. Shrinking models typically results in worse performance,
but the company claims its technique for miniaturizing AI model sizes doesn't hinder their performance.
Prism ML has compressed the size of Kwen 3.6 to less than 4 gigabytes down from around 54.
PRISMML plans to continue shrinking larger AI models even at the scale of a trillion
parameters which will bring it into the realm of cutting-edge models such as OpenAIs GPD and Anthropics
Claude, said Hasebe.
PrismML's approach may particularly appeal to Apple at the company's worldwide developers'
conference in June.
It announced its long-delayed Siri overhaul based on Google's Gemini models.
The most advanced parts of Siri are still so big that they require Apple to tap into
Nvidia chips running in Google Cloud.
Apple is currently on the hunt for acquisitions of companies that can help it run more AI on device.
The information previously reported, Apple has held meetings with Prison ML about ways it could use its technology for people familiar with the talks said, end quote.
And Character AI has launched three human-written AI-generated microdramas whose characters users can chat with and aims to eventually let users make their own shows.
Quoting TechCrunch, microdramas are such a rage these days,
that nearly every kind of company in the attention economy space,
be they dedicated microdrama apps, social media giants, TikTok, and Instagram,
or streaming services like Peacock, Amazon Prime, and India's Geo Hot Star,
is building a product to tap the opportunity.
Character AI, which lets people chat with customized AI avatars,
is also tapping this budding market by producing its own microdramas using AI characters.
But there's an interesting twist that takes advantage of the company's core product.
Users older than 18 can chat with the show's characters,
ask them questions and even roleplay different storylines.
The startup is launching three microdramas to start with a romance series dubbed last summer,
a horror show titled The Nighttime Game and a Hunger Games-like survival microdrama called Eden Fall.
Character AI says these dramas were created using AI production tools,
and in the long term, it aims to help users create their own characters and series.
This is the latest in a slew of recent features from the startup following its shift
toward entertainment-focused features last year.
In April, it teased a tool called The Lorraine.
book that users can employ to create world-building information that characters can reference
and launched another feature called books that lets users insert themselves into select classic
literature titles or roleplay as characters from them. The company said on Thursday that it is
also testing a feature dubbed C.A.fm that will let users put together audio series and
another that lets you create fiction called C.A.I. Reads. The audio series feature is
currently available to select users under its experimental C.A.A.
program, which the company says professional writers are using to create serialized audio dramas, end
quote.
The most exciting times I've seen in tech in my lifetime has been when the chessboard is
thrown up in the air and the pieces still haven't landed yet.
Think of the early 80s.
What do you use?
PC, Mac, other.
What software do you use?
Word, WordPerfect, Lotus Notes.
What browser do you use?
What search engine do you use?
Are you going iOS? Are you going Android? Maybe Palm Pre.
Times when people were just, you know, trying things out.
Switching by the day, we're in a moment like that right now, I think.
I downloaded Codex on the plane ride home last night,
and I'm more than happy to jump over there if GPT 5.6 proves to be better.
And it's so funny how granularly these things can be in terms of being better or worse.
Shifting back and forth between Opus and Fable is like switching,
conversations between a drunk at a bar and a no-nonsense junior executive. If you use them
occasionally, you can't tell the difference, but if you use them every day, it's weird how
varied these things are at the moment. Talk to you tomorrow.
