The AI Daily Brief: Artificial Intelligence News and Analysis - What to Use the Latest AI Tools For
Episode Date: September 11, 2026GPT‑Live 1 opens up a much wider range of real-time voice and vision applications, from customer service and sales to education, healthcare, and hands-free work. NLW breaks down who should be using ...the latest AI tools and the specific jobs each one is best suited to perform. In the headlines: ChatGPT for Financial Services, Cognition’s SWE‑2 model, and Projects in Cursor.Multiplayer AI Sprint - https://multiplayerai.ai/Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at https://kpmg.com/us/SophisticatedHarbor - Invest in the AI ecosystem. https://www.harborcapital.com/aidailyHyperagent - Hire a team of always-on agents. New users get $100 in free credits. hyperagent.com/aidailybriefRackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack https://www.rackspace.com/Section - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/Robots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/The AI Daily Brief helps you understand the most important news and discussions in AI. Newsletter: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? sponsors@aidailybrief.ai
Transcript
Discussion (0)
This week had so many new AI and model product releases that I've had to do not one,
but two episodes just to capture everything that's been going on.
What's interesting about this second set is that many of them show some pretty distinct
and important trends about where we're headed.
For example, it is very clear that we are going to be managing more and more of our interactions
with computers via our voice.
It won't happen all at once, but now that we have things like ChatGPT Live available to
developers to build around with a significantly
increased capability set, better ability to distinguish who's talking, background noise,
you're just going to see more and more applications that involve voice.
Another example of a trend shown off in these announcements is the continued push towards
complex model architectures where people can optimize for cost and good enough capability,
as opposed to just always seeking the highest capability.
The point is that sometimes new product releases are about what you can do with them,
and sometimes they're about what they say about what we're all going to be doing soon.
The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, quick announcements before we dive in.
First of all, thank you to today's sponsors, KPMG, Blitzy, robots and pencils, and hyperagent.
For an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts.
And if you want to learn more about sponsoring the show, send us a note at sponsors at aidailydlybreef.a.
Welcome back to the AI Daily Brief Headlines edition.
all the daily AI news you need in around five minutes.
We kick off today with a new report from Anthropic that details their efforts to detect and
counter AI misuse.
There were a lot of juicy nuggets in this thing, and were it not for the larger AI safety
conversation happening this week, it probably would have been a bigger topic of conversation.
Going through some of the highlights, Anthropics said that they had disrupted major distillation
attacks from Alibaba, Deep Seek, and Xiaomi. Each company used networks of fraudulent accounts
to extract reasoning traces from Anthropic models for use's training data.
However, Anthropic also claimed that they had detected both DeepSeek and Kimmy K3 Meeker Moonshot,
routing requests to Claude and serving the responses to their users.
Anthropic wrote,
In one instance, over a 10-day period, Moonshot relayed almost 300,000 customer request to Anthropic,
the vast majority of which were routed to Opus.
Based on the report, it doesn't seem like Moonshot was using this method to spoof their models,
using an Anthropic backend to make them appear more capable.
Instead, it seems like this was a way to gather realistic user queries for their distillation pipeline.
Still, Anthropic notes that this method exposed sensitive information
from Chinese government and corporate users on multiple occasions.
In the biology section, Anthropic discussed an incident in May
when their classifier blocked work on a grant application for scientific funding.
The application discussed gain of function research on a mosquito-borne virus called Chick-Gunia,
Anthropic wrote,
Because chicken gunya circulates naturally, a deliberate release as part of a bioweapon would be
difficult to distinguish from a natural outbreak. The grant sought to identify enhancing mutations
of the chicken gunya virus, engineer them into infectious clones, and select for virulence in vivo.
In other words, the virus would become progressively more harmful as it repeatedly infected live
animals, with researchers keeping the most disease-causing variants in each round. Anthropic noted
that this research could have innocent applications in vaccine development, but they were
concerned it might have been malicious because research was intended to be conducted at a military
research institute. The report did not identify which nation was sponsoring the research, but did note
the AI prompts continued through gray market resellers after the researchers were banned.
The report also contained dozens of other examples of blocked misuse, including Russian actors
using Claude for espionage and propaganda, Chinese, Russian, and humaneing actors attempting to write
software for conventional weaponry and a host of attempted cyber attacks and financial scams.
Anthropic noted the report was put together with intention.
writing, the cases we share here aren't typical misuse, but rather examples of the most notable
and novel threat activity we've identified to date. We're publishing this work because we believe
we have a responsibility to disclose malicious misuse of our services. As models become increasingly
capable, their risks will increase unless AI developers and society's defenders act to make them
safer. Notably, the report was entirely focused on misuse of haiku, sonnet, and opus models,
with only a single case of distillation related to fable and mythos.
Now, speaking of the changing wins around AI risk, Sam Altman has told staff he's open to an
AI slowdown. During an all-hands meeting this week, Altman told his team that OpenAI could pace
the development of new AI models. However, he said this would likely happen in coordination with
other AI labs and expressed concerns that not all rivals would agree. These comments come after a
wave of discussion about slowing the pace of AI development that started even before Jacob Coxon's
viral tweet earlier this week. Many Open AI staffers signed on to the pacing the frontier
open letter in July, and chief scientist Jacob Pachaki recently warned that recursive self-improvement
amplified the case for a slowdown. In a blog post published last weekend, Pachaki wrote that he hopes
for industry coordination and, quote, voluntary slowdowns to become commonplace until shared
safety bars are established. This appears to be the first time that Altman has indicated a
willingness to slow down. Unlike his counterpart, Dario Amadehap, at Anthropic, he did not sign the pacing
the frontier open letter and had been non-committal in public statements until very recently. On Tuesday,
remarking on OpenAI finding a solution to a Millennium Prize problem, Altman wrote,
The world has extremely capable models now. I did not expect a result of this magnitude to happen so
soon. We've been talking a lot about the need to pace progress to ensure safety. For me, this is the
strongest evidence yet of the urgency. Alongside the addition of AI safety researcher Paul
Cristiano to the board of OpenAI earlier this week, Altman's comments could indicate a coordinated
industry slowdown is becoming more likely. Now, also in OpenAI news, the company appears to
hit a compute wall and will pause top-end subscriptions as a result. We talked about this back when
Anthropic was having all of its trouble, that this was going to be a problem coming for everyone,
and it is apparently now here. Now, shortly after the release of GPT6 Astra, OpenAI product leader
Tibo Sotiao posted, demand for Astra is really unprecedented. We're pulling all the levers
possible to sustain the demand, but I've not seen anything like it until now, and we went
through very steep growth before. Priority will always be to keep excellent service for existing users,
might have to pause new pro subscriptions for a bit if this continues. Sam Altman reposted that
saying, this would suck, but we will prioritize great service for customers until we can get back on
top of things. Well, it seems like they have now run out of levers to pull with Tebow posting
on Thursday. To make sure our current users have an incredible experience and continued access to Astra,
we are going to pause subscriptions to our $200 pro plan. These put the most strain on our systems,
and we wanted to take the smallest step that allows us to continue giving the broadest access possible,
All other plans and the API remain available.
There is no impact to existing accounts,
and we are working on adding more capacity as fast as we can.
Stating it bluntly, Matt Schumer wrote,
The era of subsidized tokens is ending, prepare accordingly.
And Alex Beresh thinks,
The $200 plan will not be returning.
We will get a post talking about right-size-fit,
and they will tell us there was a large split
between people on the $200 plan,
where most would have been better off on the $100 plan,
and a small percentage of users that were constantly maxing out limits.
Now, speaking of compute, Microsoft plans to triple their data center capacity after battling
their own compute shortages over recent years.
Sources told Bloomberg that Microsoft has turbocharged their AI buildout strategy under a new
plan.
They now intend to have 38 gigawatts of global capacity by 2032, up from their current 12 gigawatts.
Only 2 gigawatts of their current fleet is dedicated to AI compute, but they will now aim to
grow the share to a third of total capacity.
In addition, the plan will see them add CPUs to power agentic workflows.
Importantly, throughout the AI buildout, Microsoft has been the most conservative of the hyperscalers.
They famously scaled back plans in early 2025, canceling some large-scale leases, which contributed to a major market pullback.
But over the past year, Microsoft has repeatedly said that compute constraints are limiting growth and causing them to turn away customers.
Microsoft's new plan reinforces that AI infrastructure is nowhere near overbuilt, and the companies driving the buildout expect at least five more years of elevated KAPX spend before they can meet demand.
Speaking of the buildout, continuing a pace,
NVIDIA CEO Jensen Huang has doubled down on forecasts of 70% growth at a Goldman Sachs conference.
Speaking on Thursday, Jensen said that growth is limited only by the supply chain,
commenting,
even though our demand is much greater than 70%,
our supply allows us to confidently deliver 70%.
Huang is never shy about self-promotion,
but this is the first time he's discussed forward estimates over the past year,
suggesting an added layer of confidence.
Before the crowd of investors, Huang explained that many of them haven't kept up
with how NVIDIA's product has evolved over recent years.
He said, most people think NVIDIA builds a chip.
I mean, you need airplanes to ship what we build.
One GPU now is not $399.
It's $8.5 million.
That's one GPU, all connected with NVLink, 2 million parts, right?
250,000 kilowatts.
That's a GPU, and we ship thousands of them.
Huang added that orders of their rack scale systems
are currently growing, unastonishing 27% month over month.
And overall, the core idea that Huang tried to communicate
was that NVIDIA is still at the center of the center
of the AI buildout, despite headlines about new competition every other week.
Still, one new thing they have to contend with is the fact that the DOJ is investigating their
$20 billion-dollar GROC deal to determine whether it was structured to circumvent antitrust
laws. When the deal was announced in December, chipmaking startup GROC framed it as a non-exclusive
licensing agreement that gave NVIDIA access to their technology for integration into future
products. As part of the deal, CEO John Ross and C.O. Suni Madra also joined NVIDIA.
The New York Times reports that the DOJ has been investigating the deal since shortly after it was
announced.
The deal was one of many not acquisitions in the AI industry over recent years,
including Google's deals with Character AI and WinSurf, as well as Meta's deal with
scale AI. By structuring the deals as non-exclusive licensing agreements, they don't trigger
automatic review by the FTC and can be blocked in the same way as traditional mergers and acquisitions.
Still, that structure came under fire in February, with a number of senators calling on the
FTC and the DOJ to investigate this particular issue. In a letter, the senators wrote that the deals,
quote, function as de facto mergers, allowing the companies to consolidate talent,
information and resources, all while apparently attempting to bypass the scrutiny typically
applied to mergers and acquisitions. Now, at this stage, it is just an investigation and it's
unclear whether anything will come of it. But when they still started happening, it was always
inevitable that this sort of investigation was going to happen. Lastly today, Meta's new personal
agent muse seems to be finding some decent traction with consumers. And beyond that, it is certainly
a big hit on Wall Street. New data from Censor Tower shows the agent was downloaded 83,000 times on
launch day in the iOS App Store for the U.S. Now, that's not particularly impressive for a new
meta product. Threads achieved 4.3 million downloads on launch day, while the meta AI app reached
108,000 downloads. Still, the debut was good enough to push Muse to second place in the app store.
However, on Wall Street, Muse is being viewed as evidence that meta can still compete in AI. The stock
jumps 6% on release day and is largely holding onto the gains. On Thursday, JP Morgan analysts
upgraded the stock to a buy viewing Muse as a sign of more to come.
In a research note, they wrote, Meta is well positioned to deliver consumer-driven AI products
to its base of around 4 billion users, and that scale distribution is a significant competitive
advantage.
There's still meaningful upside potential as meta is in the early stages of releasing frontier
models and AI-driven products beyond advertising.
Good progress for what is an important new product category, but for now, that is going
to do it for the headlines.
Next up, the main episode.
A new study from KPMG in the University of Texas at Austin found that when people
work with AI, similar skills don't guarantee similar outcomes. Researchers studied more than 500
early career professionals and found that the best performers consistently amplified the value of
AI by guiding, evaluating, and refining its outputs. These top performers, called AI amplifiers,
weren't defined by what they knew alone, but by how they worked with AI. Learn more about what
separates AI amplifiers from everyone else at KPMG.com slash US slash AI amplifiers.
Blitzie's understanding of massive codebases unlocks autonomous security fixes, modernization, and new features.
So what happens when there's no legacy code at all? Greenfield is supposed to be the easy part.
Clean slate, no technical debt. But even Greenfield moves at human speed one sprint at a time.
Blitzie changes the unit of work from the developer to the project, autonomously planning, building, testing, and validating entire applications from scratch.
Hundreds of thousands of lines of production ready code. One Blitzy customer stood up a brand new application, 534,000,
lines of code, compressing a 65-week roadmap into two weeks. Another shipped an entire application
with no front-end engineer. Legacy or Greenfield, the answer is the same. Software at the speed of
compute. Build what's next at blitzie.com. That's BLITZY.com. At this point, it's no longer a question
of whether companies are actively using AI. Using it well, on the other hand, is a whole different
story. Robots and pencils, though, is a company that I can point to that is actually built for this time.
They're an applied AI engineering firm working directly with clients on problems that matter to the business,
not experiments that live in a slide deck.
Every engagement starts by working backwards from the outcome a client actually needs.
If you're trying to tell real AI engineering apart from noise in this space, that's the difference maker.
Head to Robots &Pensals.com.
This episode of the AI Daily Brief is brought to you by Hyperagent, where you run fleets of agents your team can manage together.
Forget local agents and chat workflows waiting on your laptop to be prompted.
Hyperagent deploys always on agents.
agents in the cloud, doing real work across the tools your team already uses.
Marketing agents turn competitor moves into landing pages.
Sales agents enrich leads, draft emails, and updates the CRM.
Ops agent chases the paperwork and tracks the budget.
Every agent has access to shared context and follows your rules about scope and approvals.
It's time you add agents that feel like teammates.
Hire yours at Hyperagent.
Get $100 in credits at hyperagent.com slash AI Daily Brief.
Welcome back to the AI Daily Brief.
Today we are back to the practical.
because even as the latest round of the AI risk conversation has exploded into the mainstream,
new products that are useful for you right now, and which basically no one is arguing,
are actively dangerous, have continued to come out and create new opportunities.
And so today we're going to talk through what to use the latest AI tools for.
I'll be going through eight or nine product launches from the last couple of days with quick notes
on what the launch is, who should be considering using it, and what they should be thinking about
using them to do.
First up, OpenAI has released GPT Live 1 in the API.
This is the voice model that OpenAI released back in July with that memorable grandma's video.
This is a live voice model, meaning it can talk and listen at the same time without needing to take turns.
It boasts better handling of noisy environments and interruptions, as well as architecture that hands off tasks to a back-end reasoning model.
This allows the model to complete tasks in the background while still carrying on the conversation.
The API release came with another nifty launch video demonstrating why this is a new thing.
model could be useful for developers. OpenAI showed off the model's ability to detect voice,
even with a basketball being dribbled a few feet away from the user. It also shows the model
can control a robot and a display while carrying on the conversation demonstrating multiple
simultaneous tool use. The model is now available for five cents per minute for the audio model
and standard API pricing for the backend model on top. Wrote OpenAIs Tibo, you can now build on
top of the same voice system that we shipped to a billion users in chat GPT. So many fun applications
of a full duplex system with solid tool calling.
Going back to text-only experiences after this
feels harder than I had expected.
AI experimenter Alex Finn writes,
This is actually the release I've been waiting for for a long time,
a hacker's dream come true.
You can now build the Johnny I of OpenAI device
before they even release it.
Going to be buying a small dev screen slash mic on Amazon.
Plug it into my computer, tell Astro to hack it to put Live One API on it.
Now anywhere you go, you can just talk to the device
and have it do work for you.
Basically, chat GPT's voice everywhere.
you go. Give it commands, it starts doing work on your home computer. You can basically now build
your own personal assistant. Tons of cheap ways to do this on Amazon with small dev devices, just search for
them. Things like this just make AI fun, and I don't think we focus enough on having fun.
So when it comes to who should be using this product, in a lot of ways, this is just going to be
a strict upgrade to many existing voice-based experiences. Obviously, this is going to be extremely
valuable for customer support leaders, contact centers, small service businesses that have a lot of
miscalls because they can't be talking on the phone while they're doing their work. But I also think
that people should be experimenting a lot more with bringing voice into new types of experiences.
My general sense is that as voice technology gets better, people are going to shift more and more
from typing to their devices to talking to their devices. This is a shift that's already been
happening, it's generational. And honestly, to some extent, the only reason it hasn't happened more
is that the native voice recognition technology
and places like Siri have been so bad historically.
As more different types of users get comfortable with voice,
it opens up a lot of possibilities.
For example, for a B2B sales and inbound marketing teams,
so many of those experiences start with forcing people to schedule a demo,
but maybe instead you could just let interested prospects talk through their needs
and receive a relevant product explanation
before they even get to the demo stage.
It obviously feels like there's a ton of opportunity here for language learning
and for tutoring more broadly.
It creates some really interesting opportunities
for realistic practice and simulation type of scenarios
to the extent you're building experiences
where people are practicing a particular type of interaction
that's going to come up in their jobs.
And these are just the tip of the iceberg.
One company that has already built on GPT Live
is Cognition, who have just announced Devin Voice
with a very simple pitch.
Your favorite AI software engineer just got a landline.
You say it, Devin ships it.
The feature is powered by OpenAI's GPT,
live on the voice layer and cognition's new suite two model on the back end. And as you might have
guessed, we will get to suite two in just a minute. Now, these sort of updates while using voice for
coding are huge quality of life upgrades over turn-based voice modes. The ability to interrupt,
ramble, and correct yourself just makes the experience a lot more natural and native.
Nader dabbit from the cognition team says 90% of my day-to-day work is now done with my voice.
It only makes sense to make it a first-class citizen in Devon. writes OpenAI's Kova.
Voice is the most magical way to interact with AI.
There's something beautiful about a future where bringing an idea to life starts with saying it out loud.
So in this case, the who should be using this and what they should be using it for is a little bit more simple.
The people who should be using this are, of course, Devon users.
And what they should be doing with it is shifting their behavior from typing to Prompt and Interact to Talking to Prompt and Interact.
And even if you are not using Devon, I'll use this chance once again to implore you to start using voice if you haven't yet.
A really great way to do that is the native chat GPT voice feature if you are a chat GPT user,
but even if you're using other tools and other LLMs,
I highly recommend setting up something like whisper flow on your computer
and starting to very intentionally try to shift your behavior over to voice.
Once you do my guess is that you'll not only find your speed improved,
but you'll also find you can do more of your work while doing other things like taking a walk.
Look, it's pretty clear from these first couple announcements that the industry is pushing you there anyways.
You might as well go surf this new opportunity.
Next new launch to discuss is the previously mentioned Sway 2.
It is Cognition's latest in-house coding model
and their first major iteration of their model
which replaces Sway 1.7.
This version is a post-trained version of Kimmy K3
optimized for coding and nothing else.
We've seen a few iterations of these types of models
from Cognition and Cursor over the past year.
Generally, they've been aimed at delivering a good enough model
at a cheap price point to efficiently handle everyday tasks.
In fact, Cursor and Cognition were both a little ahead of the curve
and understanding that we had reached a certain inflection point in capability, as well as a magnitude
of tasks that are being undertaken by AI, where optimizing not just for pure capability, but also
for efficiency, was going to start to really matter. What's interesting is that with Sweet 2,
it's clear that the definition of good enough has dramatically increased in recent months.
Given that the model is built on Kimmy K3, it unsurprisingly offers near frontier performance. For example,
the model scored 50% on frontier code 1.1, putting it slightly ahead of GPT-56 sole, and just behind
Fable 51, but at a 64% reduction in cost.
The model is even cheaper than Sway 1.7, which was built on the much smaller Kimi K2.7.
Cognition is also added effort levels so you can turn down the juice for even cheaper usage.
Now, one of the interesting things about the timing for this is that even as Cursor gets further
integrated into SpaceX AI, leading to consequences like OpenAI cutting off access to its
models inside Cursor, cognition is signaled that they want to stay independent, just raising $2 billion
at a $48 billion valuation.
My guess is that that means some number of developers who prize independence and the ability
to select the models they want without constraint might go check out Cognitions products like
Devon.
And this new model release, of course, creates a good context to go see what's possible with
this cheaper native model that lives inside the Devon desktop and CLI.
Basically, when it comes to who should try this and what they should use it for, anyone who
is already in the midst of trying to build a complete model stack that can better align
capability with need to keep cost down might be interested in checking this one out.
And speaking of models to keep cost down, our next release comes from Deepseek, who have dropped
a new version of their Flash model with a release of V4.1 Flash. The model has 552 billion
parameters and is designed for high latency work. The benchmarks look pretty solid with a score of 74.2
on DeepSwi, which puts it right in line with GPT56O and Opus 5, on Terminal Bench 4.0, which is so far,
much less bench-maxable, V4.1 Flash scored 31.2% compared to 39.9 for Seoul and 74% for Opus.
Still, the model is priced at the bottom of the market, charging 30 cents per million input tokens
and a buck 20 per million output tokens. Interestingly, artificial analysis found that the
flash model actually outperforms DeepSeek's full-size pro model on the intelligence index
at a quarter of the cost. Now, this is artificial analysis's new version 4.3 index,
where the top score is Fable 5.1 and GPT6 Astra at 53, and Deepseek 4.1 is still a way off the
frontier with an overall score of 40. That puts it roughly in line with GPT-56 Luna and Gemini
3-8 Flash, and again means that this is likely a model for teams that are looking to build a more
holistic architecture. Especially for teams that have basic but highly recurring compute-intensive
types of tasks, this is one that's likely going to be worth a test. Next, we move once again
back over into OpenAI world with the introduction of the Small Business plugin collection.
Now, this is not particularly complex. OpenAI has collected some of their most useful plugins
as a pack for small business. It includes things like Dropbox, HubSpot, Canva, Figma, Shopify,
DocuSign, PayPal, Quickbook, Stripe, Gusto, Slack, Wix, Mercury, etc.
Now, this is really just trying to make things a little bit simpler for small business owners
who are using ChatGPT by putting everything all in one easily and accessible place.
What's interesting, though, is that some people are recognizing that the entire idea of plugins
gets a little bit different when it comes to the agentic era.
Responding to OpenAI President Greg Brockman posting about the plugin collection,
Abdul Wasi writes,
Plugins V1 died because they were API wrappers behind a chat box.
This time, the agent can run the whole workflow.
However, Abdul also points out, the unsolved part is still distribution.
SMB owners don't browse plugin directories.
Although given that they are now collected, that is exactly the time.
type of user who should check this out. If you are a small business owner or have a very small
team and you are using these apps already, seeing what the sort of native integrations in
chat GPT can do could be a significant upgrade to your experience. OpenAI has also released
ChatGPT for Finance with a big update for GPT6 Astra. The new version of this product is still
a bundle of data connectors and skills built for financial professionals, but it's been
updated for OpenAI's new product lineup. The product now lives within ChatGPT work.
and includes built-in access to premium data feeds, including DeLupa, Pitchbook, LSC News, and CrunchPace.
This means that users don't need to figure out MCP connections and can get started right away with minimal
configuration. The product is also designed for ChatGPT's Enterprise Security and Governance
controls to allow for organizational level control of data access. OpenAI partnered with Morgan
Stanley and Evercore for the redesign, trying to closely align it to the needs of financial professionals.
It's designed for tasks, including building financial models.
developing research and building client materials, and during a press briefing, VP of product
Nick Turley said, we're effectively teaching ChatGBTGPT to research like an analyst and back up its
conclusions like an analyst as well. It uses GPT6 astranatively and OpenAI intends to keep it
updated for new models as they become available. OpenAI's Ryan Brewer wrote, with ChatGPT for
financial services, we've done the hard work of indexing the data that you need and making it easily
accessible to the model for financial analysis. We've created a custom charting experience,
better citations, an SEC filing viewer, and much more.
Sundip Shrivestava writes,
it's not a general chatbot with a finance skin.
It's aimed squarely at what junior investment bankers spend their weeks doing,
company research and equity analysis,
LBO modeling and buyer screening,
pitch books and client decks formatted to the firm's own templates.
Now, as you might expect, this is opening up some questions
about whether those junior associates are going to be replaced.
But I think it's important to note
that it's not particularly realistic to think that high-level bankers
are going to be making their own models and slide decks,
just because the agents have gotten much better at that core work.
There's still questions of accountability, iteration,
and all these things which more senior-level bankers are not going to want to do.
Instead, what we're seeing in the finance industry
is that junior bankers are starting to be required to demonstrate AI proficiency.
UBS, for example, recently started demanding
that prospective junior bankers show proficiency in AI as a hiring requirement.
In other words, I don't think this product is a replacement for junior bankers.
I think it is a power tool for junior bankers,
which is going to allow them to be much more efficient and much better at their jobs.
The last one from OpenAI is their new data agent for ChatGPT work.
The agent is designed for handling proprietary data within an organization,
with connections to data providers like Amazon Redshift, Datadog, Google BigQuery,
Clickhouse Databricks, MongoDB, and Snowflake.
OpenAI says the features can be used to ingest sales data and generate insights around core metrics
like sales conversions and retention.
The goal is to realize the promise of being able to talk to your data and perform real analysis
without needing to touch additional tools.
Souther Jones, the chief product officer at Tableau said,
connecting Tablo with ChatGBTBT work
brings trusted business semantics into a place employees already work,
so the answers they get are grounded in the same data model
their teams rely on.
Users can easily transform insights into action
by asking questions, exploring evidence,
and publishing new views directly to Tablo
using built-in design and analytics best practices.
Taking a step back then, if you look at the aggregate
of these feature updates from OpenAI,
they are all about honing in
on specific types of business users, better connecting both the tools and context they need to do
their work, and through that better tool access and better context connections, make it vastly
easier and frankly more inviting to move more of their workflows into chat GPT.
This is not a new trend, but it's certainly one that I think we're going to see a lot more of.
In other words, verticalization that really understands how very specific types of people work.
Another small example of that trend is Grockbot, who just announced this week that they are now
more powerful for sales teams because of native integrations with Salesforce, HubSpot, Gong,
clay, Grinola, and other go-to-market tools. And in this, we see the verticalization,
tool access, and context management trends are not unique to OpenAI, but are of course shared
across the industry. Two more quick ones before we get out of here. The first is projects in
cursor. Now, this is a little interesting because while we're all used to using projects in
chatGBT and Claude, this feature takes it to slightly new places. Rather than just being
a place to accumulate context around a particular project, this feature is also designed to
house scheduled tasks, long-running agents, and orchestration. In other words, projects aren't just
a folder, they are an actual project architecture. The project's tab and cursor uses a similar
approach to Grockbot, allowing users to spin up a persistent agent that performs work is required.
In their launch video cursor said, it's graduated from turn-based chat. Instead of you micromanaging
every agent, you're working with a much more autonomous colleague. For example,
example, if you have a standard workflow after opening a PR, the project can figure out how to do
and carry out the tasks automatically the next time you trigger the workflow.
Cursors said that testers have merged six times as many PRs using projects, and merge rates have
increased by 30%.
Now, this is one where the obvious users initially are going to be developers, but the patterns
are worth paying attention to for non-developer knowledge workers as well.
Prasinjit Sarkar writes, the move from chat per task to persistent coordinator is, in my read,
the more important shift in cursor projects than the raw sub-agent count.
Every coding agent I've watched before this worked the same way.
Open a session, describe a task, agent-execute, session ends.
Cursor projects inverts the model.
One coordinator thread stays open for the life of a project.
The coordinator doesn't write code, it plans, delegates the sub-agents, and brings results back to check.
Now, as Presenge points out, this isn't a unique architecture.
You've got Devin by Cognition, Open AI Codex, and others,
but argues that the persistent thread and the proactive trigger system are different.
Now, the persistent thread pattern is one that I think is interesting.
A lot of our build projects involved having to hand off context between different instances of
something like Claude Code because the context window in one was filled and he needed to start fresh.
That has changed a lot in 26.
Kansar describes what changed.
He wrote, in Q4 of 2025, I led a big push to teach block engineers advanced context engineering
to get the most out of ClaudeCodeco. In March and April, we started switching to Codex in large numbers.
With 5-4 and especially 5-5, it could solve problems and build things that other tools could not.
But my favorite was how good it was at compaction. Instead of all the fancy sub-agent tricks and intentional
compaction, you could just keep going and going. I had one thread that built a whole sync protocol,
server, and client across multiple repos and languages and also deployed them. Since then,
then compaction continues to get better and better. Mono threads are a common topic now,
but I'm still surprised by how much you can do in one thread without any issues.
This is a pattern I found myself using a lot more in both Claude and Codex.
I have, for example, everything related to the AI Daily Brief website,
pretty much in one long thread, that has done a ton of work and has been going for a very
long time and hasn't really ever run into any context issues.
So the reminder is that sometimes a new product or feature launch is less about whether
you are going to use that particular feature, and more about what it says about where things are
headed in general.
Anyways, folks, that is a quick rip through the products and features that have launched in the last
couple of days, who I think they're for, what I think they're useful to do, and hopefully
you have some new ideas heading into the weekend.
For now, that's going to do it for today's AI Daily Brief.
Appreciate you listening or watching, as always, and until next time, peace.
