The AI Daily Brief: Artificial Intelligence News and Analysis - The Most Important New AI Tools from OpenAI DevDay
Episode Date: September 30, 2026OpenAI unveiled more than 20 launches at Dev Day, from always-on Dots agents and the shared workspace Space to cheaper models and new ways to use your ChatGPT subscription across other apps. NLW break...s down the most important announcements, the early reactions, and what they reveal about AI’s shift toward persistent agents, team collaboration, and more affordable intelligence.Register for our Free Webinar: Build Your Personal AI Benchmark - https://aidailybrief.ai/webinar/personal-ai-benchmarkNext Cohort - Learn How to Build Agents - https://register.besuper.ai/register?program=atiAIDB Fall Listener Survey - https://aidailybrief.ai/surveyMultiplayer AI Sprint - https://multiplayerai.ai/Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at https://kpmg.com/us/SophisticatedHarbor - Invest in the AI ecosystem. https://www.harborcapital.com/aidailyHyperagent - Hire a team of always-on agents. New users get $100 in free credits. hyperagent.com/aidailybriefRackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack https://www.rackspace.com/Section - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/Robots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/The AI Daily Brief helps you understand the most important news and discussions in AI. Newsletter: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? sponsors@aidailybrief.ai
Transcript
Discussion (0)
The team at OpenAI said that one of the big changes that had happened internally was that the power of the latest generation of models like Astra had increased their speed of development so significantly that many things that they thought were only going to come in 2027 were actually coming as part of this Dev Day announcement.
The result of that was more than 20 different launches and announcements, including some big headliners like OpenAI's answer to Muse and Grockbot, and some sleeper hits, like the fact that Enterprise accounts can now use OpenAI credits on an open market.
place to buy access to open source models as well. Overall, what we got at OpenAI Dev Day does
not change the big patterns and trends that we've been seeing in the industry, a move to more
cost-efficient models, those models moving to more persistent work, and some amount of that
persistent work moving from a solo to a multiplayer experience. Instead, what Dev Day reinforced
is that these are the trends to pay attention to and the ones that will reshape how we all
use AI in the months to come. The AI Daily Brief is a daily podcast and video about the most
important news and discussions in AI. All right, friends, quick announcements before we dive in.
First of all, thank you to today's sponsors, KPMG, Blitzy, robots and pencils, and hyperagent.
To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe
on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors at AIdailybrief.com.
There are also all sorts of other goodies at AIDDailybrief.A.I. On our website, you can get
access to all of our free AI training programs. Those include free, multi-dialial
multi-week self-directed programs like the multiplayer AI Sprint, which is live right now,
but also free live webinars like the one that is coming up on Thursday, October 1st at noon,
that is all about building your personal AI benchmark.
You can also find links to our paid trainings, such as the Super Intelligent Executive Ketchup
and the Super Intelligent Executive Agent Leadership Program, both of which are registering
new cohorts that start over the next couple of weeks.
Now, last note on today's episode,
originally I was going to try to jam in a full accounting of everything that happened at the
White House yesterday on top of Open AI Dev Day, but there was
just too much to fit, so I've decided to go all Dev Day recap today, and then we will dig much
deeper into what came out of Trump's meeting with the leaders of all the major frontier labs,
and if and how it changes anything about the future development of AI. For now, though,
let's get into Dev Day and find out what it says, not just about Open AI, but about where we are
with AI in general. It was an absolutely monster day of announcements from OpenAI. In fact,
there were so many that on today's episode, we are just going to get through all of the big ones
discussing what the announcement was, whether it was expected or not, people's first impressions,
and how it shifts the AI race, if at all.
Later on in the week, we'll go deeper on some of the most important ones, but for now,
it is going to take all that we have just to get through it.
First up is Dots.
Open A.I's answer to Muse, Grockbot, and the wave of personal agents that are taking
over the AI space.
Sale Maltman presented Dots as, quote, remarkably capable, always-on agents built to handle,
really anything you can think of.
They bring AI to a whole new form factor.
With dots, users can create a persistent agent, communicate with it through a text message-style interface,
and then set them to work on tasks using their own cloud computer with support for 40,000 different apps.
Dots can also communicate via voice call or across Microsoft Teams and Slack,
with persistent context carried over between services and sessions.
Functionally, it shares a lot with Muse, right down to the QT mascot.
However, at least at the time of release, there are a few caveats worth noting.
For now, users can only create one dot at a time, but OpenAI,
plans to expand teams of agents in the future. The other big difference from Mews is that OpenAI
is only offering Dots to pro-business and enterprise customers. One of the big reasons Mews has been
so successful is making it completely free, meaning OpenAI is naturally limiting their ability
to compete on that front. CFO, Sarah Friar, said that the vision is to bring Dots to the
whole consumer base, but for now, this is a pro-sumer product. OpenAI's big selling point is that
they're using GPT6 Astra to power the agents, so this is the first time we're seeing a truly
frontier model pilot a first-party personal agent. Now, this was in many ways the least surprising
of the announcements. This personal agent form factor has been basically inevitable since the launch of
OpenClaw, and even the interest in business-focused versions like Grockbot, OpenAI getting into
this particular area was pretty much inevitable. Now, how ready for primetime dots is remains to be seen.
The live demo had some issues, although I would put way less in that than the chattering classes on X.
Demos going wrong is basically just a way for you to know that they are actually live.
Unfortunately, once people got their hands on dots, they ran into a number of other teething problems as well.
Chase Browser posted a session where Dot lost all of its work.
And Tennebrass said that they were excited to try it out but found that they couldn't connect multiple computers at once to share their full context.
Others had a better experience.
Analyst Max Weinbach wrote,
My basic test for how well one of these works is can it autonomously do my expenses with access to my email in Google Drive?
Grockbot made a mess and didn't do what I told it.
It did the first like three days and then started to mess up.
Dots did it after the first time and has been good.
Justin Schroeder said that he thought that Dots was even more work-focused than Grockbott or Muse.
Think a bit less personal assistant, he writes, and a bit more Codex orchestrator and Slack collaborator.
In fact, Justin writes, Slack is where it really shines.
Your team can message your bot directly and your bot can take actions, like spin up new pains.
Also, getting at a debate, which I think will be everywhere, Justin suggests that he thinks that this should have been a separate app.
ChatGPT is great, he writes.
Codex is great.
They are very different products, in my opinion, and dots are ways.
more codex than ChatGBT. The Every Vibe Check almost leaves it as a TBD, with their users
reporting flashes where it was truly excellent, but then other times where it was just extremely
frustrating. When Brandon Chu suggested that this was the form factor that everyone was converging
on, Nate B. Jones wrote, I think we need to distinguish between form factor and utility
here. Yes, form factor is converging at the moment, that may change, but utility is not. Mews is
good at specific things, like phone calls and practical email work. Instinct is good at specific
different things like travel. Dots is good at AI context as a work surface, regardless of positioning
that's distinct utility. The question with a T for trillions is, did you pick the right utility to get right?
We'll come back to dots later in the week, but I think the big takeaway is continue to watch this
space. While acknowledging that it's a bit buggy right now, Ali K. Miller argues that there are some big
differences here that are fairly significant. Things like each dot getting its own dedicated
virtual machine, it being always on, which allows it to be more proactive, and some other changes
that are worth watching. In any case, number two on the list is the Decisions API, which is
OpenAI's answer to Jev. Now remember, Jev is a judgment model. It's not good at outputting text,
it's good at classifying things, giving confidence scores between zero and one, making judgments,
in other words, which are the precursors to decisions, which presumably is where this feature got its
name. The Decisions API allows users to call a version of Luna that mimics Jeff's quick classification
and decision-making abilities. Users define a list of questions and possible answers, and the model
outputs fast responses. OpenAI said that the API could be used for things like classification,
routing requests, or choosing an agent's next action. They claimed 10 times faster decision-making
compared to the responses API. Still, one of the big questions was how is this better than just using
Jev? One difference in actual use cases comes from Luna's support for visual inputs, which aren't
possible with Jev. Over the past couple weeks of Jev mania, many users have shown off Jev making rapid
classification for things like visual marketing, but that requires an image-to-text transformation
under the hood. The Decisions API removes that step, and that way expands its set of possible use
cases. Now, right now, the Decisions API is just in preview for testing with a limited group. And in terms of
significance, I think more than anything, this is a recognition that what Jev represents is more than just a
new model. It is an extremely useful, and dare I might say, soon to be fundamental, primitive to have
in the ecosystem. Part of why Jev has hit wasn't that it was flashy or sexy. It's just that as soon as
you see all the things that it does better than a traditional LLM, in fact, where you see that
LLMs were being rammed like square pegs into round holes into use cases that they weren't great at,
it just seems obvious that having that sort of judgment model to sit alongside your generative
model is pretty obvious in retrospect. Once again, it's too early to really do comparisons,
but I tend to think that this is a space where there is room for more than one model available.
In fact, I think that pretty much every frontier lab will have some version of this available
very, very soon. Next up on our list is OpenAI Space, their new shared.
document workspace that gives human teams and agents a place to collaborate. You can think of the
feature sort of like an AI-enhanced Google Drive. Teams can work on shared spreadsheets, slide decks,
and other documents, but space also gives teams a place to house workplace automations. Dots can
natively work on documents in space, but teams can also set up scheduled tasks to produce a deliverable
in their space. Many focused on OpenAI going after Microsoft 365, Google Drive, or Notion with
space, and while this is certainly generally in those space, each of those services
have felt increasingly like a bad fit for agentic work, and it was only a matter of time before
OpenAI launched a truly AI-native productivity suite. Now, whereas almost every other announcement
had a pretty wide diversity of opinions, space was one where especially the power users were
in love right away. Ray Fernando wrote, I've had early access to OpenAI dots and I'm not here to join in
on the Glaze Fest. Is this a Grockbot or Hermes killer right now? No. But spaces, and allowing
agents to thrive in apps, feels like the right directions for these agents. How IAI's Claire Vow wrote,
In my opinion, dots slightly overhyped and space underhyped.
Every company I know wants an AI-native collaboration workspace.
We'll be watching to see if this pulls more enterprises to the OpenAI ecosystem.
Dan Shipper from Every wrote,
The obvious win is that you stop switching windows between chat GPT and another app like Notion while you write.
The less obvious one is speed.
When a dot builds a document through its browser in Google Docs, the edits practically crawl in,
because Docs wasn't built for agents.
Native documents don't have that lag.
You can tag your dot inside the document itself instead of going back to the chat,
and tell it to check the file.
Dan concludes,
I've long expected the company to do this.
In 2024, I wrote about the potential for document slides and sheets inside ChatGPT,
and I'm glad the product is finally here.
I'm already doing most of my work in ChatGPT's in-app browser,
and this makes that process smoother.
I can tag my dot in a comment on a document
and get a revision in line as I'm working.
The back and forth feels more like collaboration.
The big model launch for the day was GPT6-1 Sol,
which OpenAI pitched as Near Astra Intelligence for a,
of the price. The release comes just a week after GPT's 6-Sole and provides some fairly notable upgrades.
Coding benchmarks are up significantly across all effort levels, meaning that 6-1-sole on medium
settings outperforms six-sole on max settings. The model's top score on Deep Sui was 75.2% on high
settings, with extra high and max actually seeing a degradation in performance. We first saw this
phenomenon with Opus 5, and it looks like top-effort settings are starting to force models to
overthink and second-guess-correct responses to their detriment. Now, that high-setting score on
on Deep Swee was actually a touch higher than Astra's best performance, supporting OpenAI's claim of
near Astra Intelligence. The pattern was similar across most major benchmarks, a big improvement over
6-Sole that landed 6-1-Sole in the same ballpark as Astra at a much cheaper price. One of the notable
benchmarks was OS World, which tests Long Horizon computer use. 6-1 Sol's best performance scored 71.4%
on max settings, beating 6-soul at 64.4%, and coming close to Astra's best at 73.5%. 6-1-Sole was also
significantly cheaper than the others, at a 3-3-1-1-soules.
third, the cost of 6-Soul and 13% the cost of Astra. What's more, increases in the effort level
barely changed the cost, suggesting that OpenAI has made some big breakthroughs in the efficiency
of computer use with this model. Artificial analysis scored the model at 52, one point shy of Astra,
and also behind Opus, Sonnet, and Fable coming in in fifth place. They also found incredible
cost efficiency that pushes the Pareto frontier, with 6-1-Soul's benchmark run completed at a quarter
the cost of Astra and 31% cheaper than 6-Sol. If you don't need to run it at max settings,
6-1-sole seems capable of hitting smaller costs like GLM-53 Flash, which is the cheaper version of
ZAI's open-source model. OpenAI called it the most cost-efficient model for its performance
available today. Now, as to whether this one was expected or unexpected, on the one hand,
it's never all that surprising to get a new model from one of these labs, especially on a big
day like Dev Day. Surprising in that we only got Six Sol last week. So when it comes to first
impressions, let's just say we're going to have to wait to get our hands on it ourselves,
Because for every post you can find like this one from dropout layer on X,
GPT's 6'1 sole just made Opus 55 look expensive.
You also get one like this one from Bridge Branch.
We ran it through our Sunset Ocean test,
same cost as GPT's 6-Sole, twice as slow, and it barely rendered an ocean.
Hello, everyone.
One big change around AI is we've shifted our thinking from how we rank our pages
to how do we become the source that AI trusts enough to answer with.
At KPMG, they're seeing this firsthand.
AI generated results now surface answers directly.
often without a single click. That's why they are increasingly focused on generative engine optimization
or GEO, structuring content so AI systems can retrieve it, understand it, and cite it as trusted
authority. This is not just an SEO evolution, but a visibility mandate. And indeed, the GEO
mandate from KPMG is simple. If AI is shaping decisions, your expertise needs to show up inside the answer.
Read all about it at KPMG.com slash us slash geo. Again, that is KPMG.com slash us
slash GEO.
Here's why most legacy modernization projects fail.
The AI doing the work can't understand codebases at scale.
It sees a small slice of context, examine syntax, and misses years of decisions distributed
across the global application ecosystem.
Blitzy solves this the way it solves everything.
Grounded in your code before any migration begins, Blitzie's agents reverse engineer the
entire legacy system into a persistent knowledge graph, every dependency, every constraint,
every piece of tribal knowledge that used to live in one engineer's head.
From that understanding, Blitzy autonomously executes language migrations, framework upgrades, and monolith to microservices transformations, all validated end to end.
One Blitzy customer modernized a $10 million monolithic insurance stack in 16 weeks against a 137 week baseline with coding agents.
That's 9x compression.
Retire technical debt while accelerating your roadmap.
See how at blitzy.com.
That's BLITZY.com.
The best teams don't have a single star carrying everyone else.
They know their own strengths than each other's weaknesses and play to both.
That's the team robots and pencils has built on purpose.
Nobody there is grinding through busy work to pat a headcount number.
People come for the hard problems and they stay because everyone around them is leveling up at the same time.
In a market full of companies that are just trying to hire fast, that's worth a look.
Check out robots and pencils.com slash careers.
This episode of the AI Daily Brief is brought to you by HyperAgent,
where you run fleets of agents your team can manage together.
Forget local agents and chat workflows waiting.
on your laptop to be prompted. Hyperagent
deploys always-on agents in the cloud, doing
real work across the tools your team already
uses. Marketing agents turn competitor
moves into landing pages. Sales
agents enrich leads, draft emails, and updates
the CRM. Ops agent chases the
paperwork and tracks the budget. Every agent
has access to shared context and
follows your rules about scope and approvals.
It's time you add agents that feel like teammates.
Hire yours at Hyperagent.
Get $100 in credits at hyperagent.com
slash AI Daily Brief.
Now, hopefully these sole class models are good enough, though, because we may never see GPT61 Astra.
The Wall Street Journal reports that OpenAI have scrapped plans to release the next version of their flagship model over safety concerns.
Sachi Jane, OpenAI's head of safety systems, said in a statement,
for anything regarding safety and alignment, there's a trade-off.
You really do need to find what's the right line between staying within scope,
but also avoiding laziness in terms of how the model actually pursues tasks, even when it hits friction.
It sounds like OpenAI tried to turn down the tenaciousness that had led to multiple incidents over recent months,
but couldn't find a happy medium.
Jane added that although the model was less lazy than its predecessors,
it, quote, didn't quite meet the bar in terms of staying within scope and authorization,
and how it communicates back to the user about the type of work it's done.
Now, obviously, I'm being a bit hyperbolic when I say that we'll never see 6-1 Astra,
but it actually does seem like they're going to need to hit some technical breakthroughs
before we get to some of the next levels of the frontier in a way that they deem safe enough to release.
Now, from here we get into what are considered the smaller announcements from the event,
many of which people are still noting as quietly significant.
OpenAI showed us the latest version of its platform plans, opening ChatGBT
BT to outside developers.
Tebow from the OpenAI team wrote,
You can now build full native apps with plugin extensions and ship them right in ChatGBT.
We have over 1.2 billion weekly users and will surface relevant plugins right in the
conversations.
This feature is technically called Plugin extensions.
Now, OpenAI has been trying to get the balance right on the internet.
this feature for more than a year, testing native integrations, plugins, and MCP as a way to expand
the ecosystem. So could this finally be the feature that lets chat GPT function as an app store,
which always seemed like one of the big goals for OpenAI? Certainly there is going to be a rush
to experiment here. I'm already seeing things like this personal stylist app from Yana Wellander,
but others are wondering if the tradeoffs for the app developer are really worth it. MIT's Christian
Catalini writes, You bring the app, OpenAI brings the intelligence. Does OpenAI also get the user
traces? An interesting partnership if your contribution is teaching your partner how to do your job.
To which Signal responded, long tail of apps might opt in, but otherwise, this is a terrible
idea for anyone else. A flip side of this is sign in with chat GPT, which is exactly what it
sounds like. It allows you to use chat CBT to sign into other accounts. OpenAI's head of
applied research, Boris Power wrote, Sign in with ChatGBTT lets you use any app built with our API,
makes it much better for app developers not needing to jump a huge hill justifying paying a separate
subscription. Chatchapit is your one-stop superintelligence juice. The idea is basically that developers
allow people to carry their intelligence subscription with them, lowering the barrier to entry for
their apps and allowing them to take advantage of the fact that so many people have already opted
into using Chatchapit for their super intelligence solution. Jackie Lua writes,
sign in with Chatchapit is a huge deal, and I don't understand why they framed it as off
when it's really OpenAI expanding favored pricing to third parties and bundling services into the
subscription. It finally starts to align their incentives with customers so that individuals stop
double paying for tokens and apps stop having to price everything on top of API costs. I'd imagine
that expands to many more, if not all apps, and Anthropic will need to follow to add value to their
subscription. In that world, apps could bypass token costs and charge for the app player alone again,
which is probably good for everyone. In other words, the idea here is that instead of apps having
to charge what looks like huge prices to get access to intelligence, they can just charge their $20
or whatever they wanted to price their app experience at
and let the API costs flow directly to the underlying.
While yes, that means they might not get to scalp those costs
and add a little premium to them,
the benefits of not having to convince people to pay those additional fees
when they're already paying for a chat GPT subscription
likely outweighs the money that they could make otherwise.
For my power users out there,
OpenAI has introduced a few new options for their subscriptions.
The first is Ultra Fast Mode,
which offers 8x faster token production in Codex and 6x speed in the app.
Right now, Ultra Fast mode is only a very important.
available for Astra and Codex and ChatGBTGBT work, but OpenAI say they will add support for
6-1 Seoul soon. In addition, OpenAI is introducing a new $500 subscription tier that offers 25 times
the usage of the plus tier. It's also the only tier that gets access to ultra-fast mode for the time
being. Heading into the event, Tebow announced that OpenAI would reopen access to their $200
a month pro tier, which they had turned off a few weeks ago, but with a few tweaks. Usage
calculations have been changed netting out to a 50% reduction in terms of API cost. Tebow
explained that this was the best of a bad set of choices. OpenAI prioritized not having to
introduce the five-hour usage limit and argue that API cost reductions and more efficient models
would mean users can still get roughly the same amount of work done. Now, on the one hand,
people were sort of expecting something like this, but at the same time, you can imagine how well
it went over to have the same model coming back with reduced value. For our purposes here of
trying to understand what it says about the state of AI, look, man, compute constraints are
real, present, and permanent. Frontier models are running up against the walls of what they can do
with the compute that we have, and new compute isn't coming online at anywhere near the speed,
that people are increasing their use of this digital intelligence. The upside of that, though,
is that we're going to see a big push around efficiency, which in the long run should net out
to more cost-effective better experiences for all of us, but it's going to have some bumps along
the way. Over in developer and VibeCoderland, Codex now has a dedicated cloud environment,
allowing users to keep working after they close their laptop. This one was always coming after
people walking around with their thumbs jammed into laptops, became so common across San Francisco
that it became a meme. But looking for broader patterns, starting with Grockbot, we've seen the
entire industry moved towards a cloud instance being a necessary part of all agendic products.
This is a continuation in the big shift from AI being something you engage with like software
to becoming a more persistent always-on application regardless of how you access it.
The Codex CLI is also getting a major refresh with a new look and new capabilities.
History goes back further as you scroll, supporting the monothread maxis, of which I am absolutely one,
while the composer stays pinned so you can always type your next line.
wrote OpenAI,
The Codex CLI now gives you a better way to manage parallel work.
Use slash agents to see what's running, check progress, and jump between tasks.
On the enterprise side, OpenAI has launched private intelligence,
which guarantees zero data retention even at inference time.
Over the past few months, privacy, which was always a huge issue for enterprises,
has become a huge issue for the companies supplying the enterprises,
with growing concerns that OpenAI and Anthropic are skimming data from their customers.
This is basically the subtext or the main text of every communication from
Microsoft these days about why you should be not trusting their competitors. All that means that
a stronger, clearer, end-to-end data privacy guarantee is a big deal for those who need it.
OpenAI has also launched a model marketplace. It allows users to buy open-weight model inference from
Base 10 through OpenAI's responses API and Codex. Now, this is one that on another day we could go
way deep into because it has some big implications for OpenAI's mode. It protects them from
open source disruption and gives customers a lot more choice. It also means enterprises can feel
comfortable making large spending commitments with Open AI, knowing that they can easily use
that spend on a multi-model strategy that includes open-weight models.
Trust me, when I say that we are going to come back to that one because I think it is bigger
than people are giving it credit for it at first blush.
So we have barely scratched the surface here, but let's try to quickly sum up what all of this
amounts to and what OpenAI Dev Day revealed about the next phase of AI.
We didn't get something blisteringly new.
What we got was confirmation of a lot of the trends that we've been cataloging on this show
over the past several months.
First, with Sol 6.1 in the Decisions API, we got a continuation and a deepening of the trend
of models that are good enough and cheap enough that they can be used to do everything.
I.e. you don't have to turn them off, you don't have to make decisions about what they're used for.
They can just do it all. With dots, we have persistent and proactive agents that take advantage of
that to actually do everything. And when it comes to doing everything, plug-in extensions
shows that that means bringing everything in, as well as going everywhere, which is chat-chip-t sign-on
for other apps. Finally, with space, we have more confirmation that the next generation of AI will
not just be single-player mode, but will also be team and multiplayer mode in a native way, even if
that remains nascent so far. Lots and lots more to explore here, but that is going to do it for today's
AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace.
