The AI Daily Brief: Artificial Intelligence News and Analysis - OpenAI Agent "Operator" Coming In January?
Episode Date: November 15, 2024OpenAI is set to release its autonomous agent "Operator" in January, a tool that aims to execute tasks like coding and booking autonomously, marking a potential breakthrough in AI's role in practical ...applications. Alongside this, OpenAI has unveiled an ambitious proposal to advance U.S. AI infrastructure, advocating for energy expansion and streamlined AI project approvals to maintain national competitiveness. This "Manhattan Project" for AI could reshape the future of tech and policy in the U.S. Brought to you by: Vanta - Simplify compliance - vanta.com/nlwThe AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614 Subscribe to the newsletter: https://aidailybrief.beehiiv.com/ Join our Discord: https://bit.ly/aibreakdown
Transcript
Discussion (0)
Today on the AI Daily Brief, we are apparently getting an Open AI agent as early as January,
which is a good thing because in the headlines, we talk about how Google is also dealing with the same AI slowdown that we talked about in the context of OpenAI a little bit earlier this week.
The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
To join the conversation, follow the Discord link in our show notes.
Welcome back to the AI Daily Brief Headlines edition, all the daily AI news you need in around five minutes.
One of the big discussions we've been having this week is whether there is an AI model slowdown.
According to reports, OpenAI's Orion model is not showing the same jump in performance that was observed between GPT3 and GBT4.
This apparently has led OpenAI to double down on reasoning and fine-tuning as a potential way to obtain the performance boost they expect from the next generation of frontier models.
And now it appears that Google is joining OpenAI and exploring new avenues to tackle some of these challenges.
According to the information sources at Google, their models are demonstrating the same lack of improvement.
The information writes, past versions of Google's flagship Gemini large language model improved at a faster
rate when researchers used more data and computing power to train them. Google's experience is another
indication that a core assumption about how to improve models known as scaling laws is being tested.
Many researchers believe that models would improve at the same rate as long as they processed more data
while using more specialized AI chips, but those two factors don't seem to be enough.
This is particularly troubling for Google whose models have failed to see the same level of adoption
as open AIs. There was a belief perhaps that Google,
Google could leapfrog OpenAI in this generation purely due to their advantage in computing resources.
That appears less and less to be likely. So, following in the footsteps of OpenAI, Google is
looking to develop new methods of improving performance. Over recent weeks, Google DeepMind has put
together a team to work on the development of reasoning models. That team is being led by
principal research scientist Jack Ray and former character.a.I. Founder Nome Shazir. To give a sense of
how important they consider the work to be, other DeepMind researchers are working on making
manual improvements to the model, including changing so-called hyperparameters, which are the
variables that determine how the model processes information and how quickly it draws connections
between different concepts. Another problem Google has run into is duplicate copies of information
within training data which could hurt performance. Google has also experimented with synthetic
training data, essentially feeding data generated by an LLM back into the corpus of training data.
They've also added audio and video, and while it was believed that these steps would lead to
significant improvements, sources at Google say they didn't make a major difference.
Meta Chief AI scientist and Turing Award winner Jan Lacoon has been predicting these diminishing
returns from model scaling for years. Yesterday, he posted on threads,
I don't want to say I told you so, but I told you so. He referenced the statement from
former OpenAI chief scientist Ilya Sutskever from earlier in the week who said,
The 2010s were the age of scaling. Now we're back in the age of wonder and discovery again.
Everyone is looking for the next thing. Scaling the right thing matters now even more than ever.
Now, what makes Ilius comments more significant is that he was basically the chief proponent of the
idea that you could just add more compute and data to keep scaling higher and higher.
So the fact that he is moving away from that suggests that there is a changing understanding
of the technical capabilities here.
Lacoon for his part commented,
We've been working on the next thing for a while.
Lacoon is referring to the fundamental AI research team at Meta,
who are pursuing new architectures as a path towards AGI.
Their focus is currently on world models which seek to train an AI on how objects
and environments interact rather than just focusing on the connection between words.
Still not everyone is convinced that this is a real thing.
Vindu Ready writes,
The AI slowdown is a non-story.
The biggest reason AI is slowing down is that there's nowhere else to go.
If you begin to saturate on benchmarks, nothing is left to do.
100 out of 100 is the highest score you can get.
Now, moving over into the world of business models,
perplexity says they will begin experimenting with advertising on their platform.
Starting this week, US users will see ads in the format of sponsored follow-up questions.
These ads will be placed to the side of generated answers and labeled as sponsored.
The initial brands and agency partners for the launch,
include Indeed, Whole Foods, Universal Mechanan, and PMG. As an example, Perplexity showed a search
for information about looking for a job with a sponsored follow-up that says, how can I use Indeed to
enhance my job search? In a blog post, Perplexity explained, ad programs like this help us generate
revenue to share with our publisher partners. Experiences taught us that subscriptions alone do not
generate enough revenue to create a sustainable revenue sharing program. Advertising is the best
way to ensure a steady and scalable revenue stream. Perplexity said the ads themselves will be
generated by AI rather than pre-written or edited by sponsors.
Advertisers also won't get access to users' personal information.
Regarding the choice of format, Perplexity wrote,
We intentionally chose these formats because it integrates advertising in a way that
still protects the utility, accuracy, and objectivity of answers.
These ads will not change our commitment to maintaining a trusted service that provides
you with direct, unbiased answers to your questions.
Obviously, right now, how Perplexity-style AI summaries influence the core business model
of the web, which is basically search ads on Google, is one of the
big open questions. So this will be really interesting to watch these experiments. Basically,
I think there are a lot more consequential than just for perplexity as a company itself.
Lastly, today, speaking of business models, Salesforce CEO Mark Benioff thinks it's crazy
talk that AI could hurt his company's bottom line. Beniof has very publicly planted his flag on the
idea of AI agents over recent months, and we're about to find out whether it will save Salesforce
from disruption. During a recent appearance on TechCrunch's equity podcast, he said,
what if your workforce had no limits? As far as being disrupted, Beniof believes that his moat
is access to client data. He said, we manage 230 petabytes of data for our customers. You could say
that might be one of the main things we do for them, and we do it with a security and a sharing model.
Anyways, for me, it's interesting to note that Beniof feels like he has to justify the potential
disruption to the SaaS business model from agents. He can say it's crazy talk all he wants, but there is
absolutely no doubt in my experience and in my conversations with enterprises. While Beniof may be
right that agents are Salesforce's future and that they're an even bigger deal and the company grows to
lofty new heights, there is an incredible pressure being put on the sort of traditional per seat
model of SaaS companies that will not be resolved easily or quickly. Certainly something that we're
really interested in and thinking about a lot as we price super intelligent, but that is a conversation
for another time and place. For now, that is going to do it for today's AI Daily Brief Headlines
edition. Next up, the main episode. Today's episode is brought to you by Plum. Want to use AI to automate your
but don't know where to start, Plum lets you create AI workflows by simply describing what you want.
No coding or API keys required. Imagine typing out AI, analyze my Zoom meetings and send me your insights in
Notion and watching it come to life before your eyes. Whether you're an operations leader, marketer,
or even a non-technical founder, Plum gives you the power of AI without the technical hassle.
Get instant access to top models like GPT40, Claude Sonnet 3.5, assembly AI, and many more.
Don't let technology hold you back. Check out Use Plum, that's Plum with a B, for early access to the future of
workflow automation.
Today's episode is brought to you by Vanta.
Whether you're starting or scaling your company's security program,
demonstrating top-notch security practices, and establishing trust is more important than ever.
Venta automates compliance for ISO-2701, SOC2, GDPR, and leading AI frameworks like ISO-402,
and NIST AI Risk Management Framework, saving you time and money while helping you build
customer trust.
Plus, you can streamline security reviews by automating questionnaires and demonstrating your
security posture with a customer-facing trust center all powered by Vanta AI.
Over 8,000 global companies like Langchain, Lila AI, and factory AI use Vanta to demonstrate
AI trust and prove security in real time.
Learn more at Vanta.com slash NLW.
That's Vanta.com slash NLW.
Today's episode is brought to you as always by Superintelligent.
Have you ever wanted an AI Daily Brief but totally focused on how AI relates to your company?
Is your company struggling with AI?
either because you're getting stalled figuring out what use cases will drive value or because
the AI transformation that is happening is siloated individual teams, departments, and employees, and not
able to change the company as a whole. Super Intelligence has developed a new custom internal
podcast product that inspires your teams by sharing the best AI use cases from inside and outside
your company. Think of it as an AI daily brief, but just for your company's AI use cases.
If you'd like to learn more, go to Bsuper.A.I. slash partner and fill out
the information request form. I am really excited about this product, so I will personally get right
back to you. Again, that's besuper.a.ai slash partner. Welcome back to the AI Daily Brief. A couple
interesting pieces of news out of OpenAI today, starting with an agent story, OpenAI is reportedly
planning to release an autonomous agent next year. Now, if you spend any time around the AI space,
you'll know that basically since ChatGBT BT launched, we've been on the verge of the agent era.
The idea of moving from just the super powerful assistance to agents actually doing human
replacement style work is something with such dramatic implications for the structure of business
in society and what we can accomplish that it captures a huge amount of energy.
In fact, probably a disproportionate amount of energy relative to how far the technology
actually is.
And yet it's clear that the major labs have been making nudges in this direction.
And so what have we learned so far about this theoretical agent from open AI?
The agent, which they've codenamed operator, can control a computer to complete tasks independently,
including coding, shopping, and booking flights.
According to Bloomberg sources, staff were told in a meeting on Wednesday that the tool
would be released as a research preview in January.
That would mean that by early next year, we could have competing computer use agents from
Anthropic, Google, and OpenAI.
There are also already more limited agents available from companies like Microsoft Salesforce
and a host of startups.
So far, we've seen two different approaches to fully-fledged computer use.
Google's agent is sandboxed in the browser window, making it more limited but potentially more performant.
Anthropics agent, which is the only one generally available, was trained to control a mouse in a full computer interface,
so it can theoretically carry out a much broader variety of tasks.
In practice, the experience is still rather limited, with the company admitting the agent is slow, cumbersome, and error-prone.
Bloomberg sources said that OpenAI is working on several agent-related products,
and the one that is nearest to completion is a general-purpose tool that executes tasks in a web browser.
Sam Alvin has been hyping agents as the next big thing over the past few months.
In October during a Reddit AMA, he said,
we will have better and better models,
but I think the thing that will feel like the next giant breakthrough will be agents.
At OpenAI's Dev Day, chief product officer Kevin Wheel said,
I think 2025 is going to be the year that agentic systems finally hit the mainstream.
The Verge writes,
AI labs face mounting pressure to monetize their costly models,
especially as incremental improvements may not justify higher prices for users.
The hope is that autonomous agents are the next breakthrough product,
a chat GPT scale innovation that validates the massive investment in AI development.
So what do people think about this?
Well, Elvis on X writes,
computer uses the kind of capability I expected OpenAI to launch first.
This time, it looks like they will be following Anthropic.
I'm still hoping for a unique twist and huge improvements.
I think there's a lot to learn from custom GPs in search.
One thing is certain, 2025 will be the year of AI agents.
I never bet against OpenAI on these things,
especially because 01 for data generation, as was used in chat chb-tsearch, will play a key role here.
They have a huge advantage.
My asks?
Make it easy to launch agents.
Make it easy to use and provide feedback.
Versatile with tools and integrations.
Reduce latency and reduce costs.
Others shared their skepticism.
I hate how companies always flex their AI agents that can quote book a flight for you as if that was not the worst use case for AI automation ever.
This, by the way, is something that is a personal pet peeve of mine as well.
I do not need an agent to book me a flight or to order me food.
I know that it is just demonstration of capabilities,
but I do think it shows how early we are that those are the things that people always point to.
Others are thinking about the implications for various domains of the world.
Callum McClark writes,
For the learning world, click and complete means we must shift from click and complete courses
to gathering meaningful learning metrics,
a shift that should have happened a long time ago.
But now hopefully these agents will be the final nail in the coffin of bad L&D courses.
There are also questions of the business model of agents,
which are swirling.
Sully Omar writes, how do we properly price agents,
especially when they keep getting more capable,
they do days of work in one hour, and they work 24-7.
We have no market reference.
Is it compute per hour per task?
Adam Silverman of Agent Ops writes,
I think as agents scale over the next five years,
pricing will be compute plus 10% margin for specific use cases.
When OpenAI and others release agents,
they will only charge compute.
There's a huge opportunity for startups to capitalize
on charging significantly more in the interim.
Now, going back to this Verge quote,
the mounting pressure to monetize costly models, and the hope that autonomous agents are the next
breakthrough product, I think whereas chat GPT and Claude and the like have been easily as
much a consumer innovation as they have been in enterprise innovation, really transforming and hitting
both equally, and by some measurements, consumers more, I believe that where we're going to see
the value from agents is absolutely in vertical, highly specific enterprise tasks.
I think that getting agents very good at very specific repetitive tasks that happen over and over and
over again all the time inside specific companies is much easier than getting good at very
general purpose use. And I think that that's where we're going to see a lot of benefit.
Now it is still very early. What we have available and what's production ready is still very
nascent. But if I were a betting man, that is where I would be placing my chips that the impact
will be in very specific verticals within the enterprise.
Second OpenAI story, the company has outlined a grand policy proposal to bolster artificial
intelligence in the U.S. in an effort to stay ahead of China.
Yesterday at a think tank event in Washington, OpenAI's head of global affairs, Chris Lehane,
presented what they're calling their official blueprint for U.S. infrastructure.
The company said the plan was, quote, as ambitious as the 1956 National Interstate and Defense
Highways Act. They outlined AI economic zones, co-created between state and federal governments,
which would give the states an incentive to speed up permitting and approval for AI infrastructure.
They envisioned constructing solar and wind power, as well as gaining clearance to restart unused
nuclear plants. OpenAI wrote, states that provide subsidies or other support for companies
launching infrastructure projects could require that a share of the new compute be made available to
their public universities to create AI research labs and developer hubs aligned with their key
commercial sectors. OpenAI also wrote a bill called the National Transmission Highway Act.
The legislation would expand power, fiber, and natural gas pipeline connectivity across the nation.
The company argues that we need, quote, new authority and funding to unlock the planning,
permitting and payment for transmission, and they noted that existing procedures aren't keeping
up with AI-driven demand. The document noted that, quote, the government can encourage private investors
to fund high-cost energy infrastructure projects by committing to purchase energy and other means
that lessen credit risk. OpenAI notably views the Midwest and Southwest as key areas for
infrastructure expansion given the plentiful land for construction. Focusing on these areas would also
ensure that the jobs and prosperity of this wave of technology aren't just concentrated on the coasts.
Now, the premise of all of this is that without government investment and the removal of red tape,
the U.S. will lose its lead in AI to China. The Open AI proposal stated,
given the stakes, we need to think big, act big, and build big. These decisions determine
whether a nation leads or lags in technological innovation, often with far-reaching consequences
for economic competitiveness and national security. The history of the U.S., they wrote,
is one of iconic infrastructure projects that move the country forward, the auto industry,
the Tennessee Valley Authority, the Manhattan Project, the interstate highway system.
One of the things that we've been tracking a lot recently is the growing push to nuclear driven
by the AI industry.
And it's notable in that regard, and something that Open AI points out that China has built
more nuclear power capacity over the last 10 years than the U.S. is built over the last 40.
Speaking to that, rapid deployment, OpenAI's head of global policy, Chris Lehane said,
we don't have a choice. We do have to compete with that.
Now, the big remaining question is whether these kinds of policies would be adopted by the
Trump administration. OpenAI says they plan to work with the Trump White House on this agenda.
So far, all we know about Trump's AI policy, however, is that the president-elect has pledged to
repeal the Biden-A-I executive order, stating that, quote, in its place, Republican support AI development
rooted in free speech and human flourishing. That's about the extent of the details we have so far.
Then again, slashing red tape and building out a ton of energy production seems in line with campaign
promises. During an appearance in July, Trump said, we will be creating so much electricity that you'll be
saying, please, please, President, we don't want any more electricity, we can't stand it.
You'll be begging me, no more electricity, sir, we have.
enough, we have enough. So who knows? I think what's pretty clear is that this is coming into the vacuum
of whatever the repeal of the executive order looks like. As Andrew Curran points out, this blueprint is
being presented to influence whatever form the new regulations will take. Still, the earnings nugget
may have summed it up best when they wrote, this is the new Manhattan Project, interesting times
ahead. But with that, we will wrap today's AI Daily Brief. Appreciate you listening as always,
and until next time, peace.
