The AI Daily Brief: Artificial Intelligence News and Analysis - Google Losing to Open Source AI? "We Have No Moat, And Neither Does OpenAI"
Episode Date: May 5, 2023This week, a Google researcher had a note leaked that argued that Google (and OpenAI) were going to lose AI to the open source community. In this episode, NLW explores the arguments and asks what the ...implications might be. Read the original note: https://www.semianalysis.com/p/google-we-have-no-moat-and-neither
Transcript
Discussion (0)
Today's AI breakdown focuses on the recent leak of a Google researcher's letter about how open source is beating Google and basically every closed source approach to AI.
Before that, we discussed today's headlines, including Microsoft teaming up with AMD on a new AI chip, Slack adding their GPT features, and much more.
Welcome back to the AI breakdown brief, all the AI headlines you need in five minutes or less.
We start with what's happening all around the tech world.
which is, of course, companies adapting to this new era, the new capabilities of artificial intelligence.
The latest company to announce their big new AI suite is Slack.
They have a whole new set of experiences that will integrate GPT natively into Slack's platforms.
So that might include editing messages, turning messages into emails, attending huddles on your behalf, summarizing them, and much, much more.
Earlier this week, we saw new features from Box that were similarly AI Integrated.
And of course, last week, Dropbox announced that they were laying off 500 people as they went to develop new AI tools.
Bloomberg is reporting that Microsoft is working with AMD on an expansion into AI processors.
So apparently these companies are teaming up in order to offer an alternative to InVIDIA, which obviously dominates the markets for AI-focused chips.
The code name for Microsoft's homegrown AI processor is Athena, but there is some question about whether AMD is actually involved.
Frank Shaw, a Microsoft spokesperson, denied it, saying AMD is a great partner, however, they are not involved in Athena.
Regardless of that denial, AMD's shares jumped on the possibility that they were collaborating with Microsoft on this new AI chip.
Speaking of companies in AI, the information is reporting that OpenAI's losses doubled last year to $540 million in the midst of developing chat GPT.
Now, before we lose our minds on how crazy it seems to lose a half billion dollars in a year, you have to remember, one, that OpenAI has a lot.
a ton of money and that this type of heavy tech venture often loses a huge amount of money on the
path to making revenue. And two, that this happened primarily before chat GPT was available and
before they started to get on a revenue run rate that is now in the hundreds of millions of
dollars. I think the major thing here is just a reminder that this is an extraordinarily
expensive industry to really try to compete and lead in. One piece of news from earlier in the
week that I didn't really have a chance to talk about yet is what happened when Chegg
talked about ChatGPT during its earnings call. Basically, the CEO of Chegg said that ChatGPT was
hurting its business. It was cutting down on the number of new subscriptions, and the stock
absolutely plummeted. You can see in this chart the exact moment when Chegg released its first
quarter earnings and had this call, and the stock price goes from 18 down to like nine,
basically a 50% loss in just a very short amount of time. Now, importantly, it wasn't just
Chegg. It was a set of education-based stocks.
So Pearson, which is one of the biggest education companies in the world, was down 15%.
Doolingo was down 10%. Udeme was also down 5%.
Effectively, people are pretty sure that chat DPT and AI more broadly are going to have a transformative effect on education businesses, even if they don't know how.
Now, one news story that we're going to get into in more depth on the main AI breakdown is the White House's meeting with AI CEOs that happened yesterday and what it suggests about the U.S.'s involvement in an approach to AI.
But where I want to conclude the brief is with these comments from Snoop at the Milken Institute conference yesterday.
This clip has been flying around the internet is something that basically gets at exactly how confused and surprised and awed and excited but also nervous so many of us feel.
One note, this is Snoop, so watch out for some colorful language.
I'm looking AI right now, they didn't make for me. This nigga can talk to me. I'm like, me and this nigga can hold a real conversation.
Like for real for real like this is blowing my mind because I watch movies on this as a kid years ago when I
See this shit and I'm like what is going on then I heard to do that the old dude that created AI someone this is not safe because the AI has got their own minds
And these motherfuckers gonna start doing their own shit I'm like is we in a fucking movie right now?
The fuck man so do I need to invest in the AI so I can have one with me up?
Like, do y'all know?
Shit, what the fuck?
That's it for this AI News headline brief.
See you back here soon for the main AI breakdown.
Today on the AI breakdown, we review a leaked letter from a researcher at Google that argues the company has no moat and is being out-competed by a legion of open-source developers.
One of the most important discussions in AI is the role of open-source development, how open-shores,
Should these models be? Are there risks to that openness? Have the nature of the risks changed
over the last few months? And does that mean something different? These are conversations that
swirl around constantly. And this week we got a really interesting document that adds to that
discussion in the form of a leaked note from a researcher at Google. So what we're going to do today
is talk about the TLDR of that note, the context for it, and then get into what it might mean
and what it argues. So the TLDR of this note is an argument that Google,
Google OpenAI and basically any closed model is going to be out-competed in AI by tinkers, by open-source.
And the reason for that is that these open-source tinkers, individuals, are advancing in ways that the big companies didn't think was possible,
and certainly at a much greater speed, even if they had thought it was possible.
Now, again, the context for this is a leaked internal document.
It is the perspective of one researcher.
It is not the perspective of Google as a whole.
Before we get into the argument for why open source is eating the lunch.
of the Open AIs in Googles of the world. We need to discuss what happened. And the way that this
piece frames it is that effectively over the last one to two months, really, there has been a sea
change in the LLM space. Individual tinkers, developers, creators have driven massive innovation
after getting access to their first, quote, really capable foundation model in Mata's Lama.
The anonymous author gives the timeline of the important events as they see them. In February 24th,
meta launches Lama. The code is open source, but the weights are not.
March 3rd, the entire Lama model is leaked, and while of course it's a non-commercial license, everyone is still able to experiment with it.
By March 12th, someone gets it working on a Raspberry Pi, which begins this larger minification effort onslaught, as they put it.
March 13th, Stanford releases Alpaca to add instructions tuning and quote,
suddenly anyone could fine-tune the model to do anything, kicking off a race to the bottom on low-budget fine-tuning projects.
Papers proudly describe their total spend of a few hundred dollars.
What's more, the low-rank updates can be distributed easily and separately from the original
weights, making them independent of the original license from meta.
Anyone can share and apply them.
March 18th is the first time someone gets it up and running on a regular MacBook CPU,
i.e. no GPU.
March 19th, Fukuna is released, which is called a 13 billion parameter model achieving
parity with Bard.
And the next few updates are all just showing how fast things are moving.
March 25th, choose your own model.
March 28th, open source GPT and multimodal training in one hour.
and then April 3rd, a milestone.
Real humans can't tell the difference between koala,
a Berkeley open-13 billion parameter model and chat GPT,
leading to April 15th when there is an open source RLHF
from Open Assistant at ChatGBT levels.
The net effect, well the author says plainly put,
open source is lapping us.
Things we consider major open problems are solved
and in people's hands today.
Just to name a few, LLMs on a phone.
People are running foundation models on a pixel 6 at 5 tokens per second.
Scalable personal AI.
You can fine-tune a personalized AI on your laptop in an evening.
Responsible release.
This one isn't solved so much as obviated.
There are entire websites full of art models with no restrictions whatsoever, and text is not far behind.
Multimodality, the current multimodal science QA SOTA was trained in an hour.
While our models still hold a slight edge in terms of quality, the gap is closing astonishingly quickly.
open source models are faster, more customizable, more private, and pound-for-pound more capable.
They are doing things with $100 and $13 billion that we struggle with at $10 million and $540 billion.
And they are doing so in weeks, not months.
Now, the author also argues that careful observers at places like Google could have seen this coming,
and that in many ways this was a stable diffusion moment, quote-unquote, for LLMs.
What he's referring to is the idea that the open-stable diffusion model,
has just absolutely crushed the closed open AI doll e model.
The author writes,
In both cases, low-cost public involvement
was enabled by a vastly cheaper mechanism
for fine tuning called Low-Rank Adaptation or Laura
combined with a significant breakthrough in scale.
Latent diffusion for image synthesis,
Chinchilla for LLMs.
In both cases, access to a sufficiently high-quality model
kicked off a flurry of ideas and iteration
from individuals and institutions around the world.
In both cases, this quickly outpaced the large place.
These contributions were pivotal in the image generation space, setting stable diffusion
on a different path from Dali, having an open model led to product integrations, marketplaces,
user interfaces, and innovations that didn't happen for Dali.
The effect was palpable.
Rapid domination in terms of cultural impact versus the open AI solution, which became increasingly
irrelevant.
Whether the same thing will happen for LLMs remains to be seen, but the broad structural elements
are the same.
Now, according to the author, there are a number of lessons.
things that it's taught them from a technical perspective.
One, this idea of low-rank adoption is powerful and they should be paying attention to it.
The stackability of Laura is one of the most important pieces of it because it allows so many
different people to build on the work of each other.
Three, thinking about large models versus iterating and training might be wrong.
The author writes, focusing on maintaining some of the largest models on the planet actually
puts us at a disadvantage.
Giant models are slowing us down.
In the long run, the best models are those which can be iterated upon quickly.
We should make small variants more than an afterthought now that we know what is possible in the
less than 20 billion parameter regime.
Now what about what has taught them in terms of a market perspective?
Well, one, the reality is that information is going to get out and it's just going to be
very hard to keep secret these types of advances.
Number two, because there are so many individuals working on this, it changes the entire
nature of the game.
Individuals, as they write, are not constrained by licenses to the same degree as corporations.
Three, this author argues that people will not pay for a restricted model when free unrestricted
alternatives are comparable and quality.
Four, this author also writes, we have no secret sauce.
Our best hope is to learn from and collaborate with what others are doing outside Google.
And I think an implication of the piece, although they never use this terminology, is effectively
that just in the world of AI development, this thing that is happening so rapidly, open networks
are just kicking the crap out of closed companies.
Now, one side implication of this is that they shouldn't really worry about open AI because
they are as well a closed company, not an open network.
But that's the gist of it.
I anticipate that we'll see a lot of counterpoints from others in Google and around Google
over the next couple days.
But there has been a ton of response to this on Twitter and in different discussion communities,
by and large recognizing a lot of truth underlying it.
Ahmad, the CEO of Stability AI, which put out stable diffusion, says,
While this article fits with much of our thesis, I think it has a misunderstanding of what moats actually are.
It's very difficult to build a business with innovation as a moat.
Base requirement is too high.
Data, distribution, great products are moats.
Robert Scoble says something similar.
This is wrong.
Developers and community are moats.
You think Khan Academy is going to rip out GPT underneath its new AI-centric education system?
And it is one of many building on top of GPT.
You are insane if you think that.
Google will never get Khan to switch.
that is a moat.
Nick Dobo says much the same thing and even more colorful language.
L.M.A.O.
Google says we have no moat and neither does OpenAI.
Holy shit Google doesn't get it.
Absolutely effing clueless and asleep at the wheel.
Google already lost.
OpenAI's moat isn't tech.
It's dev community.
75% of tech is building on OpenAI and no one is building on Bard.
Developer Nate Chan reinforces that point, saying,
The clues to why developers are so tied to OpenAI's GPT models live in GitHub and Discord
conversations happening every day.
Just look at any open source AI projects.
There is so much prompt engineering happening to improve the robustness and intelligence of these
AI products as they talk specifically to OpenAI's GPT models.
Khan Academy recently said they spent six months on prompt engineering to get their math tutoring
AI bot good enough for release.
They're not switching from OpenAI models anytime soon.
And I thought this analogy was really great.
Nate writes,
switching a model is like replacing a brilliant founding engineer with an equally brilliant
new engineer. The new engineer may be extremely capable, but you've lost a ton of context and
team synergy that will have to be reestablished over a long period of time. You would never
volunteer to replace your brilliant founding engineer if that relationship was going well. OpenAI has a
moat. So so far, a lot of these counterpoints have been around how OpenAI does in fact have
a moat, not Google. But by using a touch of irony, Sergei Keraev points out that Google certainly has a moat.
Sergei writes, Google has no moat. They don't have over 90% of search traffic. They don't have
everyone's emails and the most used email client. Their OS is not powering 70% of smartphones.
They will never be able to deploy LLM features into these products. Instead, people will run
open source software LLMs. But regardless of our debates around whether Google has any sort
of mode or at least a moat that they can take advantage of, or OpenAI has a mode, the truth that
open source development is moving extremely fast, that it's taking on a force of its own, is
absolutely true. Amad from Stability AI again writes,
We are moving to fully open on language model development over the coming weeks at
Stability AI. It's all a bit of a black box, so we will move to more continuous releases
and open discussion about various things as we are trying to go to full releases. Also share
the mistakes. Oops. Now this brings up something interesting that came out of the White House
yesterday. In advance of the meeting between Vice President Kamala Harris and a bunch of AI
CEOs, the Biden administration announced their quote, new actions to promote
responsible AI innovation that protects Americans' rights and safety.
There were a couple notable things on this fact sheet.
One was a $140 million investment to launch seven new national AI research institutes,
but maybe the most relevant piece was this bullet public assessments of existing
generative AI systems.
The fact sheet says,
The administration is announcing an independent commitment from leading AI developers,
including Anthropic, Google, Hugging Face, Microsoft, Invita, Microsoft,
Invita, OpenAI, and Stability AI, to participate in a public event.
evaluation of AI systems, consistent with responsible disclosure principles.
This will allow these models to be evaluated thoroughly by thousands of community partners and
AI experts to explore how the models align with the principles and practices outlined in the
Biden-Harris administration's blueprint for an AI Bill of Rights and AI Risk Management Framework.
Now, while the Biden administration may want there to be alignment with their own goals,
creating a mechanism to review these generative AI systems is obviously going to have a lot more
implications than just that. But the fact that there is such a push to open source and more open
development makes the set of people in this White House meeting room a little more suspect or at least
incomplete. President Biden writes, artificial intelligence is one of the most powerful tools of our time,
but to seize its opportunities, we must first mitigate its risks. Today, I drop by a meeting with
AI leaders to touch on the importance of innovating responsibly and protecting people's rights and safety.
Now, some people notice the notable absence of Meta's Mark Zuckerberg and Elon Musk.
Others asked why there weren't any critical voices or AI safety experts.
But another broader question, given this Google researcher's memo, is how do the voices of this
wider open community come to be a part of this dialogue?
Who represents them and the networks that they represent when it comes to the policy conversation?
I think it's a really fascinating question and one that we should be asking more.
For now, I will turn it back to you guys.
What do you think about this argument made by the author of this memo, that open source is just
eating the lunch of Google and everyone else like them?
Is it correct?
Is it partially correct because it applies to Google but not OpenAI?
If it's not correct because OpenAI has a moat, is it good that one company has a moat when
it comes to this technology?
Regardless of one's answers to those questions, it's pretty undeniable how fast this is changing,
and it certainly shows no signs of slowing down.
Thanks for watching the AI breakdown.
If you're enjoying, please subscribe or please go.
follow the podcast. Until next time, peace.
