The AI Daily Brief: Artificial Intelligence News and Analysis - Google Losing to Open Source AI? "We Have No Moat, And Neither Does OpenAI"

Episode Date: May 5, 2023

This week, a Google researcher had a note leaked that argued that Google (and OpenAI) were going to lose AI to the open source community. In this episode, NLW explores the arguments and asks what the ...implications might be.  Read the original note: https://www.semianalysis.com/p/google-we-have-no-moat-and-neither

Transcript
Discussion (0)
Starting point is 00:00:00 Today's AI breakdown focuses on the recent leak of a Google researcher's letter about how open source is beating Google and basically every closed source approach to AI. Before that, we discussed today's headlines, including Microsoft teaming up with AMD on a new AI chip, Slack adding their GPT features, and much more. Welcome back to the AI breakdown brief, all the AI headlines you need in five minutes or less. We start with what's happening all around the tech world. which is, of course, companies adapting to this new era, the new capabilities of artificial intelligence. The latest company to announce their big new AI suite is Slack. They have a whole new set of experiences that will integrate GPT natively into Slack's platforms. So that might include editing messages, turning messages into emails, attending huddles on your behalf, summarizing them, and much, much more.
Starting point is 00:00:58 Earlier this week, we saw new features from Box that were similarly AI Integrated. And of course, last week, Dropbox announced that they were laying off 500 people as they went to develop new AI tools. Bloomberg is reporting that Microsoft is working with AMD on an expansion into AI processors. So apparently these companies are teaming up in order to offer an alternative to InVIDIA, which obviously dominates the markets for AI-focused chips. The code name for Microsoft's homegrown AI processor is Athena, but there is some question about whether AMD is actually involved. Frank Shaw, a Microsoft spokesperson, denied it, saying AMD is a great partner, however, they are not involved in Athena. Regardless of that denial, AMD's shares jumped on the possibility that they were collaborating with Microsoft on this new AI chip. Speaking of companies in AI, the information is reporting that OpenAI's losses doubled last year to $540 million in the midst of developing chat GPT.
Starting point is 00:01:53 Now, before we lose our minds on how crazy it seems to lose a half billion dollars in a year, you have to remember, one, that OpenAI has a lot. a ton of money and that this type of heavy tech venture often loses a huge amount of money on the path to making revenue. And two, that this happened primarily before chat GPT was available and before they started to get on a revenue run rate that is now in the hundreds of millions of dollars. I think the major thing here is just a reminder that this is an extraordinarily expensive industry to really try to compete and lead in. One piece of news from earlier in the week that I didn't really have a chance to talk about yet is what happened when Chegg talked about ChatGPT during its earnings call. Basically, the CEO of Chegg said that ChatGPT was
Starting point is 00:02:32 hurting its business. It was cutting down on the number of new subscriptions, and the stock absolutely plummeted. You can see in this chart the exact moment when Chegg released its first quarter earnings and had this call, and the stock price goes from 18 down to like nine, basically a 50% loss in just a very short amount of time. Now, importantly, it wasn't just Chegg. It was a set of education-based stocks. So Pearson, which is one of the biggest education companies in the world, was down 15%. Doolingo was down 10%. Udeme was also down 5%. Effectively, people are pretty sure that chat DPT and AI more broadly are going to have a transformative effect on education businesses, even if they don't know how.
Starting point is 00:03:14 Now, one news story that we're going to get into in more depth on the main AI breakdown is the White House's meeting with AI CEOs that happened yesterday and what it suggests about the U.S.'s involvement in an approach to AI. But where I want to conclude the brief is with these comments from Snoop at the Milken Institute conference yesterday. This clip has been flying around the internet is something that basically gets at exactly how confused and surprised and awed and excited but also nervous so many of us feel. One note, this is Snoop, so watch out for some colorful language. I'm looking AI right now, they didn't make for me. This nigga can talk to me. I'm like, me and this nigga can hold a real conversation. Like for real for real like this is blowing my mind because I watch movies on this as a kid years ago when I See this shit and I'm like what is going on then I heard to do that the old dude that created AI someone this is not safe because the AI has got their own minds And these motherfuckers gonna start doing their own shit I'm like is we in a fucking movie right now?
Starting point is 00:04:15 The fuck man so do I need to invest in the AI so I can have one with me up? Like, do y'all know? Shit, what the fuck? That's it for this AI News headline brief. See you back here soon for the main AI breakdown. Today on the AI breakdown, we review a leaked letter from a researcher at Google that argues the company has no moat and is being out-competed by a legion of open-source developers. One of the most important discussions in AI is the role of open-source development, how open-shores, Should these models be? Are there risks to that openness? Have the nature of the risks changed
Starting point is 00:04:56 over the last few months? And does that mean something different? These are conversations that swirl around constantly. And this week we got a really interesting document that adds to that discussion in the form of a leaked note from a researcher at Google. So what we're going to do today is talk about the TLDR of that note, the context for it, and then get into what it might mean and what it argues. So the TLDR of this note is an argument that Google, Google OpenAI and basically any closed model is going to be out-competed in AI by tinkers, by open-source. And the reason for that is that these open-source tinkers, individuals, are advancing in ways that the big companies didn't think was possible, and certainly at a much greater speed, even if they had thought it was possible.
Starting point is 00:05:40 Now, again, the context for this is a leaked internal document. It is the perspective of one researcher. It is not the perspective of Google as a whole. Before we get into the argument for why open source is eating the lunch. of the Open AIs in Googles of the world. We need to discuss what happened. And the way that this piece frames it is that effectively over the last one to two months, really, there has been a sea change in the LLM space. Individual tinkers, developers, creators have driven massive innovation after getting access to their first, quote, really capable foundation model in Mata's Lama.
Starting point is 00:06:12 The anonymous author gives the timeline of the important events as they see them. In February 24th, meta launches Lama. The code is open source, but the weights are not. March 3rd, the entire Lama model is leaked, and while of course it's a non-commercial license, everyone is still able to experiment with it. By March 12th, someone gets it working on a Raspberry Pi, which begins this larger minification effort onslaught, as they put it. March 13th, Stanford releases Alpaca to add instructions tuning and quote, suddenly anyone could fine-tune the model to do anything, kicking off a race to the bottom on low-budget fine-tuning projects. Papers proudly describe their total spend of a few hundred dollars. What's more, the low-rank updates can be distributed easily and separately from the original
Starting point is 00:06:54 weights, making them independent of the original license from meta. Anyone can share and apply them. March 18th is the first time someone gets it up and running on a regular MacBook CPU, i.e. no GPU. March 19th, Fukuna is released, which is called a 13 billion parameter model achieving parity with Bard. And the next few updates are all just showing how fast things are moving. March 25th, choose your own model.
Starting point is 00:07:17 March 28th, open source GPT and multimodal training in one hour. and then April 3rd, a milestone. Real humans can't tell the difference between koala, a Berkeley open-13 billion parameter model and chat GPT, leading to April 15th when there is an open source RLHF from Open Assistant at ChatGBT levels. The net effect, well the author says plainly put, open source is lapping us.
Starting point is 00:07:42 Things we consider major open problems are solved and in people's hands today. Just to name a few, LLMs on a phone. People are running foundation models on a pixel 6 at 5 tokens per second. Scalable personal AI. You can fine-tune a personalized AI on your laptop in an evening. Responsible release. This one isn't solved so much as obviated.
Starting point is 00:08:04 There are entire websites full of art models with no restrictions whatsoever, and text is not far behind. Multimodality, the current multimodal science QA SOTA was trained in an hour. While our models still hold a slight edge in terms of quality, the gap is closing astonishingly quickly. open source models are faster, more customizable, more private, and pound-for-pound more capable. They are doing things with $100 and $13 billion that we struggle with at $10 million and $540 billion. And they are doing so in weeks, not months. Now, the author also argues that careful observers at places like Google could have seen this coming, and that in many ways this was a stable diffusion moment, quote-unquote, for LLMs.
Starting point is 00:08:45 What he's referring to is the idea that the open-stable diffusion model, has just absolutely crushed the closed open AI doll e model. The author writes, In both cases, low-cost public involvement was enabled by a vastly cheaper mechanism for fine tuning called Low-Rank Adaptation or Laura combined with a significant breakthrough in scale. Latent diffusion for image synthesis,
Starting point is 00:09:06 Chinchilla for LLMs. In both cases, access to a sufficiently high-quality model kicked off a flurry of ideas and iteration from individuals and institutions around the world. In both cases, this quickly outpaced the large place. These contributions were pivotal in the image generation space, setting stable diffusion on a different path from Dali, having an open model led to product integrations, marketplaces, user interfaces, and innovations that didn't happen for Dali.
Starting point is 00:09:32 The effect was palpable. Rapid domination in terms of cultural impact versus the open AI solution, which became increasingly irrelevant. Whether the same thing will happen for LLMs remains to be seen, but the broad structural elements are the same. Now, according to the author, there are a number of lessons. things that it's taught them from a technical perspective. One, this idea of low-rank adoption is powerful and they should be paying attention to it.
Starting point is 00:09:55 The stackability of Laura is one of the most important pieces of it because it allows so many different people to build on the work of each other. Three, thinking about large models versus iterating and training might be wrong. The author writes, focusing on maintaining some of the largest models on the planet actually puts us at a disadvantage. Giant models are slowing us down. In the long run, the best models are those which can be iterated upon quickly. We should make small variants more than an afterthought now that we know what is possible in the
Starting point is 00:10:22 less than 20 billion parameter regime. Now what about what has taught them in terms of a market perspective? Well, one, the reality is that information is going to get out and it's just going to be very hard to keep secret these types of advances. Number two, because there are so many individuals working on this, it changes the entire nature of the game. Individuals, as they write, are not constrained by licenses to the same degree as corporations. Three, this author argues that people will not pay for a restricted model when free unrestricted
Starting point is 00:10:50 alternatives are comparable and quality. Four, this author also writes, we have no secret sauce. Our best hope is to learn from and collaborate with what others are doing outside Google. And I think an implication of the piece, although they never use this terminology, is effectively that just in the world of AI development, this thing that is happening so rapidly, open networks are just kicking the crap out of closed companies. Now, one side implication of this is that they shouldn't really worry about open AI because they are as well a closed company, not an open network.
Starting point is 00:11:22 But that's the gist of it. I anticipate that we'll see a lot of counterpoints from others in Google and around Google over the next couple days. But there has been a ton of response to this on Twitter and in different discussion communities, by and large recognizing a lot of truth underlying it. Ahmad, the CEO of Stability AI, which put out stable diffusion, says, While this article fits with much of our thesis, I think it has a misunderstanding of what moats actually are. It's very difficult to build a business with innovation as a moat.
Starting point is 00:11:50 Base requirement is too high. Data, distribution, great products are moats. Robert Scoble says something similar. This is wrong. Developers and community are moats. You think Khan Academy is going to rip out GPT underneath its new AI-centric education system? And it is one of many building on top of GPT. You are insane if you think that.
Starting point is 00:12:10 Google will never get Khan to switch. that is a moat. Nick Dobo says much the same thing and even more colorful language. L.M.A.O. Google says we have no moat and neither does OpenAI. Holy shit Google doesn't get it. Absolutely effing clueless and asleep at the wheel. Google already lost.
Starting point is 00:12:26 OpenAI's moat isn't tech. It's dev community. 75% of tech is building on OpenAI and no one is building on Bard. Developer Nate Chan reinforces that point, saying, The clues to why developers are so tied to OpenAI's GPT models live in GitHub and Discord conversations happening every day. Just look at any open source AI projects. There is so much prompt engineering happening to improve the robustness and intelligence of these
Starting point is 00:12:49 AI products as they talk specifically to OpenAI's GPT models. Khan Academy recently said they spent six months on prompt engineering to get their math tutoring AI bot good enough for release. They're not switching from OpenAI models anytime soon. And I thought this analogy was really great. Nate writes, switching a model is like replacing a brilliant founding engineer with an equally brilliant new engineer. The new engineer may be extremely capable, but you've lost a ton of context and
Starting point is 00:13:14 team synergy that will have to be reestablished over a long period of time. You would never volunteer to replace your brilliant founding engineer if that relationship was going well. OpenAI has a moat. So so far, a lot of these counterpoints have been around how OpenAI does in fact have a moat, not Google. But by using a touch of irony, Sergei Keraev points out that Google certainly has a moat. Sergei writes, Google has no moat. They don't have over 90% of search traffic. They don't have everyone's emails and the most used email client. Their OS is not powering 70% of smartphones. They will never be able to deploy LLM features into these products. Instead, people will run open source software LLMs. But regardless of our debates around whether Google has any sort
Starting point is 00:13:58 of mode or at least a moat that they can take advantage of, or OpenAI has a mode, the truth that open source development is moving extremely fast, that it's taking on a force of its own, is absolutely true. Amad from Stability AI again writes, We are moving to fully open on language model development over the coming weeks at Stability AI. It's all a bit of a black box, so we will move to more continuous releases and open discussion about various things as we are trying to go to full releases. Also share the mistakes. Oops. Now this brings up something interesting that came out of the White House yesterday. In advance of the meeting between Vice President Kamala Harris and a bunch of AI
Starting point is 00:14:34 CEOs, the Biden administration announced their quote, new actions to promote responsible AI innovation that protects Americans' rights and safety. There were a couple notable things on this fact sheet. One was a $140 million investment to launch seven new national AI research institutes, but maybe the most relevant piece was this bullet public assessments of existing generative AI systems. The fact sheet says, The administration is announcing an independent commitment from leading AI developers,
Starting point is 00:15:00 including Anthropic, Google, Hugging Face, Microsoft, Invita, Microsoft, Invita, OpenAI, and Stability AI, to participate in a public event. evaluation of AI systems, consistent with responsible disclosure principles. This will allow these models to be evaluated thoroughly by thousands of community partners and AI experts to explore how the models align with the principles and practices outlined in the Biden-Harris administration's blueprint for an AI Bill of Rights and AI Risk Management Framework. Now, while the Biden administration may want there to be alignment with their own goals, creating a mechanism to review these generative AI systems is obviously going to have a lot more
Starting point is 00:15:34 implications than just that. But the fact that there is such a push to open source and more open development makes the set of people in this White House meeting room a little more suspect or at least incomplete. President Biden writes, artificial intelligence is one of the most powerful tools of our time, but to seize its opportunities, we must first mitigate its risks. Today, I drop by a meeting with AI leaders to touch on the importance of innovating responsibly and protecting people's rights and safety. Now, some people notice the notable absence of Meta's Mark Zuckerberg and Elon Musk. Others asked why there weren't any critical voices or AI safety experts. But another broader question, given this Google researcher's memo, is how do the voices of this
Starting point is 00:16:15 wider open community come to be a part of this dialogue? Who represents them and the networks that they represent when it comes to the policy conversation? I think it's a really fascinating question and one that we should be asking more. For now, I will turn it back to you guys. What do you think about this argument made by the author of this memo, that open source is just eating the lunch of Google and everyone else like them? Is it correct? Is it partially correct because it applies to Google but not OpenAI?
Starting point is 00:16:43 If it's not correct because OpenAI has a moat, is it good that one company has a moat when it comes to this technology? Regardless of one's answers to those questions, it's pretty undeniable how fast this is changing, and it certainly shows no signs of slowing down. Thanks for watching the AI breakdown. If you're enjoying, please subscribe or please go. follow the podcast. Until next time, peace.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.