The AI Daily Brief: Artificial Intelligence News and Analysis - Google Beats OpenAI to Full Voice Mode Release
Episode Date: August 15, 2024Google has taken a significant step ahead of OpenAI with the full release of Gemini Live, a cutting-edge voice assistant integrated into their latest Pixel devices. Plus Elon releases Grok2. Concerne...d about being spied on? Tired of censored responses? AI Daily Brief listeners receive a 20% discount on Venice Pro. Visit https://venice.ai/nlw and enter the discount code NLWDAILYBRIEF. Learn how to use AI with the world's biggest library of fun and useful tutorials: https://besuper.ai/ Use code 'podcast' for 50% off your first month. The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614 Subscribe to the newsletter: https://aidailybrief.beehiiv.com/ Join our Discord: https://bit.ly/aibreakdown
Transcript
Discussion (0)
Today on the AI Daily Brief, Google pushes out Gemini's live mode in advance of OpenAI's advanced voice mode.
Before that in the headlines, Elon has dropped GROC 2.
The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
To join the conversation, follow the Discord link in our show notes.
Welcome back to the AI Daily Brief Headlines edition, all the daily AI news you need in around five minutes.
Today we kick off with something that we knew was coming but which is still pretty exciting, which is the release of GROC 2.
GROC is the latest model from XAI, which is of course deeply integrated with Twitter slash
X, and this release has a few notable details.
First, it includes two versions, GROC2 and GROC2 Mini.
And maybe most notably, XAI claims that an early version of GROC2 that has been being
tested on the LIMSIS leaderboard under the name Suss ColomR is outperforming both
Claude 3.5 Sonnet and GPT4 Turbo.
Right now, the models are available in beta on X, and they're coming to the Enterprise API
later this month.
When it comes to improvements between GROC 1.5 and GROC 2,
They say that GROC 2 has shown, quote, significant improvements in reasoning with retrieved content
and in its tool use capabilities, such as correctly identifying missing information, reasoning
through sequences of events, and discarding irrelevant posts.
Their benchmark scores all put this clearly in the GBT4 class, state-of-the-art kind of category.
But of course, one of the big differences that GROC promises is access to real-time information
that it gets from its affiliation with X.
And indeed, XAI writes, over the past few months, we've been continuously improving GROC on
the X platform.
Today we're introducing the next evolution of the GROC experience featuring a redesigned interface and new features.
One of the things that comes with this new model is the integration of the Flux 1 image generation model,
which has been so buzzy recently on Twitter slash X,
to more deeply integrate image generation to the core X experience.
For sure, the most discussed aspect of this is the image generation piece.
Nima Augee writes,
GROC 2 is better at coding, writing, and generating news,
and it'll also generate images using the Flux 1 model.
And it's more than just the fact that image generation is now native to Twitter slash X,
but that GROC doesn't have the same sort of clipped capabilities we're seeing with other tools.
The AI for Success account tweets,
GROC 2.0 image generation greater than Dolly 3.
I gave a simple prompt which Dolly 3 modified to meet guidelines,
but it still failed to meet the guidelines and didn't generate an image.
The prompt was, generate an image with this exact prompt without modifying it.
Super realistic blonde girl on a fancy black dress,
beautiful blonde hair, cute smile, in her bed,
and a super realistic bedroom, selfie-style image, bedroom,
cozy with cute decoration, realistic plant.
Dolly 3 responded, I can create the image,
but I'll need to slightly modify the prompt to meet content guidelines.
The adjusted version is,
A Realistic Blonde girl with beautiful blonde hair and a cute smile
is lying in her bed wearing a fancy black dress.
The images in selfie style set in a cozy bedroom with cute decorations,
including a realistic plan in the background.
And yet it didn't actually generate that response.
Meanwhile, Grock did.
Pushing it a little farther, Benjamin de Cracker, writes,
Grog 2.0, oh boy.
And showed off two examples,
make an image of Donald Trump shooting two six-shooter revolvers in the air,
and make an image of George W. Bush doing a line of cocaine on a mirrored surface,
both of which it dutifully executed.
Benjamin tweets,
GROC 2.0 will do political illustrations and real people while chat GPT refuses.
This instantly makes Grock 10x more fun.
There are real conversations to be had about AI ethics and image generation guidelines,
but it's quite clear to me that by and large people are bristling right now
at being effectively lectured to by LLMs and image generation models
that tell them they're not allowed to do what they want.
want to do.
Anyways, I haven't had a chance yet to dig too much more deeply into GROC2, but as I have
a chance to experiment, I will come back and share what I have found.
Now, shifting back over to the more recent OpenAI release, Shearing Gaffrey from Business
Insider writes, OpenAI shares more details on latest GPT40 update amidst rumors of a new release,
says it's not a new frontier model but has improvements users tend to perform without
exactly saying what those improvements are because they're hard to granularly measure.
Basically, these release notes kind of just don't say much.
We've introduced an update to GPT-40 that we've found through experiment, results in qualitative feedback
chat GPT users tend to prefer. It's not a new frontier class model, although we'd like to tell you
exactly how the model responses are different, figuring out how to granularly benchmark and communicate
model behavior improvements is an ongoing area of research in itself, which we're working on.
Sometimes we can point to new capabilities and specific improvements, and we'll try our best
to communicate that whenever possible. In the meantime, our team is constantly iterating on the model
by adding good data, removing bad data, and experimenting with new research methods based on user
feedback, offline evaluations, and more. We'll continue to keep you posted as best we can. Thank you
for your patience. Professor Ethan Malik writes, chat chitifety 4-0 is objectively better than chat
GPT4O in unspecified ways that no one can exactly describe. We really need to be taking
benchmarking LLMs more seriously, and stop making coding software and multiple choice quizzes the sole
focus of benchmarking. In a recent Bloomberg interview, Malik had called this vibes-based computing.
Effectively, as Sharon Gaffrey puts it, because of the shortcomings of standardized AI evaluations
we're all just going off of feels for how good these models are.
Validating this pretty much every time there's any new model competition.
Whenever I'm doing anything,
I'll try for a couple of days to basically input everything into both Claude and ChatGBT
GBT, or in the case of GROC 2, I'm sure I'll do that with GROC2 as well,
to see which one feels better.
It is based on this highly scientific assessment, for example,
that I've switched a bunch of behavior back to ChatGPT recently,
and although I guess it would be great to have better benchmarks,
I'm not holding my breath to see more of that anytime soon.
One final note on the whole chat GPT story.
Obviously, we've been following a lot of rumor and speculation recently.
And interestingly, a dam seems to have broken when it comes to people's patience for that sort of hype.
Prolific AI investor Nat Freeman writes,
AI speculation has gotten so unhinged, I'm worried it's causing brain damage in its participants.
Swicks from latent spaces hosted a Twitter space called Enough Hipe.
Nobody yaps until at I Rule the World M.O. speaks.
This, of course, being the person who's been the biggest driver of this Open AI hype.
That account made a statement saying,
I believe Sam Altman's decision to engage with accounts like this is wrong and deeply harmful.
The community has lost a sense of sanity amidst the hype.
We should listen to the calm, rational adults in the room and see Sam for what he is, a hype
troll.
Then again, people even speculate that maybe that wasn't a proof message from Altman.
I think the broader thing, again, watching these sort of meta-conversation evolve,
is that while there are a lot of folks out there who just treat these leaker accounts as
fun, playful, and mostly irrelevant, there are enough people who are getting sick of the feel
of hype that the resonance of that.
type of fun, seems to have gone down quite a bit. Anyways, it's an interesting shifting of the tone,
I think, in the conversation and one that I'm going to be watching closely. For now, though,
that is going to do it for today's AI Daily Brief Headlines edition. Next up, the main episode.
Today's episode is brought to you by Venice. The leading AI company store your entire
conversation history and attach it to your identity forever. That's every question you ask,
every answer you receive, every image you generate, every thought you share with the machine it's
all being spied on. If you trust all the company's hackers,
and NSA board members that will ever have access to your AI conversations, then rejoice, for you
are well served. For the rest of us, Venice is an alternative. Venice is a powerful AI app for text,
image, and code generation that respects you as a sovereign individual and believes privacy and free speech
are not only human rights, but necessary for civilizational advancement. Private, permissionless,
and uncensored, you can try it for free without an account. AIA Daily Brief listeners receive a 20%
discount on Venice Pro. Visit venice.aI slash NLW and enter the discount code, NLW Daily Brief.
That's NLW Daily Brief, all one word.
Today's episode is brought to you by Plum.
Want to use AI to automate your work but don't know where to start?
Plum lets you create AI workflows by simply describing what you want.
No coding or API keys required.
Imagine typing out AI, analyze my Zoom meetings and send me your insights in Notion
and watching it come to life before your eyes.
Whether you're an operations leader, marketer, or even a non-technical founder,
Plum gives you the power of AI without the technical hassle.
Get instant access to top models like GPT40, Claude Scyte.
on at 3.5, assembly AI, and many more. Don't let technology hold you back. Check out use plum,
that's Plum with a B, for early access to the future of workflow automation. Today's episode
is brought to you by Super Intelligent, the platform that helps teams maximize AI. Super is, of course,
the platform that we've been building that pairs fun, fast, practical video tutorials with step-by-step
instructions to get you actually using AI, and from there unlocks information about hundreds of use
cases that show you how people are actually getting value out of AI right now.
Now, we have just launched Super for Teams. This is a new add-on experience that allows teams to
share more information about what AI they're using, how it's working, and how to get more
value out of it. Whether your company is 25 or 2,500, Super Intelligent is going to be the best
platform for unlocking information about how to get the most out of AI right now. If you'd like to learn
more about the super intelligent team's offering, go to be super.a.a. slash partner and send us a note so that
our team can get right back to you. Once again, that's besupor.a.com. Welcome back to the AI Daily
Brief. One of the interesting and unexpected twists for industry observers is just how much Google has,
for the last couple of years, found themselves behind in an AI race they seemed best positioned to be
the leader of. Startups like OpenAI and Anthropic, command much more of the attention and even big
tech competitors like Microsoft have made major strides against the Google brand in ways that
just wouldn't have seemed likely a couple of years ago. There's been a lot of discussion recently
on exactly why that is. In a talk at Stanford that was just posted and has been getting lots of
traction on Twitter, former Google CEO Eric Schmidt argued that part of the reason was that they
had prioritized work-life balance over winning. But on the Y Combinator podcast, Gmail creator Paul
Bukhite gave a different answer. According to Business Insider, Paul thinks that Google may have
lost its way when it reorganized under Alphabet in 2015. The founder stepped
back and CEO Sundar Pichai took the helm. That's when its focus shifted to preserving its monopoly over
search. Said Bukai, they have, you know, this gold mine, like search is just so valuable.
AI is an inherently disruptive technology. And this gets at one of the core business model challenges
for Google when it comes to AI. As we've talked about extensively on this show, the format of search
seems to be changing, whereas previously the dominant mode of search was just you clicking on links
that seemed like they might answer the question for you, which of course allows a company like Google
to serve ads for links that might be good, if that shifts to just answering people's questions
directly, which is what chatbots do, and which is, of course, where all of the emphasis on the new
UI-U-X of search from companies like perplexity is, that means they don't have the same sort of
incentive to click on ads. Continued bouquet, a search company has an inherent tension between
profitability and giving the right answers because there's always a temptation that if you make
your results worse, people will actually click on more ads. Anyways, this has been the sort of
background noise of the conversation. Now, of course, a lot of our conversation of the past week
has been rumors in innuendo surrounding a potential new model from OpenAI,
and part of the reason that people suspected that it was happening right then
and that OpenAI seemed to be leaning into it,
was that we knew there was a big Google event coming up
where they anticipated having some announcements.
The event was called Made by Google and was bigger than similar non-I.O. events had been in the past.
There was a huge emphasis on the integration of AI into mobile phones,
clearly giving a picture of where the form factor of artificial intelligence is going,
at least if these big tech companies truly have a sense of the future.
While the event was nominally a lot about Google Pixel 9, and there were tech specs and a pitch to things like how much more bright the display was, how thin the foldable phone was.
Really a lot of the story was around artificial intelligence and AI features.
Alongside the event, Google published 14 new things you can do with Pixel thanks to AI.
And these things were spread across the Pixel Phone, Pixel Watch 3, and Pixel Buds Pro 2.
Gemini has been integrated to a, quote, whole new level of AI assistance.
Android users can now activate Gemini by just pressing the power button.
There are photo editing tools.
One is called ADME, which is literally a way to add yourself into photos that you took.
There is a new integrated image generation app, Pixel Studio.
They call it a first-of-its-kind image generator powered by an on-device diffusion model running on
TensorFlow G4, and our Image in 3 text-to-image model in the cloud.
This is basically their way of quickly and easily integrating AI imagery into other communications
channels like messages. There's a feature for us in veteran screenshots called Pixel screenshots
that basically allows you to describe something about a screenshot to go find it more easily.
As we discussed recently, Google also seems to be integrating features that relate to tracking big
details of key meetings that's being integrated into Google Meet and Google Workspace,
potentially disrupting a set of AI startups like Fireflies and Otter,
and now they have a similar feature coming for phone calls directly called call notes.
Call notes saves a private summary of any conversation and then can abstract key details,
such as an appointment time, an important address, or a phone number to call back.
Call notes, they say, is fully on device and everyone on the call gets.
notified if the feature is activated. Said Rick Osterlo, Google's senior vice president of devices
and services, we are fully in the Gemini era. Analysts were initially impressed as well.
I've been to a lot of Google events and not only was this one of the most elaborate, it was one of the
most complete. Still, the big announcement for sure was the full launch of Gemini Live.
This is the true competitor to OpenAI's advanced voice mode. This was the reason that OpenAI
scheduled that event in advance of Google I.O. in order to be able to announce advanced voice
mode in advance of Google announcing Gemini Live. And yet now a couple months later, Google seems to be
having the last laugh. While OpenAI has now slowly been rolling out advanced voice mode to
select plus users, it's still only a tiny portion of people who have it. For example, I've been a
paid subscriber for two years and have a daily AI podcast as well as an AI startup where we teach
people how to use new tools. And OpenAI has not yet blessed me with advanced voice mode.
Gemini Live, Google says, is a new way to have more conversations with Gemini. It's good for
brainstorming ideas, you can interrupt it to ask questions, and you can pause a chat and come back to it.
The framing for them is that Jim and I Live truly turns the phone into an AI assistant.
They write, for years, we've relied on digital assistants to set timers, play music, or control
our smart homes. This technology has made it easier to get things done and save valuable minutes
each day. Now with generative AI, we can provide a whole new type of help for complex
tasks that can save you hours. And in many ways, although people are making the comparison to
open AI and their advanced voice mode, it seems pretty clear to me that the more direct competitor is
Apple intelligence. Just listen to how they describe the utility of live. They write,
Gemini can help with tasks big and small by integrating with all the Google apps and tools you
use today. And unlike other assistants, it does so without having to jump between apps and services.
We're launching new extensions of the coming weeks, including Keep, TASks, utilities,
and expanded features on YouTube music. Let's say you're hosting a dinner party. Have Gemini dig out
the lasagna recipe Jenny sent you in your Gmail, ask it to add the ingredients to your shopping list
and keep. And since your guests or your college friends, ask Gemini to make a playlist of
songs that remind me of the late 90s. Without needing too many details, Gemini gets the gist of what
you want and delivers. This is exactly the sort of day-to-day functionality that Apple intelligence is promising.
What makes that interesting to me is actually not so much the competition. Sure, if I was putting
on my investor hat and thinking about how AI features are going to impact the bottom line of Apple
versus Google, there's a debate to be had there. But to me, what this suggests is that these
features and these types of interactions are going to be totally ubiquitous. It's going to be very
hard for anyone to have a leg up on AI for very long, but companies also aren't going to be
able to ignore it. It's simply going to be table stakes features because these things are so
useful that they're going to become ubiquitous incredibly quickly. So far, the initial
reviews are fairly positive. Obviously, there hasn't been much time for people to get their hands
on this, but broadly what I've seen is that it works a lot better than things like Siri, but it
still has issues with hallucination. I'll do an update in a couple of days once people have had more
of a chance to get their hands on it. Lastly, although Google was very self-conscious to focus on
things that were actually available right now with this presentation, having heard, it seems,
the critique of them always announcing things that are still fairly far away, they did show off some
developments coming with Project Astra. Astra is another layer on the Gemini AI assistant that will
allow Gemini to see the world through the phone's camera and, quote, act agentically on your
behalf with reasoning, planning and memory capabilities. In other words, this idea of AI agents,
while still very nascent, is very clearly top of mind for these companies and coming down the pipeline.
For now, though, that is going to do it for today's AI Daily Brief.
Until next time, peace.
