The AI Daily Brief: Artificial Intelligence News and Analysis - 7 AI Use Cases Unlocked By Nano Banana
Episode Date: August 28, 2025Today's AI Daily Brief covers the groundbreaking release of Google's Nano Banana image generation model, which has taken the AI community by storm over the past few weeks. Google officially re...vealed that Nano Banana is actually Gemini 2.5 Flash, now available as a free preview in Google AI Studio, offering unprecedented image editing capabilities with perfect object consistency and incredible prompt adherence. The model dominates benchmarks, scoring 17% higher than competitors like Flux, and opens up seven transformative use cases from professional photo editing to 3D mesh generation. This represents a major leap forward in multimodal AI that could reshape entire industries from photography to game development.Brought to you by:KPMG – Discover how AI is transforming possibility into reality. Tune into the new KPMG 'You Can with AI' podcast and unlock insights that will inform smarter decisions inside your enterprise. Listen now and start shaping your future with every episode. https://www.kpmg.us/AIpodcastsBlitzy.com - Go to https://blitzy.com/ to build enterprise software in days, not months Vanta - Simplify compliance - https://vanta.com/nlwPlumb - The automation platform for AI experts and consultants https://useplumb.com/The Agent Readiness Audit from Superintelligent - Go to https://besuper.ai/ to request your company's agent readiness score.The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Subscribe to the newsletter: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? nlw@breakdown.network
Transcript
Discussion (0)
This podcast is supported by Google.
Hey everyone, Shrehta here from Google DeepMind.
The Geminae 2.5 family of models is now generally available.
2.5 Pro, our most advanced model, is great for reasoning over complex tasks.
2.5 Flash finds the sweet spot between performance and price.
And 2.5 Flashlight is ideal for low-latency, high-volume tasks.
Start building in Google AI Studio at AI.dev.
Today on the AI Daily Brief, seven new use cases opened up by Google's new
nano-banana image generation model.
The AI Daily Brief is a daily podcast and video about the most important news and discussions
in AI.
All right, friends, quick announcements before we dive in.
First of all, thank you to today's sponsors, Gemini, High Touch, Blitzian, Superintelligent.
And to get an ad-free version of the show, go to patreon.com slash AI Daily Brief.
Now, two quick things about today's show.
one, as sometimes happens when we have an exciting new model, the main episode got a little bit
long, so this will be a main only type of episode. There really are kind of two parts. The first half
or so is about the background context and the announcement of the new nanobanana model. And then
the second half is the seven new use cases that it opens up. So you can kind of think about it
in that way when it comes to how to divide the episode. And in the case, our normal format headlines,
and then the main episode will be back tomorrow. Second thing to note, after recording for 40 minutes,
I discovered that it had not been recording with my actual podcast mic and instead was recording directly into the laptop.
However, I am in Vegas for a keynote today and so am not able to go back and re-record the full episode.
We've done our best to use a variety of tools, AI, and otherwise to make the sound quality as good as it can be,
but apologies that it's a little bit lower than normal.
With all that out of the way, let's dive into this very exciting new model.
Welcome back to the AI Daily Brief.
Today we are talking about the image model that has had people in the image model that has had people
incredibly excited for the last couple of weeks. It is finally out. We got confirmation who it was from.
And already people are discovering new use cases that were not possible, at least in the way
that they are now, before this release came. It's a great reminder of the fact that we still really
do live right on the horizon between possible and impossible. And that every time there's a new model,
even if it's not obvious at first, there are some set of new things that move over into the realm of
the possible that simply were not before. But we're getting a little bit ahead.
of ourselves with that. A couple of weeks ago, a new image generation model showed up on Elam Arena
that was called Nanobanana. I've talked about this previously on the show. If you spend any time on
AI Twitter, you've probably seen people talking about it or sharing their generations with it.
The model quickly rose to the top of the leaderboard and gathered a ton of praise for its
state-of-the-art performance. Now, interestingly, unlike the last few leaps in image model quality,
people weren't just talking about how realistic the images were. Instead, the real standout was
the model's ability to edit images. You could generate a single shot and then modify it with
seemingly perfect object consistency and incredible prompt adherence. Now, this is one of those
things that I think people have always imagined as what they wanted out of an AI generated image
model, but hadn't really been there yet. If you spend any time with Mid Journey or the tools
inside ChatGabit, you'll know that it's pretty frequently a crapshoot of trying to get the model
to do what you want to do. In fact, for some, like Mid Journey, I don't even recommend trying to get
the model to do what you want to do, I basically recommend you winding it up, pointing in a direction,
and letting it do what it wants. Now, there have been some big updates recently. For example,
ideogram, which I use every day for thumbnails, added a new updated character consistency patch,
which made a huge difference. But still, the idea of being able to take a base image and turn it
into whatever you want in a way that was coherent is something that people have both wanted and
simply haven't had access to. Didi Das of Menlo Ventures said, this is the next generation of filters
that we've been promised forever.
In any case, once the hype started building,
Google folks began winking at the crowd on X
implying that they were, in fact, the ones testing in stealth.
Those hints, however, backfired,
as they led many to think that the Google Pixel 10 launch last week
was going to double as the nanobanana rollout.
Many were disappointed then when the event came and went with
no mention of the new model.
Now, in retrospect, keeping them separated makes sense.
These are two very different launches for very different audiences.
The Pixel 10 rollout was about selling a new AI-enabled
to Normies. And as much as it focused on some of the image features, it wasn't about advanced
AI workflows, it was about how they could do things like zoom in better. In fact, although they
didn't say anything about Nanobanana, some of the image generation features on that phone
were interesting enough that people speculated that Google had actually put the model onto the
phone without really drawing attention to it. In any case, after weeks of anticipation,
Google confirmed that they were in fact the ones behind Nanobanana and we got the full reveal.
Officially, the model is called Gemini 2.5 Flash.
So despite the resonance of the Neme version of the name,
we haven't quite escaped AI model naming convention hell just yet.
The model is now available as a free preview in Google AI Studio,
in the Gemini API API and in the Vertex AI platforms.
Now, in terms of the features that Google showcased,
the model can edit images to change backgrounds, clothing settings,
basically whatever you like using a plain English prompt.
The model is also capable of blending two images together,
for example, combining two different character images and having them interact.
One powerful way to use the model is a workflow that Google calls multi-turn editing.
Once you test the model, you'll realize it can easily fall over if you ask for multiple edits at once,
so it's far better to go one step at a time.
Google demonstrated this feature by filling out an empty room with a bookshelf, a couch, and a coffee table, one at a time.
The final feature that Google showcased was the ability to take a style or a theme from one image
and apply it to a different context.
They showed the model taking the pattern on a butterfly's wing and a powerful.
applying it to address. And at this point, I'm realizing that I should have given you a caveat
earlier, that while all of these episodes are a little bit better using video usually,
this one being about an image model is one where you're really going to benefit from going
and watching it on YouTube or on Spotify. And of course, as with any new model, the company
had to share the benchmarks. Logan Kilpatrick, the product lead for Google AI Studio, shared the
benchmarking, which was very impressive. Nanobanana was unsurprisingly the best model on Elm Arena
and during the testing, but the gap it opened up with rivals was frankly huge.
The new model was about 17% better than the next ranked model.
Flux won context according to their ELO rankings.
GBT40 Image, the model that stormed the world with Ghibli's was ranked slightly below Flux,
suggesting, and this is something we'll talk about a little bit more in just a minute,
that Google now has a big advantage over OpenAI, at least when it comes to this sort of image
generation.
Breaking down the categories, Nanobanana top the rankings in character, creative, infographics,
object and environment, and product recontextualization.
The only category where it didn't outperform was stylization,
where it fell behind GPT4O image and Quinn Image Edit.
Now, as always, big, big caveats and grains of salt when it comes to benchmarks.
Although, frankly, in some ways, I'm more interested in these subjective user preference style of benchmarks
like you get in Elm Arena as opposed to just random tests, which are A, all saturated,
and B, conducted in labs by the company putting out the model and wanting it to do well in the first place.
Now, the other notable technical point about this model is that it's based on Jeveni 2.5 Flash.
That means that it has some of the same limitations.
OpenAI leaker Jimmy Apples wrote,
Very good for what it does, but Flash is a dumb model and it's annoying for complex ideas.
The flip side, of course, is that the model is absurdly cheap and fast for what it does.
Access through the API costs around $0 per image,
which is a quarter of the price of GPT40 image on high detail settings.
Now, for most of our purposes, if you are listening here,
cost doesn't really matter that much at the moment due to the free preview.
The architecture also means that the model inherits the reasoning features of Flash, which can be used during image generation.
As you'd expect, the model has immediately taken over the discourse.
X loves new model releases, and that goes double for a visual model that really hits the sweet spot for engagement.
In fact, frankly, you kind of think that Elon must be fuming over there,
given how hard he has tried to turn Grox Image Generation capabilities into a meme recently,
without getting anywhere near the viral resonance of the Nana-Banana model.
Pretty universally, people are incredibly impressed.
Kevin Olivier edited a series of iconic sports scenes, commenting,
testing Gemini 2.5 Flash image, aka Nanobanana,
standout feature precise localized edits with context preserved.
I took iconic sports photos and anime-fied the athletes,
all while maintaining the rest of the image lighting, etc.
My favorite is the image of Jordan on his way up to dunk.
AI consultant Hadi Khan made a four-paint image of different styles of platypus writing,
Gemini Nano generating multiple styles in the same image in parallel,
one shot, one prompt, one result.
was not possible nor imaginable a couple of years ago.
Princeton CS major Chrissy Cat posted,
The best thing about Gemini ImageGen is that the original inputs are preserved,
no shiny cartoonification.
She shared an example of using the prompt to make a rom-com movie poster with these two characters,
adding Edward Cullen from Twilight,
and a shot of the lead from the summer I turned pretty,
and said, you can just imagine what suss things people are going to make.
To the extent that anyone is disappointed, it seems to be on edge cases.
Prince wrote, testing nanobanata today,
not impressed with anything other than some of its image editing capabilities.
First, world knowledge, with the prompt, picture of Shakespeare writing the famous opening line
from Mark Caesar's speech at Caesar's funeral, and picture of Nabokov writing the famous line about
arics and angels. The text that generates is slop. It doesn't have internal world knowledge.
He continues, though. However, unlike Chatchip-T, this model knows to keep the text facing the author.
Chat-GPT will often have it easily readable by the viewer instead, so it looks like the writer is writing
the text upside down. As to the portrait of Nabokov himself, Dear Nanobanana, who is
this man? Chat Chachybtee gets him perfectly. Prins ran through a series of other nitpicks
noting in particular that the model doesn't handle large quantities of text very well. However,
he did point out that it perfectly handled editing a hypnotic specter for Magic the Gathering
to be holding a modern weapon perfectly, including all the card text. And if you've ever seen a
turn-one swamp into a dark ritual, into a hippie, the version where he's carrying a massive
machine gun is even more intimidating. By the way, kudos to any of you who got that reference.
We'll come back at the end to some of the things that the model can't do really well yet,
but it is notable that even the so-called critiques that I could find still had nice things to say
about some parts of this. Now, one of the big takeaways for many is that Google is now firmly in
the lead when it comes to multimodal LLMs. AI engineer Mark Kretschman wrote, Google is becoming the
clear leader in multimodal AI. Other labs like OpenAI will have a hard time catching up. Google has
the hardware and TPUs and data in YouTube advantage. I don't see anyone keeping up with Google in the
near future. Investor Mark Turk wrote, somehow Google went from being perceived as an AI loser
a year or two ago to releasing the most exciting AI products in 2025, V-O-3, Genie 3,
and now Banana, aka Gemini 2.5 Flash Image.
Marketers spend endless hours building segments and journeys for campaigns.
High Touch just announced a new round of funding to change that with their latest product,
AI decisioning.
Instead of manually deciding which message goes to which customer,
AI decisioning deploys agents that act like a personal marketing specialist for every single
user. These agents learn from your own data warehouse, work across your existing tools like
Salesforce Marketing Cloud, Brays, and Iterable, and continuously optimize for the outcomes you care
about, like maximizing lifetime value. It's fully transparent, enterprise secure, and already
in use by teams at brands like PetSmart, Whoop, Weight Watchers, and Funrise. So if you're
ready to move past static campaigns and into the age of AI-driven marketing, check out Hightouch's
AI decisioning at hightouch.com. That's hightouch.com.
This episode is brought to you by Blitzy, the Enterprise Autonomous Software Development Platform with infinite code context.
Blitzy uses thousands of specialized AI agents that think for hours to understand enterprise-scale codebases with millions of lines of code.
Enterprise engineering leaders start every development sprint with the Blitzy platform, bringing in their development requirements.
The Blitzy platform provides a plan, then generates and pre-compiles code for each task.
Blitzy delivers 80% plus of the development work autonomously while providing a guide for the final 20% of human development work required to
the sprint. Public companies are achieving a 5x engineering velocity increase when
incorporating Blitzie as their pre-IDE development tool, pairing it with their coding
co-pilot of choice to bring an AI-Native STLC into their org. Blitzy is providing a limited
time, 30-day free proof-of-concept for qualifying enterprises. The team will provide a 5x
velocity increase on a real development project in your org. Visit blitzy.com and press book
demo to learn how Blitzie transforms your STLC from AI-assisted to AI Native. That's BLYTZY
com. If you are a regular listener, you will have heard about Super Intelligence Agent Readiness Audits at
this point. But I wanted to tell you today about the full suite of Agent Readiness products that
go beyond just the initial readiness report. Over the last six months, Super Intelligence has built
out an entire Agent Planning Suite. We help you move from discovery to planning to implementation.
After you've completed your Agent Readiness Audits, we help you double-click on your most
important use cases with what we call our use case planning reports. These reports are going to
help you understand what sort of technical preparation you need to do to be ready for a use case,
what challenges you might face in implementation, and whether you should be thinking about building,
buying, partnering, or some combination. After that, you can even get a spec document in what we
call our technical blueprint that gives either your developers or the developers of the partner
you work with what they need to build exactly the agent that you're looking for.
If you want to learn more about superintelligence agent planning suite, we've built a custom
GBT to answer your questions. Just go to bit.ly slash super super agent. That's bit.l.ly slash super super
agent, all one word. And if you have any questions, the agent can even help you book an appointment
with our team. So let's talk now about seven new use cases that Nanobanana opens up.
Like I said at the beginning, every time we get a new model, it tips into reality some set of
use cases that were not possible for. Sometimes the changes.
very small and incremental. Sometimes it's a little bit more dramatic. The first thing that jumped
to most people's mind when it came to Nanobanana was that this model absolutely kills Photoshop.
It is, of course, not the first model that's been capable of editing images, but it's the first
one that has this level of quality and consistency. Ethan Malik writes, it's impressive, crossing a
threshold that goes beyond toy, although it's a pretty fun toy too. And indeed, if you had to
summarize the general sentiment, outside, of course, the hype boys who are just going to say,
this changes everything no matter what was released, this notion that there has been a precipice
into professional possibility crossed seems to be where a lot of people are.
One of the funniest post that I saw was from AI writer Andre Berkov, who basically tried to do
a takedown, mostly seeming annoyed at those quote-unquote AI influencers who were screaming
that this was a wild model. He shared a grainy black and white photo and said, transform this black
and white photo into a color photo where the top half of my body is visible and I'm in a nice
office background wearing something casual but looking professional, not a suit. And he said,
They screamed that the model perfectly preserves your face while changing everything around it.
I tested it and what they say is a lie. The model has clear problems with removing the background.
The person with a changed background looks photoshopped into it.
I don't know, man. Yes, the resulting generated image looks photoshopped. It doesn't have the sort
of fidelity to reality that some of the other examples that we've seen are.
But frankly, given that Andre gave it an incredibly grainy black and white picture of his face
and didn't ask it to preserve that style, it does a pretty impressive job.
And certainly to the extent that we are talking about whether or not this displaces,
taking 15 to 30 minutes in Photoshop to do something,
or even much longer when it comes to a comprehensive thing like this,
you've got to think that it's going to challenge those traditional methodologies.
Sticking with the theme of this model killing entire categories of products,
our next use case that this enables is about try-on startups,
where basically every single one of them is now faced with a very big problem.
This model can natively do what those apps have been stringing together prompts and frameworks to achieve,
which is not to say that the tweaks that they can do to improve the process or the U.S. they put around it to make it more performant for a specific professional use case, can't carry them.
But this is a feature that's now going to show up in a native Google app for free by the end of the year.
writes AI for success RIP 379 startups.
He tested the feature by placing a chat-chipy t-shirt on Sam Altman and noticed the incredible attention to detail, saying,
I can't believe it replaced the entire t-shirt
and still kept that tiny microphone intact from the original image.
Now, of course, this is a much broader trend that we've been seeing
ever since the beginning of ChatGBT's release.
Capabilities that once required a ton of scaffolding
are just getting baked into foundation models.
And of course, in this particular case, what that means
is that this use case is obviously going to become rapidly commoditized
and just total table stakes when it comes to basically any sort of shopping platform.
AI Warper wrote,
hard to find motivation to build anything right now
and Nanobanana will just obliterate you in a week or two. This is another great reminder of what
Sam Altman was talking about when he said that the best builders will not be building in directions
where advancements in the foundation models upend their model, but instead where new updates
in the underlying models actually improve what they're doing. Next use case is one that people have
been exploring again ever since the early days of mid-jurney and stable diffusion. Restoring old photos
isn't necessarily a hyper-commercial use case, but it is beloved, and Nanobanana really takes it to the
next level. Rodrigo Broussaint, a professional photographer and consultant at FreePick, wrote,
Nothing has ever been like this before. I spent countless hours working on photo restoration and
colorization solutions long before AI was a thing. Nothing compares to this. Truly remarkable.
He ran through a series of incredibly impressive examples which are worth checking out if you're
interested in this niche. One of my favorites was this classic photo of Winston Churchill.
The model not only nailed a realistic colorization, but it also managed to keep the brooding
intensity of the photo based on its choice of lighting and saturation. In other words, the point
isn't just that colorization is now possible when it wasn't before, if that we're now at the point
where photo restoration can be a one-click feature with near perfect results. Now, so far,
these examples have all been about doing things that were already possible 10 times faster,
10 times cheaper, and 10 times better. Where things start to require a little more imagination
is when we start to look at features that simply were not possible before. Nanobanana
inherits Gemini's world understanding, so it has a strong understanding of real-world facts
it can use in its generation.
Bill of Al-Badhu writes,
Since Nanobanana has Gemini's world knowledge, you can just upload screenshots of the real world
and ask it to annotate stuff for you.
He included his prompt.
You're a location-based AR experience generator.
Highlight the point of interest in this image and annotate relevant information about it.
The outputs were various San Francisco landmarks, including the ferry building, the Transamerica
pyramid, and the Palace of the Fine Arts, all annotated with key stats.
There also seems to be a kind of world model embedded in the training as well.
This makes the model state of the art in doing perspective transitions.
Peter Levels demonstrated the model flipping the perspective on a first-person image of a person
holding a cup to show the man lying in bed with the cup.
He also showed that the bottle can take an image of a face and generate a full-body image
from every possible perspective.
One push to the limits, this capability is unlike anything we've seen before.
Benjamin DeKracker, a former X-AI engineer, generated an image of a city street.
Nanobanana was then able to change the perspective to a top-down view and point at the location
and orientation of the cameraman from the first photo.
Now the specific ways in which people will use this, I'm not totally sure, but it's such a powerful
capability it's hard not to imagine that people will find ways to take advantage of it and
pretty quickly.
Moving on to our fifth use case, one of the implications of having world knowledge is the capability
to think about 3D shapes.
Linus Eccanstand noted that Nanobanana is capable of generating really capable 3D meshes.
He commented,
Yeah, we were already there, there are many image-to-3D models.
What I like here is if we can accurately allow for multiple images of an object to be uploaded,
we have more control over the final output.
Now it should also be noted that this is the first image-to-3-mesh model that combines reasoning
and prompt adherence to the level the Gemini Flash is capable of.
One obvious use for this is generating game assets, although there doesn't seem to be an ability
to export meshes so far, for the moment the best you can do is combine the image
with other tools to create 3D assets for games and animations.
Still, the incredible consistency means you can generate a huge variety of poses, variations,
and angles, which makes a big difference for this actual production-style use case.
Here's another wild example taking a pretty low-quality image of a building at night and turning
it into a production-quality isometric game asset.
D-Das, again in Menlo Ventures wrote,
The best use case I heard so far is taking objects out of pictures and creating 3D models
from them for games.
Anything from a movie can be put into a game.
Now, one of the interesting things that popped up looking at all these use cases was the
extent to which Nanobanana is just taking out entire workflows.
Rather than just an edit here or there, some people are using the model multiple times to
carry their ideas all the way through.
A filmmaker called Kevin shared his use of Nanobanana to block out a scene, fiddle with elements,
and then run it through an image-to-video model.
He wrote, I've been using Nanobanana or Gemini 2.5 Flash image, as it's called, quite a
lot over the weekend.
It's the best way for me to achieve planned shots in a more direct-end.
and faster manner with greater control. Others demonstrated the workflow for product photos.
Because the model can do perspective changes, text and context shifts so easily that you can transform
a single product shot into as many different ones as you need. Nathan Snell and AI retention
marketing specialist posted, Gemini 2.5 flash image is really, really good, one-shot variation of a
hero. This might be the breakthrough we were waiting for to get the statics over the line. One of the big
issues with AI product images so far has been generating natural-looking shots of the product in someone's
hands or of a person wearing it, but the improvement with Nanobanana is a big jump up in that area.
The same is true for other applications outside advertising. VFX artist Paige Piscan noted
that you can now take one photo of a model and generate an entire photo shoot. Flowers commented
on how big these changes could be, writing, AI image generation isn't yet able to replace
fashion and editorial photographers and retouchers because resolution, detail, body coherence, and control
just aren't there yet in production quality. But the first tool that nails this will wipe out
not one job but ten at once. Photographer, creative director, art director, stylist, hair, makeup,
model, retoucher, producer, set designer, all gone. AI images are 1,000 to 5,000 times cheaper.
That's the thing with AI. Your job might be safe for now, but when it hits, it takes your whole
industry with it. Now, of course, we don't actually know how this is going to play out,
and two things can be true at once. On the one hand, directionally, it seems like there is some
amount of inevitability to what Flowers is saying there. The cost differential is going to mean
that for many types of photoshoot use cases, AI is just going to become the default option.
However, what we don't know yet is one, what sort of skills are going to be necessary to work
with the AI? We have a tendency when a new model shows up to share all of the random generations
of random people and be so impressed with what it can do without remembering that when it comes
to actual professional use cases that go into production, there's going to be a wild gap
between the people who are actually good at this good enough to pay to do it, and those who are just
doing it for fun. In other words, just because the floor comes up for everyone in terms of their
ability to create, it doesn't mean that the ceiling comes down, and companies are always going to
want to go with people who can use it at that ceiling level. The second thing is we don't know
how many photoshoots to continue with this example right now are not happening because of cost.
In other words, it's really hard to predict anything other than the simple fact that things
will change. Now, our last of seven new use cases is really more of a combination.
than an individual one on its own, when you bring all of these capabilities together and combine
them with other tools, it creates for just some wild new possibilities.
One of the big use cases for GPT40 Image when it was released, for example, was making infographics
and posters. Given that it was one of the first models that did a good job on text and was attached
to a foundation LLM, it meant that it had enough general knowledge to create believable infographics
and the actual ability to do so. Moving back to Nanobanana, AI educator Zane Shaw, posted,
Wow, I asked Gemini 2.5 image, aka Nanobanana, for interleaved text and images,
using TTS to narrate the text, animated the images,
and in minutes I had this whole animated explainer video,
full of complex 3D graphics and diagrams explaining the science end-to-end.
He showed a short video explaining how water freezes complete with molecule animations.
Now, obviously, that isn't strictly about nanobanana,
but it shows how this big improvement in image-gen quality can be stitched together
with various other tools to produce professional quality work.
Now, I should point out that we're at the stage with this model where mostly people are just
focused on discovering what it does really well.
It won't be long before we also find out where its limitations are.
A.I. Consultant Newfar Gaspar, for example, ran it through a set of three knowledge
worker tasks, infographic manipulation and data editing, a slide visual fixed and edit, and a
complex infographic generation. And while it did better than many previous models, she still
found that some of the text generation was problematic. She wrote bottom line with many text
captions the model struggles. It's better to generate blank placeholders and add the text than another
app. So yes, even with all this excitement, the banana is still not perfect yet, but overall,
hard not to be excited about this new update. I certainly can't wait to dig in there and try
things out. For now that, that's going to do it for today's AI Daily Brief. Appreciate you
listening or watching as always, and until next time, peace.
