The AI Daily Brief: Artificial Intelligence News and Analysis - 7 AI Use Cases Unlocked By Nano Banana

Episode Date: August 28, 2025

Today's AI Daily Brief covers the groundbreaking release of Google's Nano Banana image generation model, which has taken the AI community by storm over the past few weeks. Google officially re...vealed that Nano Banana is actually Gemini 2.5 Flash, now available as a free preview in Google AI Studio, offering unprecedented image editing capabilities with perfect object consistency and incredible prompt adherence. The model dominates benchmarks, scoring 17% higher than competitors like Flux, and opens up seven transformative use cases from professional photo editing to 3D mesh generation. This represents a major leap forward in multimodal AI that could reshape entire industries from photography to game development.Brought to you by:KPMG – Discover how AI is transforming possibility into reality. Tune into the new KPMG 'You Can with AI' podcast and unlock insights that will inform smarter decisions inside your enterprise. Listen now and start shaping your future with every episode. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.kpmg.us/AIpodcasts⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Blitzy.com - Go to ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ to build enterprise software in days, not months Vanta - Simplify compliance - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://vanta.com/nlw⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Plumb - The automation platform for AI experts and consultants ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://useplumb.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The Agent Readiness Audit from Superintelligent - Go to ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://besuper.ai/ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠to request your company's agent readiness score.The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Subscribe to the newsletter: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? nlw@breakdown.network

Transcript
Discussion (0)
Starting point is 00:00:00 This podcast is supported by Google. Hey everyone, Shrehta here from Google DeepMind. The Geminae 2.5 family of models is now generally available. 2.5 Pro, our most advanced model, is great for reasoning over complex tasks. 2.5 Flash finds the sweet spot between performance and price. And 2.5 Flashlight is ideal for low-latency, high-volume tasks. Start building in Google AI Studio at AI.dev. Today on the AI Daily Brief, seven new use cases opened up by Google's new
Starting point is 00:00:35 nano-banana image generation model. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, Gemini, High Touch, Blitzian, Superintelligent. And to get an ad-free version of the show, go to patreon.com slash AI Daily Brief. Now, two quick things about today's show. one, as sometimes happens when we have an exciting new model, the main episode got a little bit
Starting point is 00:01:05 long, so this will be a main only type of episode. There really are kind of two parts. The first half or so is about the background context and the announcement of the new nanobanana model. And then the second half is the seven new use cases that it opens up. So you can kind of think about it in that way when it comes to how to divide the episode. And in the case, our normal format headlines, and then the main episode will be back tomorrow. Second thing to note, after recording for 40 minutes, I discovered that it had not been recording with my actual podcast mic and instead was recording directly into the laptop. However, I am in Vegas for a keynote today and so am not able to go back and re-record the full episode. We've done our best to use a variety of tools, AI, and otherwise to make the sound quality as good as it can be,
Starting point is 00:01:48 but apologies that it's a little bit lower than normal. With all that out of the way, let's dive into this very exciting new model. Welcome back to the AI Daily Brief. Today we are talking about the image model that has had people in the image model that has had people incredibly excited for the last couple of weeks. It is finally out. We got confirmation who it was from. And already people are discovering new use cases that were not possible, at least in the way that they are now, before this release came. It's a great reminder of the fact that we still really do live right on the horizon between possible and impossible. And that every time there's a new model,
Starting point is 00:02:21 even if it's not obvious at first, there are some set of new things that move over into the realm of the possible that simply were not before. But we're getting a little bit ahead. of ourselves with that. A couple of weeks ago, a new image generation model showed up on Elam Arena that was called Nanobanana. I've talked about this previously on the show. If you spend any time on AI Twitter, you've probably seen people talking about it or sharing their generations with it. The model quickly rose to the top of the leaderboard and gathered a ton of praise for its state-of-the-art performance. Now, interestingly, unlike the last few leaps in image model quality, people weren't just talking about how realistic the images were. Instead, the real standout was
Starting point is 00:03:00 the model's ability to edit images. You could generate a single shot and then modify it with seemingly perfect object consistency and incredible prompt adherence. Now, this is one of those things that I think people have always imagined as what they wanted out of an AI generated image model, but hadn't really been there yet. If you spend any time with Mid Journey or the tools inside ChatGabit, you'll know that it's pretty frequently a crapshoot of trying to get the model to do what you want to do. In fact, for some, like Mid Journey, I don't even recommend trying to get the model to do what you want to do, I basically recommend you winding it up, pointing in a direction, and letting it do what it wants. Now, there have been some big updates recently. For example,
Starting point is 00:03:38 ideogram, which I use every day for thumbnails, added a new updated character consistency patch, which made a huge difference. But still, the idea of being able to take a base image and turn it into whatever you want in a way that was coherent is something that people have both wanted and simply haven't had access to. Didi Das of Menlo Ventures said, this is the next generation of filters that we've been promised forever. In any case, once the hype started building, Google folks began winking at the crowd on X implying that they were, in fact, the ones testing in stealth.
Starting point is 00:04:08 Those hints, however, backfired, as they led many to think that the Google Pixel 10 launch last week was going to double as the nanobanana rollout. Many were disappointed then when the event came and went with no mention of the new model. Now, in retrospect, keeping them separated makes sense. These are two very different launches for very different audiences. The Pixel 10 rollout was about selling a new AI-enabled
Starting point is 00:04:28 to Normies. And as much as it focused on some of the image features, it wasn't about advanced AI workflows, it was about how they could do things like zoom in better. In fact, although they didn't say anything about Nanobanana, some of the image generation features on that phone were interesting enough that people speculated that Google had actually put the model onto the phone without really drawing attention to it. In any case, after weeks of anticipation, Google confirmed that they were in fact the ones behind Nanobanana and we got the full reveal. Officially, the model is called Gemini 2.5 Flash. So despite the resonance of the Neme version of the name,
Starting point is 00:05:00 we haven't quite escaped AI model naming convention hell just yet. The model is now available as a free preview in Google AI Studio, in the Gemini API API and in the Vertex AI platforms. Now, in terms of the features that Google showcased, the model can edit images to change backgrounds, clothing settings, basically whatever you like using a plain English prompt. The model is also capable of blending two images together, for example, combining two different character images and having them interact.
Starting point is 00:05:26 One powerful way to use the model is a workflow that Google calls multi-turn editing. Once you test the model, you'll realize it can easily fall over if you ask for multiple edits at once, so it's far better to go one step at a time. Google demonstrated this feature by filling out an empty room with a bookshelf, a couch, and a coffee table, one at a time. The final feature that Google showcased was the ability to take a style or a theme from one image and apply it to a different context. They showed the model taking the pattern on a butterfly's wing and a powerful. applying it to address. And at this point, I'm realizing that I should have given you a caveat
Starting point is 00:05:56 earlier, that while all of these episodes are a little bit better using video usually, this one being about an image model is one where you're really going to benefit from going and watching it on YouTube or on Spotify. And of course, as with any new model, the company had to share the benchmarks. Logan Kilpatrick, the product lead for Google AI Studio, shared the benchmarking, which was very impressive. Nanobanana was unsurprisingly the best model on Elm Arena and during the testing, but the gap it opened up with rivals was frankly huge. The new model was about 17% better than the next ranked model. Flux won context according to their ELO rankings.
Starting point is 00:06:31 GBT40 Image, the model that stormed the world with Ghibli's was ranked slightly below Flux, suggesting, and this is something we'll talk about a little bit more in just a minute, that Google now has a big advantage over OpenAI, at least when it comes to this sort of image generation. Breaking down the categories, Nanobanana top the rankings in character, creative, infographics, object and environment, and product recontextualization. The only category where it didn't outperform was stylization, where it fell behind GPT4O image and Quinn Image Edit.
Starting point is 00:06:57 Now, as always, big, big caveats and grains of salt when it comes to benchmarks. Although, frankly, in some ways, I'm more interested in these subjective user preference style of benchmarks like you get in Elm Arena as opposed to just random tests, which are A, all saturated, and B, conducted in labs by the company putting out the model and wanting it to do well in the first place. Now, the other notable technical point about this model is that it's based on Jeveni 2.5 Flash. That means that it has some of the same limitations. OpenAI leaker Jimmy Apples wrote, Very good for what it does, but Flash is a dumb model and it's annoying for complex ideas.
Starting point is 00:07:29 The flip side, of course, is that the model is absurdly cheap and fast for what it does. Access through the API costs around $0 per image, which is a quarter of the price of GPT40 image on high detail settings. Now, for most of our purposes, if you are listening here, cost doesn't really matter that much at the moment due to the free preview. The architecture also means that the model inherits the reasoning features of Flash, which can be used during image generation. As you'd expect, the model has immediately taken over the discourse. X loves new model releases, and that goes double for a visual model that really hits the sweet spot for engagement.
Starting point is 00:08:01 In fact, frankly, you kind of think that Elon must be fuming over there, given how hard he has tried to turn Grox Image Generation capabilities into a meme recently, without getting anywhere near the viral resonance of the Nana-Banana model. Pretty universally, people are incredibly impressed. Kevin Olivier edited a series of iconic sports scenes, commenting, testing Gemini 2.5 Flash image, aka Nanobanana, standout feature precise localized edits with context preserved. I took iconic sports photos and anime-fied the athletes,
Starting point is 00:08:28 all while maintaining the rest of the image lighting, etc. My favorite is the image of Jordan on his way up to dunk. AI consultant Hadi Khan made a four-paint image of different styles of platypus writing, Gemini Nano generating multiple styles in the same image in parallel, one shot, one prompt, one result. was not possible nor imaginable a couple of years ago. Princeton CS major Chrissy Cat posted, The best thing about Gemini ImageGen is that the original inputs are preserved,
Starting point is 00:08:52 no shiny cartoonification. She shared an example of using the prompt to make a rom-com movie poster with these two characters, adding Edward Cullen from Twilight, and a shot of the lead from the summer I turned pretty, and said, you can just imagine what suss things people are going to make. To the extent that anyone is disappointed, it seems to be on edge cases. Prince wrote, testing nanobanata today, not impressed with anything other than some of its image editing capabilities.
Starting point is 00:09:16 First, world knowledge, with the prompt, picture of Shakespeare writing the famous opening line from Mark Caesar's speech at Caesar's funeral, and picture of Nabokov writing the famous line about arics and angels. The text that generates is slop. It doesn't have internal world knowledge. He continues, though. However, unlike Chatchip-T, this model knows to keep the text facing the author. Chat-GPT will often have it easily readable by the viewer instead, so it looks like the writer is writing the text upside down. As to the portrait of Nabokov himself, Dear Nanobanana, who is this man? Chat Chachybtee gets him perfectly. Prins ran through a series of other nitpicks noting in particular that the model doesn't handle large quantities of text very well. However,
Starting point is 00:09:50 he did point out that it perfectly handled editing a hypnotic specter for Magic the Gathering to be holding a modern weapon perfectly, including all the card text. And if you've ever seen a turn-one swamp into a dark ritual, into a hippie, the version where he's carrying a massive machine gun is even more intimidating. By the way, kudos to any of you who got that reference. We'll come back at the end to some of the things that the model can't do really well yet, but it is notable that even the so-called critiques that I could find still had nice things to say about some parts of this. Now, one of the big takeaways for many is that Google is now firmly in the lead when it comes to multimodal LLMs. AI engineer Mark Kretschman wrote, Google is becoming the
Starting point is 00:10:26 clear leader in multimodal AI. Other labs like OpenAI will have a hard time catching up. Google has the hardware and TPUs and data in YouTube advantage. I don't see anyone keeping up with Google in the near future. Investor Mark Turk wrote, somehow Google went from being perceived as an AI loser a year or two ago to releasing the most exciting AI products in 2025, V-O-3, Genie 3, and now Banana, aka Gemini 2.5 Flash Image. Marketers spend endless hours building segments and journeys for campaigns. High Touch just announced a new round of funding to change that with their latest product, AI decisioning.
Starting point is 00:11:02 Instead of manually deciding which message goes to which customer, AI decisioning deploys agents that act like a personal marketing specialist for every single user. These agents learn from your own data warehouse, work across your existing tools like Salesforce Marketing Cloud, Brays, and Iterable, and continuously optimize for the outcomes you care about, like maximizing lifetime value. It's fully transparent, enterprise secure, and already in use by teams at brands like PetSmart, Whoop, Weight Watchers, and Funrise. So if you're ready to move past static campaigns and into the age of AI-driven marketing, check out Hightouch's AI decisioning at hightouch.com. That's hightouch.com.
Starting point is 00:11:40 This episode is brought to you by Blitzy, the Enterprise Autonomous Software Development Platform with infinite code context. Blitzy uses thousands of specialized AI agents that think for hours to understand enterprise-scale codebases with millions of lines of code. Enterprise engineering leaders start every development sprint with the Blitzy platform, bringing in their development requirements. The Blitzy platform provides a plan, then generates and pre-compiles code for each task. Blitzy delivers 80% plus of the development work autonomously while providing a guide for the final 20% of human development work required to the sprint. Public companies are achieving a 5x engineering velocity increase when incorporating Blitzie as their pre-IDE development tool, pairing it with their coding co-pilot of choice to bring an AI-Native STLC into their org. Blitzy is providing a limited
Starting point is 00:12:23 time, 30-day free proof-of-concept for qualifying enterprises. The team will provide a 5x velocity increase on a real development project in your org. Visit blitzy.com and press book demo to learn how Blitzie transforms your STLC from AI-assisted to AI Native. That's BLYTZY com. If you are a regular listener, you will have heard about Super Intelligence Agent Readiness Audits at this point. But I wanted to tell you today about the full suite of Agent Readiness products that go beyond just the initial readiness report. Over the last six months, Super Intelligence has built out an entire Agent Planning Suite. We help you move from discovery to planning to implementation. After you've completed your Agent Readiness Audits, we help you double-click on your most
Starting point is 00:13:07 important use cases with what we call our use case planning reports. These reports are going to help you understand what sort of technical preparation you need to do to be ready for a use case, what challenges you might face in implementation, and whether you should be thinking about building, buying, partnering, or some combination. After that, you can even get a spec document in what we call our technical blueprint that gives either your developers or the developers of the partner you work with what they need to build exactly the agent that you're looking for. If you want to learn more about superintelligence agent planning suite, we've built a custom GBT to answer your questions. Just go to bit.ly slash super super agent. That's bit.l.ly slash super super
Starting point is 00:13:46 agent, all one word. And if you have any questions, the agent can even help you book an appointment with our team. So let's talk now about seven new use cases that Nanobanana opens up. Like I said at the beginning, every time we get a new model, it tips into reality some set of use cases that were not possible for. Sometimes the changes. very small and incremental. Sometimes it's a little bit more dramatic. The first thing that jumped to most people's mind when it came to Nanobanana was that this model absolutely kills Photoshop. It is, of course, not the first model that's been capable of editing images, but it's the first one that has this level of quality and consistency. Ethan Malik writes, it's impressive, crossing a
Starting point is 00:14:27 threshold that goes beyond toy, although it's a pretty fun toy too. And indeed, if you had to summarize the general sentiment, outside, of course, the hype boys who are just going to say, this changes everything no matter what was released, this notion that there has been a precipice into professional possibility crossed seems to be where a lot of people are. One of the funniest post that I saw was from AI writer Andre Berkov, who basically tried to do a takedown, mostly seeming annoyed at those quote-unquote AI influencers who were screaming that this was a wild model. He shared a grainy black and white photo and said, transform this black and white photo into a color photo where the top half of my body is visible and I'm in a nice
Starting point is 00:15:04 office background wearing something casual but looking professional, not a suit. And he said, They screamed that the model perfectly preserves your face while changing everything around it. I tested it and what they say is a lie. The model has clear problems with removing the background. The person with a changed background looks photoshopped into it. I don't know, man. Yes, the resulting generated image looks photoshopped. It doesn't have the sort of fidelity to reality that some of the other examples that we've seen are. But frankly, given that Andre gave it an incredibly grainy black and white picture of his face and didn't ask it to preserve that style, it does a pretty impressive job.
Starting point is 00:15:35 And certainly to the extent that we are talking about whether or not this displaces, taking 15 to 30 minutes in Photoshop to do something, or even much longer when it comes to a comprehensive thing like this, you've got to think that it's going to challenge those traditional methodologies. Sticking with the theme of this model killing entire categories of products, our next use case that this enables is about try-on startups, where basically every single one of them is now faced with a very big problem. This model can natively do what those apps have been stringing together prompts and frameworks to achieve,
Starting point is 00:16:05 which is not to say that the tweaks that they can do to improve the process or the U.S. they put around it to make it more performant for a specific professional use case, can't carry them. But this is a feature that's now going to show up in a native Google app for free by the end of the year. writes AI for success RIP 379 startups. He tested the feature by placing a chat-chipy t-shirt on Sam Altman and noticed the incredible attention to detail, saying, I can't believe it replaced the entire t-shirt and still kept that tiny microphone intact from the original image. Now, of course, this is a much broader trend that we've been seeing ever since the beginning of ChatGBT's release.
Starting point is 00:16:39 Capabilities that once required a ton of scaffolding are just getting baked into foundation models. And of course, in this particular case, what that means is that this use case is obviously going to become rapidly commoditized and just total table stakes when it comes to basically any sort of shopping platform. AI Warper wrote, hard to find motivation to build anything right now and Nanobanana will just obliterate you in a week or two. This is another great reminder of what
Starting point is 00:17:01 Sam Altman was talking about when he said that the best builders will not be building in directions where advancements in the foundation models upend their model, but instead where new updates in the underlying models actually improve what they're doing. Next use case is one that people have been exploring again ever since the early days of mid-jurney and stable diffusion. Restoring old photos isn't necessarily a hyper-commercial use case, but it is beloved, and Nanobanana really takes it to the next level. Rodrigo Broussaint, a professional photographer and consultant at FreePick, wrote, Nothing has ever been like this before. I spent countless hours working on photo restoration and colorization solutions long before AI was a thing. Nothing compares to this. Truly remarkable.
Starting point is 00:17:40 He ran through a series of incredibly impressive examples which are worth checking out if you're interested in this niche. One of my favorites was this classic photo of Winston Churchill. The model not only nailed a realistic colorization, but it also managed to keep the brooding intensity of the photo based on its choice of lighting and saturation. In other words, the point isn't just that colorization is now possible when it wasn't before, if that we're now at the point where photo restoration can be a one-click feature with near perfect results. Now, so far, these examples have all been about doing things that were already possible 10 times faster, 10 times cheaper, and 10 times better. Where things start to require a little more imagination
Starting point is 00:18:14 is when we start to look at features that simply were not possible before. Nanobanana inherits Gemini's world understanding, so it has a strong understanding of real-world facts it can use in its generation. Bill of Al-Badhu writes, Since Nanobanana has Gemini's world knowledge, you can just upload screenshots of the real world and ask it to annotate stuff for you. He included his prompt. You're a location-based AR experience generator.
Starting point is 00:18:36 Highlight the point of interest in this image and annotate relevant information about it. The outputs were various San Francisco landmarks, including the ferry building, the Transamerica pyramid, and the Palace of the Fine Arts, all annotated with key stats. There also seems to be a kind of world model embedded in the training as well. This makes the model state of the art in doing perspective transitions. Peter Levels demonstrated the model flipping the perspective on a first-person image of a person holding a cup to show the man lying in bed with the cup. He also showed that the bottle can take an image of a face and generate a full-body image
Starting point is 00:19:05 from every possible perspective. One push to the limits, this capability is unlike anything we've seen before. Benjamin DeKracker, a former X-AI engineer, generated an image of a city street. Nanobanana was then able to change the perspective to a top-down view and point at the location and orientation of the cameraman from the first photo. Now the specific ways in which people will use this, I'm not totally sure, but it's such a powerful capability it's hard not to imagine that people will find ways to take advantage of it and pretty quickly.
Starting point is 00:19:31 Moving on to our fifth use case, one of the implications of having world knowledge is the capability to think about 3D shapes. Linus Eccanstand noted that Nanobanana is capable of generating really capable 3D meshes. He commented, Yeah, we were already there, there are many image-to-3D models. What I like here is if we can accurately allow for multiple images of an object to be uploaded, we have more control over the final output. Now it should also be noted that this is the first image-to-3-mesh model that combines reasoning
Starting point is 00:19:57 and prompt adherence to the level the Gemini Flash is capable of. One obvious use for this is generating game assets, although there doesn't seem to be an ability to export meshes so far, for the moment the best you can do is combine the image with other tools to create 3D assets for games and animations. Still, the incredible consistency means you can generate a huge variety of poses, variations, and angles, which makes a big difference for this actual production-style use case. Here's another wild example taking a pretty low-quality image of a building at night and turning it into a production-quality isometric game asset.
Starting point is 00:20:28 D-Das, again in Menlo Ventures wrote, The best use case I heard so far is taking objects out of pictures and creating 3D models from them for games. Anything from a movie can be put into a game. Now, one of the interesting things that popped up looking at all these use cases was the extent to which Nanobanana is just taking out entire workflows. Rather than just an edit here or there, some people are using the model multiple times to carry their ideas all the way through.
Starting point is 00:20:52 A filmmaker called Kevin shared his use of Nanobanana to block out a scene, fiddle with elements, and then run it through an image-to-video model. He wrote, I've been using Nanobanana or Gemini 2.5 Flash image, as it's called, quite a lot over the weekend. It's the best way for me to achieve planned shots in a more direct-end. and faster manner with greater control. Others demonstrated the workflow for product photos. Because the model can do perspective changes, text and context shifts so easily that you can transform a single product shot into as many different ones as you need. Nathan Snell and AI retention
Starting point is 00:21:21 marketing specialist posted, Gemini 2.5 flash image is really, really good, one-shot variation of a hero. This might be the breakthrough we were waiting for to get the statics over the line. One of the big issues with AI product images so far has been generating natural-looking shots of the product in someone's hands or of a person wearing it, but the improvement with Nanobanana is a big jump up in that area. The same is true for other applications outside advertising. VFX artist Paige Piscan noted that you can now take one photo of a model and generate an entire photo shoot. Flowers commented on how big these changes could be, writing, AI image generation isn't yet able to replace fashion and editorial photographers and retouchers because resolution, detail, body coherence, and control
Starting point is 00:22:00 just aren't there yet in production quality. But the first tool that nails this will wipe out not one job but ten at once. Photographer, creative director, art director, stylist, hair, makeup, model, retoucher, producer, set designer, all gone. AI images are 1,000 to 5,000 times cheaper. That's the thing with AI. Your job might be safe for now, but when it hits, it takes your whole industry with it. Now, of course, we don't actually know how this is going to play out, and two things can be true at once. On the one hand, directionally, it seems like there is some amount of inevitability to what Flowers is saying there. The cost differential is going to mean that for many types of photoshoot use cases, AI is just going to become the default option.
Starting point is 00:22:39 However, what we don't know yet is one, what sort of skills are going to be necessary to work with the AI? We have a tendency when a new model shows up to share all of the random generations of random people and be so impressed with what it can do without remembering that when it comes to actual professional use cases that go into production, there's going to be a wild gap between the people who are actually good at this good enough to pay to do it, and those who are just doing it for fun. In other words, just because the floor comes up for everyone in terms of their ability to create, it doesn't mean that the ceiling comes down, and companies are always going to want to go with people who can use it at that ceiling level. The second thing is we don't know
Starting point is 00:23:13 how many photoshoots to continue with this example right now are not happening because of cost. In other words, it's really hard to predict anything other than the simple fact that things will change. Now, our last of seven new use cases is really more of a combination. than an individual one on its own, when you bring all of these capabilities together and combine them with other tools, it creates for just some wild new possibilities. One of the big use cases for GPT40 Image when it was released, for example, was making infographics and posters. Given that it was one of the first models that did a good job on text and was attached to a foundation LLM, it meant that it had enough general knowledge to create believable infographics
Starting point is 00:23:49 and the actual ability to do so. Moving back to Nanobanana, AI educator Zane Shaw, posted, Wow, I asked Gemini 2.5 image, aka Nanobanana, for interleaved text and images, using TTS to narrate the text, animated the images, and in minutes I had this whole animated explainer video, full of complex 3D graphics and diagrams explaining the science end-to-end. He showed a short video explaining how water freezes complete with molecule animations. Now, obviously, that isn't strictly about nanobanana, but it shows how this big improvement in image-gen quality can be stitched together
Starting point is 00:24:19 with various other tools to produce professional quality work. Now, I should point out that we're at the stage with this model where mostly people are just focused on discovering what it does really well. It won't be long before we also find out where its limitations are. A.I. Consultant Newfar Gaspar, for example, ran it through a set of three knowledge worker tasks, infographic manipulation and data editing, a slide visual fixed and edit, and a complex infographic generation. And while it did better than many previous models, she still found that some of the text generation was problematic. She wrote bottom line with many text
Starting point is 00:24:49 captions the model struggles. It's better to generate blank placeholders and add the text than another app. So yes, even with all this excitement, the banana is still not perfect yet, but overall, hard not to be excited about this new update. I certainly can't wait to dig in there and try things out. For now that, that's going to do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.