The AI Daily Brief: Artificial Intelligence News and Analysis - Mind-Reading AI Generates Videos from Brainwaves

Episode Date: May 24, 2023

Researchers have developed a new method called Mind-Video for generating semantically accurate video from fMRI scans. Additionally, on todays Brief, NLW covers: The most important announcements from... Microsoft Build (including AI coming to Windows 11 and ChatGPT + Bing) Google Bard images  Anthropic $450m Series C Naval Ravikant's Airchat  The AI Breakdown helps you understand the most important news and discussions in AI.  Subscribe to The AI Breakdown newsletter: https://theaibreakdown.beehiiv.com/subscribe Subscribe to The AI Breakdown on YouTube: https://www.youtube.com/@TheAIBreakdown Join the community: bit.ly/aibreakdown Learn more: http://breakdown.network/

Transcript
Discussion (0)
Starting point is 00:00:00 On today's AI breakdown, new research that generates video from thoughts. Before that on the brief, a review of Microsoft's AI-focused build event, Google adds images to Bard, and Anthropic raises $450 million. The AI breakdown is a daily video and podcast about the most important news and discussions in AI. If you're enjoying it, please like, subscribe, and share, and learn more at breakdown.network. Welcome back to the AI breakdown brief. All the AI headline news you need in five minutes or less. and today we are kicking off with Microsoft Build. This is their annual developer conference,
Starting point is 00:00:35 so their equivalent of Google's I.O., which was obviously a couple weeks ago, and not surprisingly, AI was the absolute centerpiece of pretty much everything they announced. And let's start right at the top with the biggest announcement, which is that Windows 11 is going completely, fully AI, the type of co-pilot experience that Microsoft debuted with GitHub, is now coming across the entire Windows 11 experience.
Starting point is 00:01:05 Day right, we're introducing Windows co-pilot, making Windows 11 the first PC platform to announce centralized AI assistance. We are extending BingChat plugins to Windows. We are introducing new hybrid AI loop to support AI development across platforms. We're introducing a new dev home designed to help every developer become more productive. And we're introducing new AI features and experiences in the Microsoft store on Windows. Now, the thing that many people are pointing out as the most significant is this idea of copilot now just being a centralized feature of an operating system, right?
Starting point is 00:01:41 This is the first experience that people have when they interact with their computers, is whatever the operating system puts in front of them. So by putting AI and copilot directly in the Windows 11 experience, Microsoft is betting that this is going to transform how not just developers and not just how people on the edge, but how everyone interacts with computers in the future. You probably need a robot who are admittedly, sometimes prone to hyperbole, says,
Starting point is 00:02:08 the biggest news in AI all year, famous last words, Microsoft unveils an entire operating system. Say hello to Windows 11, the ultimate productivity tool, and here's why that matters. With Windows co-pilot, a centralized AI assistant is native on every Windows computer at the operating system level. Access all local files, universally available across the system, lowers the barriers for adoption for onboarding new users to AI, game changer.
Starting point is 00:02:37 Now, the other big thing I think from this announcement is that the partnership between Microsoft and OpenAI is getting even closer in terms of how the average person experiences it. And what that looks like now is that OpenAI's chat GPT is now going to have Bing as the default search experience. Now, you might say, well, of course they will. They've had this big partnership with Microsoft, but the important detail here is that the browsing feature of chat GPT, which is obviously a huge, huge update for chat GPT, now giving it access to the current internet, not just the data set that it was trained on up to 2021, but the current internet is now going to be available, not just for plus users, which it just came to over the last two weeks, but for all users.
Starting point is 00:03:23 So Microsoft is giving the ability of regular free chat GPT users to, browse the web with Bing as part of their chat GPT experience. So what were the other important announcements from the Build Conference? Well, as we've seen over and over and over, plugins are the big theme. And plugins, I think, are just shorthand for this period in AI computing in which developers are retrofitting their applications for an AI world, building the AI powered version of their experiences, and effectively putting them in a new type of app store but for the experiences where people are interacting with AI, such as Bing or ChatGPT.
Starting point is 00:03:59 That's now being brought into the main PC experience alongside Windows 11's general AI integration. Microsoft also announced this new dev home, which is meant to be a better starting point for developers who are interacting with AI and put Microsoft right in the center of that development flow. Some smaller but still noteworthy announcements, they announced a new AI tool for building websites. This is one of a thousand AI website builders. I'm thinking of doing.
Starting point is 00:04:25 doing a video, by the way, comparing AI website building tools. If that's something that interests you, that type of comparison, please comment. And Microsoft has also launched something that's called Fabric, which people are comparing to the code interpreter plugin from chat GPT that they are getting so excited about. Now, by and large, people have been very stoked in this announcement. The one counterpoint that I've seen comes from Sully Omar, who writes, Just saw the Windows 11 AI co-pilot. I am now 100% convinced that Apple will eat Microsoft's lunch.
Starting point is 00:04:55 He says threat on this later, which has not come yet, but when it comes, if there are interesting points, I will be sure to share them. But of course, the big tech AI battle is hardly a one company race. And even on the same day that Microsoft was making all these announcements, Google had its own slate of announcements as well. A big emphasis for Google is bringing AI into their ad tools. And at their marketing live event today, they announced that they're launching Product Studio, which is a new tool that allows people to use generative AI to build better, better product shots for their brands and for the products that they're trying to sell. Some of the tools that will be available is enhancing images, increasing the resolution,
Starting point is 00:05:34 changing different backgrounds. But then there's also going to be copywriting assistance as well, with this sort of chat GPT style experience to help people write better copy that's going to presumably perform better with Google Ads. They even have a set of tools for video enhancement that, again, are all designed to help people perform better. And this is important. important because advertisers are probably the group that's the most worried about these tools in some ways, given how much it's likely to change about the experiences that they've designed their businesses around for the last 20 years. The other big Google announcement was that Bard is starting to include images as responses. This is something that we had seen promised, but it shows that Bard is kind of first out of the gate for these mainstream chatbots in terms of becoming more multimodal, not just responding in text, but figuring out what media is the right way to answer a particular. query or question. Now one could easily be forgiven for thinking that with all of these announcements
Starting point is 00:06:29 from the big tech players and this absolute AI arms race going on, that there is really no room for startup innovation, but the startups in this space are not content to let that be true. Anthropic has announced a $450 million series C raise led by Spark Capital that does have participation from Google, Salesforce, and others. And when I said at the beginning that they're trying to build the best chatbot, that's not exactly an accurate representation. They're trying with Claude to build not only the best chatbot, but the one that's designed to be the most ethical and the most safe. You'll remember that Anthropic is taking an approach that they call constitutional AI rather than reinforcement learning from human feedback, where they train AI on a set
Starting point is 00:07:09 of ethical principles underlying the entire system so that it can make better decisions about how to respond in certain complicated or morally questionable situations. They say the funding will support our efforts to continue building AI products that people can rely on and generate new research about the opportunities and risks of AI. And lastly, one more startup experience that I am noticing more and more people get very excited about is called air chat. Umjad, the CEO of Replit says, I've had multi-day, multi-party nuanced conversations on air chat, and it feels really different as a format from anything else out there. It unlocks something new. One of the company's co-founders is Naval Ravikant, and he writes, social media isn't social, it's performative.
Starting point is 00:07:53 Where's the chit-chat, the banter, the easy laughter? We made something new. Push-to-talk whenever you want with perfect transcripts and AI art. A dinner party in your pocket. Never feel alone again. Early though it is, the reminder of things like air chat is that there is so much still to be discovered about what experiences feel like in this new world and the new world that generative AI enables. Even with as many resources as they're throwing into it, it's madness to think that only Google
Starting point is 00:08:22 or Microsoft or meta, we're going to create the AI experiences of the future. So I called this video the way the next 100 million people come to AI. But was I talking about Windows 11? Or was I talking about something like AirChat? I guess we'll just have to wait and see. That's it for today's AI breakdown brief. I'll be back soon for the main AI breakdown. Mind-reading AI that can generate videos from brainwaves.
Starting point is 00:08:47 Welcome back to the AI breakdown. Today we are talking about some very, very, very, frontier research. It's called mind video, and it is all about video reconstruction from brain activity. So what we're going to do is learn a little bit about this current research, the other research that it came from, and try to get an understanding of what the implications of this type of technology might be in the future. What you're seeing right now is here on the left are the ground truth videos, the videos that were shown to test subjects under fMRI, and then on the right, the videos that were reconstructed from the brain activity that was measured by the fMRI.
Starting point is 00:09:22 I came across this research after reading a thread from Zijiao Chen, one of the PhD students involved in the study. They write, We're witnessing incredible scientific progress in image and text reconstruction from fMRI nowadays. But what about reconstructing video from fMRI? Allow me to introduce our recent preprint, Mind Video. Zijiao goes on. Our work extends previous research on FMRI image reconstruction, Mindy-dash-Viz.
Starting point is 00:09:51 While much success has been, seen in static image reconstruction from non-invasive brain recordings, the field is yet to explore the dynamic realm of video reconstruction. So to get a sense, the two papers in question here, the first came out last November and is called Seeing Beyond the Brain, conditional diffusion model with sparse mask modeling for vision decoding. The more recent research is called cinematic mindscapes, high-quality video reconstruction from brain activity. These are about six months apart. Let's look at that first paper first, seeing beyond the brain. The researchers involved come from the National University of Singapore, the Chinese University of Hong Kong,
Starting point is 00:10:27 and Stanford University. They write, for the first time, our proposed Mindy Viz is capable of decoding fMRI-based brain activities and reconstructing images with not only plausible details, but also accurate semantics and image features, i.e. color, shape, etc. So the way that the Mindy Viz model worked in reconstructing images from brain recordings was a a two-step process. The first is a self-supervised representation learning of fMRI data. This is done using what they call mask modeling in a large latent space. Mask modeling in the context of this paper refers to a technique in which certain parts of the input data are masked or hidden during training. The model is then tasked with predicting the
Starting point is 00:11:07 masked parts from the unmasked parts. This forces the model to learn useful representations of the data in order to successfully predict the masked parts. After that self-supervised representation of the data is learned, the model then uses a latent diffusion process, which is a type of generative model that transforms simple noise distributions into complex data distributions. This is a similar type of approach used in things like stable diffusion, the text image generator. The study reports that their model outperformed other state-of-the-art approaches in both semantic mapping and generation quality by 66% and 41% respectively. Now, if this is just super dense and complex, have no fear. Let's try to simplify at core what was actually happening.
Starting point is 00:11:47 Basically, subjects are shown photos or in the case of the newer study videos while in an fMRI machine. From there, brain activity is continuously recorded. That brain activity is used to create a map of how the brain responds to different types of visual stimuli. From there, whatever sort of proprietary process the researchers are testing is used to reconstruct images, or, again, in the case of the new research video, from that map of the brain's responses to the initial visual stimuli. The model is trying to decode patterns of brain activity and use that to reconstruct the images or video. But let's return to Zijow's thread and talk about the shift from image to video. Again, they write, in moving from image to video, we identified three key gaps.
Starting point is 00:12:28 One, hemodynamic response delay. Two, lack of pixel level and semantic level guidance. Three, need for generation consistency while preserving scene dynamics. To address these, we created mind video. Mind video is a two-module pipeline, with each module trained separately and then fine-tuned together. Our model progressively learns from brain signals, deeply understanding the semantic space through multiple stages in the first model. In the first step, we leverage large-scale unsupervised learning with masked brain modeling to learn general visual fMRI features. We also utilize a spatial temporal attention mechanism to process multiple fMRI in a sliding window.
Starting point is 00:13:06 In the second step, we distill semantic-related features using the multimodality of the annotated. data set. Here we train the fMRI encoder in the clip space with contrastive learning. In the final step, we fine-tune the learn features through co-training with an augmented stable diffusion model, tailored specifically for video generation under fMRI guidance. Our work, she sums up, presents a new method of brain decoding with several contributions. One, flexible two module pipeline. Two, progressive learning of brain features. Three, recovery of high-quality video with accurate semantics. So in terms of their results, they say they approached 85% accuracy in semantic metrics.
Starting point is 00:13:41 Semantic classification tasks are effectively tasks where the model has to understand the meaning or the context behind the video. So if it's a video of a dog running in a park, the model should be able to identify that it's a dog running in a park. So this idea of 85% accuracy in semantic metrics means that 85% of the time, the model correctly identified the context or content of the video. They also said that they had a 0.19 SSIM score. SSIM stands for structural, similar.
Starting point is 00:14:07 similarity index, and this is a metrics that measures the similarity between two pictures. The values in this scale range from negative one to one, with one being identical, and negative one being they are completely different. So the 0.19 score suggests that they are more similar than they are different. Zijiao also says our attention analysis revealed critical insights about how transformers decode fMRI data, dominance of the visual cortex, layer-dependent hierarchy, and progressive semantic learning. This shows our models' biological plausibility and interpretability.
Starting point is 00:14:37 One of the key things here is that with this research, we're not just understanding the AI technology needed to reconstruct images and videos, but we're actually learning about how the brain works at the same time. Now, on their website, they gave a ton of background information and even a Google Drive full of resources that you can see to actually dig into what the results were. But let's wrap up by talking about what the potential implications of this type of study are. Seism and Stoma tweets, one of the examples of technology where the quality is not yet ready, same as neural networks in the 90s. However, when the quality rises above a certain threshold, it will explode and change the world we know. There are a few different ways that that could happen.
Starting point is 00:15:16 One is BCIs or brain computer interfaces. BCIs are devices that allow direct communication between the brain and an external device. So a use case might be someone who's experiencing paralysis being able to actually interact with the physical world again. Another field that's interested in this technology is neuroprosthetics. Think, for example, about visual prosthetics where someone with vision loss could bypass the damaged parts of the visual system and stimulate brain activity with signals that correspond to the images or videos that they're seeing. As mentioned, this research is also really valuable for
Starting point is 00:15:46 neuroscience and cognitive research, which is helping us better understand how the brain processes information. Although there's some pretty serious ethical debates around it, some people are interested in the potential for forensic and law enforcement use cases. Think, for example, reconstructing images or scenes from a witness's memory. And then, of course, there is the entertainment and gaming use case, the favorite of science. in which people could control virtual worlds or interact with game experiences simply by thinking thoughts. In any case, this remains pretty frontier technology, but given how much attention there is on this space, it's not hard to imagine that we're going to see major, major advances
Starting point is 00:16:22 in this field of mind image and mind to video in very short order. Thanks to Zijiao Chen and her fellow researchers for sharing these results. They're exciting to see. And thanks to you guys for hanging out and watching the AI breakdown. As always, if you're liking it, please like subscribe and share and check out the podcast and newsletter. And until next time, peace.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.