Invest Like the Best with Patrick O'Shaughnessy - Jeremiah Lowin: Explaining the New AI Paradigm - [Invest Like the Best, EP.307]

Episode Date: December 13, 2022

My guest this week is Jeremiah Lowin. Jeremiah has been on the podcast a number of times over the years. He’s one of my oldest friends who has been a sounding board for me throughout my career. Toda...y he is the founder and CEO of Prefect, which helps companies automate and orchestrate their dataflows. In full disclosure, Positive Sum is an investor in Prefect. We didn’t plan this conversation, but when OpenAI released ChatGPT, I called Jeremiah for a primer on what’s happening under the hood and how best to contextualize this product amidst the growing AI movement. We have these conversations often, but this time I decided to record it so we can all learn from someone I consider to be a leading mind in the fields of data science and machine learning. We start off in the weeds and zoom out as the discussion unfolds. Please enjoy this conversation with my friend, Jeremiah Lowin.   Listen to Founders podcast   Founders Episode #136 A Success Story: Estee Lauder    Invest Like the Best with David Senra: Passion & Pain   For the full show notes, transcript, and links to mentioned content, check out the episode page here.   -----   This episode is brought to you by Tegus. Tegus streamlines the investment research process so you can get up to speed and find answers to critical questions on companies faster and more efficiently. The Tegus platform surfaces the hard-to-get qualitative insights, gives instant access to critical public financial data through BamSEC, and helps you set up customized expert calls. It’s all done on a single, modern SaaS platform that offers 360-degree insight into any public or private company. As a listener, you can take Tegus for a free test drive by visiting tegus.co/patrick.   -----   Today's episode is brought to you by Brex. Brex is the integrated financial platform trusted by the world's most innovative entrepreneurs and fastest-growing companies. With Brex, you can move money fast for instant impact with high-limit corporate cards, payments, venture debt, and spend management software all in one place. Ready to accelerate your business? Learn more at brex.com/best.   -----   Invest Like the Best is a property of Colossus, LLC. For more episodes of Invest Like the Best, visit joincolossus.com/episodes.    Stay up to date on all our podcasts by signing up to Colossus Weekly, our quick dive every Sunday highlighting the top business and investing concepts from our podcasts and the best of what we read that week. Sign up here.   Follow us on Twitter: @patrick_oshag | @JoinColossus   Show Notes [00:03:38] - [First question] - What a pre-trained transformer is  [00:06:12] - What latent representation means in the context of AI models  [00:09:57] - Models using math to interpret input data and generate images accurately  [00:11:43] - Whether or not understanding AI complexity in light of the results they arrive at will become a black box scenario  [00:14:13] - A high level history of the companies involved in generative AI [00:17:51] - The precursory technology that makes generative AI art possible [00:21:01] - What people are doing to improve AI models in between versions  [00:26:39] - Things that are literally happening during AI training [00:33:38] - Whether or not AI models might one day function as a utility like electricity [00:36:01] - Coding using GitHub Copilot and what it’s felt like to use it  [00:40:30] - How he’d approach starting an AI company from scratch  [00:44:40] - Developing this technology beyond general and into specific use cases [00:49:44] - The secret sauce for defensibility in the AI model space  [00:53:02] - What he’s watching more closely as the story unfolds  [00:56:32] - Whether or not he thinks that these toolkits will eventually learn how to use other systems like Unreal Engine on our behalf 

Transcript
Discussion (0)
Starting point is 00:00:00 This episode is brought to you by Teegas. Over the years of our partnership with Teegas, they have evolved from a pure expert network into a full company intelligence platform. I've been so impressed by the platform that my firm, positive sum, recently made an investment in Teegis. We did so because we feel that Teegis will be the gold standard platform for investing research for decades to come. Teaguez streamlines the investment research process so you can get up to speed and find
Starting point is 00:00:23 answers to critical questions on companies faster and more efficiently. The Teagis platform surfaces the hard-to-get qualitative issues. insights, gives instant access to critical public financial data through BAM SEC, and helps you set up customized expert calls. It's all done on a single modern SaaS platform that offers 360-degree insight into any public or private company. As a listener, you can take Tegas for a free test drive by visiting tigis.co slash Patrick. Today's episode is sponsored by Brex, the integrated financial platform trusted by the world's most innovative entrepreneurs and fastest growing companies. With Brex, you can move money fast for instant impact with high limit corporate cards, payments, venture
Starting point is 00:01:03 debt, and spend management software all in one place. Ready to accelerate your business, learn more at brex.com slash best. That's bR-ex.com slash best. Hello and welcome, everyone. I'm Patrick O'Shaughnessy and this is Invest Like the Best. This show is an open-ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. Invest Like the Best is part of the Colossus family of podcasts, and you can access all our podcasts, including edited transcripts, show notes, and other resources to keep learning at join colossus.com. Patrick O'Shaughnessy is the CEO and founding partner of Positive Sum and the CEO of O'Shaunacy Asset Management. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect
Starting point is 00:01:54 the opinion of Positive Sum or O'Shaunacy Asset Management. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of positive sum or Oshonnessy asset management may maintain positions in the securities discussed in this podcast. My guest this week is Jeremiah Loan. Jeremiah has been on the podcast a number of time over the years. He's one of my oldest friends who has been a sounding board for me throughout my career. Today he is the founder and CEO of Prefect, which helps companies automate and orchestrate their data flows. In full disclosure, positive sum is an investor in Prefect.
Starting point is 00:02:32 We didn't plan this conversation, but when OpenAI released ChatGPT, I called Jeremiah for a primer on what's happening under the hood and how best to contextualize this product amidst the growing AI movement. We have these conversations often, but this time I decided to record it so we can all learn from someone I consider to be a leading mind in the fields of data science and machine learning. We start off in the weeds and zoom out as the discussion unfolds. Before we transition to the episode, I want to highlight the Founders podcast, which is part of our Colossus network. David Senra, who hosts Founders, has devoted his life to learning from history's greatest entrepreneurs, and every week he distills the lessons of a different founder. If you want an entry point, I highly recommend starting with episode 136 on Estée Lauder. I hosted David on Invest Like the Best Like the Best this summer, and it's hard not to walk away insanely energized after listening to any episode with him.
Starting point is 00:03:23 You can find a link to founders and that episode in the show notes of this conversation. And you can search all past transcripts on our website, join colossus.com. Please enjoy this conversation with my friend Jeremiah Loewan. Can you start by explaining a pre-trained transformer? What is that? You really want to jump right in here. You know me a long time. I don't want to ask about the weather.
Starting point is 00:03:49 Oh my God, there's no warm-up question, man. Okay. There's this new class of models, transformer. models, some people call them foundational models because they're so important. They're called transformers because what they fundamentally do is transform some sequence of inputs into some other sequence of inputs following a very, very complex series of rules or heuristics. I think the most common way that we're familiar with these is translation. So you have a string of words in one language and you have a string of words in another language and the transformer sits there and it can
Starting point is 00:04:21 go from one to the other. I think language is a really good place. for us to start talking about this because it becomes very clear, very fast, that it's not just take each word, run it through a dictionary, and find the equivalent word on the other side and just do them in order. Orders change. For example, some languages will have genders, some do not. You need to do the translation appropriately. The adjectives go in different place. Syntax is different, the grammar. Ediums. You don't want to just translate them literally. I'm sure we could find a million and one reasons why translating is just horror. The complexity explodes really quickly. And so these transformers are a way of working through this sequence and applying
Starting point is 00:05:00 different techniques like paying attention to things that are more important, literally like focusing on things that are more important, or applying context in the way that lets you make a decision later in the stream based on what you saw earlier. This class of models is just incredibly incredibly powerful. And there are many of them now that are not necessarily publicly available, but they're sitting there waiting to take in some known input could be text, could be English, could be French, could be an image, could be anything really, could be a time series. And even if we're not ready to necessarily output it into a specific known form, like we take English in, we don't yet know if we're going to put it out in French or we're going to put it out
Starting point is 00:05:38 in German or what. It can take it in and build what we're called a latent representation, which becomes this really powerful information set that in a sense fully describes the input sequence. I guess it isn't necessarily restricted by the fact that it's English. It's amorphous. We can't point to it. We can't interpret it as humans. It's just all the information and it's ready to now be squeezed out and transformed
Starting point is 00:06:00 into some other form. So these large language models are in many ways the basis for a lot of the really neat things that we are seeing in the AI space today. Can you say just a bit more about what latent representation means? Because I'm sure almost everyone listening, which they'll be doing because they're interested in this explosion of AI technologies has played with Mid Journey or Dolly or stable diffusion or the GPT3 chatbot or something. Like, they've done something. What is most magical about the experience, I think, is the specificity that's possible of the outcome.
Starting point is 00:06:36 You might naively think, well, to create a great English to French translator, you should just have experts that know how to actually translate and have them encode lots of rules. it's so obvious that that's not what's happened when you see some of these outputs. These systems seem so general purpose. You can get whatever you want, despite it being clear that no one on the other side was like, let's make sure this thing can program Spider-Man smashing watermelons or something. Say a bit more about what these models are doing when they're quote-unquote training on a given data set.
Starting point is 00:07:06 When we talk about the latent representation, we're talking about intermediate steps in how we process data, how we get from A to B, how we get from Spider-Man smashing a watermelon to an actual set of pixels and colors and coordinates that represent that with no text at all. Our goal at every moment is to take the information that's represented at the beginning and end of this pipeline and keep that information intact, even as we dramatically change the form of that information from text, which could be a few bytes of information. Literally, it's the characters and the sequence of characters into potentially thousands or even tens of thousands of color pixel coordinates. This is like a wild degree of transformation.
Starting point is 00:07:51 And if we try to go immediately from one to the other, we'll almost certainly fail. That mapping is just impossible with precision. And again, as we talked about earlier, context starts to matter, especially when we're talking about images where if I've drawn an eye, the position of the people really matters. You can't just say it's kind of in the center of the eye. It's got to be very in the exact right place. and it's got to be lined up with the other eyes pupil so that the person looks like they're looking in the right direction.
Starting point is 00:08:15 Famously, these models are bad of drawing hands. Why? Because it's really hard to keep track of how many fingers you've drawn unless you have a really, really intricate sense of context around what you're doing. What we need to do is we need to get the fact that you wanted to picture a Spider-Man. We need to map that into a space that explains to us in the sense that Spider-Man is human figure, has hands, he's red, he's got a mask, there's a watermelon involved, he's proximate to it. We need to show the action.
Starting point is 00:08:39 And we'd sort of extract all that stuff, and we don't do it by going straight to the pixel representation. Instead, we go to what's called the latent space. It's a place where we take all of these concepts, map them out numerically, and I'm being a little bit vague here, because I don't know how indebt we want to get, but we map them into this latent space. What that does is it leaves each concept isolated. If we did a perfect job of this, then we'd actually have a string of numbers where the first number represents, I'm making this up, of course, but how human the character is, how many hands he has, what color he is, what the color of the watermelon is. And we could actually tweak those numbers in the latent space. We could take the watermelon from green and then dial it up to 17, which means orange, and then the picture would eventually show Spider-Man smashing in orange watermelon.
Starting point is 00:09:24 The latent space is where the power of these models come from. It's where we take the input. We map it into this place where each concept is as separately identifiable as possible. And that makes the job of the output transformation a lot easier because it can grab these concepts and work them all into their final form in just a much easier way than if we handed it the number 17. It was like, this number encodes everything there is to know about this model. Is it fair to say that in each of these things, there's input, there's output, the thing being tapped in the middle is like a mathematical representation of as granular a set of concepts as possible. that effectively what's happening is it's like a universal translation layer because we've taken
Starting point is 00:10:10 stock of everything possible in the world. And I want to talk about where that data comes from to make that core representation possible. And then we're effectively just tapping that core thing and the models themselves are getting better and smarter about how to combine those disparate elements. That's exactly right. If we go back a generation of models, actually, it becomes maybe a little bit easier to see this tangibly. So there were this models called GANS, which sort of predated the diffusion models that are very popular right now. But one of the cool things about them is that you could actually explore the latent space very, very, very easily and very quickly. When you see these models and you'd see them in sort of UIs, you'd see a dial that's
Starting point is 00:10:49 like smile, do you want more smile or less smile? Hair lengths, do you want more or less? And what that's really doing is it's going into the latent space where these models were very, very, very good at identifying latent vectors or latent dimensions, we would say, that correspond very, very, very closely to observable things in the output. I was giving sort of a facetious example before, but in this case, it's very real, by literally changing the number along that dimension of the mathematical representation of the information, the image you would get of the person on the other side would have longer or shorter hair, darker or lighter hair, glasses or no glasses, they'd be smiling more or not, they'd be looking left or looking right.
Starting point is 00:11:25 And those are all playing with the latent space, which helps us map all of these different very, very, very qualitative characteristics of the output into and separately identifiable mathematical space. That's absolutely what it is. In those examples, each variable I picture with like a slider or like a, okay, so you're saying in this output, a certain variable is relevant and I can tweak it. Am I right in saying that at a certain level as these get more advanced, the ability of the human mind to understand or interact with the variables or just the ability period starts to go away. And we just have to say, We don't really know what's going on because it's gotten so detailed. It's ever more a black box.
Starting point is 00:12:05 Here's where I think this starts getting really interesting. I wouldn't go as far as you there. I think we understand exactly as much as we did in the example that I just gave with some of these newer models. I think that it's a nice feature of those models that in that latent space, we were able to identify dimensions that gave us specific sliders for a smile. Whereas in more modern diffusion model, the latent space is so enormous that picking one number that represents the smile is silly. However, that doesn't mean that the model is less interpretable. It just means that it's more complicated. It works in exactly the same way. I think one of
Starting point is 00:12:39 the cool things that I expect to see in the near future is the reemergence of exactly that. So just because we as a human can't identify a single number that changes the smile and gives us the smile slider, that doesn't mean that a machine can't look at that vector and figure out what combination of numbers make someone smile more or less. And I think the proof of that is that if you took a prompt into an image-generating AI and said, show me spider ban smashing watermelon's comma with a big smile, you'll get it. And if the AI is capable, you'll actually get very similar to the original image just with a smile. And I think that what that means is that it is possible for us to go in and tweak these things. It just might be that the manner that we do it isn't as explicit
Starting point is 00:13:19 as finding the single number in the vector and raising it. One of the best ways for people to think about what this is and how interpretable it is, to me, is looking at clouds. So if you and I are looking at a cloud and I'm like, that cloud looks like a duck, might take you a second, but then you're like, oh, yeah, I see how it looks like a duck. And then someone else comes along like, no, no, no, it's a rabbit. They're probably just as right because the truth is cloud looks like nothing. But that's kind of what this is all about. You're taking something that's kind of resemble something and you're mapping it into an understandable space. And in fact, what these diffusion models do in a very concrete way is they do that many, many, many, many times over.
Starting point is 00:13:54 say it's a duck, okay, now it looks like a duck. You don't make it look more like a duck if the beak was a little more pronounced. And they make that change. Okay, now it's even more like a duck. Oh, it's swimming. Okay, maybe that other cloud is water. And you emerge from the latent space this very tangible characteristic. How do you think about the history here? Like this isn't all happening in a vacuum. This is all human control. There's big companies involved. There's names like Open AI and Stability AI and DeepMind and Google and these are very tangible things that are happening. If you were drawing a single PowerPoint page timeline, the big companies that have mattered and when, how would you tell the high-level version of the history here?
Starting point is 00:14:32 I think you just mentioned most of the companies. I would throw a lot of universities in as well. When I started getting involved in this, which was probably in 2009, you'd be looking at the University of Toronto, University of Montreal, you'd be looking at NYU where the O'Connor's based. You'd be looking at Stanford as well. So a lot of this research began there. And then as it went through commercial iterations, became more and more pronounced at first to large cloud providers and thinking of Google in particular, Facebook as well. And now you're seeing dedicated AI companies, like the ones that you mentioned. So I think that we go back, what, 15, maybe even 20 years since this more modern version of machine learning has come along. And what's really cool to see is the core idea has not dramatically changed.
Starting point is 00:15:20 we take an input, we run it through some sort of latent representation transformation of that input to make it more malleable and more interesting to work with, and then we output it somewhere. The ingenuity that's been applied here is, first of all, to make these models bigger. It's not as easy as just saying, I want a bigger model. Now we have to talk about memory and moving stuff around and GPUs. There's a real infrastructure challenge there. That's one thing that had to be solved in speed as well. But there's also the transformations that we're applying.
Starting point is 00:15:49 Take stability AI's model, stable diffusion as an example. The light bulb that makes stable diffusion different than, say, Dali, which was released in close proximity to it, I think just a month or two before, is that both of them are these diffusion models, which means they do the thing I describe with the clouds. They start with randomness and they iteratively pull an image out of it over many, many, many stages until they have something real. Dali operates in the actual pixel space. It starts with literally a random image, literally doing the cloud thing.
Starting point is 00:16:17 It's watching it and it's refining it over. time. Stable diffusion does the exact same thing, but it doesn't in that latent space. That's very difficult for us to conceptualize as humans. It turns out that that gave it much more expressive power because as it's doing its diffusion process, as it's learning and figuring out what's in this image, it is not bound by the same spatial constraints that the image has. For example, having two eyes and having five fingers. Now, I know I'm saying that and sometimes it's not good at doing five fingers. There's complexity here still, but Dali has to emerge, figure out what's in the image, and generate five fingers at the same time.
Starting point is 00:16:55 Stable diffusion emerges what's in the image and then actually figures out that it needs to generate five fingers at the end. That little, little, little insight about moving the diffusion from the pixel space vector to the latent space vector is, for all intents and purposes, the chief difference between these two models that has profound impact in quality and express. That is the type of thing, if you look back in time that you see over the last 20 years, you see these, I don't want to call them small changes because I think that belies the energy it takes to figure them out, test them, that they'll work, dedicate resources to them.
Starting point is 00:17:32 So they're not small changes by any means, but they are not necessarily see changes in architecture as much as they are refinements and enhancements of this neural net transformation modern machine learning paradigm. Can you say just a bit about the technology precursors that have made this possible? Because I think there was a narrative for a long time that AI was an aborted set of innovations. Everyone was expecting so much of it. There were long periods where we didn't see the sort of explosion of advancements that we've seen in the last. I mean, it's literally only been like six months.
Starting point is 00:18:06 It's crazy to think about how quickly it's going. One of these bamboo examples where I just think the roots are being laid for like months and months and the whole thing grows in a couple days. it seems to me like the internet's an important precursor for connectivity, mobile is for data gathering, cloud is for storage and compute and all these things. What else is required? Like without what else besides those three things, would this not be possible? Because I think the foundations are important here. I've been thinking about this a lot myself. I'm not sure if I have a great answer that explains exactly what the catalyst is. I'm curious for your take if I can
Starting point is 00:18:41 turn this around on you for a second. What I would throw out is that. these generative models are really, really, really powerful because they feel more powerful. Because I as a user can control what's happening in a way that's meaningful to me and interesting to me and it's somehow mine. And I think they therefore get a degree of engagement that, for example, a model that can only generate faces, which we've had very good versions of that for about five years, or a model that can only read addresses on an envelope, which we've had for probably 30 or 40 years. your earliest convolutional realm that's for handling stuff like that, it's not as interesting.
Starting point is 00:19:16 It's extremely useful, but it's very purpose-driven. It's very specific. And so imagining its application in a wide range of fields and its delivery as an API, it's like, yeah, it's great, but how many uses are there for this? And I think one of the things that these generative AIs have done is open up our imaginations of how we can apply this software in a really, really, really interesting way. Obviously, there's the availability of GPUs and the data set collection, which is critical. There are absolutely some insights that made it possible to build this type of model. But I think part of what we're seeing is a very special connection with an opportunity for people to interact with this in a very, very, very unique way. Maybe it's like how when movies
Starting point is 00:19:56 are released, there's like bugs life and there's ants. When a good idea is out there, like a lot of people are going to find out and glomp onto it. And I think that this is just one of those ideas where it's like, Yeah, these generative things are amazing. Now, one of the thing is these large language models that we have mentioned earlier in this conversation. Well, they've been sitting there for a long time, measured in years. It turns out that they encode so much information that we are still figuring out how to leverage that fact.
Starting point is 00:20:22 In a funny way, these image generation technologies would be nothing if they couldn't figure out what to do with the words that are being put into them. So that's also one of the huge unlocks here, is that we've gotten so good at processing text into a latent space where it can be transformed into something else, whether it's a different language, a voice that synthesized or an image, that's been a huge unlock. And I think it would be wrong of us to ignore that fact, because without those models, sure, I can generate a lot of images, but guiding them to something that's semantically aligned with the input, I don't know if we'd be able to do that without this advancement in text understanding. I'm curious about two pieces of this
Starting point is 00:20:59 and what is actually going on. One is training. And the other thing, is the models themselves. And the best way I can ask the second question or frame up the second question about the models is we've all seen Mid Journey v, one, two, three, and now four. We've seen Dali, we're about to see GBT4, which by all accounts is hard to believe. I mean, GPD3 is hard to believe. GPD4 apparently is 10x harder to believe. There's teams behind these things. What is happening between model iterations? What are people doing to make these things better? Maybe I'll start with that question. When I see GBT4 come out and GBT theory has been out for what, two and a half years or something, what is happening? Who is doing what? Is it human directed?
Starting point is 00:21:40 Is the model itself creating the improvement in it in model iteration? What is going on? It's a tough question to answer because I think it's different in every case. What constitutes an improvement could be wildly different. For a lot of these companies and a lot of these models that we have visibility into, we actually don't necessarily know exactly what's going on. But if you look at the releases and you look at the associated papers or blog posts, you will usually see some claim of why is this better? Like, what have we done? There's two easy ones, more data and more training. Just look at it for longer. And I think if you want to see that happening, I guess in real time isn't quite true anymore. But if you wanted to see that happening in real time,
Starting point is 00:22:18 there was an open source version of Dali that was released. Dali 2 was announced, and very shortly after an open source version of it, which I think at the time was called Dali Mini, and now it's called crane with an AI in the middle. And you could watch this model train over time because literally it was training over time. And it is significantly better today than it was in the beginning. The beginning was like noticeably worse than the full Dolly 2 model. But you could watch it and you can see it continue to train and continue to improve. And in fact, if I remember right, for anyone curious, there's a notebook where you can actually see each checkpointed model. I think they do it each week maybe. And it's the same set of prompts. And you can watch the
Starting point is 00:22:55 model improving over time on a given prompt. It was like hundreds of them. And it's just fascinating to watch. You're watching the latest space develop and you're also watching the Alpha Transformer develop and it's really, really cool to see. So those are the two easy ones, more training, more data. But that's not what leads us to announce a new model. That's what makes these models better or possibly a new version of these models. I believe I'm not sure, for example, stable diffusion 1-4 to 1-5 was principally a new weights update, which is to say a new training update, not an architecture update. The architecture updates are more interesting. And chat GPT, which was released just a few days ago, is blowing my mind in all kinds of ways. I think it's really
Starting point is 00:23:33 incredible. GPT 3 and a half might not be a terrible way to think about it. It involves a new form of reinforcement learning. Reinforcement learning is when you take a model and you train it in a way where it takes a series of actions and then you tell it at the end how good its actions were. We often talk about two different ways of training machine learning models, supervised, which means You give it an input and you tell it what the output is supposed to be, and it's got to figure that out. Unsupervised, where you give it an input and you have some way of judging the output, but we don't necessarily know in advance and it's got to just work out on its own. And then there's third one, which a lot of times gets ignored, but is really, really, really tricky, which is reinforcement learning. And why is it so tricky?
Starting point is 00:24:11 Well, because how do you tell something if it did a good job of walking? Do you tell it every time it takes a stumble, that it's screwed up and then it should start over? Do we allow the stumble and we only judge it on how quickly it got to an endpoint? How do you actually reinforce a series of actions over time? So one of the ways that chat GPD got so good is that OpenAI came up with a new, more efficient, much faster to converge and requiring less data form of reinforcement learning. So for them, I don't know about the underlying architecture exactly, but in this case, it's the training of the model and their ability to deliver quality
Starting point is 00:24:44 results in a short amount of time thanks to this new form of training. And they also leverage other AIs to help train it. So there's other AIs that are judging its output that are assigning the reinforcement and scores that are then going into the training. It's like this fascinating way of setting up the training process. So that's one way that you can see this advancement. Another way is through new architectures entirely. So Dolly and Dolly II are completely different architectures. I'm forgetting as we speak right now what Dali I was. It was not a diffusion model. But Dolly 2 is a diffusion model. It's very different than Dolly 1. And it's capable of a much higher level of fidelity.
Starting point is 00:25:17 But that's a massive risk to take. If you think about it, someone's got to go do that research. There's a lot of different choices you could make. So I think a lot of the expertise is actually somewhere in the intuition of being able to look at a model and understanding, if I nudge this in a certain direction, or if I take our architecture in a certain direction, that will deliver results because you understand what the limitations of each model is. And you are trying to advance them in a concrete way with a new architecture, with a new model. I mentioned earlier, stable diffusion comes along and says, yeah, we can do diffusion models, but instead of operating on the pixel space, we're going to operate on the late space. That turns out to be on its own, this revolution. Mid Journey, we don't know exactly, but what we do know about Mid Journey is that the aesthetic quality of its images is extraordinary and always has been.
Starting point is 00:25:58 It's very hard to get Mid Journey to put out something bad. I don't know if that's because they didn't show it any bad images in the training set. I actually think that that's not necessarily the case because if all yet to do to get quality outputs was show something a bunch of quality inputs that would belie the complexity of the model. I suspect that there's something that they've done there to guide always to high quality. they've found some way to build, whether it's an architecture or some sort of guiding along the way, to guide things to that output, and that's their secret solve. The reason that's such a hard question to answer is you can make advancements in these things
Starting point is 00:26:29 in so many different ways right now that generalizing that is probably impossible at this time. Can you say a little bit more about what is actually happening during training? It's become a word that everyone's kind of says, yeah, the model's training, it's training, it's training. GPUs, GPUs, GPUs. What is literally going on in a generalized... sense. When we think about the model, the monotonous two components, principally. It's got the architecture itself, which is to say, okay, you give us this vector, we're going to transform it in the following way, 17 times. We're going to run a bunch of diffusion steps on it. We're going to give you an
Starting point is 00:27:01 image. That's the architecture. And then there's the parameters of that architecture, which are usually called the weights. And this could be 10 or 100 billion numbers. And these numbers go into that architecture and they tell that architecture how to behave. They literally weight different parts of it. When we're training the model, we have the architecture, but the weights are all random. They're all garbage. And if you ask it for Spider-Man and a watermelon, you're going to get literally just garbage noise, nothing, you can't see anything. Because it has no idea how to take that input. I mean, don't get me wrong. It knows how to take that input and literally run it through the architecture, but the architecture has no knowledge embedded in it. It doesn't do anything interesting.
Starting point is 00:27:39 Training is the process of changing those weights, literally the numbers, those potentially billions of numbers changing them so that they encode information in a way that when you run an input through the model, you get an output that is sensible. There's many ways that we do that. I just reference a couple of the high level like supervised training, unsupervised training. But what all of them have in common is that we run something to the model. We look at it. We have some way of deciding if it was good or not good. And we have some other way of deciding what could make it better. I'm vastly generalizing here, but this is really what happens. And then part of what makes these model so interesting is we can actually go back to each of those 100 billion numbers and
Starting point is 00:28:19 figure out which one would impact the output in what way. Now, I'm not saying that we end up with Jackson Pollock painting on the other end and we're like, oh, if we make this 17 and 18, it'll be Spider-Man. But we are saying if we get this complete random noise at the other end, we can go back in and we can say, okay, literally mathematically, we know that if we nudge these numbers up and these other numbers down, we'll end up with still noise, to be clear, but noise that's slightly closer against whatever our objective function is to what the person wanted. And we do that hundreds of times, thousands of times, millions of times, millions of times. Training is a very, very, very slow iterative process. And for a long time,
Starting point is 00:28:58 you don't even know if it's working. The math tells you it's working. If we're talking about an image generator, you may not actually see a sense of output for a very long time after training. And if we can dig up that notebook on the Dolly Mini project, you can watch at training. The outputs are kind of nonsensical in the beginning. And it takes a few weeks until you even start to see shapes even really emerging, but you can see clear progress. So training is the way of taking a mathematical expression of the fidelity of the model and figuring out, as the model is looking at data, how to tweak all of the weights to get better and better and better results. So that's when we're waiting for something to train, when we're deploying GPUs to train it. That is the process
Starting point is 00:29:36 that's happening. It's extremely compute intense and it takes a long time. Can you talk about the role of scale in all of this, just in terms of the real world impact it will have of who controls these models and who can create a marginal model that's interesting or useful. If you'd asked me six months ago, I would have said, it seems as though Open AI is the hegemon here. They've spent all this money, and money is a major barrier to entry, money and time, to train these things. All the benefits are going to accrue to the scale players, whether that's an open AI or whether that's Google that has all this great data and also the resources to spend on training. And then all these alternatives have started to come out, which begs the question,
Starting point is 00:30:18 what is going on? Is capital really scale constraint here for new entrants? Will we have a hundred dollies competitors or will there always just be a couple? How do you think about the power dynamics of who can build these things and how that will play out? I think it's very early to make predictions on that dimension because we don't really know yet what the constraints are. Here's something I feel very confident about. The things that we're calling state of the art today will be obsolete 12 months from now. I don't mean because more training will happen. I mean, new architectures, new applications, completely new things. So in a funny way, I don't think capital is a constraint on any one iteration of this. So if you need to spin up and say 4,000 to 800 GPUs,
Starting point is 00:31:01 order of magnitude that's $40 million a year of cost. It's a lot of money. It's a lot of money. for some guy, but for a well-funded company that knows what it's doing, that's a very reasonable, in fact, maybe even a very small cost. And it's a retail cost, by the way. Just for reference, the number I'm giving is the number that Stable Defusion trained their original model on. And the reason it's interesting is I think at the time it made it the 10th most powerful supercomputer in the world. We're not talking about astronomical sums. It probably costs more to staff the team, in a sense, that's going to come up with all this research. Because remember, it's not like we take a group of people and they instantly come up with the right idea.
Starting point is 00:31:36 We have a group, they make a bet. Months later, we have an answer. We know if it work. We go back. It's a very iterative process. Time, energy, and intuition and ingenuity are, I think, really the constraints on this, really wanting to improve this and wanting to move forward with it more so than just raw capital, especially depending on what happens with the price of Bitcoin.
Starting point is 00:31:53 It could be a lot of cheap GPUs on the market, as a matter of fact. Maybe not A100s, but processing power is probably not the limiting factor here or access to it, I think, at the end of the day. As for who controls it and how many of these things will see, I think that's one of the big and fascinating open questions that we'll be watching. And yes, that's kind of me saying, I don't know, but also in a more realistic sense, do we need 17 different diffusion models in the world? Or is stable diffusions open source version, the solid base that we then build a million and one fine-tuned, industry-specific or personalized models on top of? Because remember, these models are very much a common. combination of models. There's the text processing part, there's the image processing part, this is the latent part, it's a diffusion part. So what a lot of people are doing right now is they're taking, for example, the stable diffusion model, which is open source, and they're tacking on to it, some other transformation they want to do, or they're going in and they're fine-tuning
Starting point is 00:32:48 the weights themselves. So stability AI has done the hard work, if you will, the month-long process of generating the weights that characterize the stable diffusion model in its canonical form. And then folks are going in and they're saying, you know what, I really want this to only output things that look like a Van Gogh painting. And they go in and they fine tune it. They tweak the weights a little bit more and they train it for a little bit. And now the thing's only capable of putting out Van Gogh. And it probably does that better than if you just tack on comma in the style of Van Gogh to a prompt the original model. So I can easily see that as being the accessible way of extending these models into specifics. And you have this canonical base, whether it's open source
Starting point is 00:33:24 or via API. And you're doing fine tuning for whatever your application is, I would be shocked if that's not a major, major, major way that this shows up in the world. Where does an electricity analogy fall apart? I've always liked the way of describing APIs that it's like an electrical outlet, that all you know is that whatever you plug into that thing is going to get a certain kind of power to it. Whatever application you want to design on the other side of that outlet, do whatever you want. This is what's going to come through the outlet.
Starting point is 00:33:51 Is it fair to think about this almost like utilities 10 years from now where on the other side of that outlet is just these models? Maybe there's a diffusion model. There's a couple forms of electricity, but basically like it's just a utility that's being built and maybe that makes Open AI the biggest company in the world. Is that analogy oversimplified? I think it depends on what actually the utility is producing. So is the utility producing an API that takes in text and returns images, then probably not
Starting point is 00:34:21 because that feels too close to the end product. That would be like if the outlets in my house show light. It's too close to what I'm actually trying to do. do with it. It's not a raw material. Do they produce weights? Do they produce latent representations? I think that would be incredibly interesting, but also incredibly difficult to actually do anything with. I don't know at this time what the right utility metaphor is. One thing that I could believe is that there are a few versions of this where, yeah, we could believe that the utility metaphor applies and it is producing images or a chatbot, for example, like an AI that you can converse
Starting point is 00:34:52 with. A search engine would be a good example of this. Is Google a utility that provides links? If so, then absolutely I see the metaphor to these. But I'm hesitant to equate an API. APIs are more interactive than power outlets. Power outlets are one way. I know what I'm getting out. The best I can do is I know this has different voltage or damperage or whatever, and I have to mix and match appropriately.
Starting point is 00:35:16 APIs, I think, are more two-way because typically with the API, I send information, and then I get something back that's useful to me. And really what I'm doing with an API is I'm outsourcing the business logic or some form of processing. And I think that these companies are going to be well-suited to do that because I'm going to send it text, I'm going to get back an image. I'm going to send it a text, and I'm going to get back a response. I'm going to send it a search query. I'm going to get something back. But I think it's too early to say, like, what role do the large players? It could be that they just, huge, air-quets or unjust, that they just provide the training and the canonical models over time, whether
Starting point is 00:35:48 that these foundational models or what we would just call foundational models. And then every company, every industry has their personalized version of it for whatever their purposes. You told me that you've been coding using GitHub co-pilot for a little while, and I think it would be neat to have you described that experience because it's one thing to create Spider-Man smashing the watermelon, which I did as we were talking, and it's pretty incredible what you can get. And that's a play thing, and it's silly, and it's funny, and you do it every so often. But building source code for something, as you were doing,
Starting point is 00:36:18 in a way that's much faster and better as a result of this technology, is a nice, interesting, tangible example of why it's so powerful. Maybe just describe what Copilot is and how you've been using it and what you've noticed about it, what it's felt like to use it. So Copilot is a GitHub product, which essentially writes code for you. That's what it does, and it does it surprisingly well. It was certainly the earliest of these generative AIs to hit the public, just because of the demographic that it's approachable to,
Starting point is 00:36:47 I think didn't get quite the attention that Spider-Man, smashing a watermelon does. but it's very much in the same vein. What it does is you write some code, and for many, many, many years, the software that we used to write code has been able to suggest what it thinks the next word is or auto-complete, basically,
Starting point is 00:37:03 whatever your variable name you're typing. What Copilot does is it actually just goes and writes lines and lines and lines of code that it believes are what you're trying to do. I recently had a project and I was like, you know, I'm going to try this thing and let's give it a shot. And I was absolutely shocked at the quality of the code that it produced, but also how useful it was. I was very tempted to view it in the way that I think we all are,
Starting point is 00:37:26 which is how will this handle my edge cases? How will this possibly be able to understand that this one route that I need to call has a very unique way of being called? And it's definitely not going to do that. Therefore, this is a waste of time. I have a similar argument with my father all the time about my car, which drives itself,
Starting point is 00:37:43 and he's constantly worried about, well, how will it know? The other day, he was like, What if a school bus is stopped, but the stop signs are all the way out? Will it even know? Will it be able to recognize it? What would happen if I wrote on a sign myself? Don't go on this road. It's icy. Well, the car just go right there. And I'm like, probably it will. But you're looking for the AI to perform in the situations where actually I don't want the AI to perform. I want the AI to perform when I'm driving on the New Jersey Turnpike. And it's a straight line. And I just need to hold the car in that line. I'm looking for the AI to perform when I'm stuck in traffic on the way to, my kid's school. And similarly in code, I'm looking for the AI to perform when once again, I'm writing a CRUD API and I just need to pass arguments from my user back into my server. Or I'm writing whatever boilerplate I'm writing, I'm writing tests, I'm writing documentation. These are things where every single time you do it, it's a little bit different.
Starting point is 00:38:35 It's easy to characterize New Jersey Turnpike as a straight shot for the entire state. But every time you do, it's a little bit different. Some guy cuts you off in a little bit of a different way. There's a puddle on the road. It's raining. There's a little bit of traffic. There's not. The exit's closed. The lane is close. It's always something a little bit different. This is, I think, where AI's at the stage we are today in history excel. They are really, really, really good at taking well-defined situations that have some degree of randomness or idios and crusade to them and navigating them in a really, really, really, solid way. I don't think we're at a point yet where we have any reasonable expectation of the AI actually going and figuring out every possible edge case. And I don't think that's what we're looking for now. And in fact, I think that people, who are looking for that are just going to be disappointed for a long time. When we talk about
Starting point is 00:39:20 AGI, these generalized intelligence that are going to take over the world, that's the sort of stuff that they're going to be able to do in theory. If we ever see them, we are so far from that. And that's why when I think people point in some of these AIs and are like, well, this is the one that's going to take over the world. I'm like, this thing could barely power a toaster if it hadn't seen one before. This is not really the threat that you think it is. It's going to be humans who misuse this that pose a threat. I'm sure we can talk about those sort of ethical concerns in a moment. What Copilot has really done is it's just freed up my time to focus on the aspects of my code and the APIs that I'm writing, which are uniquely interesting and
Starting point is 00:39:56 idiosyncratic and deliver a core value. Like, someone's got to define what exactly this thing does. And then I get to hand off to it filling in the blank spaces around that. And I've gotten to the point where a couple times I've been writing code and I've waited a second for it to do the auto complete and it hasn't for one reason or another. And I get frustrated. And I'm like, come on, Copilot. Like, you know this. You did this on the last function. Why aren't you filling it in. And every single time I've done that, I'm like, wow, this is a partner to me right now. I'm pair programming with the AI in a way that I absolutely didn't expect to at this stage. It's been a really pleasant surprise. I was joking with you that you're not allowed to do what I'm
Starting point is 00:40:29 about to ask about. But if you were not running Prefect, which itself is infrastructure to this whole world, which I think is great, and you were forced to go start an AI company, as cliche and ridiculous as that sounds, how would you approach that problem? What would be the things that you would be thinking about, and I'm thinking about all the entrepreneurs out there that are interested in this space and have the technical capabilities to attack something, but would want to create something that's actually useful to users. How would you think about the problem of, okay, what would be a great set of models or set of thinking for how to start a valuable AI-based company that's built on top of these new enabling technologies? Because so much of quote-unquote startup history
Starting point is 00:41:10 is explosions of applications or systems on the back of new enabling technologies. This appears to be one, how would you approach this particular one as an entrepreneur? So we need to find the problem to solve. And I think when you have any technology like this, it's extremely tempting to add to the solution that you already see. And we're already seeing that in the AI space. First of all, it's amazing that the stable diffusion model is open source, because it means that you really can just build a full stack thing and control it and own in. We're seeing this massive amount of experimentation. And we've seen some early business successes. avatars, AI generate avatars for Twitter.
Starting point is 00:41:46 There are all these things that people are paying money for because it is a valuable product. But none of these things are going to be startups. First of all, they're indefensible in a lot of ways. They don't represent any gathering of information or data or anything that would lead a business to be dominant over time. Are they really solving a problem or are they just delivering cooler versions of the solution that these foundational models already represent in the forms that they've been
Starting point is 00:42:09 delivered to us? So I think instead, what you need to go look for is a problem. I think that the first class of problems that will be subsumed are the ones which generating stuff is hard. So generating text is hard. Of course we're going to use AIs to do text generation for marketing, for copy, for whatever it is. Image generation is hard.
Starting point is 00:42:28 There are lots of people who need to communicate through image, because text is hard as well. Architects, visualizations, theme park designers, all this stuff where you need to be able to communicate something to somebody rapidly, iteratively, and these models actually make it possible to do that storytelling in a new way. I think that given how many companies and industries there are already today that are engaged in some form of storytelling, Disney probably being the king of all of them, I wouldn't go put my startup in their way. I think the problem to be solved here in what these models are capable of
Starting point is 00:43:01 removing frictions is the interactivity. So I'd go look at very boring places where a lot of time is spent figuring things out. And the first thing that comes to my mind is customer support and customer success where we don't even have to come up with a problem because literally the point of the company is to solve the problem that the user has. The problem is almost certainly going to be idiosyncratic and unique to that user and going to require some degree of customized response. The person is probably frustrated, which means the faster we can get them in answer, the better, but also they're not going to have a lot of patience for low-quality responses. There's a two-way friction there. That person actually doesn't care how we get them the answer. They just want the
Starting point is 00:43:38 answer. We, speaking from Prefect's own knowledge, we have an extraordinary customer success team. Fortunately, we don't have a lot of unhappy users, but we definitely just need assistance or need help or are doing onboarding or need to deploy stuff. And I see the value that can have, and I can also imagine how awful it would be to not engage with that. And one of the things we spend a lot of time thinking about is how will we scale this? Because we can't possibly grow that team in a linear fashion as our customers are growing. Our customers are skyrocketing. We've already experimented with putting a GPT power bot in our Slack to handle, let's call it, that first 80% of questions that are probably somewhere in our documentation or somewhere in our knowledge bank or somewhere in our own collective knowledge. But we don't necessarily want to task a human with taking the time to repeatedly identify what the person's asking, go off into the knowledge base and deliver the answer.
Starting point is 00:44:23 And that's a perfect place to deploy one of these technologies that can do the two-way understanding of remove the friction on the person's part of actually having to communicate information appropriately and report. the friction on our side of actually doing the knowledge of discovery. Use that as an example for how these things could be made turned from general to particular. So you said something important there, a GPT3 powered bot, meaning it's a bot. It's something performing a function in a very particular specific environment, this case, a Prefect user that's got a question. What literally is happening for you and your team to take a generalized model and use it to, quote unquote, power a specific use case?
Starting point is 00:45:00 There's two different technologies that we can deploy. The first is we've got this massive generalized model that's capable of taking a question, producing an answer. You first need to tell what kind of answer are we looking for? You and I were joking earlier with a chat GPT that I asked a question and I threw an instruction at the end. Answer is if you were a pirate. Yeah, it does.
Starting point is 00:45:19 Pretends it's a pirate, but it answers a question still, which is incredible and mind-blowing. We want our AI, our chatbot, to have a specific way of answering questions, a personality, if that's not too loaded of a word. We want it to be brief and direct. We want it to have links to documentation where possible. So the first part of the powered by is actually teaching you what sort of answers we want. What we don't want is we don't want something to say, hey, how do I write a Prefect Workflow and for it to tell a joke?
Starting point is 00:45:46 That would be bad, even if the joke was about Prefect Workflows. That would be bad. And almost certainly wouldn't do that, to be clear. But I'm using it as a counter example of the fact that we do want to answer that in a specific way. We wanted to deliver the message. And we also wanted to deliver a link to documentation. so that this user can go and serve themselves and continue their own knowledge journey. So we have to teach you that.
Starting point is 00:46:05 Now, that's not training. That's like through examples. That doesn't change the original model. It's through examples. We can also, if we choose to, if we want to take a step further, we can actually do the fine-tuning thing that we talked about before. We used image generation and in the style of Van Gogh as the example. You can actually go train it to do that.
Starting point is 00:46:20 We can also fine-tune these models, these language models, on our own data store. We have a proprietary data store, which is extremely, extremely rich. of user questions, our own responses, our internal discussions, our internal design documents, which we would never make it public. They got all kinds of stuff in them, but we can absolutely take sections of them, send them off to one of these bots and let it synthesize that information in a clean and accessible way for our users. That's what it means to take one of these licensed models, these generalized models, and apply it specifically for ourselves. Can I ask the question one more time just to make sure I understand it? And I'm going to make it
Starting point is 00:46:58 really simple and selfish. Let's say I wanted to create a little chat bot that was in the front of the Colossus website where we've got, I don't know, 500 very long transcripts of conversation. Let's say we have a thousand. We might be getting close to a thousand now. And I want to have the chatbot say, well, what are you interested in? I say, I'm interested in enterprise software. Oh, like, what about enterprise software you're interested in? These three things. And have it be like, oh, well, I suggest you You should go listen to this conversation with Chathin Putagunta. What would I do to take the proprietary thing I have similar to the resources you just described, the prefect has?
Starting point is 00:47:32 I've got a thousand transcripts, let's say. It's like literally what are the steps that I would take? I'm going to do this. If you tell me how to do it, I'll tell Joe, our engineer, to go build this. What would he have to do? How hard is it? What are the steps? There are two steps here.
Starting point is 00:47:46 One, we need to point it at your data. We need to show it that this is the corpus that we're interested in. And the two, we need to literally just show it. What does it mean to answer a question? So I would actually say we would go a step further than that. We wouldn't just say to it, go listen to this. We would have it say, oh, Chathen says the following on open source business models. And by the way, he did this podcast with Jeremiah, and they disagreed on this point.
Starting point is 00:48:06 And so this is the point of contention in open source business models. And yeah, I could link out to them. But I think the principal thing here is not just the search engine replacement, not just the natural language search engine replacement. I know there's a podcast over here and I'm here with a question, get me over there. it's actually the retrieval mechanism. I don't have to go to the podcast in a sense to find out if my answer's in it. The AI will have already read all the information,
Starting point is 00:48:30 and it's going to deliver a synthesis of that information to me that I can make use of. And now if I choose, I can go further into the system. And I think that is, when we talk about those latent space and those representations of information, the real power here is about getting us to information. You think going back to, like, what was the amazing thing that Google did? It wasn't that I type in a thing.
Starting point is 00:48:50 thing and it goes to the web page. That's how I experienced the technology of Google. But the real value ad was that they figured out how to know which web page should take me to. It's not that I type in the word car and I go to the page that has the word car on it the most times. It's that I go to the page that other pages were the word car pointed to the most. And that little lightball, which I'm sure is a million times more complex today than it was at the beginning, but that was the insight. And when we look at these models and how they're doing the information synthesis, there's still this core idea of it's less interesting for them just to take you somewhere. And it's more interesting to think, how do we teach it what is interesting to me?
Starting point is 00:49:27 And so that's where your energy would be as you set up examples for the spot. You would literally ask a question. You'd give some example answers of what you're looking for. You'd point it at your data. And you teach it what is a high quality response to a question like this? Do you think then that the most defensible products or businesses that might be built all have a component of incorporating existing and ongoing customer data or feedback. Is that the quote-unquote moat that might exist? There has to be two places that the secret sauce comes from. One comes from
Starting point is 00:49:56 a proprietary data set that you simply can't access otherwise. That could be that the data is not accessible publicly. And using Google as an example, that could be that no one can collect all of the data or in the right form as possible, since the web is obviously public. And anyone in theory could go and index every page on it. I think one of the place. places is having this mode of proprietary data, which you are then making available to people in this consumable way. But I do think the other one is going to be some secret cells in the models themselves. We're theorizing here that Mid Journey is doing something incredible in the model to guide to these high-quality answers. Not only do I think that that's going to be a critical
Starting point is 00:50:34 part of this, I think that that's going to be one of the major steps forward here. Because right now, I think that we have a little bit of a wrong way precision in these image generation technologies. Like we're doing all this stuff called prompt engineering, which for anyone who has not seen this term before, it's the idea that if you just type Spider-Man smashing a watermelon into one of these things, you probably get something that very literally looks like that. But then the meme, I think, is that you throw in a bunch of comma, comma, trending on art station, comma 8K, comma 4K, comma HD, comma, whatever. And then you have to put all these what are called negative prompts, things you don't want the model to do.
Starting point is 00:51:08 comma, no bad anatomy, comma, get the right number of fingers, comma, like all this stuff. and you really have to be very, very, very, very precise. Now, the model is still doing an incredible thing. It's taking that mishmash and words and producing a very coherent image at the other side, but that is not the future that I want to live in, where every single time you talk to one of these things, you have to say all that to it. That is going to be a major opportunity that is as of yet unsolved of how you guide to a high-quality outcome without requiring users to jump through all these hoops.
Starting point is 00:51:39 Because if you think about using Google today, if you're looking over someone's shoulder, what would make them a good user of Google or a bad user of Google? A good user of Google probably uses the smallest number of words that bring their result to the top. They're precise, they're efficient, they know how to create their query in a way that Google understands,
Starting point is 00:51:59 and Google's gotten really good at making that possible. A good user of, for example, stable diffusion today is exactly the opposite. It's someone who writes a paragraph into the thing to get what they want out. Now, that may be necessary. It may be that using something of this degree of power will always require some proportionate degree of precision on the input. But what I don't expect to see is these hacks, basically, comma trending on art station, comma 8K, comma 4K, become necessary. That's not a good prompt engineer. That's literally someone hacking the system to go into the latent part of the system that they want to be in, the high quality latent area of the system.
Starting point is 00:52:36 These models are as capable of generating garbage as they are generating high quality. all of anything, so that's just a different area of their latency space. They've learned what bad stuff is, so you tell them to stay away from it. That's another dimension of defensibility here is how is the user interface, so to speak, that you expose to users and how easy you make it for them to get what they are searching for. What are you watching most closely? You and I've been texting back and forth on the chat TPT here the last few days, and every time one of these new things comes out, we talk a lot about it.
Starting point is 00:53:07 But if you zoom out a little bit and look forward, and I know some of this is unknowable, but we don't know what's going on with perfect clarity. But six, 12 months, that kind of horizon, where do you have alerts set up so that if something happens, you want to be the first to know about it? There's a couple things that I'm looking for right now. I'm looking for video generation, which is we're just starting to see the first little examples of it come out.
Starting point is 00:53:28 I think that's fascinating. I'm looking for text generation, which is more interactive and interesting. So if you do play with this chat, GPT thing, and by the way, I'm in love with the fact that these are all accessible to the public or to bank customer. in some way. It's just incredible to have access to this technology. If you look at ChatGBT, it does amazing things. If you ask you the same question over and over, you don't just get the same
Starting point is 00:53:49 answer. You get the same answer. It's the same. You're very clearly keying into a very specific node and it's like understanding. And that is a good thing from the point of view of information retrieval, but that's a bad thing from the point of view of, am I having a conversation, or is it just a really good search engine? That's something that I'm hoping we'll see in the near future. By the way, I say that not to take away from the just incredible achievement that this is, but that's what I would look to see improve on that side. The thing that I'm most interesting in, though, is neither of these things. I am most interested at this moment in time in something called neural rendering, which is a representation of basically 3D scenes, like letting you
Starting point is 00:54:29 explore them. And so there's a company, for example, Luma Labs, which is commercialized something called NERF, a neural rendering field. It's a fantastic video made by Renf from Corridor Digital. recently that made this available to a lot of people in a way that I think was previously inaccessible. I think that these technologies are really, really, really incredible because one of the biggest problems we have right now in terms of representing not just an image or video or text is like, how do we actually take scenes, entire scenes, and make them interactive? How do you walk around something? It doesn't have to be in VR. It can be for like video effects, for making a movie, for filming a TV show, for a presentation. The high fidelity capture of visual
Starting point is 00:55:06 information. It's still very, very, very, very hard. And you and I have talked about engines, like the new advances in Unreal Engine, for example, that make that possible, which are incredible. But there's this new class of technologies, no surprise, which are AI-based renderers that can capture a scene in extraordinary fidelity and then represent it back. For example, in your browser. I'm extremely excited about these technologies because I think that they will make possible things like video calls where I don't need this webcam staring at me. I don't need to sit in my chair. I don't need to fiddle with my mind. and everything because I can just walk around and it's just capturing everything and it's able
Starting point is 00:55:39 to transmit it in high fidelity. I've seen very simple examples of this where my web cameras are above my monitor, but here you are in the center of it. So you probably think this entire conversation, I'm not looking at you, but I swear to you, you have my full attention. And video's got a demo where they just make your eyes look at the camera with AI in a natural way. It's extraordinary. It makes such an enormous difference. And I think it's really cool. And there are all these little things that have to do with neural rendering and transformations that I think are amazing. To be clearly, I think it's very different
Starting point is 00:56:06 than the Nerf thing that I was talking about before. But that's where my attention is right now on the AI-powered, how are these technologies going to transform? Because all of us interact with generated worlds all the time, whether it's in a game, a movie, a TV show, even VR, the ability to communicate these virtual worlds
Starting point is 00:56:25 into the real world is, I'm very excited with. My last question is, do you think that these old tool kits, it's funny to call them old, given how I'm advanced there, but something like Unreal Engine 5, let's say, which if you look at the outputs, it's clear it's capable of producing anything anyone could imagine. But if you said to me, go build a video of Spider-Man smashing the watermelon in Unreal Engine, I couldn't do it. Do you think that that bridge will get crossed because of these technologies where AI will learn how to use other toolkits and allow a simpleton like me to engineer using the toolkit like Unreal
Starting point is 00:56:59 engine 5 to create the outcome that I want? Is that relevant? Absolutely. I think the question is, will be in the form of, say, Unreal Engine 5? For example, will you use Photoshop to make images that are collages? Or will you literally go to a stable diffusion UI or a stable diffusion going to be baked into Photoshop in an interesting way? I think we need to separate the technology from the interface in order to answer that question. I have no doubt that you'll be able to say, I want projected onto my tabletop, a castle with a dragon so that my kids can play with. No doubt that that is a very, very, very possible thing and it'll be interactive and they'll use their own little action figures to interact with
Starting point is 00:57:34 100%. The fascinating thing to me in terms of watching just those technologies progress is with what fidelity can we generate these scenes? So right now, we're on the verge, I think, of a major achievement in image generation, which is that all hands will have five fingers. I know that sounds ridiculous. That's where we are as far as like what are we trying to achieve. So when I think about the world itself and capturing the world and all of its beautiful detail, reflections are really hard. Light bounces a million times and captures all this detail and gives rise to how interesting the world is to look at. That's really hard to program into something like Unreal Engine, which does some very expensive computations to produce.
Starting point is 00:58:10 But a neural renderer uses an AI to encode that all in a latent space and can do it more easily and more quickly. And so this is one of the reasons I'm very interested in these neural fields is they advance exactly the areas that our current technologies, struggle with by using heuristics to skip having to do all the explicit computation. My main conclusion from all this is I want to fund a company called Latent Space. It seems like at the end of the day, that is this incredibly powerful substrate that is merging into the world and we're figuring out how to both create the substrate but then tap it. Also, this is just a killer of friction between our minds and outcomes or outputs. And that is incredibly exciting.
Starting point is 00:58:50 I can't remember a time when I've been as glued to my screen around changing technology in my whole life and how cool it is and how lucky we are that we get to watch it. And I really appreciate, as always, when there's something technical to understand my ability to call you and have you explain it to me. Thank you. I totally agree with you. The latent space, it's like if I ask you, hey, Patrick, what are you thinking about? Don't tell me what you're thinking about.
Starting point is 00:59:15 Show me your thoughts. Show me what you're thinking about. That's the latent space. And that's why it's so powerful because you can choose anyone of, a thousand different ways to turn that thought into words. You could draw it for me. By the way, you absolutely could draw it, Spider-Man, smashing a water mountain, which is take you a really long time
Starting point is 00:59:28 and you'd have to learn technique to do it. So what if we could take that latent idea that you have and use a machine to accelerate that? And I think that's why we're so fascinated by this is we're seeing that play out in real time. You know, anytime you want to train a chatbot on me, just make it ramble for a long time, throwing a couple puns.
Starting point is 00:59:44 You don't actually conclude anything, and that's what I'm here to do anytime. So thank you. Thanks, my friend. Anytime. If you enjoy this episode, check out joincolossus.com. There you'll find every episode of this podcast complete with transcripts, show notes, and resources to keep learning.
Starting point is 01:00:01 You can also sign up for our newsletter, Colossus Weekly, where we condense episodes to the big ideas, quotations, and more, as well as share the best content we find on the internet every week.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.