Latent Space: The AI Engineer Podcast - [Practical AI] AI Trends: a Latent Space x Practical AI crossover pod!

Episode Date: July 2, 2023

Part 2 of our podcast feed swap weekend! Check out Cognitive Revolution as well."Data" Dan Whitenack has been co-host of the Practical AI podcast for the past 5 years, covering full journey of the mod...ern AI wave post Transformers. He joined us in studio to talk about their origin story and highlight key learnings from past episodes, riff on the AI trends we are all seeing as AI practitioner-podcasters, and his passion for low-resource-everything!Subscribe on the Changelog, RSS, Apple Podcasts, Twitter, Mastodon, and wherever fine podcasts are sold!Show notes* Daniel Whitenack – Twitter, GitHub, Website* Featured Latent Space episodes:* Benchmarks* Reza Shabani* MosaicML and MPT* Segment Anything* Mike Conover* Featured Practical AI episodes:* From notebooks to Netflix scale with Metaflow* Capabilities of LLMs 🤯* ML at small organizations* Prediction Guard* Data DanTimestamps* 00:00 Welcome to Practical AI* 01:16 Latent Space Podcast* 04:00 Practical AI Podcast* 06:20 Prediction Guard* 08:05 Daniel's favorite episodes* 10:21 Alessio's favorite episode* 10:54 Swyx's favorite episode* 12:44 Listener favorites* 15:14 LLMOps* 17:06 Reza Shabani* 19:06 Benchmarks 101* 20:06 Roboflow* 21:38 Mode collapse* 26:21 Rajiv Shah* 28:01 Staying on top of things* 33:11 Kirsten Lum* 34:31 datadan.io* 38:48 Prompt engineering* 40:38 Unique challenges engineers face* 42:51 AI-UX* 45:31 NLP data sets* 50:49 Unlabeled data sets* 55:07 Lightning round!* 55:20 What's already happened in AI?* 56:27 Unsolved questions in AI* 58:01 Get hands on* 58:53 OutroTranscriptFull transcript is over at the Changelog site! This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

Transcript
Discussion (0)
Starting point is 00:00:00 Hello again, it's Swix back again with the second of our crossover episodes, this time with the practical AI folks who have been covering the AI space since before it was cool, as we often like to say. Something that Alessio and I are pretty mindful of every time we take to the mic is that we're pretty new podcasters to this space. And there is a longer history and heritage behind all this, as well as a lot of areas and niches that will probably never even touch. So I always like recommending these podcasts with a much longer backlog that actually give you a way to travel back in time for kind of an oral history of AI, which is very rare and very valuable. These are tokens that are generated in real time, just like we're doing it now for this current wave of AI trends. And so what we did podcast to a podcaster, I always love talking to other podcasters because they know what's up is we went through our respective stats and picked out the crowd favorite episodes. as well as some personal picks. I do have to note that this was recorded just before our George Hott's episode,
Starting point is 00:01:04 which obviously is now our top downloaded episode across a podcast, as well as our newly launched YouTube. Go subscribe, please, if you haven't seen it. But anyway, if you're looking for good AI podcasts, these guys have been covering it for a long time, very, very consistently, and covered a lot of topics. Definitely check them out, and please enjoy our special crossover episode with Practical AI in the studio.
Starting point is 00:01:28 Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at Fly.io. Well, hello, we have a very special episode for you today. I got the chance to sit down with the guys from Latent Space, Swix, and Alessio out in San Francisco. They were kind enough to let me into their podcast recording studio. And we got a chance to talk about our favorite episodes of both of our shows and some of the overall takeaways we've had from those discussions. We cover some of the trends that we've been seeing in AI. And they even get a chance to grill me on my opinion.
Starting point is 00:02:37 about prompt engineering. So enjoy the show. Hey everyone. Welcome to the Latenspace podcast. This is Alessio, partner, and CTO and residence at Decibel Partners. I'm joined by my co-host is Swix, writer and editor of Latenspace. And today we're very excited to welcome Dan Whiteneck to the studio. Welcome, Dan. What's up, guys? It's great to be here. This is a podcast crossover. If you recognize this voice, Dan is the host of Practical FM. He's been in my ear on an offer the past five years, covering the latest and greatest in AI, before it was cool. Yeah, yeah, yeah, before the AI hype back in these like weird data science times,
Starting point is 00:03:14 whatever that is now. Yes, everything is merging and converging. So I'll give a little bit of your background, and we can go into a little bit on your personal side. You got your PhD in mathematical and computational physics, and then you spent 10 years of data scientists, most recently at SIL International, which I actually thought it was like an agri-tech thing, and then I went to the website. It's actually a nonprofit. It's a international NGO. Yeah. So they do language-related work all around the world. So I spent the last five years building up a team that's been working on kind of low-resource scenarios for AI if people are familiar with that. So doing like machine translation or speech recognition, that sort of thing in languages that aren't yet supported.
Starting point is 00:03:53 Yeah. And we'll talk about this later, but I think episode three on Presco AI was already featuring the global community that AI has and addresses. Yeah, yeah, yeah. It's been an important theme. throughout the whole time, throughout over 200 episodes. Yeah, yeah. And you recently left SIL to work on prediction guard, which we can talk a little bit more about that. You are also interim senior operations development director at T. Candleco.
Starting point is 00:04:18 Yeah, yeah. And yeah, what else do people know about you? Yeah, I mean, I, as probably can be noted from the intro, I love working on various projects and having my hands in a lot of things.
Starting point is 00:04:31 But yeah, I've code on the side for fun. and that's how I usually get into these side projects and that sort of thing. But outside of that, yeah, I live in Indiana. I was telling you guys that I'm trying to coin the term cerebral prairie. So we'll see if that catches on. Probably not. Your second guest in a row from Indiana, Linus from Notion was Indiana.
Starting point is 00:04:53 And we're talking about how there's surprising a number of international students in there. Yes, very true. Purdue is a strong university. Yeah, yeah, a very strong university. It's a great place to spend time. And there's a lot of fun things that happen around that area, too. So I'm also very into music, but not any sort of popular music. I play, like, mandolin and banjo and guitar and play folk music.
Starting point is 00:05:16 Low resource. Yeah, low resource music. Low resource languages. Yeah, all those things. Anything low resource is in my territory, for sure. And maybe we can cover the story of Practical AI. How did you start it? Tell us what the early days were like and just fill everyone in.
Starting point is 00:05:31 Yeah. It was kind of a winding journey. Some people might be familiar with the change log podcast, which I think they've been going now for like 11 or 12 years. It's pretty prolific in around, I think originally around more open source. Now it's kind of software development in general. But they have a network of podcasts now. And at a Go conference actually, so I'm a fan of the Go programming language. That's another fun fact. But at GoferCon, I think it was in 2016, maybe. I met Adam's. Stokoviac, who is one of the hosts of the change log. At the time, I was giving a talk about data science, something, I forget. But he kind of pitched me. He's like, we've been thinking a lot about doing like a data or data science podcast. And at the time he had a name, it was like, I think it was hard data or something like that, which never caught on for obvious reasons. But I kind of stored that away. It didn't really do anything with it. But then over the next couple of years,
Starting point is 00:06:33 I met Chris Benson, who's my co-host on Practical AI, and helped him with a couple of talks at conferences. We met through the Go community as well. And eventually, he was working at a different company at the time. Now he's a strategist with Lockheed Martin working in AI stuff. But he reached out to me and said, hey, would you ever consider doing kind of like a co-host podcast thing? And at that point, I remembered my conversation with Adam. So I reached back out to Adam with the change log. and then we kind of started working on the idea. We wanted it to be practical. So at the time, well, there's a lot of people doing things now with AI, like hands-on.
Starting point is 00:07:10 Back then, there were kind of some podcasts that were really hyped AI, like not practical at all, which is why we kind of came to practical AI, something that would actually benefit people. And that's like a great thing to hear from people that when they listen to the show, they do actually learn something that's useful for their day-to-day. That's kind of the goal. Yeah. Nice. And I think that's one of the things in common with our podcast. You know, there's a lot of content out there that can get a lot of clicks with fear of AI, you know, and all these different things. And I think we're all focused on more, yeah, practical and day-to-day usage. Yeah, tell us more about prediction guard, you know, that kind of fits into making AI practical and usable.
Starting point is 00:07:50 Yeah, yeah, sure. Appreciate that. So, yeah, prediction guard is what I've been working on since about Christmas time-ish. Originally, I was thinking a lot about large language model evaluation and model selection, but it's kind of morphed into something else. What I've realized is that there's this market pressure, there's internal company pressure for people to implement these kind of generative AI version of models into their workflows because enterprises realize the benefits that they could have. But in practice, when they go and they go from like chat GPT, they type in something, it's amazing.
Starting point is 00:08:26 And then like, how do we do this in our enterprise where we have maybe rules around data privacy or compliance issues? And also, like, we want to automate things maybe or we want to do data extraction. But I just get like text vomit out of these models. Like, what do I do with that? It was a unstructured text. How do I build a robust system out of inconsistent text vomit? So prediction guard is really focused on those two things. One is kind of compliance and running kind of state-of-the-art AI models in a compliant way,
Starting point is 00:08:58 and then layering on top of that layers of control for structuring output and validating output. So some people might be familiar with projects like guardrails or guidance or these things. So we've integrated kind of some of the best of those things into the platform, plus some ways to easily do like self-consistency checks and factuality checks and other things on top of large language. model output. Nice. We did have a Sharia Rajpal from Garel's as a guest. Yeah.
Starting point is 00:09:26 Yeah. So, yeah, that's another episode that people really like. Yeah, maybe, you know, just to give people a sense of what practical AIS as a podcast, you want to talk about maybe like the two, three favorite episodes that we have. And we can go maybe alternate, you know, like our favorites. We've done some prep for this episode. Yes, yes. So this is kind of like, I think our conception of this is kind of like a review for listeners
Starting point is 00:09:48 who are new to us, either of our podcasts, to go back in. revisit the favorites. Yeah, yeah. I think I can talk about some personal favorites of mine and then maybe like favorites from the audience. I think some of my personal favorites have actually been, we call them like fully connected episodes where Chris and I actually talk through a subject in detail together without a guest. To be honest, those are great episodes just for like me to learn something, like have an excuse to learn something. And we've done that recently like with chat GPT and instruction tuned models. We did it with stable diffusion and diffusion models. We did it with alpha fold. So all of those are episodes with with us too and just talking through like how practically
Starting point is 00:10:33 can you like form a mental model for how these models were trained and how they work and what they output. Those are some of my favorites just because I learn a lot because I do a little bit of prep. We talk through all the details of those and it helps me form my own sort of intuition around those things. Another personal favorite for us was that we did a series about AI and Africa. That was really cool. You mentioned the global AI community. We did actually a series of those. They're all labeled AI for Africa highlighting things like Masacane.
Starting point is 00:11:06 So people don't realize that like some of the models that we develop here like in the West Coast or wherever, they don't work great for all use cases around the world. And there's a lot of thriving grassroots communities like Masacane and Turkey. interlingua and other communities that are really building models for themselves, machine translation, speech recognition, models that work for their languages around the world, or agriculture, you know, computer vision models that work for their use cases around the world. So those are a couple of highlights on my end. Do we go with our personal highlights? Go ahead. I think you already picked one out. Yeah, I think mine is definitely the episode with Mike Conover from Databricks, who's the bird.
Starting point is 00:11:51 and leading the Dahlia first there. I think obviously the content is great, and Mike is extremely smart and prepared, but I think the passion that he had about these things, you know, the red pajama data set came out the morning that we recorded, and we're all kind of like nerding out or like, yeah, why is that so interesting? Like, he was so excited about it,
Starting point is 00:12:09 and it's great to see people that have so much excitement about things that they work on. You know, it's kind of like an inspiration in a way to do the same for us. I think personally, so I tend to drive, the news-driven episode once, like the event-driven ones, where something will happen in AI, and I'll make a snap decision that we'll have an episode recording on Twitter spaces, and we'll have just a bunch of people tune in. I think the one that stood out was the Chat-T-Up Store, the Chat-T-Ptplugns release,
Starting point is 00:12:38 where like 4,000 people tuned in to that one. That's crazy. And we did, like, an hour of prep, right? And I think it's important for me as a quote-unquote journalist to be the first to report in something major and to provide a perspective of something major, but also capture an audio history of how people react at the time, because this is something that we're talking about in the prep, chat to beat plugins have become a disappointment compared to our expectations then, but we captured it. We captured the excitement back then, and we can consider compare
Starting point is 00:13:08 and contrasts where we thought things were going and where things have actually ended up. It's really nice piece of, I guess, audio journalism. Yeah. Yeah, I mean, it was just last year. I mentioned stable diffusion and all that. We were talking about this. I was had in my mind, oh, everything's going to image generation. Like, should I quit doing NLP and start thinking about image?
Starting point is 00:13:30 And now all I do is NLP and language models. But at the time, that was, you know, that's what was on our mind. Same thing. I was working on WebUI for Stable Diffusion just like a thousand other funnen developers were. And yesterday was the first time I opened Stable Diffusion in six months.
Starting point is 00:13:48 And a lot has changed. And it's still an area that's developing, but it's not, yeah, it's not driving thought process at the moment. Yeah. Well, especially because I think just it depends on what you're, what you think you want to do. And I'm definitely less visual. I'm more, more of a text-driven person. So I naturally lean towards LLMs anyway, like NLP. Yeah.
Starting point is 00:14:09 I can hit some listener favorites. Yeah, proud favorites. So we have like one clear favorite, which is actually, I would say it's a surprise to me. Not because the guest wasn't good or anything, but just the, so the topic was Metaflow. So I don't know if you've heard of Metaflow. It's a Python package for kind of full stack data science modeling work developed at Netflix. And we had a Villatoulos on, who was the creator of that package. And that has had so, it's like maybe a 30% more listens than any other episode. And I, I, I, I, I think the title, so we titled it from, I think from notebooks to production or something like that.
Starting point is 00:14:56 Yeah. So it's like this idea of from notebooks to production, there's all sorts of things that prevent you from getting the value out of these sorts of methodologies. And my guess would be that talking about that is probably like the key feature of that episode. And Metaflow is like really cool. People should check it out. it is one way to kind of do this both versioning and orchestration and deployment and all these things that are really important. But I think a takeaway for me was like practically bringing into the some people might call it like full stack data science or like model life cycle things. Like the model life cycle things interests people so much.
Starting point is 00:15:42 So beyond making like a single inference or beyond doing like a single fine tuning, what is. What is the life cycle around a machine learning or an AI project? I think that really fascinates people because it's like the struggle of everyday life in actual practical usage of these models. So it's one thing to go to hugging face, try out like a hugging face space and like create some cool output or even just pull down a model and get output. But how do I handle like model versioning and orchestration of that in my own infrastructure? How do I tie in my own data set to that and do.
Starting point is 00:16:18 do it in a way that is fairly robust. How do I, you know, take these data scientists who like use all this weird tooling and like mash them into an organization that deals with like DevOps and like non-AI software and all of that? I think those are questions people are just wrestling with all the time. Yeah. It feels a little bit in conflicts with the trends of foundation models where you, the primary appeal is you train once. Yeah. And then you never touch it. again or you'd release it as a version and people kind of just prompts based off of that. And I feel like this, I feel this evolution moving from essentially the MLOPS era into, for lack of a better word, LLMOps.
Starting point is 00:17:01 Yeah. How do you feel about that? No, I think you're, I think you're completely right. I think there will always be a place for like these models in organizations that are task specific models like psychet learn models or whatever that solve a particular problem. because organizations like finance organizations or whatever will always have like a need for explainability or whatever it might be. But I do think we're moving into a period where like I've had to rebuild a lot of my own intuition as a data scientist from thinking about gather my data, create like my code for training, output my model, serialize it, push it to some hub or something, deploy it, you know, handle orchestration,
Starting point is 00:17:45 to now thinking about, okay, which of these pre-trained models do I select and how do I engineer my prompting and my chain? Maybe going to fine-tuning. Like, that is still, like, a really relevant topic. But some of these things that, like, I've been working on with Prediction Guard, I think, are the things that have a parallel in MLOPS, but they're slightly, like, there's just a slightly different flavor. I think it's how MLOPS is graduating to something else, versus like people are still concerned about ops. It's just like you say, it's kind of a different kind of ops.
Starting point is 00:18:22 Yeah, and I think that's reflected in our most popular episodes too. So I think all three of our most popular episodes are model-based. They're not more like infrastructure based. So I think number one is the one with Reza Shabani, where we talked about how they trained the Replit code model and the Amjud vibes that they used to figure out whether or not the model was good. And I think, you know,
Starting point is 00:18:42 that makes sense for our community's most, mostly software engineers and AI engineers. So code models are obviously a hot topic. Yeah. Yeah, that was really good. And I think like it was one of the first times where we kind of went beyond just listening traditional benchmarks, you know,
Starting point is 00:18:57 which is why we did a whole thing about I'm Judd Eval. It's like a lot of companies are using these models and they're using off-the-shelf benchmarks to do it. And what, you know, in other episodes that we'll talk about is like the one with Jonathan Frankel from Moshek-ML. And he also mentioned a lot of the benchmarks are multiple choice. But most production workloads are like open-ended text generation questions.
Starting point is 00:19:19 So how do you kind of reconcile the two? Yeah. Did you all get into it all, you know, the whole space of LLMs evaluating LLMs. And sort of this was something on a recent episode we talked to Jerry from Lama Index about in terms of on the one hand generating questions like you're talking about to evaluate LLMs or using an LLM to look at the question. a context and a response and provide an evaluation. I think that's definitely something that I think is interesting and has come up in a few of our episodes recently, where people are struggling to
Starting point is 00:19:55 evaluate these things. And so, yeah, we've seen a similar trend in one direction thinking about benchmarks and in another direction thinking about this sort of on the fly or model-based evaluation, which has existed for some time. Like in machine translation, it's very common. So like unbabel uses a model called comet. That's a... like one of the most popular highest performing machine translation evaluators as a model. It's not a metric and that sort of thing like blue. So yeah, that's a trend that we've seen is is evaluation and specifically evaluation for LLMs, which can kind of get dicey. Yeah, we did a benchmarks one-on episode that is also well liked. And we talked about this concept of like a benchmark driven
Starting point is 00:20:38 development, you know, like the benchmarks used to evolve every like three, four years. And now the models are catching up every like six months. So there's kind of this race between the benchmarks creators and like the models developers to find, okay, the state of the art benchmarks is here. And GPT4 on a lot of them gets like, you know, 98 percentile results. So, you know, GDP4 is not a GI. Therefore, to get to a GI, we need better evils for these models to start pushing the boundaries. And yeah, I think a lot of people are experimenting with using models to generate these things. but I don't think there's a clear answer yet.
Starting point is 00:21:15 Something that I think we were quite surprised to find was specifically in Hello Swag, where the benchmarks, instead of being manually generated, were adversarially generated. And then I was very interested in our, I mean, this is kind of like segwaying, we're not really going in sequence here, segueing into our second most popular episode,
Starting point is 00:21:34 which was one RoboFlow, which covered segment anything from meta. I think you guys had a discussion about that too. Yeah, it's been mentioned on the show. I don't think we've had a show devoted to it. Well, the most surprising finding when you read the paper is that something like less than 1% of the data of the mass that they released were actually human generated, a lot of it was AI assisted. So you have essentially models, evaluating models, and the models themselves trained on model generated data. We're very many layers
Starting point is 00:22:07 in at this point. Yeah, yeah. And I know that there's been a few papers, recently about the sort of things that were done with Lama and other models around model-generated output and data sets. It'll be interesting to see. I think it's still early days for that. So I think at the very minimum, what all of these cases show is that models, either evaluating models or using simulated data, I think back a few years ago, we would probably call this simulated data, right? I don't think that term is quite as popular now. Augmented? Yeah. Or augmentation data. augmentation, simulated data.
Starting point is 00:22:44 So I think this has been a topic for some time, but the scale at which we're seeing this done is kind of shocking now. And encouraging that we can do quite flexible things by combining models together, both at inference time, but also for training purposes. Well, have you ever come across this term of mode collapse? What I fear is, especially as someone who cares about low-resourced stuff,
Starting point is 00:23:09 is that stacking models on top of models and top of models, you just optimize for the median use case or the modal use case. Yeah. Yeah, I think that one maybe, so yeah, that is a concern. I would say it's a valid concern. I do think that these sort of larger models, and this gets, I guess, more into like multilingualism and the makeup of various data sets of these LLMs, the more that we can have linguistic diversity represented in these LLMs, which I know, I think Cohere for AI, just,
Starting point is 00:23:41 announced like a community-driven effort to increase multilinguality in LLM data sets. But I think the more we do that, I think it does benefit the downstream lower resource languages and lower resource scenarios more because we can still do fine-tuning. I mean, we all love to use pre-trained models now. But like in my previous work, when you were looking at maybe an Arabic vernacular language, rather than standard Arabic. There's so much standard Arabic in data sets. Making that leap to an Arabic vernacular is much, much easier if that Arabic is included in LLM data sets because you can fine-tune from those. So that is encouraging that that can happen more and more. There's still
Starting point is 00:24:29 some major challenges there. And especially because most of the content that's being generated out of models is not in, you know, Central Siberian Yupik or one of these. languages, right? So we can't purely rely on those. But I think my hope would be that the larger foundation models see more linguistic diversity over time. And then there's these sort of grassroots organizations, grassroots efforts like Masacane and others that rise up kind of on the other end and say, okay, well, we'll work with our language community to develop a data set that can fine-tune off of these models. And hopefully there's benefit both ways in that sense. Since you mentioned Masa Kani a couple times, we'll drop the link in the show notes so people can find it.
Starting point is 00:25:17 But what exactly do they do? How big of an impact have they had? Yeah, I would say so if people aren't familiar, if you go to the link, you'll see it. They talk about themselves as a grassroots organizations of African NLP researchers creating technology for Africa. So we have our own kind of biases as people in an English-driven. sort of literate world of what technology would be useful for everyone else. Like it probably makes sense for maybe listeners to say, well, wouldn't it be great if we could translate Wikipedia into all languages?
Starting point is 00:25:55 Well, maybe, but actually the reality on the ground is that many language communities don't want Wikipedia translated into their language. That's not how they use their language. Or they're not literate and they're in oral culture. So they need speech, right? texts won't do them any good. So that's why Masakane has started as a sort of grassroots organization of NLP practitioners who understand the context of the domain that they work in and are able to create models
Starting point is 00:26:26 and systems that work in those contexts. There's others, you can hear them on like the AI for Africa episodes that we have that talk about like agriculture use cases. Agriculture use cases in the U.S. might look like. you know, John Deere Tractor with a camera. Like, I don't know if people know this, but like John Deer tractors are these big tractors, they literally, they have a Kubernetes cluster on, like, some of them have a Kubernetes cluster on the tractor.
Starting point is 00:26:53 It's like a at the edge Kubernetes cluster that runs these models. And like when you're laying down pesticide, there's cameras that will actually identify and spray like individual weeds rather than like spraying the whole field. So that's like at the level that, you know, maybe is useful. here in Africa, maybe the more useful thing is around disease or drought identification or disaster relief or other things like that. And so there's people working in those environments or in those domains that know those domains that are producing technology for those cases. And I think that's really important. So yeah, I'd encourage people to check out Masacane. And there's other groups
Starting point is 00:27:33 like that. And if you're in like the U.S. or Europe or wherever and you want to get involved, there's open arms to say, hey, come help us do these things. So yeah, get involved too. What else is in your top three? Oh, yeah. So one recent one from Raj Shah from Hugging Face. Some people might have seen his really cool videos on LinkedIn or other places. He makes TikTok videos about AI models, which is awesome. And his episode is called the capabilities of LLMs. And I thought it was really a good way to help me understand like the landscape of large language models and the various features or axes that they're kind of situated in. So one axis is for example, closed or open, right? Can I download the model? But then on top of that, there's another there's another axes which is, is it
Starting point is 00:28:31 available for commercial use or is it not? And then there's other axes like we already talked about multilinguality, but then there's like task specificity, right? Like there's code gen models and there's language generation models and there's, of course, image generation models and all of those as well. So yeah, I think that episode really helps set a good foundation, no pun intended, for language models to understand where they're situated. So you can kind of, when you go to Hugging Face and there's, what is there like 200,000 models now? Maybe there's, I don't know how many models there are. How do I like navigate that space and understand what I could pull down, or do I fit into one of those use cases where it makes sense for me to just connect to Open AI or Cohere or Anthropic helps kind of
Starting point is 00:29:19 situate yourself. So I think that's why that episode was so popular as he kind of lays all of that out in an understandable way. How do you personally stay on top of models? You know, there's leaderboards, there's Twitter, there's LinkedIn. Yeah, I think it's a little bit spread out for me between the sources that you mentioned as podcasters. I think that's one of the, yeah, it's one, well, it's also a benefit for us. I think like if I didn't have every week on Wednesday, like, I'm going to talk about this topic, whether like I'm planning to think about a certain thing or not, it kind of helps you prompt and look at what's going on. So I think that is an advantage of like content creators is it is kind of a responsibility, but it's all.
Starting point is 00:30:04 also an advantage that we can have to, like, have the excuse to have great conversations with people every week. But yeah, I think Twitter's a little bit weird now, as everybody knows, but it's still a good place to find out that information. And then sometimes, too, like, to be honest, I go to Hugging Face and, like, I'll search for models, but I also search and I look at the statistics around the downloads of models, because generally when people find something useful, then they'll download it and download it over and over. So sometimes when I hear about like a family of models, I'll go there and then I'll look at some of the statistics on Hugging Face and like try a few things. Yeah. And some of these forks, I see the download numbers, but I've
Starting point is 00:30:49 never heard of them outside of Hugging Face. Yeah, it's true. It's true. Yeah. And some of them, like there'll be a fork or like a fine tune or something. And it's a, you do have to do a little bit digging around licensing and that sort of thing too but it is a useful like there's tons of people doing amazing stuff out there that aren't getting recognized at the like you know falcon or mpt level but there's a lot of people doing cool stuff that are releasing models on hugging face maybe that they've just found interesting any unusual ones that you recently found well there's one that i'll highlight which i thought it was cool because i don't know if you you all saw the um meta released this, I think it was like, six modality model. Yeah, yeah. And it was interesting because we did this work
Starting point is 00:31:36 with Masacane when I was at SIL. We did this work with Masacane and Koki, which is a speech tech company to create these language models in like six African languages. And I was like, okay, that's, that's cool. Like we did that. We formed like the data sets. It was satisfying. But now I'm like learning that then Meadow went and found that data on on Hugging Face and that's kind of incorporated in these these new models that meta has released. So it's cool to see like the full cycle thing happen where there was grassroots organizations seeing a need for models gathering data, doing baselines, and now there's like extended functionality in kind of like a more influential way, I guess, at like that higher level. Yeah.
Starting point is 00:32:28 Yeah, I think, I mean, talking about open and close models, when we started a podcast, it kind of looked like a cathedral kind of market where we had coherent, anthropic, open AI, stability, and those were like the hottest companies. I think now, you know, as you mentioned, you go on the Hocking Face. Like, I just opened there right now. There's the Sutter's news research 13 billion parameters model that just got released, fine-tune on over 300,000 instructions. It's like, models are just pop. up everywhere, which is great. And yeah, we had an episode with, as I mentioned, with Jonathan Frankel and I've been off from a mosaic camel to introduce MPTE7B and some of the work that they've done there.
Starting point is 00:33:09 And I think like one of their motivation is like keeping the space as open as possible, like making it easy for anybody to go obviously ideally on Moshekamel's platform and turn their own models and whatnot. So that's one that people really liked. I thought it was really technical. So I was really a little worried at first. I was like, is it going to fly over most people's head? But it was actually super well received.
Starting point is 00:33:34 Exactly. Now that was a good learning. Leaning in. Exactly. And Jonathan is super passionate about open source. He had this rant halfway through the episode about why it's so important to keep models open. And I actually edited in a crowd applause into the podcast, which I kind of love. I love little audio bonuses for people listening along.
Starting point is 00:33:56 And I think the change log guys do that really well Especially in their newer episodes Yeah, we need to, there is a way for us to integrate some of those things Yeah, like the soundboard thing and we we've never got into it too much I need to work with Jared from the change log and see It just spices it up exactly exactly You can have you can only have so many hour long conversations about ML We're yeah I keep thinking that but then we keep going
Starting point is 00:34:22 Right right right sorry I didn't mean like it was like a No, she switches it up and makes it audio interesting to add variety. Cool. I don't know if there are any other highlights that we want to do for. I'll just highlight maybe one more. Kirsten Lum was on, she had an episode about machine learning at small organizations. I think that's a great one. Like if you're a data scientist or a practitioner or an engineer at like either a startup or a mid-sized company where I think the thing that she emphasized was these
Starting point is 00:34:56 different tasks that we think about, like, you know, whether it's curating a data set or training a model or fine-tuning a model or deploying a model, sometimes at like a larger organization, those are functions in and of themselves. But when you're in this sort of mid-range organization, that's like a task you do, right? So to think about like those tasks as tasks of your role and time box them and understand like how to do all of those things well without getting sucked down into any one of those things. That was an insight that I found quite useful in my day to day as well as to sort of start to get a little bit of like spidey sense around, hey, I'm spending a lot of time doing this, but which probably means I'm like stuck in too much. Like I'm making my ML ops
Starting point is 00:35:46 too complicated, right, to track versions and like tie all this stuff together. Maybe I should just like do a simple thing and like paste a number in a Google sheet and move on. or something. I think that's a good segue into some of the other work that you do. You run the Datadan.com.com. website, which is kind of like a different types of workshop and advising that you do. I think a lot of founders especially are curious about how are companies thinking about using this technology. There's a lot of demos on Twitter, a lot of excitement. But when founders are putting together something that they want to sell, they're like, okay, what are the real problems that enterprises have? What are like some of the limitations that they have? We talked about
Starting point is 00:36:27 commercial use cases and something like that. Can you maybe talk a bit about, you know, two, three high-level learnings that you had from these workshops on like how these models are actually being brought into companies and how they're being adopted? Yeah, I think maybe one higher level comment on this is even though we see all these demos happening, everybody's using chat GPT. The reality in enterprise is most enterprises still don't have like LLMs integrated across their technology stack, right? So that might be a bummer for some people, like, oh, it's not quite as pervasive. But I actually find it as refreshing. Maybe because some of us feel like stuff happens every week.
Starting point is 00:37:13 It's exhausting to keep up. Like, oh, if I don't keep up with this stuff, then like I'm getting left behind. But it takes time for these things to trickle down. And not everything, like we were talking about the stable diffusion use case and others, like not everything that's hyped at the moment will be a part of your like data day life forever. Right. So you can kind of take some comfort in that. I think it's really important for people to, if they're interested in these models,
Starting point is 00:37:39 to really dig into more than just kind of a single prompt into these models. The practical side of using generative text models or LLMs really comes around either what some people might call prompt engineering, but, you know, understanding things like giving examples or demonstrations in your prompt, using things like guardrails or or reject statements or prediction guard to structure output, doing like fine tuning for your your company's data. Like these things go, there's kind of a hierarchy of these things. I think I think you all know Travis Fisher. He was, who's a guest on practical AI and talked about this hierarchy from prompt engineering through like data augmentation to fine tuning to eventually like training your own generative model.
Starting point is 00:38:38 I've really tried to encourage enterprise users and those that I do workshops with to think something like that hierarchy with these models. Like get hands on, do your prompting. But then like if you don't get the answer that you want immediately. I think there's a tendency for people to say, oh, well, it doesn't work for my use case. But there's so much of a rich environment underneath that with things like Lingchain and Lama Index and, you know, data augmentation, chaining, customization, fine-tuning, like all this stuff that can be combined together. It's a fun, new experience, but I find that enterprise users just haven't explored past that very most shallow level. So I think, yeah, in terms of the trends that I've seen with the workshop, I think people have
Starting point is 00:39:27 gone to chat GPT or one of these models. They've seen like the value that's there, but they have a hard time connecting these models to a workflow that they can use to solve problems. Like before we all had intuition, like, I'm going to gather my data. It's going to have these five. features. I'm going to train my psych at learn model or whatever. I'm going to deploy it with Flask and like now I have a cool thing. Now all of that intuition has sort of been shattered a little bit. So we need to develop a new workflow around these things. And I think that's really the focus of the workshops is kind of rebuilding that intuition into a practical workflow that you can think through and solve problems with practically. You have a live prompt engineering class.
Starting point is 00:40:16 Prompt engineering, overrated or underrated? Yeah, I think prompt engineering as like a term is probably too hyped. I think engineering and ops around large language models, though, is a real thing. And it is sort of what we're transitioning to. Now, how much you want to say is like, that term gets used in all sorts of different contexts. It could mean just like, oh, I wrote a good prompt and I'm going to like sell it on Twitter or something. Prop base.
Starting point is 00:40:47 The marketplace of prompts. I wonder how they're doing, to be honest, because they get quoted in almost every article about prompts engineering. They got really, really good PR. Yeah, yeah. I mean, if people can sell their prompts, I mean, I'll offer that.
Starting point is 00:41:00 That's cool. I got prompts right here. You know. But I think it goes, like some people might just mean that. And I think that's maybe overhyped in my view. But I do think there's this whole level of engineering and operations
Starting point is 00:41:12 around prompts and chaining and data augmentation that is a real workflow that people can use to solve their problems. And that's more what I mean when I'm referring to like whatever, however you want to combine the word engineering with prompting and language models. Yeah, I've just been calling it AI engineering. AI engineering, that's good. Rangel with the AI APIs, know what to do with them. That is a skill set that is developing that is a self-specialty of software engineering.
Starting point is 00:41:40 Yeah, yeah. It is what it is. And I think part of something I'm really trying to explore is this, is this spillover of AI from the traditional ML space, like where you need a machine learning researcher or machine learning engineer. It's spilling over into the software engineering space. And there's this rising class of what I'm calling AI engineer
Starting point is 00:41:57 that is specialized in conversant in the research, the tooling, the conversations and themes and what do you think are the unique challenges that like someone coming from that latter group, like engineers that are advancing into this AI engineer position versus like probably more like, my background where I was in data science for some time and now I'm kind of like transitioning into this world. What do you think are the unique challenges for both groups of people? Oh, I mean, so I can speak to the software side and you can speak about the data science side.
Starting point is 00:42:30 It's simply that we are, for many of us, dealing with a non-deterministic system for the first time. That, by the way, we don't fully control because there's this conversation about the GPT4 regress in its quality and we don't know because model drifts is not within our control because it's a blackbox API from from open AI but beyond that there's this sense that the the latent space of capabilities yeah is not fully explored yet yeah right like there there's a hundred and seventy five billion or one trillion parameters in the model we're maybe using like 200 of them it's literally where where is that meme where like we're using 10% of our brain we are we're probably using 10% of what is capable in the model. And it takes some ingeniousness to unlock that.
Starting point is 00:43:17 Yeah, I think from the data science perspective, there's probably a desire to quickly to jump to these other things around fine-tuning or training your own models, where if you really do take this prompting, chaining, data augmentation seriously, you can do a lot with models sort of off the shelf and don't need to like jump immediately into training. So I think that is like a knee-jerk reaction on our end. And fine-tuning is going to be around for the foreseeable future as far as I can tell. But data scientists have maybe a different, because we've been dealing with the uncertainty or non-deterministic output for some time and have developed some intuition around that.
Starting point is 00:44:03 But that's mostly when we've been controlling the data sets, when we've been controlling like the model training. sort of thing. So to throw some of that out, but still deal with that, it's a separate kind of challenge for us. I just remembered another thing that we've been developing on the Latin space community, which is this concept of AIUX. Yeah. Right? That the last mile of showing something on the screen and making it consumable, easily usable by people is perhaps as valuable as the actual training of the model itself. Yes. So I don't know if that's an overstatement, to be honest. Obviously, you're spending like hundreds of millions of dollars training models.
Starting point is 00:44:38 Yeah. And like, you know, putting it in some kind of React app. It's not, it's not the biggest innovation in the world. But a lot of people from OpenEI say like chat GBT was mostly a UX innovation. Yeah, I think like leading up to chat GBT, like when I saw the output of chat GPT, it wasn't, I don't think I had the same earth shattering experience that other people had in believing like, oh, this output is coming from a model. Like that, sure, it came from a model.
Starting point is 00:45:04 but the reception to like that interface and like the human element of the dialogue like that was so maybe it's both and right like it's not like you're not going to get that experience if you don't have the innovation under the hood and the modeling and the data set curation and all of that but it can totally be ruined by like the ux i typically give the example like one day in gmail i logged in and like i was typing my email and then like had the gray auto-complete, right? I did not get like a pop-up that said, like, do you want us to start writing your emails with AI? Like, it just, like, was so smooth and it happened and it made, like, it created value for me instantly, right? So I think that there is really a sense to that, especially in this area where people have a lot of, like, misgivings or fear around the technology itself. And we're going to have Alex Gravely on in a future episode, but GitHub, when they had the initial codex model from OpenEI, they spent six months tuning the U.S. Just to get co-pilot to a point where it's not a separate pane.
Starting point is 00:46:12 It's not a separate text box. It's kind of in your code as you write the code. And to me, that's more the domain of traditional software engineering rather than ML engineers or research engineers. Yeah. Yeah. I would say that is probably, yes. To circle back to what we were talking about, like challenges that are unique to like
Starting point is 00:46:29 engineers coming into this versus like data scientists coming into this. That's something data scientists, I think, have not thought about very much at all. At the very most, it's data visualization that they've thought about, right? Whereas engineers generally, like, there's some human, I mean, unless you're just a very pure backend systems engineer, like thinking about UI, UX is maybe a little bit more natural to that group. You mentioned one thing, which is about data set curation. We're in the middle of preparing this long overdue episodes on Datasets 101.
Starting point is 00:47:02 Any reflections on the evolutions in natural NLP datasets that have been happening? Yeah, great question. I definitely like, I think, are you all familiar with Label Studio? And that is one of the most popular kind of open source frameworks for data labeling. And they've been, I think they've been on, we have them on the show, like, we try to have them on the show every year as, like, data labeling expert. Maybe it's time for that. It's just reminding me. So they just released the so Aaron McHale is in the late space discord.
Starting point is 00:47:36 I think you had her on at the ODSC. Yeah, she was at ODSC. Yeah. So they just release new tools for fine-tuning generative AI models. Exactly. Yeah. It's a location. I think maybe the that being an example of this is maybe a trend that we're seeing there is
Starting point is 00:47:53 around augmented tooling or tooling that's really geared towards an approachable way. to fine tune these models with human feedback or with customized data. So like I know with label studio, a lot of the recent releases had somewhat to do with like putting LLMs in the loop with humans during the label process similar to like I think Prodigy has been doing this for some time, which is from Spacey. So this sort of human in the loop labeling and update of a model, they brought some of that in. but now like this new kind of set of tooling around specifically instruction tuning of models. I think before maybe people, and I've seen actually this misconception, I was in an advising call with a client.
Starting point is 00:48:44 And they're really struggling to understand like, okay, our company has been training or fine tuning models. Now we want to create our own like instruction tuned model. like how is that different from what we've been doing in the past? And kind of what I tried to help them see is, yes, like some of the workflow that happened around like reinforcement learning from human feedback is unique. But reinforcement learning is not unique. There's an element of training in that. There's data set curation in that. There's pre-training that happen like before that whole process happened.
Starting point is 00:49:22 So the elements that you're familiar with are part of that. they're just not packaged in the same way that you saw them before. Now there's this clear pre-training stage and then the human feedback stage and then this reinforcement learning happens. So I think the more that we can bring that concept and that workflow into tooling, like what label studio is doing, to make it more approachable for people to where it's not like this weird like reinforcement learning from human feedback sounds very confusing to people like PPO and helping people understand like how reinforcement learning works.
Starting point is 00:50:01 It's very difficult. So the more the tooling can just have its own good UI, UX around that process, I think the better. And probably Label Studio and others are leading the front on, leading the way on that front. I was thinking like so labels are one thing. And by the way, okay, I'll take the side tangent on labels and I'll come back to the main point. Yeah, yeah.
Starting point is 00:50:21 I actually presume that scale would win everything. Yeah. And it seems like they haven't. Yeah. And, sorry, there's scale, this snorkel, there's this generation of labeling companies that came out. Like data-centric AI companies, yeah. Right. Right.
Starting point is 00:50:34 What happened? Like, how come there's still new companies coming up? There's label box, there's label studio. I don't have a sense of how to think about these companies. Like, obviously, it was important. Yeah. Yeah, I think also even before that, there was like tool, at least features even from cloud providers or whatever, like auto ML. like came before that like upload your own data create your own custom model so I think that maybe it's
Starting point is 00:51:02 that like companies that want to create this sort of custom models and this is just my own opinion I'll preface that maybe they don't want like when they're thinking about that problem they're not thinking about oh I need a whole platform to create custom models using our data they're more thinking about like how do I use these state of the art models with my data. And so it's still, if those statements are very similar, but if you notice like one is more model-centric and one is more data-centric. So I think enterprises are still thinking like model-centric and augmenting that with their data, whether that be just through augmentation or through fine-tuning or training. They're not necessarily thinking about like a data platform for AI. They're thinking,
Starting point is 00:51:52 about bringing their AI or their data to the AI system, which is why I think like APIs like cohere, Open AI that offer fine tuning as part of their API. It's sort of like people love that. It makes sense. Like okay, I can just upload some examples and it makes the model better. But it's still like model centric, right? Yeah. I get the sense that OpenEI doesn't want to encourage that anymore because they don't have
Starting point is 00:52:18 fine tuning for 3.5 and 4. And then so the last thing I'll do about datasets. and we can go into the lightning round. I was actually thinking about un-labled data sets for unsupervised learning or self-supervised learning. That is something that we are trying
Starting point is 00:52:31 to wrap our heads around. Common crawl, Stack Overflow archive, the books. I don't know if you have any perspectives on that, like the trends that are arising here, the best practices. As far as I can tell, nobody has a straight answer
Starting point is 00:52:45 as to what the data mix is and everyone just kind of experiments. Yeah, well, I think that's partly driven by the fact that like the most popular models, you don't really have a clear picture of what the data mix is, right? So the people that are trying to recreate that and they're not achieving that like level of performance, right? Then they, one of the things that they know is, well, what are all the different data mix options that I can try and try to replicate some of what's going on, right? So I think it's partly driven by that is like we don't totally know what the data mix is like sitting
Starting point is 00:53:21 behind the curtain of Open AI or others. But I think there's a couple of trends, I guess, which you've already sort of highlighted. One is, like, how can I mix up all of these public data sets and filter them in unique ways to make my model better? So some, I listened to a talk, I believe it was at last year's ACL, and they did this study of common crawl, right? And they found that actually a significant portion of common crawl was like mislabeled all over the place, right? Like trash, yeah.
Starting point is 00:53:58 So like I think it was 100% of the data that was labeled as Latin character Arabic. So Arabic written in Latin characters was not Arabic. Like 100% of it. And there was like all sorts of other problems and that sort of thing. So I think there's one side, one group of people or set of experiments that you could think about as like, how do I take these existing data sets, which I know have data quality issues or maybe other data biases or problems that I would like to filter out, like not fit for work data, that sort of thing. So how do I create my own special filtered mix of these and train a model? So that's one kind of genre. And then there's a other genre, which is like maybe take. taking those, but augmenting them with this like simulated or augmented data, right, that's out of a model, like a GPT model or something like that. So I think you could combine those in all sorts of unique ways. And I think it is a little bit of like the Wild West because we don't totally have a good grip on what is the winning strategy there.
Starting point is 00:55:05 And so I think that's where I would also encourage people to try a variety of models. So this is maybe a problem with benchmarks in general, right? You can see the open large language model benchmark on hugging face, and these models are at the top. And you could come away with that and say, well, I'm like anything below like the top three I'm not even going to use, right? But the reality is that each of those had a unique sort of flavor of this data under the hood that might actually work quite well for your use case. So one example that I've used recently in some work is the Camel $5 billion model from writer. You know, it doesn't work great for a lot of things. But there's certain things around like marketing copy and others that it does a really good job at.
Starting point is 00:55:55 And it's a bit smaller model that I can host and run. And I can get good output out of it if I put in some of that workflow and structuring around it. But I wouldn't use it for other cases. But that has a lot to do with the data. and I'm guessing writers focus on that copy generation and such. So, yeah, I would encourage people specifically on this topic to maybe think about what's going on under the hood and also give some models a try for different,
Starting point is 00:56:23 like gain your own intuition about how a model behavior might change based on how it was trained in the mix of data that went in. Awesome. Let's jump into the lightning round. We have three questions for you. It's lightning, but you can say 30. seconds to answer. All right, cool. So the first question is around acceleration. What's something that already happen in AI that Utah would take much longer? Yeah, I think the thing that I was thinking about here was like how general purpose these large language models are beyond
Starting point is 00:56:57 traditional in all P tasks. So it doesn't surprise me that maybe they could do like sentiment analysis or even like NLI or something like that. These are things that have been studied for a long time. But the fact that I can, like, at ODSC, I was in like a workshop on fraud detection. And they were using like some, I forget the models they were using, some statistical models to do fraud detection. I was like, I wonder if I just like do a bit of chaining and like insert some of the examples of these insurance transactions into my prompts if I can get the large language model to detect a fraudulent insurance client. And it seemed to like, like I got pretty. far doing that. So that fact of like you can do something like that with these models are that
Starting point is 00:57:44 generalizable beyond traditional NLP techniques I think is surprising to me. Awesome. Exploration. What are the most interesting unsolved questions in AI? Yeah. I think there is still such a focus on English and Mandarin. It's like that like you're kind of large language model wise. If you look at the drop off performance after you get past like English, Mandarin, German. It's Spanish to some degree, but German is actually better than Spanish because of how much it's been studied in NLP. And of course, Mandarin has a lot of data. Spanish still does good. But like there's languages, even in the top hundred languages of the world that are spoken by millions and millions and millions of people around the world that don't like perform well in these models. So that's like thing one,
Starting point is 00:58:38 but even modality-wise. I know there's a lot of work going on in the research community around sign language, but like there's all of these different modalities of language. Written text is not, does not equal communication, right?
Starting point is 00:58:54 Written text is a synthesis of communication into a written form that some people consume. But the combination of all of these modalities along with all of these languages, there's just so much room to explore there and so many challenges left to explore that will eventually, I think,
Starting point is 00:59:15 help us learn a lot about communication in general and the limitations of these models, but is an exciting area. It's definitely a challenge, but an exciting area. Awesome, man. So one last takeaway. What's something or a message that you want everyone to remember today? Yeah, similar to when you're asking about my workshops, I think I would just encourage people to get hands-on with these models and really dig into the new sets of tooling that are out there. There's so much good tooling out there
Starting point is 00:59:45 to go from like a simple prompt to inject your own data, to form like a query index, to create like a chain of processing, even like trying agents and all those things. Like get hands on and try it. That's the only way that you're going to build out this intuition. So yeah, that would be my encouragement.
Starting point is 01:00:04 Excellent. Well, thanks for coming. on. Yeah. Thank you guys so much. This is awesome.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.