Latent Space: The AI Engineer Podcast - AGI is Being Achieved Incrementally (DevDay Recap - cleaned audio)

Episode Date: November 8, 2023

We left a high amount of background audio in the Devday podcast, which many of you loved, but we definitely understand that some of you may have had trouble with it. Listener Klaus Breyer ran it throu...gh Auphonic with speech islolation and we figured we’d upload it as a backdated pod for people who prefer this. Of course it means that our speakers sound out of place since they now sound like they are talking loudly in a quiet room. Let us know in the comments what you think?Timestampsthe cleaned part is only part 2:* [00:55:09] Part II: Spot Interviews* [00:55:59] Jim Fan (Nvidia) - High Level Takeaways* [01:05:19] Raza Habib (Humanloop) - Foundation Model Ops* [01:13:32] Surya Dantuluri (Stealth) - RIP Plugins* [01:20:53] Reid Robinson (Zapier) - AI Actions for GPTs* [01:30:45] Div Garg (MultiOn) - GPT4V for Agents* [01:36:42] Louis Knight-Webb (Bloop.ai) - AI Code Search* [01:48:36] Shreya Rajpal (Guardrails) - Guardrails for LLMs* [01:59:00] Alex Volkov (Weights & Biases, ThursdAI) - "Keeping AI Open"* [02:09:39] Rahul Sonwalkar (Julius AI) - Advice for Founders This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

Transcript
Discussion (0)
Starting point is 00:00:04 Hey everyone, this is Swix coming at you live from the Newton, which is in the heart of the cerebral arena. It is a new AI co-working space that I and a couple of friends are working out of. There are hot desks available if you're interested. Just check the show notes. But otherwise, obviously, it's been 24 hours since the opening AI dev day. A lot of hot reactions and long-standing tradition, one of the longest traditions we've had on the latent space pod is to convene emergency sessions and record the live. thoughts of developers and founders going through and processing in real time. I think a lot of the roles of podcasts isn't as perfect information delivery channels, but really as an audio and oral history
Starting point is 00:00:46 of what's going on as it happens while it happens. So this one's a little unusual. Previously, we only just gathered on Twitter spaces and then just had a bunch of people. The last one was the code interpreter one with 22,000 people showed up. But this one is a little bit more complicated because there's an in-person element and then a online element. So this is a two-part episode. The first part is a recorded session between our Latent Space people and Simon Willison and Alex Volker from the Thursday iPod, just kind of recapping the day. But then also, as the second hour, I managed to get a bunch of interviews with previous
Starting point is 00:01:22 guests on the pod who we are still friends with and some new people that we haven't yet had on the pod, but I wanted to just get their quick reactions because most of you have known and I loved Jim Fan and Divgarg and a bunch of other folks that we interviewed. So I just want to, I'm excited to introduce to you the broader scope of what it's like to be at OpenEye Dev Day in person, bring you the audio experience, as well as give you some of the thoughts that developers are having as they process the announcements from Open AI. So first off, we have the Layton Space pod recap one hour of OpenAID DEVD. Hey, everyone. Welcome to the Latenspace podcast on Emergency Adepa. addition after opening I'd have day.
Starting point is 00:02:04 This is Alessio, partner in CETN residents and decibel partners. Then as usual, I'm joined by Spuix, founder of Smalley Eye. Hey, and today we have two special guests with us covering all the latest and greatest. We love to get our band together and recap things, especially when they're big. And it seems like that every three months, we have to do this. So Alex, welcome from Thursday, I. We've been collaborating a lot on the Twitter spaces and welcome Simon from many, many things. but also I think you're the first person to not make four appearances on our pod.
Starting point is 00:02:34 Oh, wow. I feel privileged. So welcome. Yeah, I think we're all there yesterday. How do we feel? Like, what do you want to kick off with? Maybe Simon, you want to take first and then Alex. Sure. Yeah. I mean, yesterday was quite exhausting, quite frankly. I feel like it's going to take us as a community several months just to completely absorb all of the stuff that they dropped on us in one giant batch. It's particularly impressive considering they launched a ton of features. What, three or four weeks ago, chat GPT voice and the combined mode and all of that kind of thing. And then they followed up with everything from yesterday.
Starting point is 00:03:06 That said, now that I've started digging into the stuff that they released yesterday, some of it is clearly in need of a bit more polish. You know, the reality of what they released is, I'd say about 80% of what it looked like it was yesterday, which is still impressive. You know, don't get me wrong. This is an amazing batch of stuff. But there are definitely problems and sharp edges that we need to file off. And there are things that we still need to figure out before we can take advantage of all of this.
Starting point is 00:03:31 Yeah, agreed, agreed. And we can go into those trap edges in a bit. I just want to pop over to Alex. What are you with us? So, interestingly, even folks at OpenEI, there's like several booths and a help desk. So you can go in and ask people like actual changes and people like they could follow up with like the right people in OpenEI and like answer you back, et cetera. Even some of them didn't know about all the changes. So I went to the voice and audio booth and I asked them about like, hey, is Whisper 3 that was announced by Sam Outman?
Starting point is 00:03:57 stage just like briefly, will that be open source because I love using Whisper. And there's like, oh, did we open source? Do we talk about Whisper? Like some of them didn't even know what they were releasing. But overall, I felt it was a very tightly run event. Like I was really impressed. Sean, we were sitting in the audience and you like pointed at the clock to me when they finished.
Starting point is 00:04:16 They finished like on 45 on that, I think, right? And this was after like doing some extra stuff. Very, very impressive for a first event. Like I was absolutely like, good job, guys. Good job. Yeah, apparently it was their first keynote. And someone, I think, was it you that told me that this is what happens if you have a president of White Combinator, do a proper keynote, you know, having seen many, many, many presentations by other startups. This is sort of the sort of masterstroke.
Starting point is 00:04:43 Yeah, Alessio, I think you were watching remotely. Yeah, we were at the Newton. Yeah, I think we had 60 people here at the watch party. So it was quite a big crowd. It makes a reaction from different founders and people, depending on well, always being announced on the page. But I think everybody walked away, kind of really happy with a new layer of interfaces,
Starting point is 00:05:05 big in use. I think to me the biggest takeaway was like, and I was talking with Mike Conover, another friend of the podcast about this, is they're kind of staying in the single-treaded, like synchronous use cases lane, you know? Like the GPD's announcement are all like still chat-based, one-on-one synchronous things.
Starting point is 00:05:23 I was expecting maybe something about async things. like background running agents, things like that. But it's interesting to see there was nothing of that. So I think if you're a founder in that space, you're quite excited. You know, they seem to up to pick a product lane, at least for the next year. So if you're working on async experiences, so things working in the background, things that are not copilot like, I think you're quite excited to have them be a lot cheaper now. Yeah.
Starting point is 00:05:52 As a person in itself, like I often think about this as a passing of a, a big risk in terms of uncertainty over opening as roadmap. Like, you know, they've shipped everything they're probably going to ship in the next six months. You know, they sort of marked out the territories that they're interested in. And then so now that leaves open space for everyone else to pursue. So I guess we can kind of go in order. Probably top of men, top of mind to mention is the GPT4 turbo improvements.
Starting point is 00:06:21 So longer context length, cheaper price, anything else that stood out in your viewing of the keynote and then just the commentary around it. I was waiting for Stateful. I remember they talked about Stateful API, the fact that you don't have to keep sending like the same tokens back and forth just because, you know, and they're going to manage the memory for you. So I was waiting for that. I knew it was coming at some point.
Starting point is 00:06:43 I was kind of, did not expect it to come kind of at this event. I don't know why. But when they announced Stateswell, I was like, okay, this is making it so much easier for people to manage state. The whole threads. I don't want to mix between the two things, so maybe you guys can clarify, but there's the GPD4-2B, which is the model that has new capabilities, a whopping 128K, like, a context link, right? It's huge. It's like two and a half books, but also, you know, faster, cheaper, etc.
Starting point is 00:07:12 I haven't yet tested the fasterness, but, like, everybody's excited about that. However, they also announced this new API thing, which is the assistance API, and part of it is threads, which is will manage the thread for you. I can't imagine, like, I can't imagine how many times I had to like re-implearned this myself in different languages and type script and Python, etc. And now it's like, it's so easy. You have this one thread. You're sending it to a user and you just keep sending messages there and that's it. The very interesting thing that we attended and by we, I mean, like SWIGS and I have a live space and like 200 people. So it's like me, SWIX and 200 people with us as well, they kept asking like, well, how's the price happening? If you're sending just the
Starting point is 00:07:48 tokens like the Delta, like what the new user just sent, what are you paying for? And I went to open AI people and I was like, hey, how do we get paid for this? And nobody knew, nobody knew. I finally got an answer. You still pay for the whole context that you have inside the thread. They still pay for all this. But now it's a little bit more complex for you to kind of count with TikTok. So you have to hit another API endpoint to get the whole thread of what the context is.
Starting point is 00:08:12 Then TikTokanize this, run this to TikTok, and then calculate. This is now the new way officially from Open the Eye. But I really did have to go and find this. They didn't know a lot of how the pricing is going to have. Ouch. Yeah. Yeah. Does the API at least tell you how many tokens you used,
Starting point is 00:08:30 or is it entirely up to you to do the accounting? Because that would be a real pain if you have to account for everything. So in my head, the question I was asking is, like, if you want to know in advance before hitting the API, like with the library hook token, if you want to count in advance or like make a decision like advanced on that, how would you do this now? And they said, well, yeah, there's a way.
Starting point is 00:08:49 If you hit the API, get the whole thread back, then count the tokens. But I think the API still really sends you back the number of tokens. But isn't there a feature of this new API where they actually do, they claim it has, does it have infinite length threads because it's doing some form of condensation or summarization of your previous conversation for you? I heard that from somewhere, but I haven't confirmed it yet. So I have a source from Dave Waldman.
Starting point is 00:09:17 I actually don't know what his affiliation is, but he usually has pretty accurate takes on AI. So I think he works in AI circles in some capacity. So I'll feature this in the show notes, but he said, some not mention interesting bits from opening eye dev day. One unlimited context window and chat threads from opening eye docs. It says, once the size of messages exceeds the context window of the model, the thread smartly truncates them to fit. I'm not sure I want that intelligence.
Starting point is 00:09:43 I want to chime in here. Just real quick, the not want this intelligence. I heard this for multiple people over the next conversation that I had. Some people said, hey, even though they're giving us like a content understanding and rag, we are doing different things. Some people said this with vision as well. And so that's an interesting point that like people who did implement custom stuff, they would like to continue keeping implementing custom stuff. That's also like an additional point that I've heard to talk about. Yeah.
Starting point is 00:10:11 So what opening us doing is providing good defaults and then, well, good is questionable. We'll talk about that. I think the existing sort of like chain and llama indexes of the world are not very threatened by this because there's a lot more customization that they want to offer. So, frustration is that OpenAI, they're providing new defaults, but they're not documented defaults. Like they haven't told us how their rag implementation works. Like, how are they chunking the documents?
Starting point is 00:10:36 How are they doing retrieval? Which means we can't use it as software engineers because it's this weird thing that we don't understand. And there's no reason not to tell us that. Giving us that information helps us decide how to write good software on top of it. So that's kind of frustrating. I want them to have a lot more documentation about just some of the internals of what this stuff is doing. I want to highlight an additional capability that we got, which is document parsing via the API.
Starting point is 00:11:02 I was blown away by this, right? So we know that you could upload the images and vision API we got, we could talk about vision as well. But just the whole fact that they presented on stage, like the document parsing thing, where you can upload PDFs of the United Flight and then they upload like an Airbnb. That on the whole, like that's a whole category of like products that's now. open to open eyes, just like given developers to very easily build products that previously it was a pain in about for many, many people. How do you even like parse a PDF?
Starting point is 00:11:29 Then after you parse it, like what do you extract? So the smart extraction of like document parsing, I was really impressed with it. And they said, I think, yesterday that they're going to open source that demo, if you guys remember, that like friends demo with the dots on the map and like the JSON stuff. So it looks like that's going to come to open source. And many people learn new capabilities for document parsing. So I want to make sure we're very clear what we're talking about when we talk about API, when you say API, there's no actual endpoint that does this, right?
Starting point is 00:11:55 You're talking about the chat GPT's functionality. No, I'm talking about the assistance API, the assistant API that has threads now, that has agents and you can run those agents. Actually, maybe let's clarify for this point. I think I had to, somebody had to clarify for this for me. There's the GPTs, which is a UI version of running agents. We can talk about them later, but like you and I and my mom can go and like, hey, create a new GPT that like, you know, only does Technoric, jokes, like, whatever.
Starting point is 00:12:23 But there's the assistance thing, which is kind of a similar thing, but not the same. So you can't create, you cannot create an assistant via an API and have it pop up on the marketplace on the future marketplace they announced. Oh, can you not? No, no, no, not via the API. So there are like two separate things and somebody in the opinion I told me they're not, they're not exactly the same. That's so confusing because the AI looks exactly like the UI that you used to set up the GPTs.
Starting point is 00:12:49 I assumed there was an API for the same feature. And the playground, actually, if you go to the playground, it kind of looks the same. There's like the configurable thing. The configure screen also has like you can allow it browsing, you can allow it like tools. But somebody told me they didn't do the full cross-mapping. So like you won't be able to create GPPs with API.
Starting point is 00:13:07 You will be able to create assistants. And then you'll be able to have those assistants do different things, including call your external stuff. So that was pretty cool. Okay. So this API is called the assistant API. That's what we get like in addition to the model of the GPT4 Turbo.
Starting point is 00:13:21 And that has document parsing. So you can upload documents there, and it will understand the context of them and then return you like structured or unstructured input. I thought that that feature was like phenomenal just by all its own. Like just on its own, uploading a document, a PDF, a long one
Starting point is 00:13:36 and getting like structured data out of it, it's like a pain in you have to build. Let's face it, guys. Like everybody who built this before, it's like it's kind of horrible. When you say structured data, are you talking about the citations? The JSON output,
Starting point is 00:13:48 the new JSON output, the new J. Jason output that they also gave us. Finally, if you guys remember last time we talked together, I think it was like during the functions release, EmergencyPod. And back then, their answer to like, hey, everybody wants structured data was, hey, we're going to give you a function calling. And now
Starting point is 00:14:02 they did both. They gave us both like a JSON output like structure. So like you can, the models are actually going to return JSON. Haven't played with it myself, but that's what they announced. And the second thing is they improved the function calling significantly as well. So, so
Starting point is 00:14:17 I talk to a staff member there. I've got a pretty good model for what this is. Effectively, the JSON thing is they're doing the same kind of trick as Lama-grammas and JSON-forma. They're doing that thing where the tokenizer itself is modified, so it is impossible for it to output invalid JSON because it knows survive. Then on top of that, you've got functions, which actually can still, the functions can still give you the wrong JSON. They can give you JSON with keys that you didn't ask for if you're unlucky,
Starting point is 00:14:44 but at least it will be valid. At least it will pass through a JSON parser. And so they're very similar sort of things, but they're slightly different in terms of what they actually mean. And yeah, the new function stuff is super exciting because functions are one of the most powerful aspects of the API. But a lot of people haven't really started using yet. But it's amazingly powerful what you can do with it. I saw that the functions, the functionality that they now have is also plug inable as actions to those. Right.
Starting point is 00:15:11 So when you're creating assistants, you're adding those functions as like features of this assistant. and then those functions will execute in your environment, but they'll be able to call different things. Like they showcase an example of integration with, I think, Spotify or something, right? And that was like an internal function that ran. But it is confusing the kind of the online assistant, APIable agents, and the GPT's agents. So I think it's a little confusing because they demo both.
Starting point is 00:15:37 I think it's worth us talking about the difference between plugins and actions as well. Because, you know, they launched plugins back in February. And they've effectively, they've kind of deprecated plugins. They haven't said it out loud, but it's clear that they are not going to be investing further in plugins, because the new actions thing is covering the same space. But actually, I think, is a better design for it. Interestingly, a few months ago, somebody quoted Sam Altman saying that he thought that plugins hadn't achieved product market fit yet.
Starting point is 00:16:05 And I feel like that's sort of what we're seeing today. The problem with plugins is it was all a little bit messy. People would pick and mix the plugins that they needed. Nobody really knew which plugin combinations would work. With this new thing, instead of plugins, you build an assistant, and the insistent is a combination of a system prompt and a set of actions which look very much like plugins. You know, they get a JSON schema to call an API somewhere.
Starting point is 00:16:27 And I think that makes a lot more sense. You can say, okay, my product is this chatbot with this system prompt, so it knows how to use this tools. I've given it this combination of plug-in-like things that it can use. I think that's going to be a lot more, a lot easier to build reliably against. and I think it's going to make a lot more sense to people than the sort of mix and match mechanisms they had previously. So actually, maybe it will be cool to cover kind of the capabilities of an assistant, right?
Starting point is 00:16:52 So you have a custom prompt, which is akin to the system message. You have the actions thing, which is you're going to have add the existing actions, which is like browse the web and code interpreter, which we should talk about, like the assistants now can write code and execute it, which is exciting. But also you can add your own actions, which is like the functions calling thing, like V2, etc. Then I heard this incredibly quick thing that somebody told me that you can add two assistants to a thread. So you literally can mix agents within one thread with the user.
Starting point is 00:17:22 So you have one user and then you can have like this assistant and that assistant, they just glanced over this. I was like, that is very interesting. That is not very interesting. We're getting towards like, hey, you can pull in different friends into the same conversation. Everybody does the different thing. What other capabilities do we have there? Do you guys remember? Oh, like context, uploading it with context, like with the full, here's our API documentation.
Starting point is 00:17:47 Well, that one's a bit more complicated. So you've got the system prompt, you've got optional actions, you've got, you can turn on Dali-free, you can turn on code interpret, you can turn on Gras with Bing. Those can be added or removed from your assistant. And then you can upload files into it, and the files can be used in two different ways. There's this thing that they call, I think they call it the retriever, which basically does, it does rag. it does retrieval augmented generation against the content you've uploaded. But code interpreter also has access to the files that you've uploaded. And those are both in the same bucket.
Starting point is 00:18:17 So you can upload a PDF to it. And on the one hand, it's got the ability to turn that into, like chunk it up, turn it into vectors, use it to help answer questions. But then code interpreter could also fire up a Python interpreter with that PDF file in the same space and do things to it that way. And it's kind of weird that they chose to combine both of those things. Also, the limits were amazing, right? you get up to 20 files, which is a bit weird because it means you have to combine your documentation to a single file.
Starting point is 00:18:44 But each file can be 512 megabytes. They're giving us 10 gigabytes of space in each of these assistants, which is vast, right? Of course, I tested it'll handle SQLite databases. You can give it a gigabyte 12 megabyte SQLite database, and it can answer questions based on that. But yeah, it's, like I said, it's going to take us months to figure out all of the combinations that we can build with all of this. I was going to say for the storage. I saw Jeremy Howard tweeted about it's like 20 cents per gigabyte per assistant per day. Just to compare like S3 cost like two cents per month per gigabyte.
Starting point is 00:19:22 So like 300X more, something like that than just raw S3 storage. Ouch. There will still be a case for like maybe roll your own rag, depending on how much information you want to put there. but I'm curious to see what the price, the client curve looks like for the storage there. Yeah, they probably should just charge that at cost. There's no reason for them to charge so much. That is wildly expensive.
Starting point is 00:19:49 It's free until the 17th of November, so we've got 10 days of free assistance, and then it's all going to start costing us. Crikey. They gave us 500 bucks of API credit at the conference as well, which we'll burn through pretty quickly at this rate. I confirmed a very important. important question, everybody was asking. Did the five people who got the $500 first got actually $1,000? And I think somebody in OpenNine, I said, yes, there was nothing there that
Starting point is 00:20:15 prevented the five first people to not receive the second one again. Okay. I met one of them. I met one of them. He said he only got 500. Ah, interesting. Okay. So again, even Open AI, people will not necessarily know what happened on stage with Open the Eye. Simon, one clarification I wanted to do is that I don't think assistants are multimodal on input and output. So you do have vision, I believe, not confirmed by do believe that you have vision, but I don't think that Dali is an option for an option for GPs, but the guys... Oh, that's so confusing.
Starting point is 00:20:45 The assistants, the checkbox for Dali is not there. You cannot enable it. Well, you just add them as a tool, right? So, like, it's just one more... It's a little finicky. In the GPD interface. Yeah. I mean, to be honest, if assistants don't have Dali 3, we...
Starting point is 00:21:02 Does Dali 3 have an API now? I think they released one. I can't... There's so much stuff. That got lost in the pile. But yeah, so code interpreter, wow. That I was not expecting. That's huge.
Starting point is 00:21:16 Assuming, I mean, I haven't tried it yet. I need to confirm that it definitely works, because GPT is going to have code interpreter. Can assistance? Assistance will have code interpreter as well, yeah. Awesome. Huh. It's incredible.
Starting point is 00:21:29 Yeah, so I tried to make it do things that were not logical yesterday. because one of the risks of having the God model is it calls the wrong model inappropriately whenever you try to ask it to something that's kind of vaguely ambiguous. But I thought it handled the job decently well. I think there's still going to be rough edges. It's going to try to draw things. It's going to try to code when you don't actually want to. And in a sense, opening eye, it's kind of removing that capability from chaty-b-t.
Starting point is 00:22:00 It just wants you to always query the God model and always get feedback. on whether or not that was the right thing to do. Which really sucks because it runs, I like asking a question and it goes, oh, searching Bing. And I'm like, no, don't search Bing. I know that the first 10 results on Bing will not solve this question. I know you know the answer. So I had to build my own custom GPT that just turns off Bing because I was getting frustrated with it always going to Bing when I didn't want it to.
Starting point is 00:22:27 Okay. So this is a topic that we discussed, which is the UI changes to chat GPT. So we're moving on from the Assistance API. and talking just about the upgrades to chat GBT and maybe the GBT store. You did not like it. And I love this. I mean, both sides of this. Okay.
Starting point is 00:22:45 So my problem with it, I've got the two things I don't like. Firstly, it can do Bing when I don't want it to. And that's just, just irritating because the reason I'm using GPT to answer a question is that I know that I can't do a Google search for it because I've got a pretty good feeling for what's going to work and what isn't. And then the other thing that's annoying is it's just a little thing. interpreter doesn't show you the code that it's running as it's typing it out now. Like it'll churn away for a while doing something and then they'll give you an answer and you have to click a tiny little icon that shows you the code.
Starting point is 00:23:14 Whereas previously you'd see it writing the code so you could cancel it halfway through if it was getting it wrong. And okay, I'm a Python programmer so I care and most people don't, but that's been a bit annoying. Yeah, and when it errors, it doesn't tell you what the error is. It just says analysis failed and it tries again. But it's really hard for us to help it. Yeah. So what I've been doing is firing up the browser dev tools and intercepting the JSON that comes back and then pretty printing that and debugging it that way, which is stupid.
Starting point is 00:23:41 Like, why do I have to do that? It's really good feedback for OpenEI. I will tell you guys what I loved about this unified mode. I have a name for it. So we actually got a preview of this on Sunday. And one of the folks got like an early example of this. I call it MMIO, multimodal input and output, because now there's a shared context. between all of these tools together.
Starting point is 00:24:04 I think it's not only about selecting them, just selecting them. And Sam Altman on stage I said, oh yeah, we unified it for you, so you don't have to call different modes at once. And in my head, that's not all they did. They gave a shared context. So what is an example of shared context, for example?
Starting point is 00:24:18 You can upload an image using GPT for vision and eyes, and then this model understands what you kind of uploaded vision-wise. Then you can ask Dali to draw that thing. So there's no text shared in between those modes now. There's like only visual shared between in those modes and Dali will generate whatever you upload it in an image. So it's eyes to output visually. And you can mix the things as well.
Starting point is 00:24:39 So one of the things we did is, hey, use real-world real-time data from Bing, like weather, for example. Weather changes all the time. And we asked Dali to generate like an image based on weather data in a city. And it actually generated like a live almost like, you know, like snow, whatever. It was a snowing inventor. And that I think was like pretty amazing in terms of like being able to share contacts between all these different models and modalities in the same understanding.
Starting point is 00:25:03 And I think we haven't seen the end of this. I think like generating personal images, adding context to Dali, like all these things are going to be very incredible in this one mode. I think it's very, very powerful. I think that's really cool. I just want to opt in as opposed to opt out. Like I want to control when I'm using the gold model versus when I know, which I can do because I created myself a custom GPT that does what I need.
Starting point is 00:25:27 It just felt a bit silly that I had to do a hot. custom bot just to make it not do Bing searches. All solvable problems in the fullness of time. Yeah. But I think people, it seems like for the chat GPT at least, they're really going after the broadest market possible. That means simplicity comes at a premium at the expense of pro users. And the rest of us can build our own GPT rappers anyway.
Starting point is 00:25:51 So not that big of a deal. But maybe do you guys have a, oh, sorry. So the GPT rappers thing, guys, They call them GPTs because everybody's building GPDs. Like, all the rappers, whatever, they end with the word GPT. And so I think they reclaimed it. That's like, you know, instead of fighting and saying, hey, you cannot use the GPD. GPD.
Starting point is 00:26:10 It's like, we have GPTs now. This is our marketplace. Whatever everybody else builds, we have the marketplace. This is our thing. I think they did like a whole marketing move here. It's a very strong marketing moves because now it's called Canva GPD. It's called Zapier GPT. And they're basically saying don't build your own websites, build it inside of our
Starting point is 00:26:28 God app with chat GPT and that's the way that we want you to do that. In a way, it sort of makes up, it sort of makes up the fact that chat GPT is such a terrible name for a product, right? Chat GPT, what were they thinking when they came up with that name? But I guess if they lean into it, it makes a little bit more sense. It's like chat GPT is the way you chat with our GPs and GPT is a better brand. It's terrible, but it's not, it's a better brand than chat GPT was. So talking about naming, yeah. Yeah, so Simon actually, so for those listeners, we're actually going to release Simon's talk at the AI Engineer Summit where he actually proposed, you know, a better name for the sort of junior developer or code developer.
Starting point is 00:27:10 Coding, I'm coding intern. Coding intern. Yeah, coding intern was it, yeah. But did you know, did you notice that advanced data analysis is dead, you know, 2023 to 2023, you know, a sales driven decision that has been rolled back effectively because now everything just called coding. Oh, that's, I hadn't noticed that I thought they'd split the brands. And they're saying advanced age analysis is the user-facing brand and code is the developer-facing brand. But now have they ditched that from the interface then? Yeah.
Starting point is 00:27:40 So it's unified mode, yeah. Yeah. So like in the unified mode, there's no selection anymore, right? You just get all tools at once. So there's no reason to differentiate this. But also in the pop-up, when you log in, when you log in, it just says code interpreter as well. So, yeah. And then also when you make a GPT, the drop down when you create your own GPT,
Starting point is 00:28:02 it just says, call interpreter. It also doesn't say it. You're right. Yeah, they ditch the brands. Good Lord. On the UI. That's amazing. Okay.
Starting point is 00:28:11 Well, you know, I think, so I may be one of the few people who listen to AI podcasts and also Sester podcast. And so I heard the full story from the opening ice head of sales about why it was named Advanced Data Analysis. I saw that. Yeah. Yeah. There's a bit of civil resistors, I think, from the engineers in the room.
Starting point is 00:28:30 It feels like the engineers won because we got code interpreter back, and I know for sure that some people were very happy with this specific thing. I'm just glad. For the past couple of months, I've been writing code interpreter parentheses, also known as advanced data analysis. And now I don't have to anymore, so that's great. Yeah, yeah, let's back. Yeah, I did want to talk a little bit about the GPT creation process, right?
Starting point is 00:28:51 I've been basically banging the drama a little bit about how AI is a better prompt engineer than you are. And sorry, am I speaking over assignment because I'm lagging. When you create a new GPT, this is really meant for low code, such no code builders. It's really, I guess, no code at all because when you create a new GPT, there's sort of like a creation chat, and then there's a preview chat, right? And the creation chat kind of guides you through the wizard of creating a logo for it, naming a thing, describing your GPT. getting custom instructions, adding conversation structure starters, and that's about it that you can do in a sort of creation menu. But I think that is way better than filling out a form.
Starting point is 00:29:30 Like, it's just kind of have a job to fill out a form rather than fill out the form directly. And I think that's really good. And then you can sort of preview that directly. I just thought this was very well done and a big improvement from the existing system, where if you tried all the other, I guess, chat systems, particularly the ones that are done independently by this storywriting crew, they just have you fill out these very long form. It's kind of like the match.com, you know, you're trying to simulate.
Starting point is 00:29:59 Now they've just replaced all of that, which is chat. And chat is a better prompt engineer than you are. So when I... I don't know about that. I'll drop this in, which is when I was creating a chat for my book, I just copied and selects it all from my website, pasted it into the chat, and it just did the prompts from chatbot for my book.
Starting point is 00:30:18 book, right? So, like, I don't have to structurally, I don't have to structure it. I can just dump info in it and it just does the thing. It fills in the form for you. Yeah, did that come through? Yes, now it does. Yeah, I built the first one of these things using the chatbot. Literally, on
Starting point is 00:30:37 the bar, on my phone, I built a working like bot. It was very impressive. And then the next three I built using the form, because once I I've done the chat bot once, so it's just, it's a system prompt, you turn on and off the different things, you upload some files, you give it a logo. So yeah, the chat bar, it got me on boarded, but it didn't stick with me as the way that I'm working with the system now that I understand how it all works. I understand, yeah, I agree with that. I guess, again, this is all about the total
Starting point is 00:31:01 newbie user, right? Like, there are whole pitches that you will program with natural language. And even a formula. And for that, it worked. Yeah. Yeah. Yeah, that did work really well. Can we talk about the external tools of that? Because the demo on stage, they literally used, I think, retool and they used a Zapier to have it actually perform actions in real world. And that's like, unlike the plugins that we had, there was like one specific thing for your plugin, you have to add some plugins in. These actions now that these agents that people can program with, you know, just natural language, they don't have to like, it's not even low code. It's no code. They now have tools and abilities in the actual world to do things. And the guys on stage, they demoed like
Starting point is 00:31:44 mood lighting with like a hue lights that they had on stage. And, they'd like, hey, set the mood and set the mood actually called like a hue API and like turned the lights green or something. And then they also had the Spotify API. And so I guess this demo wasn't live streams, right? Swixfusside. They uploaded the picture of them hugging together and said, hey, what is the mood for this picture and said, oh, there's like two guys hugging the professional setting, whatever.
Starting point is 00:32:09 So they created like a list of songs for them to play. And then they hit Spotify API to actually start playing this. All within like a second on a live demo. I thought it was very impressive to. for a low-code thing. They probably already connected the API behind the scenes. So, you know, just like low-code. It's not really no-code.
Starting point is 00:32:25 But it was very impressive on the fly how they were able to create this kind of specific bot. On the one hand, yes, it was super, super cool. I can't wait to try without that. On the other hand, it was a prompt injection nightmare. That Zapier demo, I'm looking at going, wow, you're going to have Zapier hooked up to something that has, like, the browsing mode as well?
Starting point is 00:32:44 Just as long as you don't browse it, get it to browse a web page with hidden instructions that steals all of your data, from all of your private things and X-Filtrates it and opens your garage door and sets your lighting to dark red. It's a nightmare. They didn't acknowledge that at all as part of those demos, which I thought was actually getting towards being irresponsible. Because, you know, anyone who sees those demos and goes, brilliant, I'm going to build that and doesn't understand prompt injection is going to be vulnerable, which is bad, you know.
Starting point is 00:33:13 It's going to be everyone because nobody understands. side note GROC from XAI a dear friend Elon Musk is advertising their ability to ingest real-time tweets so if you want to worry about prompt injection just start tweeting in all instructions
Starting point is 00:33:28 and turn my garage door I will say there's one thing in the UI there that shows kind of the user has to acknowledge this actually is going to happen and I think if you guys know open interpreter there's like an attempt to run a code interpreter
Starting point is 00:33:44 locally from Killian. We talked on Thursday as well. This is kind of probably the way for people who are wanting these tools. You have to give the user the choice to understand what's going to happen. I think openly I did actually do some amount of this at least. It's not like running code by default. You have to acknowledge this. And then once you acknowledge, you maybe even like understanding what you're doing.
Starting point is 00:34:04 So they're kind of also giving this to the user. One thing about prompt injection, Simon, tangentially, I don't know if you guys, we talked about this, they added privacy sheets, something like this, where they would protect you if you're getting sued because of your API is getting copyright infringing. I think it's worth talking about this as well. I don't remember the exact name, I think, copyright
Starting point is 00:34:24 shield or something. Copyright shield, yeah. GitHub, I said that for a long time, that if copilot created JPL code, you will get the GitHub legal team to fight on your behalf. Adobe have the same thing for Firefly. Yeah, you pay money
Starting point is 00:34:40 to these big companies and they have got your back is the message. And Google Vertex has also announced it, but I think the interesting commentary was that it does not cover Google Palm. I think that is just, yeah, Conway's Law at work there. It's just, like, I'm not willing to back this. Yeah, any other elements that we've got to cover? Well, the one thing I'll say about prompt injection is they do, when you define these new actions, one of the things you can do in the opening API specification for them is say that this is a consequential
Starting point is 00:35:14 action and if you mark it as consequential then that means it's going to prompt the use of confirmation before running it that was like the one nod towards security that I saw out of all the stuff they put out yesterday yeah I was going to say to me the main takeaway with jubtis is like the funnel of action starting to become clear so the switch to like the god model I think it's like signaling the chat dcd is now the place for like long tail not repetitive tasks you know if you had like a random thing you want to do that you've never done before, just go and chat GPT. And then the GPDs are like the long tail of repetitive tasks, you know? So like, yeah, startup questions. It's like, you might have a ton of them, you know, and you have some constraints, but like you never know
Starting point is 00:35:58 what the person is going to ask. So that's like the startup mentor and the Sam demo on stage. And then the assistance API, it's like once you go away from the long tail to the specific, you know, like have you build an API that does that and becomes to focus on both non-repetitive and repetitive and repetitive things, but it seems clear to me that, like, their UI-facing products are more phase on, like, the things that nobody wants to do in the enterprise, which is, like, I don't want to solve the very specific analysis, or like the very specific question about this thing that is never going to come up again, which I think is great. Again, it's great for founders that are working to build experiences that are, like, automating the long tail before you
Starting point is 00:36:39 even have to go to a chat. So I'm really curious to see the next six months of a Starbucks coming up you know I think you know the work you've done Simon to build the guardrails for a lot of these things over the last year now a lot of them come bundle with open AI and I think it's kind of be interesting to see what what founders come up with to actually use them in a way that is not chatting you know it's like more autonomous behavior for you interesting point here with the GPs that you can deploy them you can share them with a link obviously with your friend but also for enterprises you can deploy them like within the
Starting point is 00:37:11 enterprise as well and I'll see I think you bring a very interesting point where like previously you would document a thing that nobody wants to remember. Maybe after you leave the company, whatever, it would be documented like a Nasana or the conference somewhere. And now maybe there's a there's like a piece of view that's left in the form of GPT that's going to keep living there and be able to answer questions like intelligently about this. I think it's a very interesting shift in terms of like documentation, staying behind you, like a little piece of Alessio staying behind you, sorry for the balloons, to kind of document this one thing that like people don't want to remember, don't want to like, you know,
Starting point is 00:37:44 A very interesting point. Very interesting point. We're the first immortals. We're in the training data and then we will. You'll never get rid of us. If you had a preference for what lunch got catered, you know, it'll forever be in the lunch assistant in your company. One thing I find interesting about the shareable GPTs is there's this problem at the moment with API keys,
Starting point is 00:38:07 where if I build a cool little side project that uses the GPT4 API, I don't want to release that on the internet because then people can build. through my API credits. And so the thing I've always wanted is effectively AWARF against OpenAI. So somebody can sign in with OpenAI to my little side project, and now it's burning through their credits when they're using my tool. And they didn't build that, but they've built something equivalent, which is custom GPTs.
Starting point is 00:38:31 So right now I can build a cool thing, and I can tell people, here's the GPT link. And okay, they have to be paying $20 a month to OpenAI as a subscription, but now they can use my side project. And I didn't have to have my own API key and watch the budget and cut it for people using it too much and so on. That's really interesting. I think we're going to see a huge amount of GPT side projects because it now doesn't cost me anything
Starting point is 00:38:53 to give you access to the tool that I built. It's built to you. And that's all out of my hands now. And that's something I really wanted. So I'm quite excited to see how that ends up playing out. Yeah, excellent. I fully agree with all that. And just a couple mentions on the other multimodality things.
Starting point is 00:39:10 Text to speech and speech to text just dropped out of nowhere. Go for it, go for it. You sound like you have strong. Oh, I'm so thrilled about this. So I've been playing with chat GPT voice for the past month, right? The thing where you can, you literally stick an ear pod in, and it's like the movie her without the, without the cringy, cringy phone sex bits.
Starting point is 00:39:28 But yeah, like I walk my dog and have brainstorming conversations with chat GPT. And it's incredible, mainly because the voices are so good, like the quality of voice synthesis that they have for that thing. It's, it's, it really does change. It's got a sort of emotional depth. to it. It changes its tone based on the sentence that's reading to you. And they made the whole thing available
Starting point is 00:39:50 via an API now. And so that was the thing that, I built this thing last night, which is a little command line utility called Ospeak, which you can PIPP install, and then you can pipe stuff to it, and it'll speak it in one of those voices. And it is so much fun. And it's not like, another interesting
Starting point is 00:40:06 thing about it is I got it, so I got GPD4 Turbo to write a passionate speech about why you should care about Pelicans. That was the entire prompt. because I like Pelicans. And as usual, like, if you read the text that generates, it's AI generated text, like, yeah, whatever. But when you pipe it into one of these voices,
Starting point is 00:40:23 it's kind of meaningful. Like, it elevates the material. You listen to this dumb two-minute-long speech that I just got a language-mort generating. I'm like, wow, no, that's making some really good points about why we should care about pelicans. Obviously, I'm biased because I like pelicans, but oh my goodness, you know,
Starting point is 00:40:37 it's like, who knew that just getting it to talk out loud with that little bit of additional emotional sort of clarity would elevate the content to the point that it doesn't feel like just four paragraphs have junked at the model dumped out. It's amazing. I absolutely agree that getting this multimodality and hearing things with emotion, I think it's very emotional. One of the demos they did with a pirate GPT was incredible to me. And Simon, you mentioned there's like six voices that got released over API. There's actually seven voices. There's probably more, but like there's at least one voice that's like pirate voice.
Starting point is 00:41:09 We saw it on demo. It was really impressive. It was like, it was. like an actor acting out a role. I was like, what? Is this making no sense? Like it really, and then they said, yeah, this is a private voice that would not get to release, maybe we'll release it. But also being able to talk to it, I was really, that's a modality shit for me as well, Simon, like you, when I got a voice and I put it in my air pod, I was walking around in the real world just talking to it. It was incredible mind. It was actually like a FaceTime call with an AI. And now you're able to do this yourself because they also, open source Whisper 3. they mentioned it briefly on stage
Starting point is 00:41:43 and we're now giving a year and a few months after Whisper 2 was released which is still state of the art automatic speech recognition software. We're now getting Whisper 3 I haven't yet played around bench ones but they did open sources yesterday and now you can build those interfaces
Starting point is 00:41:59 that you talk to and they answer in very very natural voice all via open AI kind of stuff the very interesting thing to me is their mobile allows you to talk to it but you were sitting together and they They typed, most of the stuff on stage they type. I was like, why are they typing?
Starting point is 00:42:15 Why not just have an input? I think they just didn't integrate that functionality into their web UI. That's all, it's not a big complaint. So if anybody in OpenEI watches this, please add parking capabilities to the web as well, not only mobile, with all benefit from this, I think. I think we just need sort of pre-built components that assume these new modalities. You know, even the way that we program front ends, you know, and I have a long, history of in the front end world. We assume text because that's the primary
Starting point is 00:42:45 modality that we want. But I think now basically every input box needs, you know, an image field, needs a file upload field, and needs a voice fields and you need to offer the option of doing it on device or in the cloud for higher accuracy. So all these things are you because you can run whisper in the browser. Like it's it's about 150 megabyte download, but I've seen that I've used demos of Whispur running entirely in WebAssembly. It's so good. like these and these days 150 megabyte well I don't know I mean
Starting point is 00:43:17 React apps are leaving in that direction these days to be honest you know no honestly it's the the the stuff that the models that run in your browsers are getting super interesting I can run language models in my browser the whisper in my browser I've done image captioning things like it's getting really good and sure like 150 megabytes is big but it's not unachievably big
Starting point is 00:43:37 you get a modern MacBook Pro 100 on a fast internet connection 150 meg takes like 15 seconds to load and now you've got full whisk you've got high quality wispy you've got stable diffusion very luckily without having to install anything it's it's kind amazing i would also say i would also say the trend there is very clear those will get smaller and faster we saw this still whisper that came like six times as smaller and like five times as fast as well so that's coming for sure i got to wonder whisper three i haven't really checked it out whether not it's even smaller than Whisper 2 as well, because Open AI does tend to make things smaller. GPD Turbo, GPD4 Turbo is faster than GPD4 and cheaper.
Starting point is 00:44:16 Like, we're getting both. Remember the laws of scaling before where you get like either cheaper by like whatever in every 60 months or 18 months or faster. Now you get both cheaper and faster. So I kind of love this like new new law, scaling law that we're on. On the multimodality point, I want to actually like bring a very significant thing that I've been waiting for, which is G54, Vision is now available of the API. You literally can send
Starting point is 00:44:39 images and it will understand. So now you have like input multimodality on voice. Voice is getting accurate related to text. So we're not getting full voice multimodality. It doesn't understand for example that you're singing. It doesn't understand intonations. It doesn't understand anger. So it's not like full voice
Starting point is 00:44:55 multimodality. It's literally just when saying to text. So I could like, it's a half modality, right? Like it's eventually. But vision is a full new modality that we're getting. I think that's incredible. I already saw some was from folks on RoboFlow that do like webcam analysis, like live webcam analysis with with GPD4 Vision. That I think is going to be a significant upgrade for many developers in their toolbox to
Starting point is 00:45:18 start playing with this. I chat with several folks yesterday, Sam from new computer and some other folks, they're like, hey, Vision is really powerful, very, really powerful because it's, I've planned the open source models, they're good, like Lava and Buck Lava from folks from News Research and from Skunkworks. So all the open source stuff is really good as well. Now, nowhere near GDP4. I don't know what they did.
Starting point is 00:45:41 It's really, I'm kidding how good this is. I saw a demo on Twitter of somebody who took a football match and sliced it up into a frame every 10 seconds and fed that and then got back commentary on what was going on the game. Like good commentary. It was, it was astounding. Like, yeah, it turns out FFMPEG slice out a frame every 10 seconds. That's enough to analyze a video. I didn't expect that at all.
Starting point is 00:46:03 I was playing with this. Oh, I think Jim Fan from InVideo was also there. And he did some math where he sliced, if you slice up a frame per second from every single Harry Potter movie, it costs like $15, $45. Oh, it costs $180 for GPC4V to ingest all eight Harry Potter movies, one frame per second, and 360P resolution. So $180 to update everything is the price of for vision. Yeah.
Starting point is 00:46:31 And yeah, actually, at our hackathon last night, I skipped it. a lot of the party and I went straight to hackathon. We actually built a vision version of V0 where used vision to correct the differences in sort of the coding output. So V0 is the hot new thing from Vsale where it's drafts front ends for you, but it doesn't have vision. And I think using vision to correct your coding actually is very useful for front end. Not surprising. I actually also interviewed Div Garg from Multion. And I said I always, I've always made that vision would be the biggest thing possible for desktop agents and web agents, because then you don't have to parse the DOM.
Starting point is 00:47:08 You can just view the screen just like a human word. And he said it was not as useful, surprisingly, because he's had access for about a month now, for specifically the Vision API. And they really wanted him to push it. But apparently it wasn't as successful for some reason. It's good at OCR, but not good at identifying things like buttons to click on. And that's the one that he needs. Right.
Starting point is 00:47:30 I find it's very important. You need to go ahead. Click here. Because I asked for coordinates and I got coordinates back. I literally upload a picture and said, hey, give me a bounding box. And it gave me a bounding box. And also I remember the first demo. Maybe it went away from that first demo.
Starting point is 00:47:43 Which you remember the first demo. Brockman on stage uploaded the Discord screenshot. And that Discord screenshot said, hey, here's all the people in this channel. Here's the active channel. So it knew like to highlight the actual channel name as well. So I find it very interesting. You said this because I saw it understand UI very well. So I guess it will find.
Starting point is 00:48:01 out. Many people will start getting these tools. Yeah, there's multiple things going on, right? We never get the full capabilities that OpenEye has internally. Like Greg was likely using the most capable version, and what div got was the one that they want to ship to everyone else. The one that can probably scale as well, which probably lower, yeah. I've got a really basic question. How do you tokenize an image?
Starting point is 00:48:23 Like, presumably an image gets turned into integer tokens that get mixed in with text? How? Like, how does that even work? work. Yeah. There's a paper on this. It's only about two years old. So it's like it's still a relatively new technique. But effectively, it's convolution networks that are reimagined for the vision transform age. But what tokens are you, because the GPT4 token vocabulary is about 30,000 integers, right? Are we reusing some of those 30,000 integers to represent what the image is? Or is there another 30,000 inches that we don't see? Like, how do you even count token?
Starting point is 00:49:00 I want tick token but for images. I've been asking this and I don't think anybody gave me a good answer. Like how do we know the context links of a thing now that like images is also part of the of the prompt? How do you count? Like how does that? I never got an answer. So folks, let's stay on this and then let's give the audience an answer after like we find it out. I think it's very important for like developers to understand like how much money this is going to cost them and what's a context link.
Starting point is 00:49:28 Okay, 128K text tokens, but how many image tokens? And what do image tokens mean? Is that resolution-based? Is that, like, megabytes-based? Like, we need a framework to understand this ourselves as well. Yeah, I think Alessio might have to go. And Simon, I know you're busy at a good at you. I've got to go in-minute.
Starting point is 00:49:49 Yeah. So I just wanted to do some in-person takes, right? A lot of people, we're going to find out a lot more online as we go about our learning journeys with OpenEye. But just like, what was, you know, any interesting conversations from you say in-person observations. I'll volunteer in mind, which is Sam Altman, came out to the after party for the conference and just stood there in Japan. No bodyguard, just him for like a few hours. And it was just really impressive how much he, I guess, personally demonstrated that he cares about meeting developers. I really liked meeting everybody in the kind of the after party, whatever it was called, reception.
Starting point is 00:50:30 It was very like buttoned up in the Young Museum in San Francisco. It was really like well organized. Actually, probably not surprising, but I know that like the whole event was like extremely well organized. We talked about this a bit in the beginning. So this was my takeaway from all this. Folks got like a hundred dollars credit for a newber because like the party was not at the same place as the event where like it usually is. And to me personally like the music was too loud. I wanted to talk to people and not scream at people.
Starting point is 00:50:59 So like I always like this happens for some reason, but like I just wanted to like talk. Networking was really powerful. It was like a self-selected event. Many people didn't get in. Like I didn't get in until I met Logan and Logan thankfully invited me. Thank you, Logan. It was amazing. But it was like a very selected event.
Starting point is 00:51:17 So I actually met a few people who are working on some incredible things. I met somebody who was working on AI for education for special needs kids, for example. And he got invited to open AI directly because like he's working. in Italy for all these types of things. So actually, like, meeting the people who are working around the world was the biggest impact. There wasn't as many as I thought there would be. And a shout out to Open AI for this, but, like, please invite more of the world. I'll back that up.
Starting point is 00:51:47 Every conversation I had, just talking to a random person, they were doing something interesting. Like, they clearly did a very good job of funneling people who are actively hands-on building stuff into this event. That was really fun. I did actually want to, one thing I'll say, the venue itself for the main conference was a multi-story car park that had been converted into an event venue. I thought it was a great venue. I just thought it was hilarious that we were walking up ramps between floors because the best thing about multi-story car parks is that you can park cars on the roof. So the roof was where they set up
Starting point is 00:52:16 the lunch and they had a big tent up and stuff. And it was great. I hung out on the roof socializing. Yeah, but what a fascinating thing. Like a multi-story car park that's turned into a top-notch event venue. I've never seen one of those before. Alessio on the ground there with Newton, any founder conversations that you liked? It was, you know, I think the thing, you know, Tab is like an office here and they're doing one of the AI. Maybe you want to introduce Tab.
Starting point is 00:52:41 You know, they were recently, yeah, it's one of your personal companions that can chat with you in real time. And for example, Avi was using it for investor pitches. So he would get notifications on his phone during a pitch and be like, hey, you forgot to mention this and whatnot. And I know you might remember like there was the room over like Johnny I working with Open AI on a on a hardware project and I think like this GPDs announcement kind of make me think of you know, maybe they're building their own hardware assistant that you can load with a bunch of GPPs and you know Alex just mentioned how good it was to talk to one and maybe they want to go further down in that direction. I think that would be quite
Starting point is 00:53:20 quite interesting but yeah I think a lot of excitement and you know we just announced the the latest space launch red so we're on the side of of the builders. We don't think Copenhagen is going to do everything. Excited to see what people will come up with. Cool.
Starting point is 00:53:33 So I will stitch up this recording. I actually recorded a bunch of interviews on site with a bunch of other founders as well. So I'll put that at the end of this chat to get perspectives
Starting point is 00:53:42 from everyone. But thanks so much for jumping on with this quick call. Very, very exciting day. And I think we'll all be having a lot more takes
Starting point is 00:53:50 as we build with these APIs. I just want to say a quick round of thanks to everyone here. It's been awesome to like, experience these changes with all of you guys. It's a personal shout-out.
Starting point is 00:54:01 It's been crazy. It's been crazy, but also, like, the fact that, like, we were, like, the only space live from the actual event, and, like, we got joined by, like, 200 people in the audience. Yeah, we got officially sanctioned as podcasters. Yeah, it was funny. Yeah, but officially, like, the only two podcasters in the opening animals. Yeah, we got press passes.
Starting point is 00:54:21 We got press passes would have had an easier time, but, yeah. Maybe they would have led you with the whiteboard inside if we had the press, We made it happen. But yeah, that's another thing. Chatubit is not even one year old, right? Like, anniversary is November 30th. So we're 11 months and a few days in. And this is the craziness that it's been kind of matching what we'll be like in the year's time.
Starting point is 00:54:44 Yeah. And I think Sam Alvin mentioned this on stage as well. Like in the year's time, this will seem like trivial. But we get some very exciting announcement for today. So, you know. Honestly, I can't predict. I can't predict four weeks ahead the way things are going. It's fascinating. Cool. I probably should let you all go, but thank you so much for jumping on.
Starting point is 00:55:05 Thank you everyone. Thanks. This was really fun. All right. That was part one of this very long opening I Dev Day episode, but I promise you will be worth it because part two is some of my favorite work that I've done in audio form. So I basically carried a microphone around, and when I ran into someone that I wanted to interview, I just paused them and ask them for five minutes. And the first is someone that we haven't yet scheduled on the pod, but we've been extremely friendly with the Jim fan, everyone. Jim fan from the landmark Voyager paper and more recently the Eureka paper, but all of which comes out of his work at Nvidia and advising at Stanford. So on top of actually leading a group of researchers, he's also very good on Twitter.
Starting point is 00:55:50 And I think that is a very useful skill to have because you can communicate the value of your work to a wide audience. and that is something that we also aspire to do at Lean SpacePod. So, yeah, it's good to see you. We're going to see with you, Sean. Yeah, so great. I always wanted to get you on the podcast, and then like never got around to scheduling you in the studio, but since we're at the events, like, this is the big one. This is the best event to have the podcast in, so thanks for having me.
Starting point is 00:56:12 Yeah, yeah. And I also saw you've been tweeting some stuff. Like, what's the most interesting to you so far? I think a couple of things. Like, one is kind of the economy of scale. Yeah. Like how cheap the GP4 and GP3 APIs. have become, I think that's going to be a game changer.
Starting point is 00:56:27 So I just did a back-off envelope calculation. Like if you feed the entire Harry Potter books, like all seven books into GD4, it's going to cost only like $15 to read all of them. And $45 to write all of them. And that is just crazy. And you can have GPD4, right? It's going to be better than 3.5.
Starting point is 00:56:48 And the other thing is GPD4V API is also available. Yeah. And if you feed all of Harry Potter's like, you know, eight movies into it, that's going to be like 20 hours, frame by frame, you know, one frame per second, it's only going to cost $180 to watch all of these movies as 260P resolution, right? Yeah. So this economy of scale is crazy, and I think that's really hard for other companies to beat. Yeah.
Starting point is 00:57:13 Yeah. Is it a surprise to you, the rates at which they have been bringing down their pricing? I am not surprised. I think, you know, the pricing is going to follow some kind of exponential annealing from now on. It's just going to be exponentially cheaper. as compute becomes cheaper, as the economy of scale is going. So that's one thing. And the second thing is I am amazed by kind of how open AI is doing the integration, right?
Starting point is 00:57:34 If we look at the Assistant API, it basically has all of the things that open-add develop in a one-stop shop. So you have like code interpreter, you have, you know, stateful API, you have browsing, and it can integrate with, I suppose, all of the plugins on the open-air store. And then it can also switch between those, right? We have seen those demos. So yeah, the API, I think it's going to be way better and way more flexible. So that's the second thing.
Starting point is 00:58:00 And the third thing is the UGC platform, right? Now everyone can build their bots and share them, you know, share not just the prompt, but actually like entire behaviors, entire GPTs. That is a huge advancement. Yeah. Yeah, it's really fascinating. And I think one of the thing is that it's interesting. This is supposed to be a dev day.
Starting point is 00:58:17 Yeah. But actually, like, I think the first half was not a dev-focused thing. It was kind of low-code or no-code programming with natural language. something that they're all saying a lot. And it's something that you've been doing a lot as well. I've been following your work somewhat. Yes. I feel that it's going to be this new programming.
Starting point is 00:58:31 Well, we'll just use natural language. And I refine it through dialogues. And I think that is the most natural way to do programming in the future. And the GPD app store is showing us a glimpse of it. Like you talked about and then you can refine the pavia and the bot can ask you like clarification questions. Yeah. That is the way. Yeah.
Starting point is 00:58:48 That is the right way. Exactly. The GBT creation pain, you're no longer filling out a form. You know, question, answer, answer question, answer question. Oh, yeah. You're having a chat, and then it prompts for you on the other pane. Yeah. And I thought that was a much better way than filling out custom instructions, because you don't know what you want.
Starting point is 00:59:04 Exactly. And also it feels very natural and intuitive because we as humans also onboard new employees in this way, right? Like, we don't send them a form. We have a dialogue with them and we tell them this is the expected behavior, and they can ask follow up questions if there are details that are not clear. So it is like just the most natural way to program. So two more questions. One is, so they mentioned the word agents.
Starting point is 00:59:27 Sam said the word agents on stage, but here they're calling it GPTs. Do you see a big gap that they still need to fulfill to become a full agent? Or is this the new direction that we should think about? I think it is the beginning. So it's kind of hard to predict what agents people will build. And also how good the base models are. Because I feel that the agent's robustness and capabilities are ultimately bottlenecked by the underlying model. So GPD4 turbo looks like it's a bit fine-tuned towards the agent use case, right?
Starting point is 00:59:59 It can do better function calling. It can do better like tool switching. These things are critical to agents. So I'm pretty optimistic. But we'll see. We'll see kind of is there like an emergent behavior once you put a UGC platform out there. Yeah. You mentioned tool switching.
Starting point is 01:00:13 Actually, I was thinking when you said tool switching. Actually, they're also doing model switching. Oh, yeah. Which is new. Like they have some kind of internal model router or like their mix sure it's good enough. they just don't care. Yes. They got rid of the model selector
Starting point is 01:00:26 and now it's the God model that does everything. Yeah, and you can also do retrieval as opposed to retrieval also has an embedding API in it that's automatically down under the hood. So yeah, very exciting. Yeah, yeah, yeah.
Starting point is 01:00:35 Okay, and then the last bit is, a lot of your work is sort of reinforcement learning plus plus. Yeah. Zero gradients, reinforcement learning. What do you think, you know, and we just went to one of the closed door sessions
Starting point is 01:00:47 where they talked a little bit about how they received their feedback. What do you think they're doing well or like might be a, You speculate a little bit, like next step, if they were to take anything from your research interests. I'm also very excited by GPD4's fine-tuning API, right? Because the rest of the APIs we see today are no gradient APIs. You cannot really fine-tune them, but you can only prompt them in different ways. But the fine-tuning on top of GVD4 with a custom data may have completely new behaviors.
Starting point is 01:01:14 And it's also a new way to program. Just it's a bit more complicated. It's not programming by dialogue. It's programming by data, right? you bring a dataset and then you have a new GPD4. So I think, you know, this year's theme is customization, customized by system API, customized by dialog, customized by data. So I see this kind of trend going into the future.
Starting point is 01:01:35 Yeah. Looking forward to it. I think there will be a lot of work in this area. I'm excited to just go hack. I am very excited. I want to skip the after party, but like, there's so many people here in person, so it's great. Jim is actually such a curious person that he does something that a podcast guest rarely does, which is turn the mics around and ask me.
Starting point is 01:01:51 So here's part two. Yeah, Sean, tell us what are you most excited about? So I'm taking over the show, man. Of course, please, please. Me personally, I was actually not even expecting them to release most of these things today. Like, a lot of people were like, I don't think they have like the Dolly 3 API ready. I don't think they have like...
Starting point is 01:02:07 Oh yeah, they actually have everything ready. I don't think you're texting speech ready. Oh yeah. It speaks volumes that when Sam Altman announced the Whisper 3 model, no claps. It's the smallest news. But it is actually going to be huge. I would love to, you know, put my hands dirty on whisper. Yeah, so honestly, I'm just overwhelmed.
Starting point is 01:02:29 I know some of the team. I know they're working extremely hard. This is their sprints until to get everything on today. Yeah. So, I mean, I think that's very important one. That now just like, they just shipped everything. They just, even though they're, even though they're doing very well, they still push themselves extremely hard to be top of it.
Starting point is 01:02:47 And they're really earning their spot for developers and for the general of the general AI market. And I hope they take some holiday after today. Yeah. Yeah. Yeah. It's too much of updates. And then so the next interesting thing to me is that they are integrating, they're
Starting point is 01:03:01 Sherlocking a lot of the startup features. So there are a lot of startups that are built on providing rag for people. Yeah. A lot of startups that are built on like maybe building agents on top of GPT. Yeah. So this is the first time where, you know, I think it's pretty common in large platform companies like AWS Reinvents often does this as well. They call this a red wedding.
Starting point is 01:03:20 Like they invite all your. customers to the same room. And then they're like, all right, let's see who survives. You know, stats, that's step. Oh, my, my God. So that is the sort of memey, funny, jokey version of this. Yes. Yes. I don't, I mean, realistically, I'm sure Harrison and Jerry and all the other rag people,
Starting point is 01:03:35 they have some heads up about all this stuff going on. But I think because it's built in so easily into the playgrounds, into the API, into the chat EPC itself. And also the tools, all the integrations, right? You don't need a lot of tooling just to set up a simple chatbot with rag. Yeah. It's like, so for example, for my conference, we did a Summit AIBot. All right.
Starting point is 01:03:56 Where we did, where we set up a land chain stack, we integrated and put it in a widget on the website. Now you can set it up with no code inside of the playground and it just let people play with it. It's great, but it's also very scary for a startup because if that was your whole mode, you don't have that mode. I agree. Yeah. Yeah.
Starting point is 01:04:14 That's going to be a problem. So it's interesting that I can sort of easily build this in and obviously the stateful API is something I was considering building. And I roughly knew that like, this would be the next thing. This is on the critical path, so I don't build it. I agree. Yeah. But then the question is like, all right, what do startups do?
Starting point is 01:04:32 Yeah. I think maybe one thing I was missing from Sam was like, hey, this is the biggest gathering of all your ecosystem developers. They're afraid of you. You have given them no assurance as to like where do you think people should build. Okay. So because like opening up just wants to do everything. I think so, right? judging from today's trend,
Starting point is 01:04:51 they literally are doing everything. Yeah. Yeah. So I feel a little bit, I mean, it's fine. Everyone who's building with the eye today opted in to cutting edge. And sometimes you work on a cutting edge,
Starting point is 01:05:02 you bleed. Yeah. And that's okay. That's right. That's right. Yeah. But I do feel like there's a lot of tension between the startups that build an opening eye
Starting point is 01:05:11 and opening eye itself. Yeah. So that's my two says. Sounds great. It's great to see you. Yeah, good to see you. Thanks for jumping on. Thanks for having me.
Starting point is 01:05:18 Is it? And next week, catch up with the former guest, Raza Habib, back for his second time on the pod. Last time we talked about Human Loop and we recorded in London and there was a pretty popular episode. I love that you guys care about Foundation Model Ops as Raza puts it. So check out the Human Loop episode if you want, but also here's Raz's take on OpenEye Dev Day. Welcome back to the pod.
Starting point is 01:05:39 Here's the second appearance. It's always a pleasure. Nice to see you again, Sean. Good to see you as well. All right. Let's just get right into it. What was most interesting to you? I mean, the sheer density of announcements.
Starting point is 01:05:49 I actually, I came with high expectations and there was a lot of stuff I was hoping to see but I think they under-promised and over-delivered, which I thought was really good. I think seeing that they're having a second run at plugins and doing it right this time and having the GPT store and like really allowing people to do that, I thought that was really cool.
Starting point is 01:06:06 Product decisions around how you design and build the GPTs like the low-code builder for these chat agents. I thought that was really nicely done, that they have this conversational interface that elicits from maybe someone who's not very expert or how to do prompting and things like that. I thought it was really thoughtful. It fills out the form for you, right?
Starting point is 01:06:23 It's a very simple thing, right? Like, ultimately, it's just filling out the system prompt and filling out what abilities it should have. Yeah. But actually, despite its simplicity, I think it's very powerful, and I was impressed by that. So, yeah, a lot of really cool things. And then all the changes to the API I'm really excited about.
Starting point is 01:06:40 I have some questions. Like, I'm not uniformly positive about all of the new API things, but I'm sure they'll get there. Okay. Anything in particular? Do you want to touch on? Yes, I think things that I'm excited about with the new assistance API or the new APIs in general, like multi-modality is really cool,
Starting point is 01:06:56 longer context windows, really cool, cheaper, faster models. I think everyone's going to be super excited about that. JSON mode is like, it seems like a small feature, but actually so many people say this is a problem for them, so I think that's going to be great. So I maybe missed the importance of this. Isn't that the same as the function-calling API? It's related, but you might want to have it in context
Starting point is 01:07:16 where it's not strictly doing function calling. Right, okay. So a little bit more general. Typically, I'll just make up a function that isn't actually a real function that is JSON. People say that for complex things, sometimes it violates the valid JSON things. I think just making that more reliable.
Starting point is 01:07:33 Okay. Some stuff that I thought was, initially I was excited about and then as I've like chewed on it a bit more, I'm a little bit less clear. So one is this like ability to jump in a bunch of documents and have it do a rag for you. Yeah.
Starting point is 01:07:44 I think like... 20 documents, man. or something. Yeah, I think that like it's a cool feature, but it feels a bit gimmicky to me. Like it feels like for serious practical applications, it's going to be hard to get that to work. If you think about what a large enterprise needs for RAG, like, it's, you know, it's rarely sufficient that you could just jump in a bunch, dump in a bunch of dollars. It's usually permissioning. Yes.
Starting point is 01:08:05 As like which users can actually access which bits of data. Yes. There's so much control that I think most developers will want to have for serious applications that I think it's cool for the like GPTs in the low code version. I'm skeptical that it'll get that much use by serious developers. And I feel the threaded, stateful, like, assistance API is really awesome, but I would like more clarity over how it's doing the statekeeping. Like, what ends up in the context?
Starting point is 01:08:31 I think for that to be really popular, they need to make that transparent. Yeah, there's an API booth downstairs. I don't know if you've spoken to them, but they wouldn't answer any of these questions for now. Okay, of course. But, you know, obviously that really affects human movement. But this is, you know, this is commentary over what I think overall was a set of really exciting announcements. Yeah.
Starting point is 01:08:48 And the last time we talked also, you were talking about, we were talking about the multimodal APS. And now you have them. It's finally here. What happens now? As I said to when I spoke to you last time, right? Like, it's a relatively straightforward addition to the email loop product. Like, everything will continue to work, but now you'll also have images in and images out and audio in and audio out. It's kind of interesting, like, seeing, you know, the assistance playground for opening eye that they just really. some things like that.
Starting point is 01:09:14 Like, it feels like they're starting to get close to supporting all of these things, but not quite yet. Yeah, yeah, excellent. And then I think the last part is I saw Human Loop, actually probably not you, probably somebody else, but also talking about the fine-tuning. There was a price drop.
Starting point is 01:09:26 I don't know how much because there was just so many announcements. I imagine that's only good things for fine-tuning. Yeah, I mean, yeah, there's so many answers. I also missed the price drop, but I know from speaking to folks at Open Eye as well that they think a lot more people should be fine-tuning. Yeah, fine-tuning is going to have, like, huge importance in the future.
Starting point is 01:09:43 That's why they're building out the EY for it. So it's something they're investing in very deeply. And, yeah, I still view fine-tuning as like an optimization stuff. Yeah. I think of it as like the compilation you do, like once you have something that's working. Which is what they said in the LLL performance session just now. Okay, cool. Yeah.
Starting point is 01:09:59 I'm glad that my tips are aligned with opening hours. I think you're very aligned. You're often leading them in what they specifically, which I think is good. Yeah, whatever I used to be, Sean? What did you think? Oh, I've said this in a previous recording. But effectively, I also thought they would do much less than they did today. I think they underpomised and overdelivered exactly like you said.
Starting point is 01:10:21 And even things like text to speech, which happened to that. It's not just text to speech. It's really good text to speech. So I think I told you last time I did like a near year-long internship at Google, and I was working on the first mural TTS team. The Takritraun team there were amazing. So what did you get from their demo? I think I need to play with it more, but I was impressed by the quality.
Starting point is 01:10:42 Yeah. Like the quality of the prosody, the variation. I think they're only releasing six voices, but... Secret Seventh voice with the Pirates. The Secret Seventh Voice with the Pirates. And then I was chatting to Andrei just now. And he was saying that internally, like, they have voice cloning set up as well. So they can do it with something like 30 seconds of speech. I'm not sure that's public. It's not public? I don't know.
Starting point is 01:11:02 He didn't tell me it wasn't public. Okay. All right. All right. But maybe filter it out when you publish this. For what it's worth, I've been talking to a lot of people in and outside of Dev Day. and a lot of people have heard about the voice customization stuff. So it's not really going to get anyone in trouble, I don't think. So I just chose to love it in there. Whatever.
Starting point is 01:11:22 I mean, it exists elsewhere in other products. And I think it's fair playing to compete with other companies. I don't know if they're going to release it for obvious reasons, right? There's a lot of safety concerns about releasing that kind of product. And for what it's worth, someone else, I think Fixi AI did a comparison of the pricing. they are severely undercutting like PlayHT and some of the other text to speech companies as well
Starting point is 01:11:46 on the pricing. They're between three to ten times cheaper per second or something than the other existing TTS companies. Yeah, I think that's very interesting. I think in general, their promise to keep cutting prices and then following through is building a lot of confidence. People who weren't previously nervous about building on them. What's interesting, I
Starting point is 01:12:04 think, is that because they have such a large economy of scale and they continue to drive down prices, The option of like self-hosting a fine-tuned model, even for smaller models, starts to be like less obviously economical because of the like spin-up and spin-down costs. So unless you have the like volume of usage to justify having it on all the time, it actually starts to become cost competitive to use one of these third-party APIs rather than having even a smaller model.
Starting point is 01:12:31 Right, because it's serverless in a way. Yeah. So what can you give people an idea of what kind of volume that is? Are you talking about concurrent requests? So if you look at most of the people who will provide you in like a served model, if you look like a replicate or a mystic AI or something like this. Fireworks.
Starting point is 01:12:48 Fireworks. There's a few of these companies. They tend to actually charge by like compute hour or compute minute. And so if you're not like going to have it on all the time, then like the reason is dollars. The reason, yeah, you end up needing it on all the time though because there's like spin up spinet. Cold starts. And so if you if you don't actually have enough,
Starting point is 01:13:08 usage to justify having it on all the time, it starts to become cost competitive to just use open I am. Yeah. So what I'm trying to get to is it's just dollars though. Like it's, if it's like $5 an hour, yeah, whatever. Yeah, I agree. I agree. Depend to new use case. But yeah. Okay, got it. All right, cool. Well, thanks so much for jumping on. I know this is last minute, but it's just nice to see people. No, no, I always love chatting with you. Yeah. Yeah. Hopefully it'll be more of the future. Yeah, for sure. The next guest is going to be a new name to many people. He hasn't done many public appearances, but he is a force to be reckoned with on Twitter. His name is Suria Danturi,
Starting point is 01:13:41 and this is a story of somebody whose startup got killed by Sam Altman. So we're here with Suria. Hey. Hello, my name, Syria. You're new on the pod, but also we've been around each other in the tech circles for a little bit.
Starting point is 01:13:53 You're a very famous developer of vector databases and of plugins. Yeah. What did the sound of the plugins that you've done? Yeah, so I worked on a few plugins. I worked in like chat with PDF, chat with like video, chat with website, chat with like Git.
Starting point is 01:14:07 I made a lot of cool plugins. Making decent money too. Yeah. I mean, they give like better functionality to like the whole GPD4 interface. Initially I wanted to do my homework with them. So I might as well make a plugin for it. So yeah,
Starting point is 01:14:23 I mean, they give, there's a lot of cool functional. I made one with called chat with like instructions, which would allow you to save more, more custom instructions and use that when you're talking to get GPD4. But yeah,
Starting point is 01:14:35 I mean, they're making revenue. It's pretty thick for, you know, people paying in 85 different countries. It's like nuts. How many people are like, or how big the scope is? How many people can use it? I think you may have shown me this before, but there was a plug-in platform that you use for monetization?
Starting point is 01:14:51 No. No? Oh, you build your own. I built my own thing. Okay. I've seen someone do like a Firebase for, you know, I'd, yeah, I don't know. RIP. No, I mean, they're doing well, but like, I just don't want to, you know, pay a 10% tax and all that stuff.
Starting point is 01:15:05 Yeah, yeah. For sure. Obviously, you're very technically savvy. Okay, so what happened today? The announced GPTs. What's going on? Yeah, so I made a tweet this morning, being like, Sam, I want to kill my startup. And a joke, okay?
Starting point is 01:15:18 I just wanted to talk. I was like, I was trying to notify people while I'm here, and I just want to meet up. I made it the joke. And then a couple hours later, my friend, Matt, he works at Julius. He showed me the new UI. I'm like, okay, cool. And he was forced me to look at it on my phone. I'm like, okay, sure, I'll pull it up.
Starting point is 01:15:35 I pulled it up on my phone and plug-as were gone flaggeds were gone you don't you can't I think you can go between models so you can go between four and three but the whole options of like code interpreter and like
Starting point is 01:15:48 Dolly 3 and all those stuff all those good stuff were gone from the UI I think this only if this only applies for people who are here at the event yeah I think they give access or like the new UI to the people here and they also but yeah plugers were gone and I'm like oh shit
Starting point is 01:16:03 and I asked the person like hey like where Where can I, like where are the plugins? Like, where does it go? They basically told me like, you have to make a new GPD as a developer and you can import your schema into the new GPD and only that way can you, you know, kind of revitalize your plugin. But your existing users will be like...
Starting point is 01:16:24 No, I think they're gone. I mean, I haven't looked at my stats today. Well, I mean, this is not widely rolled out yet, but when it rolls out... Sure. When it rolls out, I'm pretty sure all of the plugins... They have to discover you again. Yeah, they're kind of dead. I mean, there's like no way,
Starting point is 01:16:37 I don't think there's a way to link them. Yeah. Like, there's like no way for the users who were using it previously to be using the new thing at all. But I mean, it's like, a side project for me, it's like not like a full time thing for me.
Starting point is 01:16:48 It's a fun project to do and like, it's like a nice, nice thing to work on. So I'm really bullish on, you know, the old new GPDs thing. I think they're a better abstraction. Yeah, I think GPDs are, I was starting a few open engineers and I was like agreeing with them
Starting point is 01:17:01 because like, I think GPDs are a much better abstraction on what plugins with hobos to be. I think plugins kind of died on arrival. Well, Samson said they did not have PMF, right? Yeah, he said that a long, he said that, he said that like one plugin started. Yeah. So it's like pretty much.
Starting point is 01:17:16 But yeah, I think GPs are a better abstraction. And I also love their doing revenue share. So revenue share is also a good thing. Because like Jeep plugins were like a really weird way of monetizing. You do a bunch of finicky stuff. But yeah, I mean, also like just by the way for people who don't know, Poe, you know Poe, right? Yeah.
Starting point is 01:17:36 Poe did this a long time ago. They did this a couple months ago. They have these bots. They call bots. And you can, you know, make your own, like Poembot, or you can make your own, like, S-A-Bot or whatever. And then the bots have custom instructions, and also they use a very specific model that the developer specifies. And you can install these bots, or you can chat with these bot, and the bot will do whatever the developer made them to do. So I think they're just basically open-edged, made the same thing.
Starting point is 01:18:04 and they brought it over to them. But, yeah, but effectively, plug into crime dead. Oh, RIP. Yeah, I mean, RIP, but it was a fun part. I mean, it's fun. I think GP, honestly, it's good that plugins died because, like, they had a bunch of issues. So one of the issues is that you can't share them.
Starting point is 01:18:21 You can't share a link to them. GPTs, you could share a link to them. So, like, I can share my link to my GPD thing to you. So it's much better for discoverability. Because previously, the only way to discover a plugin was through the plugin store. you just search for it, you do a bunch of stuff, and it wasn't very good in that aspect.
Starting point is 01:18:38 But sharing a link to them, having revenue share, and you can also give custom instructions, custom context. So they also came out with retrieval or whatever, and that can basically give you a custom, a vector database directly in your GPT, I think. So that's all great, all good features that should have came with plugins, probably. Yeah, awesome. And then lastly, just like any of the new stuff that was launched today,
Starting point is 01:19:02 What interests you in sort of building with them? Like if you were to build on the new API or a new GPT? Yeah, totally. I have some ideas. The thing is like, this is really weird to say. But like some of my ideas that I've said before for plugins, they kind of get copied quickly. So you want to keep it to yourself?
Starting point is 01:19:24 Yeah, that's fine. Yeah. But that's one part of it. The second part of it, I don't have any good ideas regarding what you can do with all the new functionality. like that's like a that's like a good product I don't know honestly there's a text speech came out the their internal vector DB thing came out internal vector DB thing or like retrieval or whatever it's called okay yeah people
Starting point is 01:19:43 have been saying they have an internal vector DB thing but yeah it's just a retrieval yeah it's like zero non-configurable right like it's it's gonna be for a simple use cases it's fine and then after a while you're gonna need one to control over chunking yeah I'm also excited by once we need a contact window I was a big user of cloud for a while because Claude, they basically gave you 100K contacts window directly on the UI, and you can upload your PDFs to it, and everything would work very well. Yeah. But, and then Cloud had some issues regarding, I mean, actually very recently,
Starting point is 01:20:13 Claude came out with this whole bullshit thing, bullshit copywriting thing. Copywriting thing? Yeah, yeah, it's really weird. So if you upload a PDF now out of Claude, like just this week, they made this weird tweak where it doesn't answer any questions, because if there's a copyright symbol or a copyright name, anywhere. It just like blocks you out.
Starting point is 01:20:32 It's like, what? Apparently you can prompt inject that by insisting that you are the author and then it just overrides it. Oh, really? It's like, don't worry, I got this.
Starting point is 01:20:38 I'm the author of this. There's no copyright issue. Anyway, so thanks. This is a really good story and I wanted to people to share it. And excited for what you work on to become more public. Yeah, thanks works. All right.
Starting point is 01:20:53 So that's what happened to chaty PT plugins, which we covered back in March. But don't worry, that's not the full story. His startup is not. not fully dead. We actually cover what happens later on. I just wanted to capture the confusion that was happening at Dev Day. So he referred to Julius and we'll actually talk in and check in with Rahul later on in this episode. But first, we have to go to our next guest. When Open AI launched
Starting point is 01:21:16 with GPs and the Assistance API, one of the lead launch partners that they launched with was Zapier. And I managed to catch up with Reed Robinson, who is lead AIPM at Zapier, to talk about it. All right. All right. Oh, Reed. Nice to meet you. It's really great to run into you as we're leaving. So you guys had a big sort of partnership launch on stage. Yes, yeah. We launched AI actions for GPTs, which we're really excited to see out there. We also today launched an update to our chat GPT integration that supports the assistance API functionality that was announced.
Starting point is 01:21:50 And you were one of the earliest to go. In my mind, Zaprio was very, very early in the natural language actions. NLA, I don't remember. Good memory. Yeah. Yeah. Yeah, we launched our natural language action. So we were a launch partner for Chituity plugins.
Starting point is 01:22:02 Yeah. And that's when we launched our natural limit actions API. And actually the AI actions that we're calling it today, kind of a rebranding that side of thing to really focus on that functionality. Yeah. And I just interviewed Asuria, who is a pretty prominent plugins developer. Plugins are dead. You know, reborn.
Starting point is 01:22:20 Yeah. It's going to be interesting to see what happens. There's clearly a difference. I think one of the things I talk about is the fact that, you know, with TPTs, you're able to constrain the prompts quite a bit. our plug in for chat to BT, the initial one, you needed to give it access to every single action you ever wanted to have access to, which meant that the kind of content, I don't know, like, yeah, that's going to be an issue. Yeah. The common one I give is, like, you know, if you had
Starting point is 01:22:43 given it Gmail and Google Calendar and asked it like, hey, what's going on next week on my, like, agenda, it would sometimes search Gmail because it'd be like, yeah, events are in Gmail or, like, you know, calendar invites are going to go to Gmail, so I should search there. But now you can, you know, define what apps it should use. You can define, like, how it's, you know, should use those. So some really fun use cases. I mean, honestly, we've been hustling hard to get this out there. I'm really excited to see what people actually build with this. Right. And what gets released there. Yeah, we'll be monitoring, trying to listen to people really closely. And so, like, something that's interesting about Zapier is that you are a collection of actions in
Starting point is 01:23:20 and off yourself. Yeah. So there's kind of multiple layers in which to do to do this. Like, what should exist at the GPT layer? What should exist at the Zapier layer? Yeah. Well, what's nice? I mean, it's a good point. We have about 6,000, like, apps on the platform today. Really, what the A.A. actions is, is it the ability to use any of those searches and actions using kind of natural language inputs. That would be, like, the instruction that the model gives it. So it's like, you know, check this user's calendar for Monday. And, you know, it might even give the, you know, the actual date for Monday, right?
Starting point is 01:23:52 And Zabier on our side will take that natural language request and process that into an actual API, like the actual API call to a tool like a calendar. And then we all run the response. So, you know, you can't just take the entire response of a, especially like Gmail, Google Google Cloud Theory. Yeah, responses are very, very, very, very, very long and very confusing. And so we actually do a lot of work to kind of, if you will, like, massage that data so that it makes sense for an LLM on the other side that it is giving it the right information it needs and not just like the entire payload.
Starting point is 01:24:24 Right. So it really helps it just kind of deliver like a more, again, more like contain, more refined experience for leveraging integrations alongside, you know, like chat chit. So existing Zaps cannot be ported one for one over to LLM ZAPs. It's really one-off actions is the better way to think about it. And you can, you know, as you saw in today's demo, you know, was using Google calendar for the search and a Slack action. You can actually chain those together.
Starting point is 01:24:49 And so, you know, how much is that as like a one-off action versus an actual like all a sudden a zap? But in this case, it's almost more like the trigger is the human in chatypti. right? Like, you need to trigger it to run for that. But on the flip side, you know, the assistance API is extremely exciting for me as well because you look at now like the, that functionality of building a GPT. You still getting used to the name. Allows you to kind of port that over to run asynchronously. So a common one, the two examples that I love giving for that API that I love in Zapier is number one, like data export. You know, think of every tool out there like looker, mix panel, amplistically. to all so many tools are able to send these like massive exports of CSV data on a regular basis.
Starting point is 01:25:34 Like you could say, hey, every Friday, export my blog traffic content or see it as a CSV, right? Normally, someone's going to get that CSV and have no clue what they're doing, right? But now you can actually create an assistant in Zapier and you can give it instructions to say, like, hey, tell me the top 10 performing blog articles in the last week. And also, you know, tell me highlights on, you know, maybe keywords that were used or SEO tags that were used and how that impacted conversions, right? You can be pretty detailed depending on what you're providing it.
Starting point is 01:26:03 And that can now run asynchronously. That can run automatically. So every Friday, you know, 8 a.m. You could be getting the export of that data. It's going to go to an assistant. That assistant's going to reply with even charts and graphs. And those will come through
Starting point is 01:26:14 and you can then send it to Slack. And so you can have every Friday a post in your team's, you know, the blog team's Slack performance. Yeah. And that'll run automatically. And then they can even reply in Slack to that post and have a content.
Starting point is 01:26:28 continuous conversation with that assistant. Oh, my God. So it's like really everywhere. Yeah. So you really put them everywhere. And that's one of the things I like about what's released. And I think people are going to continue to learn really just how kind of wild that is. It's the fact that you can use your actions in the UI of Type TBT and a one-off action,
Starting point is 01:26:47 but you can also run these things extremely well asynchronously. And yeah, like OpenAI releasing API support for the vision model and for code interpreter and retrieval. that these assistants can use is really cool. Yeah. Is there a Zapier angle to any of that? That's what I did. They're all the same. You're doing Zappir, right?
Starting point is 01:27:06 They're all the kind. You can, the whole like creating of an assistant and running that through an assistant is today support. You could do that literally right now. Yeah. So it's really cool. And the other one is like retrieval, right? I talk about, you know, you could go in, create an assistant.
Starting point is 01:27:18 Give it, let's say, you know, I talk about our accounting team a lot, right? You could give it, like if you have a team that approves budget requests from your company, right? Everyone does, right? They can actually have. have, take their Slack channel, or create an assistant first that would have the documents of your policies of like, hey, here's what you can expense, here's how you can expense, here's eligible, right, all these sorts of things. And actually then set up something like, again, I'll pick on Slack.
Starting point is 01:27:42 It's just easy. It's like a new message in your accounting budget request channel, right? And have it trigger a, the assistant and send the user's request to the assistant with all of your documentation with retrieval. And now it'll try to understand what your policies are. what everything is and check the information against what that. And you can even like I did one internally where we have a tool called, I think it's called Stacker that tracks each employee as like software budget and home office setup budget, right? You could see how much they've spent of their budget.
Starting point is 01:28:12 You could actually include that data in the context of the user message so that the model will be able to say like, hey, I see you want to expense this webcam. It's actually over the recommended budget. But you personally do have budget left if you wanted to use it for that, right? And so autonomy there. Yeah. And that's really cool.
Starting point is 01:28:30 Yeah. So you can start to do all of those sorts of things now in ZAPs that really were never possible. Yeah. So yeah, the querying of knowledge, running of data analysis, writing code even. I think in a very real way, you are the perfect partner to Open Eye because they've sort of built a reasoning sort of glue between all these things. It's definitely been a good and fun partnership. I think, yeah, the big thing for me that I always say is like I'm really, really excited now to see what he's. people do it. That's how we can improve it.
Starting point is 01:28:58 Yeah, awesome. Is there anything, you know, you've been developing with these APIs for a while? Is there anything that you caution people not to get too excited about? Like, what, yeah. I mean, call-outs that I'll always make is like double-check accuracy, right? Like, you want to call out like, how accurate. So make sure that information is accurate. Make sure you're putting some human in the loop steps before you're putting this into like a critical.
Starting point is 01:29:19 Which they show and like confirmed and I. Yeah, that sort of thing. But even, yeah, all sorts of things. You really want to make sure that you're comfortable with like, what can go wrong, what is likely to go right, right? Like all those sorts of constraints. The other side that I often talk about is just like keep an eye on, you know, if you have free form human input somewhere in your application that is triggering these things,
Starting point is 01:29:38 now that can something to get right. Yeah, prompt injections. Those are a real thing. And I think, you know, a lot of people are still trying to do really what that means and how bad that can be. Yeah. And so I always try to caution people about that as well, right? Like you really want to be realistic on kind of how far reaching you're doing this. So, yeah, that's why I like the internal use cases, that, you know, like things like that is a great way to start.
Starting point is 01:30:01 Yeah. To get familiar with the technology, to get familiar with the constraints for that. Other than that, no, I mean, the voice model stuff, I'm really excited to try that. I really want to think, yeah, that'll be really good. I love this secret pirate mode that they demoed. I don't know if you caught that session. I didn't see that session. So they obviously they have six voices, but there's a secret seventh mode if you add in the prompts to speak like a pirate.
Starting point is 01:30:24 Love it. I love it. That was an old, I don't know if you remember Facebook way back in the day, had that if one of the languages you could select. Yes. Yeah, yes. That reminds me of that. Yeah. Yeah. Lots of fun to be had with me as well. Okay. Well, thanks so much for jumping on. I know it's very random, but also, yeah, people love to hear from builders. So that's awesome. I love hearing from builders. And most of the interviews were done as we were sort of leaving the Devday venue and going to the after party. And I caught Divgarg of Maltion, who we've been talking around and circling around a possible episode on, he's definitely one of the leading
Starting point is 01:30:59 voices and thought leaders on agents because he's building a browser agent that's a very prominent one. Unfortunately, I have to take an L on this one because the audio is not great. Div's mic wasn't working and I don't know what happened to it. I try to always check these things, but you're only going to hear the output from my mic, which is slightly worse, but I opted to leave it in because div is actually building an agent with opening eye stuff. and had access to GPT4 Vision. And I think that people building with GPT4 Vision will be surprised at his answer to me
Starting point is 01:31:31 on whether or not it's useful for agents. You're to meet everyone. I'm this founder of Multion, which is an AIVab agent that can automate browsing for you. So we can book a flight, order stuff on Amazon, order dinner, whatever you can imagine.
Starting point is 01:31:43 Yeah, and I was actually reflecting, so everyone who has this is today already knows what was announced. I was actually reflecting that they didn't have any browser-based actions. So what were your thoughts on just generally their approach to agents? So it'll be very interesting because I feel like browser actions are just so risky. So things can go wrong. So if a company or your Open Air, you won't want to build up.
Starting point is 01:32:02 And they're better of just relying on a third party to who wants to own that. And that's also the strategy we are taking with them. They're like open air and launched like a ZAPID integration for APIs. But we want multi-empt to be like the no API solution. Like I want to do things beyond APIs. I want to connect to my personal accounts where I just have my logins already or I already have the cookies and I once want to multi-e want to go and like interact with my personal accounts or personal data. very easily. And I think this is very fascinating for us where we can like launch a multi-on
Starting point is 01:32:28 integration with their new platform and then you can just go and like give it a command like, oh, like can you book this platform me on chat jibbd? And then to launch a browser and the browser you can see what's happening and then we'll do the whole thing for you. And it'll be all seamless. And then people can have a lot of fun. Just like trying out all these different capabilities and like automating their like daily hook flows. You can like save this custom integrations. So for different agents, you can have different custom like multi-on prompts that are that are already pre-saved, and then you were like, oh, I want to now go order something on like DoorDash.
Starting point is 01:32:58 I want to order my favorite burger. Then, like, chat you can go and like, suggest you order your favorite burger style. And then it's like, now order this for me, multi-on and multi-goes and like, say it does that and vice. So we saw the payment for you, we saw identity for you.
Starting point is 01:33:09 And like, we are owning all the risky, like, actions that can pay. So you're going to build a GPT version of Multian? Yeah, we'll have a multi-on GPT. Okay. Will that be like a replacement to your existing thing or just like an alternative way to use your same API
Starting point is 01:33:23 or something like that. So the direction we're going for is we want to make our AI agent embedable within existing applications. So our launching an API. And we already have a chat chad chaddbid plugin. And so this will be like sort of like I will like use the API to power this sort of like new chdbddd experience. So for us we actually don't have to like change anything.
Starting point is 01:33:41 It'll be like very streamlined just integrate our API into chat chivity and like if we can start using it. Yeah. Yeah. Awesome. What about the, I guess the vision API? I think one of the things that have always constrained browser agents is the DOM. which is very heavy.
Starting point is 01:33:54 So the alternative approach is to use Vision. Would you explore that? What are your thoughts? So for us, we actually had early access to the Vision API for more than a month.
Starting point is 01:34:02 We tried it on a bunch of websites. Maybe like 5% of the websites is actually really useful, which are more like image-heavy because 95% of the websites even if you do OCR, that's good enough. Yeah, it's not in the dataset. We have really good like parsing.
Starting point is 01:34:13 So most websites, we can compress less than 3K tokens. So we don't really have to like worry about how heavy the taxes. So we have one interesting use case about the Vision API. We had a user who got it up on Tinder. And then like the, then like Miltian.
Starting point is 01:34:27 Hot or not? Oh, okay. Left and right. And the user actually got a match. Yes. I think you have found the killer use case for Maltion. Yeah. Like this. You did it in Lacham, right? Yeah. Oh, my God. Okay. Interesting. Okay. But so, but only image heavy sites. That's surprising to me. because you know the original Vision demo they actually showed a screenshot of Discord and they have perfect OCR. It's true.
Starting point is 01:34:59 It should be good for you. It can be very interesting. But I think it's like even without mission we can just do like so much things. So like adding vision maybe like helps a bit but not it's not like really game changing for us. That's surprising. Okay.
Starting point is 01:35:13 Well good to know. Anything else that you would highlight from today? I'm just like really excited about like open air trying to become like a marketplace. Yes. an app store. Yes. So if this can take off,
Starting point is 01:35:23 they could potentially kill like Apple App Store and become like the new thing there. And it's really hard to say like how things will go. But they tried this with plugins before but this might actually work this way. But we're just really just in listening to see like how
Starting point is 01:35:35 two years from now, how a lot of the developing might like how the world looks like. Yeah. I'm very excited about like two years from now. Like everything will be so different. We might not even use computers or even like mobile phones. You just have assistant. You talk to it and the assistant goes and does everything.
Starting point is 01:35:48 Yeah. It would be a fascinating world. Yeah. So one last question before you go. You have a nice side gig teaching at Stanford. Well, you were a PhD student and you put on top. But you're still teaching or curating Transformers United. Yeah.
Starting point is 01:36:01 So I draw it out from the PhD, but I'm still a lecture at Stanford. Yeah. Okay. So like, what paper should people read to like catch up on this? Like what is like top of mind in terms of like research that is informing what we're seeing? Yeah. Definitely very, it's a good question. Just things are moving so fast and there's like hundreds of research papers coming out,
Starting point is 01:36:19 like literally like a few days. I'm really excited about like the development that are happening at like meta so a lot of this work is open source on the Lama step and all the Mistral stuff I feel like that's very interesting on the transformer side Do you believe sliding window attention was the key for Mistral? I feel so for them
Starting point is 01:36:34 but I feel like there might be other ways to do that. There's some secrets right There was probably some secret yeah okay well that's all the time we have but thank you so much thanks a lot thanks okay and our next guest is Louis Nightweb CEO and co-founder of Blupe AI and organizer of the AI meetups in London where he is a very
Starting point is 01:36:50 prominent and staunch member, unlike Raza, who has defected to San Francisco since our last conversation. Louis always has very interesting takes in person, and it was a pleasure to finally actually get him to come on the pod, but also we recorded this while inside of a Waymo on the way to our after party. So Louis, you are new to the pod, but we've been friends for a while. Maybe explain that, maybe introduce yourself and how you come to the world of AI. Yeah, I guess. So we started Bloop, me and my co-founder, three years ago in a very different era for machine learning. And we both started the company because we wanted to help engineers navigate large codebases in a much better way.
Starting point is 01:37:35 And originally that was training our own models to do natural language to code search. And today, we still do that, but obviously those language models are very small compared to The state of the art. Yes. And so they're just one part of a much bigger pipeline. I see you as a very astute technologist. You used to be a VC. You wrote the first check into Human Loop.
Starting point is 01:37:58 And you used to share an office with Human Loop to the point that I called it Human Bloop. Yes. I think you liked that. Yeah, I did. Yeah, that is good. We're considering renaming. And you also run AI Tinkers in London. I do.
Starting point is 01:38:12 Yeah. London has a kind of a slightly different mix of talent. than say San Francisco. You've got a lot of agencies, a lot of enterprises. And so, yeah, we just felt a need to start like a very startup focused event. And that's why we created AI Tinker at London. Yeah, I think Alex Gravely would be very happy to hear about other stuff that you've been doing. And I've been to one of them. And it's really good work. I might be the only one that's been to both. Yeah. I've been to both as well. Okay. So let's fast forward to today. A whole bunch of things was announced. What's top of mind for you?
Starting point is 01:38:45 Yeah, so I think like context length is something that that we spend a lot of time evaluating whenever something new drops. All of the kind of standard evals, you know, the kind of literacy tests, things like that, they generally don't do a good job of measuring whether a model can actually use the context length that it claims it has. Yeah, context utilization is what I saw will depute today call it. Exactly. So this basically started maybe five months ago over the summer when Claude II dropped. And, you know, obviously I had 100K context and we were really excited about that. So we ran an experiment to see basically if we hid 10 pieces of information in the prompt and we increased the size of the prompt, you know, so you do it at 1,000 tokens, 4,000, 8,000,
Starting point is 01:39:35 et cetera, up to 100,000, how many of the original 10 pieces of information can it retreat? and we essentially found that the accuracy drops off a cliff between one and 10,000 tokens. And we repeated the same experiment with GPT4, and we found similar results that 32K GPT4 can only find one of the 10 pieces of information. But if you are only using 1,000 tokens, it can find nine of the pieces of information. So what that tells us is that context utilization five months ago was not great with all of the state-of-the-art models. So with the announcement of 128K today, and... That's the first test you'll run? That's the first test I'll run.
Starting point is 01:40:17 Having spoken to a couple of the team members who do Eval today from OpenAI, they're pretty confident that the model's got better ability to answer questions at those context lengths. So it's time to measure. Time to measure. Any other of the API features, reproducibility, does that matter to you? I think, to me personally, no. I kind of like the creativity. I normally have my models at like, you know, temperature.
Starting point is 01:40:44 Yeah, exactly, and a bit of temperature. Yeah. But I know lots of people on the blue team, he'll be very happy, I'm sure. And then I guess the JSON features, the, there's so many, like the multimodal features, any of that appeal for you personally. Jason is definitely a big one. I think it allows you to kind of standardize how you call different models. Yeah.
Starting point is 01:41:07 So instead of having to build, you know, the, and it's quite a massive thing to build. but to build the kind of function-calling integration, and then if you want to try Anthropic, you've got to go and have a completely different way of interpreting the output. So if you can just stick with JSON across all of your different LLM providers, open-source models included, that's definitely advantage.
Starting point is 01:41:26 Which allows you to evaluate different models more easily. Yeah, very excited about that. You are, so you compete in a pretty competitive space with the code assistance, code search, code systems, right? We do. There's source graph, there's codium, there's other codium, as co-pilot and so on. You've never ventured into the agent side of things.
Starting point is 01:41:48 Yeah. Is that conscious strategy? Are you waiting for the right time, waiting for the right APIs? I think, I mean, we're seeing traction at the moment with companies that have very large code bases, right? And it's not something we hear from those users that, you know, when we listen to their problems,
Starting point is 01:42:05 it hasn't been like an obvious fit to try and build like maybe an auto-GPT type of agent. I'd still say, you know, we're very interested in agents. The pipeline we have at the moment, it's basically GPT in a big while loop with function calling, which, you know, like nine months ago definitely did count as an agent, maybe less so now. So, you know, it's just customer and problem-driven, and we don't, you know, it's not a, it's not a hammer for the nails that we've got.
Starting point is 01:42:34 Yeah. So two comments on that. One, I think opening eye has sort of put there a flag a little bit in the definition of agents. They had three things, right? They had custom knowledge. They had custom instructions, and then I forget the third one, custom tools, let's just say. So by that definition, we're doing, yeah, so we've been doing that since about February. That's the definition. Then the second observation I'll say is you talk to developers, but what if the target customer for agents is not developers? It's the PMs, right? So we definitely see a lot of PMs using
Starting point is 01:43:07 using the product or people that I define as like reading more code than they write. So, you know, could be designers trying to understand the implications of an interaction. Could be PMs trying to fact check a contentious time estimate from a developer or something like that. Low trust environment there. Talking from, I've seen some stuff. Egregious things, yes. Yeah. So basically it's still not that appealing for you, but you'll keep.
Starting point is 01:43:37 look out for it? I think based on the definition of Open AI, you know, released today, we tick all the boxes. And I think we were one of the earliest adopters of that, if that's the definition. You just don't brand yourself with the agents. I don't think it's important to users.
Starting point is 01:43:55 I don't think that's why people use the product. I mean, we're very solutions focused. I think a lot of our branding at the start of the year was about models. And, you know, we put GPT4, GP3 right there on. the front page and now, you know, we've kind of reoriented to be more about solutions. I think that that reflects kind of maturity of the ICP we're going after and where we are with sort of stage of company life. Yeah, yeah. Cool. Any other things that you personally, you personally, not blooper related, are just excited by, interested by, from today. Any interesting conversations
Starting point is 01:44:31 with others? Loads are really interesting ones. I had a fascinating talk with some safety researchers who... They were here? They... So there's a couple of people who were kind of PhD students who had kind of looked at adversarial attacks through fine-tuning of models and found that basically, like, it's such a hard problem to solve. If you enable fine-tuning, it's basically impossible or very difficult to make it so that you can't disable all the safety features. You can just train it to spit out all sorts of stuff. So that was pretty fascinating.
Starting point is 01:45:10 I'm being excited about the Waymo we're in right now. Oh, yes. So we should tell people we're recording in a Waymo. We haven't been looking at the road the whole time. Is this your first Waymo? It is my first Waymo, actually, yes. It's my first Waymo, too. Thank you for taking my Waymo virginity.
Starting point is 01:45:23 I've got to experience this together. I've been a cruise stand the whole time until they ran over someone. So my take on cruise, like, at sample size 10 cruise journeys before they got shut down, and three of them resulted in something popping up on the screen saying that I had been in a collision. Did they use the word collision? Yeah, yeah, yeah. That's surprising. I'll shoot.
Starting point is 01:45:47 After that, I got pictures of it. I take a fair amount of cruises and it didn't, yeah. And so it was the same situation almost every time, which was a car was in front trying to park. And I think they just maybe bump fenders or maybe the crash detection. Oh, there was actual contact. I think in one of the cases I think there was. In the other two, I didn't feel anything. But it came up saying, like, you've been in a collision and somebody comes over the intercom.
Starting point is 01:46:06 checks like that. So yeah, I mean, but out of a, you know, 10, 10 rides and three of them ended like that. So I think, yeah, definitely some questions there. But this way moves pretty smooth. Maybe also we're in a better neighborhood for driving because we're going to the Golden Gate. The time of day, that was, that's a really good point. I noticed that all of the ones I took at night, all of the cruises I took at night were fine. And when I took one during rush hour, it was a completely different experience because the routes it would take it had this really aggressive.
Starting point is 01:46:36 maybe traffic management, something that was going on, so take a long time to get from A to B. Yeah. Yeah. Yeah. It often puzzles me slash interest me that self-driving is almost solved. You know, we still have some bumps in the road. Sometimes the bumps are human. It's solved in San Francisco where you've got wide open roads, nobody cycles, and...
Starting point is 01:46:57 That's not true. Some people's like... I live here. Excuse me. Some people's like, okay, I mean, compared to like, okay, compare to like... Okay, fine. London where you've got, you know, roads half the size built for horse and carriage and millions of cyclists and buses and all sorts. So I think, you know, it's going to be a long
Starting point is 01:47:18 time until we have that same experience of a cruise or Waymo today, London. I understand. London's a tougher neighborhood. But still, you know, we're 80% there, 75, 80% it there. Whatever, right? But like, and it seems like the stuff that we do in the rest of our lives in terms of AI automation is so primitive compared to this, which is the car that we're sitting here right now. And I find that weird. I find like the relative ease or the relative like heerness of this technology is very disparate. Like, how come it didn't trickle down from self-driving to the rest of tech? Yeah, it's interesting, isn't it? Well, I don't know how those pipelines are built. I assume that's,
Starting point is 01:48:00 the secret source, right? But the flip side of that argument is like, maybe it's very scary that we know, like now many more people understand the mistakes that these types of systems can make because we're all getting hands on with GPT. And this system is equally as problematic and we're just oblivious to it because it's a black box. Almost at your drop off. Oh? Check the app for walking directions. Okay, way more. All right. Well, I think, yeah, that's probably our. right. But thanks so much for giving a quick review. Thanks for having me. Yeah. So that was Louis, whose opinion, I think, is very reflective of the people who are building
Starting point is 01:48:40 code generation or code search type startups based on top of GPC4. And as we headed into the Devday venue, we actually caught Shreya Rajpal from Guard Rails AI. And there was an interesting comparison here in our conversation between how she views the LLM stack versus how OpenEI I've used the LAM stack. OpenEI actually had a closed door session where they gave some thoughts on how they felt that people should start from prompting and build up into a full software system. And they actually deferred a little bit from Treya. Don't worry, all that's recorded.
Starting point is 01:49:13 The videos will come out in a week, but you can listen to TRIA's take. So, so much. We're reviewing AI Engineer Summit. Yeah, we're reviewing the AI Engineer Summit. And it was a very, very well-organized conference. And a small thing that I was thinking about is that your swag for speakers. Is it on? Okay, it's on.
Starting point is 01:49:29 Yeah. Your speaker swag was, like, not surprisingly, I guess, but like really weirdly very nice. And it just kind of like showcases this attention to detail that I think like really kind of permeated the entire, you know, conference. Okay. Like, every single decision was very well thought through and, you know, kind of like to a degree of like quality that's very rare to see. So yeah, it was amazing. I thought you guys like an absolutely fantastic job. Yeah.
Starting point is 01:49:51 This one mostly goes to Ben. So I'm definitely going to make sure that Ben understands that I really appreciate the work that you does. And this is why I couldn't do it myself. You know, I'm mostly the content guy, but I don't, he's the logistics and he's run conferences for eight years. So that's why I keep working with him. Yeah, I also kind of really enjoyed the 18 minutes, you know. Really?
Starting point is 01:50:10 Yeah. When I saw that, I was like, huh, is this going to be, you know, is this going to be enough? And like, is that? But it was like, it would be great. Yeah. Yeah, yeah, yeah. Yeah. I think the 18 minutes was actually the right kind of bite size.
Starting point is 01:50:21 It's optimized for YouTube. Yeah, I see. Interesting. Because it's not the in-person audience that matters. I see. I see. interesting. Okay. I need to promote my my video more. Yeah. Is it was yours up yet? I don't think it's up yet. It's not up yet. Yeah, we're releasing, we're dripping them out to spread it out.
Starting point is 01:50:37 Sounds good. So yours, maybe in two weeks from now. Okay. Okay. Yeah. Sounds good. Okay. So welcome back. Thank you for having me. I think you were guest number five. Yeah. You were super early. So, so we're at the after party now. How do you feel about the whole day? I'm, I'm really excited. I think it was, yeah, I think the, I think, I think, I think, I think, the excitement in the air with like everybody just like waiting with bated breath to see I guess like what gets destroyed but also like what gets really optimized I think this is like very it feels like you really part of a movement and as shaman who like you know us like early people in this space we got to stick together because like whatever happens to any of our companies you know there's
Starting point is 01:51:15 such a like there's such a transformative moment in technology that yeah so you don't care right yeah we're all going to like look back on this time but I had a I had a blast like I really really enjoyed the releases yeah What got destroyed? What got destroyed. I'm mining for hot takes here. Once again, I think my takes are... Unfortunately, very measured.
Starting point is 01:51:37 I wish I had spicier take. Your takes are within the guardrails of common behavior, yes. I think retrieval is like the big one for me. I think it's kind of really exciting to see the retrieval baked in. And that's one thing where I'm very interested to see, like, does that pattern become common by model providers? A, by commercial model providers and also by open source model providers and then how much of retrieval do you have to do yourself?
Starting point is 01:52:02 You know, and like what remains challenging about retrieval compared to just like, you know, this really easy API to just like have it done for you, right? Yeah, I think what they did was effectively build the basic patterns in. Yeah. But for the more advanced stuff, you're still going to need Langchain Lama Index, all those.
Starting point is 01:52:17 Yeah, yeah, yeah. So for the longest, I believe that in RAG, it's the retrieval. That's the hard part, right? Yeah. And then generation is really easy. as long as you're better, like, good retrieval, you can, like, get really, really far, and the generation only gets you, like, a little bit over.
Starting point is 01:52:32 And so I'm really curious to say, like, how, once again, like, how complex do you need it to be in order to start seeing good results? Yeah. Okay. Interesting. And what are your normal benchmark tests like? Do you actually have a set of tests that you run whenever you're, like, exploring something? Or some personal favorites of, like, use cases that you think are tricky for LLMs to do well? I think like a big focus of ours is on hallucinations.
Starting point is 01:52:57 So always kind of like checking out hallucination and like conflicting instructions, etc. As one. Tirst responses is another. You know, like how well is it? Like not, you know, you ask it a question and here's this 10 point list and you know, very, very verbose. Do you have a terse response as a validator? Yeah. Well, we don't have it.
Starting point is 01:53:12 Like, we don't have it publicly. But like we do kind of like check it. Yeah. So I think like those are kind of some of the things. There was one, there's one example in one of the close door sessions where the, The only answers were two tours. Yeah, yeah. Where I think everyone would laugh when they were like,
Starting point is 01:53:26 can you write a block host about this? And the guy and the GPT said, sure, I'll do it tomorrow. Yeah. Yeah, yeah, yeah. Yeah. I think like those are, I think those are, I'm really, really excited about,
Starting point is 01:53:35 I'm really, really excited about JSON generation. I'm actually kind of surprised to see how long it took them to get, like, they're probably just doing constrained decoding under the hood, right? Like, constrained generation. Okay. Because they're now saying that guaranteed correct JSON rather than, you know, more correct. Do you get what I'm saying?
Starting point is 01:53:53 I was parsing through their words. They've never had an issue producing JSON. It's just that sometimes it doesn't fit the JSON schema. Right? Am I wrong? You would know more than me.
Starting point is 01:54:04 No, I think there are also issues with like producing. I think the obvious thing is like unbalanced brackets. When it's on context land, I think that's like an obvious thing, right? But like weird things when you have like really long strings
Starting point is 01:54:14 then quotes, etc., become kind of weird. So I think those are some other ones. Schema is obviously kind of challenging. etc. Yeah. I think there are, even with function calling,
Starting point is 01:54:23 like function calling, at least I haven't played around with it yet today, but previous generations of function calling wouldn't guarantee that your schema is matched, which would be an issue. And I think they're still not guaranteeing it
Starting point is 01:54:34 because I kept waiting for them to say it. I haven't read any of the public docs or anything. Do you know if they're guaranteeing that it fits a schema or they're like... Oh, that's a good question. Yeah, that's a good point. They never said they guarantee. Yeah, they never said they got...
Starting point is 01:54:45 They guaranteed correct JSON. They didn't guarantee if the JSON matches the schema. Okay, you can call JSON loads. Yeah, yeah, yeah. I'm very curious to see, like, once again, if this is a pattern that, you know, all of the other foundation model providers adopt,
Starting point is 01:54:59 and I don't see why not, right? Like, I think for them to kind of, like, own specific decoding models is going to, like, make a lot of sense compared to, you know, like, yeah, a lot of the, a lot of the hacky stuff. Yeah, cool. Any other favorites, you know, doesn't have to be guard, guard,les related.
Starting point is 01:55:15 Any favorite conversations, favorite demos, favorite? I, oh, the GPT. and the assistants. You want to make one for yourself? I do want to make one for myself. It doesn't add, like, yeah, not very Godreels related. I do want to kind of play around
Starting point is 01:55:28 with like how well it works with like some of the things we track. But yeah, it was just so fascinating to see the marketplace. I am very, very curious to see, you know, what the marketplace looks like. Like, are people going to have like really,
Starting point is 01:55:40 really vertically specialized things on the marketplace? Like, if you have a generic, you know, sales assistant or something, right? Like how much, or SQL generator, how much, how popular does that become versus like sales assistant for X vertical at Y stage of the sales process. Oh my God.
Starting point is 01:55:56 Do you know what I mean? Like it's so easy to do this now. Yeah. That like where at what level of specialization do you need to be to kind of start seeing the results? And that is one thing I'm very excited to see like how that pans out. It scares me a little bit because it's basically, they said the future programming is natural language or something like that. Yeah. And that's great.
Starting point is 01:56:15 But like it really is a new platform, a new operating system almost that they're creating. And I don't know how to position myself. Not that I have to, because my world is very developer-oriented. Yeah, yeah, yeah. But this is a whole no-cold world that you and I don't touch. Yeah, yeah, yeah, yeah. Whoa. Yeah, yeah.
Starting point is 01:56:33 I really want to see, like, is there just going to be like assistance for everything? Like what I'm generally curious to see the impact of this on knowledge work, you know, which, yeah, like how much of my work. Like, if I'm getting annoyed by something, is my first instinct going to be like, you know, let me just, you know, spend the five minutes. is to build an assistant for this? Like, is that how everybody's now going to start thinking? You know, and that's one thing I kind of really want to see. Yeah, that's exciting. Yeah.
Starting point is 01:56:57 Okay. Last question, you spoke at AI Engers Summit. Let's advertise your talk a little bit and point people to your talk. Yeah. Yeah. Yeah, so thank you again for inviting me to the AI Engineer Summit. One of my favorite conferences that I've attended, you know, this year. My talk was about the new paradigms for working with large language models, you know,
Starting point is 01:57:15 for building really production-ready applications when the technology that you're working with is underneath all of it, you know, non-deterministic. Really fascinating thing, which was the open AIs talk about building production grade applications, talked about how essential it was to build Godreels as a way to make it to product. You're talking about the one from today. Yes, the one from today. Which people haven't seen yet, but really, really cool talk. So I think it really validates what we've been saying pretty much since the beginning of the year,
Starting point is 01:57:42 which is that you'll get like a certain, you know, you'll get to a certain point, but at that point, you need to start adding guardrails to your application. If you need to get your users to start, you know, getting value out of what you build out, right? So I have your chart and I have their chart. They put guardrails at the first layer. It's not at the end. It's actually right at the beginning for a user experience. Yeah, that's right.
Starting point is 01:58:06 Yeah. Yeah, that was kind of interesting to see that they put it as part of the UX. I'm still kind of very candidly. I'm still kind of digesting that. Like I think of it as I think of it as part of the infrastructure. and I don't know if it's as much UX as it is, you know, just like one of the components that you need in your stack. But I think the pat, like a lot of what they said today,
Starting point is 01:58:26 completely validated, you know, what we've felt for the longest time. And also what I go really in depth about, like in the talk that I gave, right? Which is that what happens once you have the bare bones application, ready, what is the process of actually adding guardrails for what you care about? Like, what does that look like? Yeah. You know, what are the risks that you care about? How do you verify that those risks are happening or not?
Starting point is 01:58:46 not happening, if they are happening, how do you quantify them, and then how do you mitigate them? That was what the talk was about, which I would really recommend people go and check out. Awesome. Well, you did a great job. We're going to post the talk soon. And thanks. It's good to see you again. Thanks again for inviting me. And that was about all I managed to get before the after party. At the after party, there was actually an after after party thrown by news research. So let's hear a little bit about open AI versus open source AI from Alex Volkov. Okay, so we are in the one day after Dev Day here with Alex. Hey.
Starting point is 01:59:19 Hey. Hey. Very, very recognizable voice right now. We don't have to introduce you. Hey, everyone. And we're here to talk about the two parties that happened yesterday. There was one official Dev Day open AI after party where I interviewed Shrea, who's just before this.
Starting point is 01:59:31 And then there's an unofficial one for keeping AI open by news. Yeah. So what was it like, just compare and contrast? So let me maybe start with like who news research is. Oh, yeah. Yeah. Most people haven't heard of this. It's written N-O-U-S, or I mispronounced it,
Starting point is 01:59:46 multiple times, like it's news research. It's one of the few organizations online that started like from a Discord and then like kept going up until like a significant amount of people are like working with them, affiliated with them, of folks who take open source model to its most extreme capability. So collect data sets from open source, open source and more close source and depending on that they released like with different licenses. And then they find to an open source models that were like released to us from like Lama for example and mistral, which is a French company that recently released a seven and B model that's the best. And they've been doing this since Lama 1, but recently it really kicked into high gear with Lama 2 releases because Lama 2 ended up being with a commercial license. So you could actually use this for actual products and services. And Mistral came out with like a full Apache 2 license with a BitTorrent link. I think you remember that. And so these organizations suddenly became like
Starting point is 02:00:36 a very very important currency in the in the world of like where the whole world of AI is going because they're lining local models. And many companies love OpenEI, but either cannot afford this or cannot risk the chance that OpenEI changes something like we saw with Dev Day. And so many people are turning on to like,
Starting point is 02:00:55 okay, if we want to run our own hardware, how do we actually do this? And you can run it, you can run a Lama 2 and all these models on your own hardware, but then you want to fine tune them for your own purposes. And so how do you actually fine tune?
Starting point is 02:01:06 And now organization like news research was probably the biggest one, alignment labs, shout out to Austin and folks from alignment labs, skunk works and many of, of these people come up and say, hey, we have the know-how. And we only started learning about this like eight months ago, six months ago themselves. But now they're like the specialized more people that find two models and actually
Starting point is 02:01:26 release the best kind of models on the Hagenface open source leaderboard. Yeah. Yeah. And in my knowledge, all two models that I keep hearing about, one is Hermes. And they recently switched the base model for Hermes from Lama to Mistral because apparently it's better. Yeah. Hermes is like an instruction dataset, 900,000 instruction.
Starting point is 02:01:43 I don't really know where it's from. Maybe I don't want to know. They also do some, like, fun models. There's, like, a mystical model that they do. Trismestus, yeah. Some stuff like that. I think it's actually a little bit weird that they keep releasing models.
Starting point is 02:01:56 Like, they release three models a week. It's insane. Right? And it's very hard to keep up. Like, I'm like, okay, which one is actually the one that I should pay attention to? Yeah.
Starting point is 02:02:04 So, first of all, you're welcome to join Thursday. And then we talk about all the models every week. It's kind of interesting to that. If I do, like, a recap for a month, the beginning of the month, most of the updates don't matter because like every month. I'm doing monthly and I feel this like. Every month or a month.
Starting point is 02:02:21 I'm doing this for historical posterity. Like five years from now, people want to look back. Then they can look at my notes because I only have 12 a year. Yeah, nobody's going to look at your notes. They're going to have a GPT train or your notes answering everything. I have, yeah, I'm doing like every week. And every week we're talking about like this model outperforms that model like significantly. And we're noticing significant changes from week to week.
Starting point is 02:02:41 Literally in the spend of a month, we went from a 30. three billion parameter model, which is big. And parameter count is not everything there is, right? You can have a smaller model with like larger, longer training that actually will perform better than whatever. But we're noticing smaller and smaller models doing outperforming bigger ones significantly. Zephyr from Hagenface outperformed Lama 70B and Zephyr is like only like a 7B model. On some things. On some things for sure. And so this is very interesting because like it's really hard to evaluate. Evaluation frameworks are bad. Everybody's saying they're not representing of anything. People can fine tune and overturn on them. And so this is very interesting. This is a very important. And so this
Starting point is 02:03:13 There's this whole kind of subculture of open source, mostly on Discord, some of them on X and Twitter spaces. And for some reason, but I find it very humbling and incredible. They also hung out on Thursday. And so that's how I got to this. That's how I got to meet like news research folks, Ticknew, Imozilla, and organized the counter party event last night, together with some other EAC people that we know from Twitter as well. Including Mark Hercid. So apparently he was supposed to, I didn't see him. I saw a photo with a bald head of a big guy,
Starting point is 02:03:46 so I was like, is that Mark? I don't know. Anyway, but the opening eye party was at an art museum, and then the news research party was at a club. It was a club, yes. Folsom Street in San Francisco, a club. 10, 15, Folsom, I think. Opening eye was a very high-brow, buttoned up event.
Starting point is 02:04:04 There was a live band, someone playing jazz. Yeah. Which I think I mentioned this once. It was too loud. We want to talk. We don't want to listen to music. No, no, no. We're just old.
Starting point is 02:04:14 Everything is too loud. And then it was like a lot of people, a lot of networking, a lot of people trying to get together, maybe do business together. And like very, very awesome. Many people from Open AI actually showed up. A lot of people. We stood in line. There was a long line for the Magnometer to step in. And then everybody, like, passing us around was like Open AI employee that passing like straight through.
Starting point is 02:04:33 Yeah. And then that ended around eight, which is like the standard San Francisco like buttoned up in. Oh, yeah. That's when you go to bed. That's where you go to bed. And that's when the other party. kind of started and I think they just seized the opportunity because everybody's in town for the open AI stuff. Why not make a splash and announcement for like for open sourcing AI. So literally
Starting point is 02:04:53 the invite was keep AI free.com yeah which was the website and the API open open.com and you had to register you had to go in there and this was to me an incredible kind of show of Twitter in real life. So all of the folks who follow Mark Andreessen, he received and they stepped into this thing with like the techno optimism stuff. He started to boost the E effective accelerism, EAC folks. And so there's a lot of like signature stuff from that like ecosystem on Twitter.
Starting point is 02:05:24 There's like don't thread on me with like you don't take away my GPUs. There's like all these signs across the club. The it's a very visual club as well. So where the DJs is a whole like a 3D projected thing. So there's like a bunch of like art and like live things about keep AI open. I found it like very, very super cool. I'll have to tell you a tidbit. I saw me and Killian were there from Open Interpreter.
Starting point is 02:05:46 We saw two people with lap codes. I was like, what's the deal with lap codes? So we went and asked. And they just said, hey, we just like, we came back from our work where we work on semiconductors. We were actually like touching chips, whatever, which is like didn't change out of it. And my head was like so incredible in the keep AI open GPU kind of poor party.
Starting point is 02:06:02 We have people who literally work on superconductors came from the work on like they're working on chips. Semiconductors or superconduble, very different things? I think semiconductors. Yeah. We had the Superconductor episode a while back. I think people were still recovering. I'm personally still recovering from that.
Starting point is 02:06:18 There was a whole thing for me. So is news research like vibes? You know, like what is the mission apart from to keep publishing open source models? I think you'll have to get some news people to actually speak like about the mission, about the actual product. But as far as I understand this, no matter how much the product site will be and there might be, There's so many people they're doing like so incredible stuff that people notice like, you know. So no matter how like how much of the the business side will be, they're like committed to fully open source as much as possible, including data sets, including models that are like trismestous, for example, their model that's like trained on the occult and the physical and metaphysical.
Starting point is 02:06:57 You can't expect open eye to let you talk with a model to answer with like these mystical questions, mystical stuff. astrology, Halloween. So you're very like easy. into the astrology in Halloween. They're talking about, like, you can ask this model about like resurrection. And stuff, right? Like all of the occult, like craziness
Starting point is 02:07:14 that they've collected, open I will not let you do that. And so there's a thing that I will not let you do by default because they have lawyers and they were doing sued. Yeah. Recently they announced the protection shield thing, so you won't get sued because of their model.
Starting point is 02:07:27 So they're damn, entropic, all these big companies. It's very important for them to protect the outputs and the models. Here, these folks are like, hey, if you want to build a model, fine-tune this. We're going to teach you how, jump on our Discord. we're going to help you with producing the biggest models.
Starting point is 02:07:41 And then if, you know, there's going to be like financial aspects as well. If your company that wants to run this, we'll also help you do that. Yeah. So it's the same as stability, basically. That's from what it's from talking to him. That's what I gather. Yeah. Cool.
Starting point is 02:07:55 Anything else that people should know about the party news? I found this whole day to be like a very singular AI day. And we don't get many of this. GPT4, I think was the biggest one previously. Yeah, March. Like a single March 14th. That's what Thursday I started. We started talking about this every week.
Starting point is 02:08:11 This was a singular day in San Francisco. This, like, started pregame party with Swiss and some other folks that I got to feel like a little bit of San Francisco. And then Dev Day was incredible. We just heard from Simon. There was like a garage that they made into a venue event. Yes. Probably a custom venue event on the fly, which like just talks to how much they can pull off.
Starting point is 02:08:31 It felt to me that like this Dev Day event and then the following party, it felt a little bit like almost like an Apple thing. where like it's going to be a yearly thing that people will like try to get in as much as possible. One thing to note that in the other party there were many people who didn't get in to this party. And so, you know, they were watching for like a party. Yeah, this office right here.
Starting point is 02:08:50 This office people watched here. And people watched in the life space that we... 8,000 people tuned in to our spaces. 8,000 people tuned in. I didn't even have a chance to look at it. I always want to know the number. Oh, it shows the relative level of interest. And, you know, like, so codens 22,000.
Starting point is 02:09:05 and this is 8,000. Just relative interest by developers. There's like two spaces as well. Robert Scobarra. He stole the thunder a little bit. Stolled some audience for us. Shout out of Robert. And I think that it was a singular day.
Starting point is 02:09:18 And I think the newsresteris keep open source, open, EAC, Mark and Driesen, like all these things together, also added to the top of this because like it happened in the same day, one on top of another, in the same place, San Francisco. I find it incredible. I would definitely come back next year to it. Yeah. Okay.
Starting point is 02:09:33 Yeah. Well, I think you'll be back sooner than that. Yeah, probably. There'll be other things going on. All right, thanks. Awesome. All right, last but not least, we go back all the way to the Newton where I started this podcast, where we checked in with Rahul Somwaka, better known as Rahul Ligma, who just celebrated his one-year
Starting point is 02:09:50 anniversary as one of the biggest memes and celebrities in San Francisco. But by day, he's also the CEO and co-founder of Julius A.I. And I'll match it up. What's up, Swicks. Hey, good to see you. It is one day after a deaf day. and we all had a chance to process. How do you feel?
Starting point is 02:10:08 What's your top takes? That would have awesome. We got to see a bunch of really smart people who are building cool things with OpenAI, GBT, Dolly. The event was very well to put together. The keynote was awesome. The energy in the room was crazy. And I could see real-time social media firing up
Starting point is 02:10:24 with all these takes. Overall, I think it was a good day. Yeah, I interviewed Suria Dantaluri. Yeah, I think you know him. He was like, Samud just killed my startup. And he was, was almost true for him because he has a bunch of plugins. Yeah.
Starting point is 02:10:39 And plugins are kind of deprecated. Yeah. Yeah. Yeah. Yeah. The plugin thing was interesting because it was, it's going to be deprecated, but they just accidentally turned it off yesterday. Yeah.
Starting point is 02:10:51 So he freaked out of it. It freaked out. And then like they bought it back up. Yeah. Yeah. So top features that you're interested in that you want to explore more. I think people are super psyched about the assistance API. But personally, if you ask me,
Starting point is 02:11:05 Two things that I am most excited about is turbo. Yeah. The speed is crazy. And have you actually, have you measured, do you know any like rough measured? Because I don't think they actually ever mentioned the speed relative difference. I started noticing the speed difference in chat GPD actually like a few weeks ago. Oh, I see. So they already slowly eased this into it.
Starting point is 02:11:25 Yeah. Yeah. And I saw like takes on Twitter that did anyone notice chat GDPD get much faster and I noticed it too? Yeah. But so it's turbo. was exciting, but the second thing that's exciting is multiple function calling and then the JSON output formatting. I think as developers are building on the dev API, so that's the thing that's super exciting to me. Of course, as version stuff, there's code interpreter as a tool
Starting point is 02:11:51 in the API. But I think what will bring the most applications is actually the speed, because there are so many things. If you look at our number. numbers on Julius. People are not patient. They want an answer and I want to answer quick. And we see clearly if you can get that answer to them a few seconds faster, there's a clear difference in the conversion. So speed is going to be paid. What is conversion for you? Is that just paying or? Oh, no, just like from first message to second message. I see. So we we do code gen and then we run the code and then the code has an output. They use it as a second message and we can just see the funnel.
Starting point is 02:12:32 Yeah. Where if it's faster, the code runs faster. And the second thing is multiple function calling. I think you're basically telling the AI that, so I think the people
Starting point is 02:12:41 misunderstand function calling it's essentially to use. And if you can tell the AI, hey, you can give me multiple tools to use at once. Yeah. I think that's going to
Starting point is 02:12:52 unlock different applications than before, because before it was just like, okay, this is a task. Tell me one tool and was the input for it. Yeah. But if the AI can now use multiple tools in parallel, you can first of all have more specialized tools.
Starting point is 02:13:07 And then the AI more specialized instructions for each tool. Yeah. It's just going to unlock a lot of cool applications that previously weren't possible. There was a practical limit in the number of tools that you can give it, right? So we had this discussion in March, February, March, April, when they released the function API, that this is subject to context window, the JSON schema itself. Yeah. Does that change at all? or I don't know if you know.
Starting point is 02:13:30 I don't. Yeah. But what I noticed, though, even before, was that more functions and more options just confused it. And that's what I want to play with next is like, okay, where's the breaking point? They see. Like, does more options, you know, confuse it? Does it make it?
Starting point is 02:13:46 Would you use multiple function calls as well? Oh, totally. Is that just theoretical? No, no. I have a direct application for it right now. One of them is oftentimes GGB rights code. And then we run that code. realize that, oh, from GPD's last knowledge update, that module in Python has changed.
Starting point is 02:14:03 It has new function, new APIs. So today, the way we do it is when the error happens, we tell GPD, okay, you can go look up new documentation and then fix that error. But with multiple function calling, the way we would do it, is like, give me the code, but then also give me a documentation look up. And then when the error happens, I can just quickly fix that without another GPD call. Yeah. And then keep moving.
Starting point is 02:14:28 Nice. But I mean, in general, it's just like multiple to use to me. It's just so exciting as a developer. And I wish people were talking more about this. Yeah, I mean, people are still coming to terms of just like the base model and prompt engineering and all that. That's still important. But for engineers, I think you should explore these other advanced features. True.
Starting point is 02:14:46 Yeah. Anything on the multimodality side that you're interested in. I mean, originally will be super interesting for sure. And we have this functionality in Julius right now. you can generate React and each team opponents. Like V0. I think Matt was showing me
Starting point is 02:15:02 a little bit of that demo. Yeah, yeah. We've been hacking on it a lot. I think the missing piece here is that well, you have an engineer who knows how to react and they probably wouldn't find this useful. But if I can allow like anyone in the world to just draw a mockup
Starting point is 02:15:17 on a piece of paper and then run that and have the vision. Yeah. Yeah. Yeah. Turn into like actual components. is I could use on a webpage, that'll be sick. And what's even more sick is like, have the feedback,
Starting point is 02:15:32 where you take a screenshot of the page generated and then feed that screenshot back into vision and then come up with more instruction and have that loop. Wow, like a self-improving web page. Isn't that crazy? Yeah, yeah. I'm so excited.
Starting point is 02:15:45 Yeah, yeah. So in my mind, Julius is very data-focused. By the way, I didn't introduce you. I was just going to do it separately. Yeah. But people know who you are. Yeah. You have a Wikipedia page.
Starting point is 02:15:57 Yeah. You just passed your one-year anniversary as Rahul Lingma. Thank you. By the way, any fun things happen on the anniversary? What are the fun thing? Ilya said, Ilya recognized you on the spot. Oh, Ilya was like, oh, my God, this is, oh, you're famous or whatever. And, no, these guys are so awesome.
Starting point is 02:16:11 Like, they're so humble. But any happen on the one-year anniversary, nothing really, like, it's, I mean, you knew about it a week before. I like to set anniversary dates. That's awesome. Because it reminds people of the passage of time. Like, it's like, wow, shit. Has that been a year?
Starting point is 02:16:24 Yeah. And then you're like, I think it motivates me, it motivates me more than like memento Mori. Like, yeah, it's case, you know, sometimes you're out of date. But it reminds me to spend my years wisely to do interesting things with the time that I have. Momentumory is kind of depressing, whereas this is like, oh yeah, did you know one year ago? Wow, it's been a year. Yeah. Okay, cool.
Starting point is 02:16:45 But Julius, you data analysis chat thing. Yeah. Basically, code interpreter plus plus is how I think about it. Exactly. And also you just across the 100,000 users. Yep. You have delivery modes across, your plugin as well as a chat box,
Starting point is 02:17:00 like a dedicated web app. Yep. Okay. Anything else that people should know? Well, our vision is, you know, writing code is super fundamental to doing things. You could not only automate a bunch of tasks in your life,
Starting point is 02:17:14 but just writing code, but also it's how you just like interact with the universe. Right? You have code that brings you a Vamo car and picks you up, drops you out somewhere. And I think allowing these language models to write code and do things for you is really powerful. And
Starting point is 02:17:30 data announces this application that we're most excited about right now because that's what it's good at immediately. But just on Friday, we launched FFMPEC support. And there were people trying to upload videos, turn the videos into JIFs, or like take a YouTube video turn it into a, you know,
Starting point is 02:17:46 short summary and all these different use cases that we didn't truly like hard code into Julius. We just told it, hey, now you can run ffmpeg and you can run idlp and movie pie and all these different things do these tasks for me and then people were just like organically discovering those things there's this guy TDM on on on twitter cTO junior and he took some meme video and put it on my own tweet over later on my own tweet and then like tweeted that and then that got a bunch of likes and i was like dude like this is the first one that
Starting point is 02:18:17 gets a lot of likes on and you know fmpeg on julius so that has a lot of meme potential it's a lot of potential, but that's not what we're going for. Yeah. It's just like letting people like do things. Your target market is like the S&D, the enterprise. It's actually individuals who have data on the hand and they just want to drum academics. A lot of academics, actually. Yeah, a lot of academics, a lot of students, researchers, any kind of CSVXL data, you can just dump into Julius and then have it analyze for you.
Starting point is 02:18:47 We have this video coming out in a few days where you can now actually train a nano-GPD on Julius. So you can give it, hey, here's the good update you for Carpeti. So you have GPUs to train it on or is just training CPU? CPU, because it takes a minute. Yeah, yeah, yeah. That's true, that. Yeah. I mean, Copathia will like that.
Starting point is 02:19:04 Yeah, yeah, yeah. Okay, cool. So the thing I really want to sort of ask you as a founder on is, you know, I think there's always this existential threat about opening eye building your features, right? Yeah. In a way, so like the number two default bot in the, in the GPT app store. Yeah. Is data analysis.
Starting point is 02:19:20 Yeah. And people can build their own by, customizing and adding code interpreter. Although I think there's also opportunities for you. So on the roadmap that they presented in the closed session, they also said you can bring your own code interpreter. Yeah. So how are you thinking about that?
Starting point is 02:19:35 I mean, as a founder or as? Founder. So who's the audience? Is it like other founders or is it? Yeah. Other founders? And then people are just interested in how you are processing this. Yeah.
Starting point is 02:19:48 I mean, I think it's a very interesting story of processing this live because the news just dropped yesterday. Yeah. Totally. Well, so the story behind Julius is that we actually launched Julius three months after Code Interpreter was announced and a few weeks after it was rolled out to everyone else in the world. So we were number two. And even then we got 100,000 users because I think there's a lot of work to do to get something to work properly.
Starting point is 02:20:14 And there's a bunch of examples of this on the Internet. So if I'm talking to founders, what I'll tell them is, man, so many people give up. before even getting started. And that happens. Don't do that. Sure, you can change your idea. You can find new things to work on. But the way I'm processing is that,
Starting point is 02:20:32 wait, we were actually, we launched after a code interpreter came out. And there's 100,000 people who think Julius is better than code interpreter. Or just tried it out. Yeah. I'll try it out. And use it over code interpreter. And there's like a lot of work to do.
Starting point is 02:20:48 Like, for example, the FFMbeck stuff we launched on Friday. Or the HTML stuff. Or React component stuff. stuff, all these different things. To get them to work, it takes some effort. How I'm processing it, I mean, you know, that's like, that's what startups are all about. It's like risk, right? If you want to build a risk-free startup, you probably don't want to work on startups.
Starting point is 02:21:09 Yeah, just go get a job. Just go get a job. Exactly. So I'm having so much fun. The way I'm thinking about this is like, whoa, there's all these new different things I could do now. I could build. That's so exciting to me. And I'm pumped.
Starting point is 02:21:21 Yeah. Awesome. That's it. Any last words, call to action? Call to action. Let's go build some cool things and get a bunch of users. Let's do it, guys. All right.
Starting point is 02:21:32 Awesome. Thanks so much. Thanks, Wix. I think that's a meme that we can all get behind. Let's go build things for a bunch of users with AI.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.