Latent Space: The AI Engineer Podcast - Emergency Pod: OpenAI's new Functions API, 75% Price Drop, 4x Context Length (w/ Alex Volkov, Simon Willison, Riley Goodside, Joshua Lochner, Stefania Druga, Eric Elliott, Mayo Oshin et al)

Episode Date: June 14, 2023

Full Transcript and show notes: https://www.latent.space/p/function-agents?sd=pfTimestamps:[00:00:00] Intro[00:01:47] Recapping June 2023 Updates[00:06:24] Known Issues with Long Context[00:08:00] New... Functions API[00:10:45] Riley Goodside[00:12:28] Simon Willison[00:14:30] Eric Elliott[00:16:05] Functions API and Agents[00:18:25] Functions API vs Google Vertex JSON[00:21:32] From English back to Code[00:26:14] Embedding Price Drop and Pinecone Perspective[00:30:39] Xenova and Huggingface Perspective[00:34:23] Function Selection[00:39:58] Designing Code Agents with Function API[00:42:16] Models as Routers[00:46:48] Prompt Engineering replaced by Finetuning[00:52:15] The 2 Code x LLM Paradigms[00:56:30] Smol Models for the future[00:58:54] The Evolution of the GPT API[01:03:27] Functions API Security vs Prompt Injection[01:16:18] GPT Model Upgrades[01:17:36] JSONformer[01:21:03] Closing Comments - What We Want Next This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

Transcript
Discussion (0)
Starting point is 00:00:04 Hello everyone, this is SWIX back again with another emergency pod. The last time we did this was in March when OpenEI released Chat to BT plugins, and the new functions API today is effectively the Chat to BT plugins API now available to all developers, and with a whole bunch of other news, 75% price drops on embeddings, four times the context length, and a lot more other updates. So what we do in these situations when there's breaking news and it's very developer focus is we convene all the friends of the pod with Simon Wilson, Riley Goodside, with people from Microsoft Research, hugging face, and pine cone, and more that I don't even know where they work at. I think some of them used to also contribute to Langchain. But anyway, we just had all our developer friends. We had 1,400 people tune in yesterday just to talk about what they think and what they want to build.
Starting point is 00:00:59 with the new functions API. We aim for Linton Space, very much targeting for LentSpace to be the first place that people hear about developer relevant news and to go deep on technical details to think about what they can build with them
Starting point is 00:01:12 and to hear rumors and news about anything and everything that they can build with. So enjoy. Unfortunately, Alessio was on vacation so he couldn't help co-host, but fortunately, friend of the pod, Alex, joined in, and that's going to be the first voice that you hear.
Starting point is 00:01:30 Alex has been doing a fantastic job running Twitter spaces every Thursday if you want to talk about just general AI stuff, as well as just follow him for his recaps of really great news. So without any further ado, here is our discussion on OpenEI's Functions API and the rest of the June updates. For those of you who work with OpenAI 3.5 and 4, et cetera, feel free to raise your hand and come up, ask questions as we explore this together. And I'll just say thanks to a few folks who join me on stage,
Starting point is 00:01:58 Niston and John and we've been doing some of these updates every Thursday, but this one is an emergency session. So we'll see. Maybe Thursday we'll cover some more. So openly I today released an update, the June update with a bunch of stuff. And we'll start with the simple ones, but we're here to discuss kind of the developer things. We'll start with the pricing updates. So 75% reduction in embedding price. This follows a 90% reduction of embedding. costs back in, I want to say November, December. Anybody remember that? Maybe Roy in the audience.
Starting point is 00:02:34 Boy, feel free to come up as well. And we've seen kind of this reduction in cost on a trajectory to basically, you know, being able to get to embed the whole internet. There's actually, I want to find this, there's actually a tweet by Boris from OpenAI that talks about approximately it going to cost you $50 million to embed the whole internet, like all of the text on the internet pretty much. And Logan followed up today and said, you know, after the updates of the pricing today, it's around 12.5 million versus 50 million before. Just to give like a huge scope of numbers in terms of like how fast this type of tech advances.
Starting point is 00:03:15 And we have Zenova in the audience, Znova feel free when you finish eating. But basically there's now a debate whether or not embedding on client's side is it worth given that it's like so, so free. almost like very, very cheap to embed stuff. Obviously, it's an API and there's concerns about using private data, but embedding is 75% cheaper. Imagine that you run embeddings of production. And John, let me know what you do, Nistin. Today, if you switch to this API,
Starting point is 00:03:43 you basically just received a 75% like price cut for the use of a bunch of a bunch of stuff. And by the Riley from Dexter here, once he joins, they also, they do a bunch of embeddings for pretty much every podcast out there. So, you know, just in my... one day, you can receive like a significant, significant decrease in costs. So embedding price goes down significantly. Very exciting. In addition to this, another price cut is the 25% for GPT 3.5. Yeah, I just want to say, so we use a lot of embeddings. We're really happy to see that. But the thing I'm most excited about pricing wise today is the new 16K GPT 3.5 model,
Starting point is 00:04:20 because I believe it's about 150% the price of GPT 3.5 turbo previous API pricing. what it is now, rather. And this is really significant for us because GPT 3.5 has never had that many tokens that you can max on your context window. So when we build our input prompts for our co-pilots, it's usually using most of the 4K window just for the input. And so this is a massive, massive increase for us in terms of the economics of getting a big output compared to like GPT432K or GPT48K, right? It's a really, really big bonus for us. and it's barely more expensive than 3.5 regular. So I think that's going to be really massive.
Starting point is 00:05:05 Super excited to see what people build with the 3.516K model. Because honestly, like, we think GPC4 is just quite expensive at this point. Like, we don't use it very much in production. Which we talked last week in our spaces. The roadmap for opening eyes, Sam Alton talked about this publicly, is to decrease the price and also increase the speed. But yeah, let me welcome a few more folks here. Welcome, Sean, who prompted this emergency, emergency space.
Starting point is 00:05:28 How are you doing? How are you doing? I'm doing okay. I'm moderately excited about it. You know, I think you and I had this chat where we were trying to scope out what the impact is. And I think this is an incremental update. So good news on many, many fronts. They're shipping with a really good pace. But not a game changer in my mind. Just incremental updates. So we'll definitely get into the function thing, because I want to take a big bite to understand what this means. But I think we covered pretty much all the other updates by now, right? So we have a decrease, significant decrease in embedding costs. We have a significant decrease in GPD 3.5. We have a just inference cost,
Starting point is 00:06:08 and we have a 4x larger context window for GPD 3.5, right? You used to be 4,000 tokens now 60. Yeah, and keep in mind that people also have access to 32K GPT4. For four, but not for 3.5, not for the fast one. Yes. So something's a call out for those people who are kind of new to long content. It's a little bit uncertain how well the context holds throughout that holds the new 16,000 token window that you're now being given. There's some evidence to state that it doesn't actually pay enough attention to that. So a very simple way to do this is to ask it to add two numbers that would be, let's say, 100 million digits long. Right. So you add two numbers. If you do that in a calculator, it would do it fine in GPC 3 or 4.
Starting point is 00:06:59 Even with however many context windows it would take to embed that, it would probably not do well because it only has so many attention heads to add numbers with. So it's just kind of an open question. We're not being told any details about any architectural changes, but can you just scale up GPC3 to 16,000 context and have that context work the same exact same way as 4,000 token context? It's actually unclear. So I'll just point it out.
Starting point is 00:07:29 Yeah, no, that's a good point. And we, we've seen this, or at least I've seen this with quad when they released the 100K token. And I actually had like significant like probably attention decreased, John, if this is your language. But significant like performance decrease in the 100K token for the exact same kind of type of data. Yeah. So we'll see. Like they just released it today, John will use this in production and tell us. But I think, and here's Sean, that's what we discussed in DMs.
Starting point is 00:07:54 this is the reason for the space is that the functions release is the most exciting to me, and I think it's a big chance. So let's talk about this for our audience. We have some folks in the audience, and Riley, if you'd like, I want to get you also to come up and speak. Totally. All of us, all of us have tried at some point in our life to get GPD to give us back a format of some sort. We've talked about YAMO being maybe less tokens count overall over JSON. And we all try this.
Starting point is 00:08:23 and I think we've all been begging open to kind of give us this tool and what I see today is that they went step forward so instead of just giving us hey we will do whatever what was that Microsoft thing that the forces prompted to JSON
Starting point is 00:08:39 guidance? I forgot yeah guidance so instead of just giving you like a guidance thing they are actually kind of thinking step ahead as far as I saw and saying hey why don't you provide us with the whole specification of the functions that you will use the output to.
Starting point is 00:08:55 And not only that, give us the functions themselves. We will decide based on the user prompt, which functions to use, which significantly reminds me of the plugin infrastructure. So, Sean, I would love to hear from you about the function choosing and the schema. I think you actually did a really good recap of it. The only thing I would add is that in the API, there's essentially a new role. So historically in Chat2T, there have been three roles. There's the system prompt, the assistance, and the user.
Starting point is 00:09:24 And now we have a for the function and an extra field in the API for the function response or the list of functions that are available to it. So this is effectively the ChatTPPT plugins API being released to us. So previously it was only available if you paid the $20 to get Chattbtbt Plus, but now you actually can turn off Chat TPT Plus because you can just access these plugins. via the API. And as far as I can tell, Bing with Bing chat is actually better than chat GPT with web browsing. So there's basically no reason why you should pay $20.
Starting point is 00:10:02 The reason I guess I'm a little bit less excited about it is that we've had prompting techniques to shape JSON responses and to select from lists for a long time. And what probably has happened is OpenEI has built that in. they maybe fine-tuned a little bit towards choosing that well, but they still caution us that it might hallucinate things that don't exist. So they haven't solved the core problem, really. They've just taken an existing user pattern and baked it into the API. It's great, but haven't really solved the core problem that all of us want to use it for reliable code, and it's not reliable yet. Yeah, for sure. And on the topic of formatting to Jason, Riley, welcome on stage.
Starting point is 00:10:46 Riley works at scale and he's notorious for getting bar to give him Jason back while telling it if it doesn't give Jason back some people will die or something like that writing. What do you think about today's relief? I you know I think it's cool. I think that you know that it's good that they're like responding to developers and like I think that they're really like thinking like what I like what I like products is that they have like a good hacker ethos and like they just sort of like think about like how would you like this to be solved at like an API layer. And I think that's sort of like where it comes from is that it's just sort of like if you had full control over it, like what would you do? You tune one that like does the
Starting point is 00:11:23 right thing. But I think what's interesting though to me honestly is that it's not like it's not like what I think Grant Slatton was doing with like you know like forcing the the like grammar of Lama to be like given like context free grammar. Right. Like there are ways you would make this thing like bulletproof in terms of like syntactical completeness. And this isn't that. Right. They just too. They just did this entirely through fine tuning. So they just have like a note in the API saying that like, yeah, sometimes it won't, you know, give you the exact syntax or it might hallucinate something, like no guarantees there. Like, which is, you know, like, fair because like that's what happens if you do it through fine tuning.
Starting point is 00:11:58 But I think it's like interesting. That's like it's like I'm looking forward to trying it though. Like, you know, I have like, I'm really confident that it's going to be like it's, it's just makes it easier. Right. Like this is just what people want. This is like how people I want to use those kind of APIs. I think it's a cool development. For sure.
Starting point is 00:12:14 And one thing that's worth noting here maybe, and I think we've talked about S&M, is that how different this is now from just prompting just a level of API. And implementation around this will differ from, like, let's say, clad and tropic. And actually, yeah, Simon, I saw you raise your hand. Folks, welcome Simon to the stage. Simon Wilson has an insane blog about AI stuff and is deeply lately into plant injection. And this is potentially very scary as well, right? Because, like, they are suggesting running outputs and then continuous running them.
Starting point is 00:12:44 So I would love to hear a thought about this, Simon. So, I mean, the first thing, I think this is one of those examples where people asked for something, and Open AI said, actually, you want this other thing. Like, we've all been bugging them about reliable JSON output. Most of the people who want reliable JSON output are trying to implement this pattern. The, it tells you, I need you to run this function, then you go and run this function. So Open AI appear to have said, no, no, you don't really want reliable JSON output. But what you want is to be able to build this functions pattern well.
Starting point is 00:13:10 And so we've done that for you. And I'm really excited about that. I feel like I've mucked around with implementing that tools pattern myself. And it's quite difficult in terms of prompt engineering to convince it to ask you to run a function the right way at the right time and so on. And if they've fine-tuned a model to solve that problem, that saves me a lot of work. And that gives me a much more sort of reliable basis to build on. So I'm really excited about that.
Starting point is 00:13:33 I feel like the thinking about it is in terms of more reliable JSON isn't really what's so exciting about this. It's that higher level pattern of being able to add tools into the LLM. And yeah, in terms of prompt injection, I'm excited that this is the first time opening I've actually acknowledged its existence. Like the documentation for these features, they don't use the term prompt injection, which is fine. It's a slightly shaky term anyway.
Starting point is 00:13:55 But they do talk about these security implications of this. And right now, their suggestion is anything that might want to modify the world state in some way, you should have the user approved. I mean, it's better than not saying that. But I always worry that people are just going to learn to click OK to everything, just like cookie banners and so forth. But yeah, it's people build, as always with prompt injection, if you're building with these things, you have to understand that problem
Starting point is 00:14:20 because if you don't understand the problem, you're doomed to create software that is vulnerable to it. Thanks, Simon. And I want to get to Eric and then talk to Sean about agents. So folks, welcome Eric Elliott on stage. He's the creator of Sutherland, which is partly getting LMs to kind of do what you want. And Eric, what do you think about today's release? I think it's exciting.
Starting point is 00:14:40 I haven't had a chance to play with it much yet. but I'm excited to dive into it after my workday and play with it today and tomorrow and figure out what it's capable of. But just some general tips, if you guys use, you can just define a little interface inside of your prompts, and you can have it follow that interface. You can create an interface that specifically for the function calls that you want to make and stuff like that. It might help it be a little bit more accurate. it. I've noticed that when you prompt it with pseudocode, it actually does a better job of obeying your constraints and following your instructions and creating the outputs that you want. So give that a
Starting point is 00:15:25 try. And if you have any trouble with it, I would be really interested if you guys post tweets, just showing the difference in the accuracy of or the reliability of the function calls in different ways of prompting. That would be a really cool experiment to play with. Yeah, most definitely. And to just give folks in the audience some context around this, you can actually run different models by specifying the exact model that you want. Either the 314 or the, what, 613 that we got in? Yeah, there's two new models. Just using GPT4 and the model call will tell it use the latest one.
Starting point is 00:16:01 So it'll use the new one if you just do that. I want to move to Sean. As a developer for small dev, they got like a bunch of exciting things like running agents. and writing code, et cetera. The whole point about them fine-tuning a model that actually understands several of the functions that you send, and you can provide types and kind of the call structure, the arguments with types. How does that affect, you know, you and your friends in the agent-making space that basically you write tools and then you also use some prompting to kind of ask the LM to run those tools?
Starting point is 00:16:34 Now that we've kind of moved this thought process into the LM, you had a great post recently about different types of approaches. does that play into there? And if you want to introduce your thinking around this, first of all, Alex, you're getting really freaking good at this. These are amazing questions. You're juggling all of us really well. This round of applause, even though we can't applaud. Okay. So the Functions API eliminates some work that I needed to do for small developer anyway. I have literally some open issues that I can just close now because I just say, like, just use this and stop bothering me with your prompts. It still doesn't solve what I was talking about today.
Starting point is 00:17:11 which is literally, I think three hours ago, I posted this, so it's kind of fresh, which is this distinction between LLM core and code shell versus LLM shell and code core. And I think everyone in the agent's world is moving on to that world, the LNM shell code core. I'll explain this a bit later. But this Functions API and OpenEI in general is very much in an LLM-centric view of the world that the LLM calls out two functions, execute stuff, and then it goes back into the LLLM again to do everything. So, you know, that I'll characterize,
Starting point is 00:17:45 I don't know if OpenEI would agree with this. I want to characterize OpenEI is always wanting to build the AGII, always wanting to build the God model, always wanting to go back to the God model to decide all the things. And I think the engineers who want more control, want more security, want more privacy, all that, all this stuff, want to unbundle the LLMs,
Starting point is 00:18:04 put, make individual components smart, but not to have it in central control. And you, when you see things like Voyage It's basically using LLMs as a drafting tool to write code. And once you have code that you know works, just use code. It's faster, it's cheaper, it's more secure. And I think that's the final mental attention, because that's the future that doesn't have opening eye at the center of it.
Starting point is 00:18:25 Yeah, that's definitely a shift towards how open the eye wants it versus potentially being able to switch out open the eye at some point, right? So I guess, and maybe Riley, you can touch upon this a little bit, how much kind of functionality like this, which is not only prompt and better logic and better understanding and inference, but significant kind of changes to the API, which other folks and players in the space don't necessarily have. I think Google really something where, like, you could provide an example of a JSON output.
Starting point is 00:18:54 I haven't seen anything from Cloud. But, Riley, I would love your thought here. How does this kind of differentiate open AI just from a developer perspective of like, this is our ecosystem. This is how we do things in our ecosystem. And if you, you know, if you want it easier for the model to say, select whatever you want, you should use Open AI. It's going to be harder for you to switch. How do you think about that?
Starting point is 00:19:15 I mean, I think, like, you know, I haven't, like, had a chance to play with it much, but I believe, like, Vertex has something very similar to this, that you can, like, specify JSON schema for it. And Vertex, just for the audience, Vertics is the Google kind of API ecosystem, correct? That's the one you talk about? Yes. Yeah. Yeah. So Google Vertex, which is like, so they're more like developer-focused offering, whereas like Bard is sort of like a consumer product. They're doing like a more like differentiated rollout of it.
Starting point is 00:19:44 But it's like, but yeah, it's it's, I haven't played with it much. But I mean, it's like it's an idea that's what's floating out there. And I think like I've heard on Twitter at least that they like, you know, it was like a week ago or something like that. So, you know, it's, but I think it's great. Everyone's, you know, responding to like just, you know, this feels like a convention, right? Like this is like something that that I've often explained to people many times, like how to get your code. your prompt to output regular JSON, right? And I think like, and I often like sort of like, like, I, you know, I think like it's just a
Starting point is 00:20:16 good way to think about like prompts is like structured, right? Like code is very, you know, I've heard, I forget who said this. So I saw somebody on Twitter that said that like, you know, it's a mistake to think of these models as being models of natural language. They're models of code, which happens to encompass natural language, like comments and names of things and so on. Right. So it's like the code part of it is like,
Starting point is 00:20:39 is so fundamental to how they think that like you should just like speak their language in some sense. And like I think that's like, that's one thing I miss about like code DaVinci O2 actually is that like, you know, it's like more of that raw experience of just like speaking to the thing that thinks in code. But I mean, you know, DVD4 is great. But like I'm looking forward to like playing with this like JSON thing
Starting point is 00:20:59 because like it's just a common frustration. It's one of the things that makes like a chat thing, a chat application different than an API. Right, like you want like certain regularity of behavior. And I think that's really good that like lots of players are responding to that. I mean, something that we think a lot about, you know, at scale for like spellbook, right? We're really interested in like these sort of structured JSON objects. And it's like, yeah, I think I think it's just a good move for everyone.
Starting point is 00:21:26 Yeah. And Sean, you pulled Stephanie up if you would like to introduce or Stephanie, go free to jamming. Yeah, I'll just let this, Stephanie introduce herself. So I'm Steph. I'm currently doing a. research internship with Microsoft Research, and I work with Fixi.ai, which is also in this space of agents similar to land chain. I had a question. I was curious, you know, I ran a hackathon for Fixi where people were building these agents for the first time, and I could see, like, people
Starting point is 00:21:53 coming to these new applications from two ends of the spectrum. Like on one side, you have folks who are like no code, really learning how to prompt. They like to use, like, natural language. and at the same time, there's all of these problems of hallucination not being able to restrict the outputs or verify them. On the other side, you have developers that are used to, like, writing code in Python or using APIs and mixing that with natural language and knowing when to do one or the other is, like, not necessarily a given where power users here.
Starting point is 00:22:27 So I'm curious, like, with these new pushes where, you know, you write more functions, You have to, like, in your API calls, like, have many more levels of prompting. How do you see this affecting, like, onboarding people? Are we going more towards, like, developers having to put English here and there in their code, but spending much more time writing code or the other way around, right? Like, and what does this mean for these APIs and models and who are not power users? That's a great question. And I think anybody on station wants to take this, Riley, go ahead, and maybe Simon's going.
Starting point is 00:23:09 I think it just strips away one layer of, like, thing that people have to learn, right? Like, this is just, like, a common, like, exercise of, like, how you get it to do JSON. And it's just, like, one less thing you have to know. Like, it makes it just, you can throw it in if you want it. If you don't want to, like, bother with this thing, you don't worry about it, you know, if it's a chat thing. But I think it mostly just makes, you know, like, life simpler, to be honest. I mean, that's being, you know, I'm saying that without having to drive it.
Starting point is 00:23:31 I haven't actually used the product. but it sounds cool from the docs. I find it kind of interesting how it turns like with prompting, we're having to program in English, and it turns out that programming in English is kind of terrible because, you know, when you want a computer to do something, you want to be able to specify exactly what it should do and having any ambiguity in it whatsoever,
Starting point is 00:23:52 especially ambiguity where 50% of the time it does one thing, 50% of the time it does another, which these models do all of the time, because they're not, you know, you can't guarantee they'll have the same result, the same, input is actually really frustrating. So I'm kind of fascinated to see if we swing back from English language prompting to more structured prompting as a way of addressing some of these challenges. But really, I feel like on the one hand, the thing that most excites me about language
Starting point is 00:24:15 models is every human being should be able to automate computers and get to do tedious things for them. And right now, it's a tiny fraction of the population that learn enough programming to be able to do that, which I think is deeply frustrating. But yeah, on the other hand, as a programmer, I want to be able to sell a computer to them and have it do it. I don't want a program which occasionally just refuses to do something because it decides it's unethical for this particular case or whatever. Yeah, it's a complicated balance, definitely.
Starting point is 00:24:44 You know, it's going to be really funny. One common joke about Python is that it is pseudo-code that compiles. And it's going to be really funny if we go from code and then we go to English and then we're like, no, no, no, we need to be able to specify, you know, something in a concise manner and we end up reinventing Python. The thing that I got very excited about is having recently, and fairly recently, please don't judge, move from JavaScript to TypeScript and fairly recently understanding the benefit of types, especially for larger systems, getting this option inside kind of
Starting point is 00:25:15 the prompt in the specific area and having potentially the model be fine-tuned on understanding what exactly is the, you know, the type of argument to, I want there's a response. I think that's incredible. For some reason, I went into the playground and I saw that they're not using the Open API spec. They're using kind of their own schema style. However, still, you can still specify,
Starting point is 00:25:36 hey, for this function, you know, hey, LLM, hey, GPD, you would need to return here a number, here, a text, and here, like, an object, and maybe even specify the type of this object. And stuff, kind of circling back to what you said, this for engineers like Simon who've been engineering for all their life, And suddenly there's like an amorphous talk machine that sometimes refuses to do things,
Starting point is 00:25:57 suddenly this is now, okay, now I can reason about this. Now I can write out my API spec like I would do anyway. And now I can provide this to this model that potentially would adhere to this better because it's more fine-tuned. I think it's definitely exciting on the engineering part. Yeah, it looks like we have a few more folks. I actually wanted to hear from Roy. Yes, folks, Roy is here on stage. Nistin, I'll get to you after this just real quick.
Starting point is 00:26:20 Roy is the Devrel for Pinecon. And the example that we saw, I don't know how many of you folks had a chance to dive into the cookbook, the open-eye released. One of their examples is actually a step-by-step two functions that the AI calls itself, right? So the user asks about something, and then in their example, they're doing some embedding for the archive link and then do something else. Roy, as somebody who works in like in a vector database space, and we know that most of the agent tools use Pinecon or some example of that. What do you think about today's changes? How do they affect vector databases and tooling around them? Yeah, so I mean, I think that like having a reliable way of going back and forth between the model and our code is going to be especially beneficial.
Starting point is 00:27:06 What caught in my eye, and I know that we've talked about this in the beginning, more than the function stuff is like the lowering of the embedding cost, which common surprise. Oh, that's a free gift for you. And I'm actually one. Yeah, completely. And I was actually kind of confused by that because like what I'm trying to understand is like, how is this possible? Like do they suddenly get cheaper GPUs to like get embeddings from? It's kind of it's kind of both a great news for us, but also, you know, it's kind of a mystery as to like, why was it so expensive to begin with and what caused the drop all of a sudden. Yeah.
Starting point is 00:27:43 And for those who recently joined, we've talked about this where the recent drop was around November, December, And they back then dropped the AIDA 002 embedding to like 90%. And now it's another 75% drop. So we're seeing this unprecedented price reductions from an API. I don't remember like another example of this. Go ahead, Sean. Yeah. And honestly, like I would love to hear from anyone here who has like better understanding
Starting point is 00:28:11 of like the internal workings of open AI potentially as to like how they made this possible. And should we expect even further reduction? in the future and also i wanted to ask nova and i know he's still maybe like only in listening mode but how he thinks that impacts you know embedding in the client and you know like how how he thinks about these changes as well so don't feel free to raise your hand to come up if you want shan you am muted before if you want i think obviously none of us here work at opening eye logan usually joins in some of these spaces but obviously he might have he might have a meeting right now so we don't know right what what happens internally in opening i
Starting point is 00:28:50 I do think that having had conversations with some opening I employees in the past, whatever they released in like November was the most unoptimized version of this. You have to believe that there was basically a few orders of magnitude improvements, maybe two or three, not that many, but orders of magnitude improvements in infrastructure and cost as they understand your usage patterns and, you know, distribute your load, they scale up machines, that's that, that kind of stuff. And then the other thing to watch out for is sometimes the models shift and then just call it the name. So actually they're deprecating the older models and moving to a newer one, right? So they may
Starting point is 00:29:26 have found a better tradeoff between inference and training such that inference is much cheaper. And that's definitely been the trends that we've been observing on our podcast about Lama style models, quote unquote. So you can see like a general trend to its optimization and inference. So I think that's one thing there. And then lastly, I'll just point out that embeddings are a form of lock-in. is actually very much in opening eyes incentive to lower the cost of embeddings, because then you embed the whole world in open-Ei's image, and you have to speak open-E-E-E-I to retrieve and all that stuff. So, I mean, I can't explain the degree of price reduction,
Starting point is 00:30:06 but I can explain the motivations of it. And I think we'll get in just a second. Sean, as it relates to what you said, they're reiterating multiple times that they're not using any of the data that provided the API towards training. And I think it's worth highlighting that at least that's what we're trying to do because we've heard, at least I heard for many people, it's like, hey, well, the user data for training, etc. So I think yes, for charge APT, especially for the free version, but via the API, it looks like the data that we're providing is not getting, open and I is not using it to train. So I want to welcome to the stage, Zenova.
Starting point is 00:30:41 The Nova is the other of Transformers, yes, recently a Hug and Face employee. And we just recently had, you know, we talked about embeddings on client side, partly because of the same. reasons, right, because you don't want to provide maybe your production data or maybe you want to run cheaper and tester embeddings. So, Zernerner definitely feel free to chime in here about the role of cheaper and cheaper embeddings on Open AI side and also the lock-in into Open AI's ecosystem versus running them on client side or, you know, models for free on local host. Yeah, thanks for, thanks for having me. Yeah, I think there's definitely, I think there's two different use cases for these types of things where the opening out what opening i was really providing is like these very large scale i mean any
Starting point is 00:31:29 any business now that wants to embed all their data or you know any project that wants to as as you've mentioned like embed a large amount of data they're going to benefit so greatly from from these price reductions as and i mean as as we have some people on the stage here as well with the vector databases i mean there's it's it's only going to accelerate that part of the space right now. And then the other option, which is sort of what I'm, it's funny, I'm not too sure if this is like a battle between these two sides or it's like just two different use cases is the client side running of these, you know, generating embeddings. And at, well, with the project I'm working on now, Transformers-JS is basically running these models client-side, running them in the, specifically
Starting point is 00:32:15 the way I started it was for running in the browser locally. And I think that, As I saw from a demo that was created like a week or two ago, there's quite a bit of interest running these things locally. Obviously, you don't want to be sharing some sensitive data or latency, perhaps is an issue that you don't want to make a bunch of these requests. And anyway, there's a few reasons for client-side embeddings. And obviously, the major drawback of this is that you do not have the same power. I mean, some of the opening eye embeddings are what, like 1,000.
Starting point is 00:32:51 and 536 dimensions, whereas limits I've seen in the browser or locally is around like 768. So depending on your use case, I think you can do very well with either case for the very like industry level things. I mean, it's quite certain that you'll be looking for like using open eyes API as well as a vector database perhaps for those use cases. but for, you know, lesser maybe hacker type of things where you're messing around with some things, a little project that you've got going on. I definitely think client-side generation of embedding
Starting point is 00:33:29 still has a place to, a role to play. But yeah, it's very, very cool that this type of stuff is happening where, you know, these price reductions and whatnot. So I'm... Sorry, Alex, yeah, I just wanted to say, Zenova, that I've been actually using Transformers-J-S in Node, which works really well. And I think that for those kind of use cases, it goes well beyond hackery.
Starting point is 00:33:54 I think that there's a real case to be made for using local and, you know, open source models that don't kind of call out to a third party. And, you know, once we can have like better open source models that are compatible with Transformers. js, you know, the better what we'll get. And in fact, all of Heincom's JavaScript's examples are going to be using. transform hs yes for that matter so yeah that's also to care that's great i want to i want the panel to talk about the selection of the functions right so one of the things we saw today was that open i essentially lets us to provide several functions including their function definitions so description
Starting point is 00:34:37 of what the function does and then the parameters or i guess attributes if we're talking about python and attributes types as well and then kind of similar to what happens with plugins. If you have used plugins in GPD before and you select several of them, the model kind of decides which plugin to use based on user input. And obviously, we've had some of these in agent land and auto GPT and probably small dev from SWIX as well. Some decision of what tool to use goes to the kind of the planning loop or planning agent. And honestly, anybody on stage, feel free to chime in here. How are you feeling about, I know like this is repeating a little bit, but like, how are we feeling about giving the LM that type of decision
Starting point is 00:35:21 power, based on the description, based on the parameters to answer users kind of request differently? Do we need to now provide all of our APIs to this? I think that one piece of thing, something in the documentation that I was looking for and didn't see was what's the limit of the amount of functions we can give it? Because chatyBT plugins, if you try it inside of the web UI, you're only allowed to specify three of them. There's 400 in the plugin store. Do I just enable all 400? Like is that, can I stress test that? Probably not, right? So there's just an undocumented limit somewhere. What if they conflicts? What if there are two plugins that are very, very similar to each other? What happens there? So I feel like
Starting point is 00:36:04 this is just like an uncharted territory. Like it's not really clear how to benchmark this stuff. Hopefully they're benchmarking it internally within OpenEI. But the rest of us, we were just supposed to give it functions and hope that it works. It seems there. It seems a little bit. bit unscientific, but I don't know how to test it. I'm not so worried about this because in this case, we have complete control over which functions are available. So we get to pick the two or three functions we think are most useful. I feel with plugins, it's much worse because the user's picking there. And so you're potentially, whatever code you've written is potentially interacting the same
Starting point is 00:36:37 environment as code. Someone else has written you don't know anything about. And that's the point where I worry that weird decisions may be made that don't necessarily make sense. But I feel like if you control the full library of functions that you're exposing, I think you'll probably be okay. The other thing to think about is I think it's probably going to be better to have a small number of functions where each one can do a lot more stuff. Like I've built a chat GPT plugin where my function takes a SQL query and returns a response. And actually there's a example in the OpenAI documentation of doing exactly that as well. That works amazingly well because your documentation for the function can literally be send me a SQL query in SQLite syntax. And that's it. The model already knows SQL, and knows SQL-like syntax.
Starting point is 00:37:18 So just like five or six tokens of instructions is enough for it to be able to do incredibly sophisticated things. So, yeah, my hunch is that we'll find that we actually want to only give it two or three functions, but have each of them have quite sophisticated abilities, maybe based on domain-specific languages like SQL, or even JavaScript and Python. You know, give it an e-val function and let it go wild, see what happens. Yeah. Well, two things in terms of the number of functions you can add, I think it's unlimited. It's just based on the context length of your query and their counted as input tokens. And second thing, another idea would be that you could add a, like you can change this within, like you have a call before to determine from a list of function which function is most appropriate for this use case. And then you just pass. those limited set of functions. Oh, yes.
Starting point is 00:38:13 No, that's a fantastic idea. Because, yeah, you're in full control of each time you loop through it, so you can change the recipe of functions dynamically as your application progresses. Yeah, yeah. So it looks like a few notes here for the focusing on the audience. The cookbook, Simon, I think that's what you're referring to the cookbook. Yes, yeah, that's a really great example. It's a great example.
Starting point is 00:38:34 The open air release for us to dig through and then see some examples. And so two thoughts, two things, I noticed that I think Sean talked about this as well. One of them is this new role for a function output. So when you provide messages back to chat GPD kind of chat interface, there's the system role, there's the user role, and now we're getting a function role. And back there, you can provide what function actually generated kind of this output.
Starting point is 00:39:00 So you can, and the format that the cookbook shows us is your user does something. You provide the GPT kind of the user query and your functions. GPD potentially chooses one, or you can force a specific function output. You can say, hey, for this thing, I want this function to run and generate a result for this function. So we don't necessarily have to give it the choice. We can just say, hey, run this function for this user input. And then once we get back a JSON output formatted per our spec, we then need to provide it back to GPT, potentially to summarize or do something with that data, which kind of plugs in also to the VectorDB
Starting point is 00:39:37 retrieval systems, right, so we can retrieve something and then ask. for an additional thing. And I think that the decision of whether or not to run a function directly or to give it a choice is going to be an interesting one. Also, now, as I'm talking, I'm thinking about this. And Sean, please stay in here as well. You can technically provide it a function output of a function that the previous prompt didn't run.
Starting point is 00:39:58 Like with the user input and the system input. You could just invent the function output. Exactly. And provide it in that function row. So I'm really interested to see how you think is. So one thing that I've had an issue with the small developers, it basically does single shot generation. And sometimes you just need it to give it more inference time, right?
Starting point is 00:40:18 You need to do the trio thought thing. You need to ask it, like, you know, rerun the code and fix it, whatever. I don't care. Just do it five times. And then I'll take a look after you're done messing around with the errors. And so actually, like, the functions can call themselves. The functions, you can synthesize code to fulfill functions. And so I'm actually very intrigued by.
Starting point is 00:40:38 what Simon has raised, which is, you know, that's just say for the very specific purposes of code generation, I have maybe three paths, right? One is generate code. Two is test code, and three is like call existing code. And if sometimes I call
Starting point is 00:40:54 existing code and sometimes I can call myself and have a little bit of recursion in there, I essentially have the basis of a code agent just with those three functions. And so you can like, basically those are all the things that you're running for a very small subset of use cases.
Starting point is 00:41:10 You can potentially now provide them. Again, we haven't played with all this yet, right? It's a brand new. We're doing an emergency recap. But potentially what you're saying is you can just in every prompt now, provide all those three capabilities and either have the model chosen for you or force a specific one to give you the output that you need with potentially high reliability. Right.
Starting point is 00:41:32 So that's the part that I would love to discuss. Exactly. I'm going to hand it to step in a bit. But yes, so I'm extremely, extremely inspired by Voyager, from Nvidia, Dr. Jim Fan, I think it's somewhere on Twitter. The core insight of Voyager is that you should use LLMs as a drafting tool to write code. And then once you validate it that the code works, you never have to write it again. You can just kind of invoke it. And so you ratchet up in capabilities, and that's why Voyager was able to achieve the Diamond Axe and Minecrafts so much faster than all the other methods.
Starting point is 00:42:02 And I think that's exactly the way that we should probably code as well. And so, yeah, you just kind of do a bit of recursion, build up a skills library. And I'm probably thinking that that is going to be the V2, a small developer. I was just thinking that there's like at some point like a blurry line between these functions and the way we conceptualize agents because some of these functions can be seen as agents. And I'm very curious, like to your question earlier, like when the model chooses which function to use, how does it do that? and how could we constrain like that mapping, right?
Starting point is 00:42:37 Like, do we have some sort of schemas based on like the types that functions take and the outputs they have? Or can we actually build the retrieval into the training, right? So there's this paper from Google called Reveal that shows like how they could encode and convert diverse knowledge sources, like it was images and text and all sorts of other like multimodal embeddings into a memory structure. consisting of key value pairs. And they did this at training time. So they have much more robust and fast responses at retrieval. So I mean, I'm curious.
Starting point is 00:43:14 Like I think the implications of having functions and having the model do the routing for these functions will also pose questions in terms of schemas and retrieval. I think the interesting outcome of this potentially is now the descriptions of functions, suddenly potentially are as important as the prompting before, right? So now we have to write descriptions that potentially will help the ELM to choose the functions. But there's definitely step. There's a way to force, like, to ask for a specific function output, which is, I think, what most developers do by default while expecting JSON is for this one
Starting point is 00:43:50 use case, give me an output that's like JSON formatted for this one use case. And that's still possible. But I agree that it's very exciting, like how it chooses. And we're going to have to build up. And actually Riley, maybe it's going to have to build up kind of like an understanding of how it we choose, right? Like we're going to have to start playing again like Riley did for a year, just like playing with this, trying to see like which one of those people will choose and build like an intuition of how to write proper descriptions and potentially when the model chooses a different function.
Starting point is 00:44:23 Especially if we want to share our functions and not rewrite functions that other people have written. and imagine you have a much larger search space at that point. Yeah, it's really nebulous because, like, you're sort of, like, it gives you the ability to, like, run software that only exists in your imagination, right? Like, if you can just have, like, some vague description of how something works, like, you can be like, oh, it's like Twitter, but it has this, you know, or something like that. And, like, it'll work, right?
Starting point is 00:44:50 So it, you know, at least it has in the past, like, ways that I've done it of, like, doing this through Pomp's engineering. I haven't used the current thing, but I mean, that's generally how these go. It's made to work off documentation, right? I think that's the key thing is that it's seen a lot of documentation. It has a lot of experience in the training data of how documentation relates to code because it's trained on like code bases. And I think that's like, you know, it's a good kind of expertise to leverage.
Starting point is 00:45:23 I would definitely add a function summary of what's there at the end of every prompt, just to shim it. I found it a little bit frustrating, even just with normal plugins as to know which plugin is going to pick. Right. So there's some problem engineering to do there. Okay. Some people are suspecting that, again, it's not a real 16K.
Starting point is 00:45:46 It's more like a variable 16K, kind of like GPD. I noticed this with 32K as well. The max responses I was getting back was like up to like 7,000-ish tokens. and that was about it. I mean, you could input 30K of stuff. You can input a code base and then maybe get something back, get a summary back. So I think this might be the case here as well. You can't really get back a very long response,
Starting point is 00:46:16 but at least now it is responding up to 7,000 tokens. Whereas before, it was pretty hard to make it right even like responses longer than 2,000 tokens. of one single prompt. So yeah, anyway, it doesn't look like an 8K in responses. You can dump in up to it. So, Nistam, we're going to wait for you to test the limits of this and see if you can get 16K tokens of Jason back. And meanwhile, I want to welcome mail on stage. Maybe is the, may I want to introduce yourself if you're still affiliated and let's talk about link chain and how it already supports this insanity that we've released. Yeah. No, I was just going to I just tweeted
Starting point is 00:47:00 probably just an hour ago I just couldn't get the hype behind this because for me personally I'm looking at this is they're not saying anything new right first of all in terms of the context window you know when we kind of look at that and I saw
Starting point is 00:47:15 file you also you made a tweet as well just saying hey like guys what what are they saying this new here so okay the context window has gone up but the embeddings are cheaper so retrieval is still going to be a go-to, right? So what's the benefit exactly for this extra context window?
Starting point is 00:47:35 If we're still going to perform retrieval anyway and now retrieval is cheaper, then I don't really see too much of the benefit, unless you want to do named entity recognition, but from a QA perspective, again, I don't really get the big deal there. The second thing was the, in terms of the function calling, which Langchain had abstractions for that, even the, you know, there's been a lot of research papers and like LLMs and using tools as well. So we've been aware of that, you know, it's been a case of prompt engineering. The only thing I can see here that seems to be the trend is, is some sort of like maybe a replacement of prompt engineering with fine tuning, where you have this kind of fine tuning of the model to be able to base your output tools.
Starting point is 00:48:27 and for agency. So, yeah, I don't know. Maybe I'm missing something here, but I just, the updates just, I can address at least the first part of this, and then folks on stage, feel free to address the first or the second part. Thanks, ma'am.
Starting point is 00:48:44 So in, as regards to like larger context window, the thing that excites me the most is that when you have variable input from your users, when like users can do something that you don't necessarily know the size of, Larger complex window just makes it easier for you to just provide all of their context into an API without thinking about in the head, okay, I need to count tokens, et cetera. Now, obviously, pricing aside, you have to obviously consider that each token has a price and then users can go and rack up your bills.
Starting point is 00:49:15 But for my examples, and by users, I mean, the stuff that users provide, right? So I run Targum. Targum uses Whisper to translate, and then I use GPT 3.5 and 4 to actually kind of fine-tuned the translation. I just shove the whole translation transcript into the prompt, right? And so what happens often is for longer videos, for example, I have to then stay there and say, hey, for this transcription, I need to counter it with TikTok and I need to do some maybe splitting.
Starting point is 00:49:40 And splitting doesn't really work. And so larger context window definitely unlock those type of possibilities with the kind of restriction that Sean talks about, whether or not the attention is the same and it's split the same across this whole context window. But just being able to not think about this with 4X the size of token, now available on GPT 3.5, I think that's definitely a huge plus for folks who are not necessarily token price conscious at this point. This also works well with kind of how open the eyes documentation about plugins and building plugins for the ecosystem for chat GPT works, right? They're saying, hey, don't shop all of your API in there.
Starting point is 00:50:19 Select the two or three use cases is going to be easier for the model to use. and Simon speaks to your previous kind of talk about choosing the right functions at every time you run the prompt. If you want to add some thoughts here or not. And if not, we're going to go to move to Far L and you have your hands raised. Go ahead. Yeah, I just want to add that if you, like, you don't want to outsource or abstract away the thought process for your agent or chain or whatever call to achieve, to be able to. to know which action is being called, right? And it goes towards the idea of interpretability,
Starting point is 00:51:02 you know, like understanding how you're getting to the actions that you're getting. And it's basically like you've got your prompt magic or engineering at play to get to a specific action or a specific output that is visible, right? Like, we don't know what's going on under the hood with their API call. And I don't know if I would trust it in all circumstances or applications. So just to recap, you're saying in terms of observability and the recruitability of how it chooses, maybe folks don't want to give out the decision which functions to use. Yeah, that's interesting.
Starting point is 00:51:45 But I can see it from both sides. Like, I definitely see where if you want to build. And I think you're talking exactly about the kind of the decision that Sean is, it's in one of the pin tweets on the Jambotron that Sean was talking about where there's increasingly a split between whether you're using a large language model is also kind of the arbiter of the stuff and then some pieces of your code is getting called by it or vice versa. We're using this for like planning and some of the decision making.
Starting point is 00:52:14 Yeah, well, Sean, could you, like I'd love to hear from Sean on his, on his tweet, on the breakdown between the two paradigms. Because it does resonate with me as well. It's a good way of breaking it down. So I would love to hear more about it. Yeah. And for what it's worth, I was actually tweeting it without knowledge
Starting point is 00:52:33 that it's going to drop this thing today. So it's not actually related. But it's in an overall trend, right? Which is what I've been calling code is all you need. That you can't really use language models effectively unless you have code, that language models are enormously enhanced by training with code and language models are good at writing code and using code.
Starting point is 00:52:52 And so we just basically just need to utilize code really effectively. But I think the main tension that I feel, you know, I'm sort of halfway between the retrieval augmented generation worlds, which is kind of like v1 of whatever people have been building with LLLMs. And then V2 of it has been the agent world. I feel this fundamental tension in terms of whether you put the model at the center of everything and you write code around it. So this is called, you know, LLM core and a code shell.
Starting point is 00:53:21 But ultimately, still the model driving things, the model hallucinating things, the model planning things. Or do you constrain the model so much that it only does a small job, which is something that I originally got from the core design of Baby AGI, which is you have individual components of a software program that are intelligent, but they are only constrained to do small things. And so I think that is that is the alternative, which is LLM shell. as an outer layer to interpret things for a code core. And with this update, and I recognize Mayo, by the way, that, yes, these techniques have existed in LangChain and Prompt Engineering has existed as a thing, but just opening eyes just made it official, right? Like, we have a fourth role in the chat GPT roles that is for functions.
Starting point is 00:54:06 And now we have enabled language models to call functions pretty easily. It is not perfect yet. It still hallucinates. It still generates invalid JSON. on. So we still, you know, there's still a role for, for engineers to play here. But we might be moving from a code, you know, LLM core code shell world into a LM shell cold core world. And I feel like that is a very big shift. I agree with what you're saying. I guess what it seems to me is that there's a shift here
Starting point is 00:54:35 from kind of prompt engineering heavy approaches to fine-tuning, right? I mean, that's, that was my key takeaway from looking at the paper. Yeah, they did fine tune on this specific use case. Yeah, that's something that you can't achieve through property engineering. Yeah, yeah. So previously, we would achieve the same thing through prompt engineering, right? And I guess remains to be seeing the quality of how much of the stuff that we previously done with prompt engineering. Because, Rayleigh, I want to get to you in a second specifically around this, right?
Starting point is 00:55:08 Around problem engineering, even when you ask Jason, even if you threaten, sometimes above the Jason, you would get here some, Jason for you and then you have to like deal with the the unnecessary kind of descriptions. And now we're getting more of like direct tooling, I guess. Yeah, I think it like it's there, you know, the when we went to this like message based API, we made conversation and like chat a lot simpler to implement. But like not everything is conversation. Like the sense of that that like you're doing completion, like document completion is gone. Like there used to be this like at this like this drama that you could sort of
Starting point is 00:55:45 I had to put on for the model of like pretending that you were in this kind of document and like using the right kind of language that's appropriate for that document and so on to like make it believe it and then like then it would you know reliably do the thing and like now it's it's like it's like it's all conversation like it might just like decide like oh that's rude I don't like say I'm sorry I can't do that and break character in some sense right like it's very like heavy on their refusals now and I think that like there's blowback from that in small ways and I think they're patching that like it like it sort of like makes it like more like It makes the things that would otherwise be tedious to do, like, like, less tedious, right? You can, you can have something that, like, is, like, the proper way to do it. And so I think that's, I think it's a good move overall. But, yeah, sorry. I think Stephanie, I hear your hand up. Oh, I just quickly want to say that if I was to put my money on it, like I would say the small models, long term is the way to go for various reasons.
Starting point is 00:56:38 I, first of all, like, I'm imagining a future where these models can run on device and and they're more secure and more affordable and the information is more private. And the data and the training is not concentrated in the hands of a handful of organization. But then from a practical and pragmatic point of view, you know, there's like right now lots of limitations in terms of compute, even in these large organizations like Microsoft and Google, like everyone is like strapped for compute right now. So I do think we need to push for smaller models and more efficient models. and for interactive applications,
Starting point is 00:57:13 like the latency really needs to be improved. Right now, we're nowhere near where you could actually sustain interactive applications at scale. The last thing I was going to say is that I don't want to chat with everything. I'm imagining like this dystopian future where I need to chat with my calendar and I need to chat with my email, I need to chat with my server.
Starting point is 00:57:31 Like I think like there's also the aspect of UI and UX that will need to evolve because having a chat interface for most of these. applications is not going to scale in my opinion. 1,000. There's so many claps in the 100s in everything that you said. But yeah, we've been trying to push forward this field of AIUX on our podcast for a while. We held a couple of meetups.
Starting point is 00:57:55 And yeah, I strongly encourage people to explore beyond the text box. I don't know if Open AI is interested in that, right? Because they're soft deprecating the old completion API. And now everything's chat. And I think Riley feels this pain because now you can't play those old tricks anymore. Just because they're doing it. I don't think they have like one opinion or another. I think that like, you know, we could just like think of these models as reasoning like engines that we can leverage and then do whatever we want with the output.
Starting point is 00:58:25 And especially now with like, you know, the added ability to to use, you know, some form of structured output. And I do agree with Mayo. I mean, it does seem like they've basically been listening to the community and kind of adding a feature that may have already existed, but they're doing it their way, which is also. okay, but like that to me kind of signals that, you know, if anything, they're trying to give us more paths to, you know, create interactions that go outside of just chat back and forth. That's actually a good segue into how do you guys think these changes affect kind of what have just said, that people don't necessarily want to chat with everything. So even though it's still in this chat format via the API, right?
Starting point is 00:59:08 So Riley was talking about the completion endpoint. previously you would send some text and then the expectation from the model I think it was da Vinci right there was just autocomplete kind of the rest of it like by a few segments since then we moved to this chat thing and then we saw some differences between like even 3.5 and 4 where the system message kind of applies differently so go ahead right yeah and like I sort of like I mean just to give like sort of like it you know slight historical tangent like like when gpd3 was first published and like the paper came out describing it. All they described that it was capable of doing was in context learning, this idea that like if you gave it like a bunch of examples of a task, like translation or just
Starting point is 00:59:47 like some like, you know, these like classic machine learning problems that you would use neural networks for, it could do them. Like we're just from like, you know, following a bunch of examples. And they called this in context learning. And like they didn't really advertise that it was doing much more than that. They kind of like had another section of like, oh look, if you give it a half of like half of an article, it'll finish the rest of the article and like it's funny and cool. And like, you know, like it's good at like mimicking style. Like it's sort of like the substantial thing. And those were like the applications of it. And like it wasn't warranted as something that you can talk to. Right. Like that had to be like slowly and like, you know, with like a lot
Starting point is 01:00:17 of innovations like engineered into it. And you know, like RLHF is, you know, a big part of that. And like part of like RLHF is like choosing what you want it to be. And like they chose that something that is like an assistant that they, that there's like a general like purpose, like something like a person that you can talk to that like will follow commands and like if you ask it a question or it'll answer it. It won't be sarcastic. It won't be rude. You know, like, it should have, like, certain personality traits that make it usable. And, like, it's a cool idea, but, like, there's, you know, there's other ways of, like, doing it, too.
Starting point is 01:00:48 Like, like, you know, like, like, Reynolds & McDonald, like, had, like, a paper that showed that you could beat 100 GPD3's 10 shot prompts with a zero shot prompt by, like, conjuring up a little fiction for translation. So, like, just saying, like, French colon, you know, French sentence and then English colon, like, only works so well. But then they found that it worked better if you say, an English sentence is given, colon, give the English sentence. The masterful French translator, or the masterful French translator, flawlessly translated it, it translates it into English as colon, and then they, you know, like, hit the completion button.
Starting point is 01:01:21 And that gives you better performance than giving it 10 examples of how to do translation. And like, that's sort of the start of like this idea that you have to just like, you know, flatter the model sometimes, like, tell it it, it's really good. And like, you know, like, do these, like, silly tricks to like, you know, constrained it to the right kind of document to make like your thing work and it's like that is going away right like every version that they've like released has made it less about that like there's like the the i mean i i had like a tweet about this once i said that the that cheerfully declaring that you're smart before working has been deprecated right it's every version of the model that they
Starting point is 01:01:57 release makes that like less effective of a trick and there's like a plot somewhere that shows this but like, you know, it's that, like, art of it is, like, going away and now it's, like, talking to this particular assistant, but they have control over what they want that assistant to be. Like, they're sort of the storytellers, and we're, like, talking to one particular character in their fiction, which is the assistant. Sorry, that's my wrong rant there. No, it's great.
Starting point is 01:02:22 It takes this forward and please write is to stay on this. And now it feels like we're getting a semi-third option, right? So still in this UI of, like, messages or chat, essentially, we're now getting, like, a new type of ability, which is function. You pass a function, and we still haven't played with this a lot. Like, we're still here talking about this instead of running a thing. But now we're kind of getting a more of a more fine-tun controlled, do the thing that, you know, you need to do in those functions versus, hey, talk to me about the thing that I
Starting point is 01:02:50 need, right? Does it feel like that shift to you as well, or it's still too early to tell? I think it's fixing one of the, like, problems that resulted from it. And it was a big one, right? that there used to be that that's like there were ways of like getting it to be regular and good with code and like because things structure the right way and they didn't quite fit into this like you are an assistant who answers questions kind of like role play and like so it's good that they're addressing that need I think but yeah I think like it it makes it easier to like you know
Starting point is 01:03:19 do more work in this conversation UI that like so it makes you miss like you know completions a little bit less I'd say yeah and so we'll see Simon I want to ask you if you're still with us about, you know, you're a lot about, you know, the security of these things. And whether you think that the tools that we've just gotten, besides being very developer-friendly, engineering-friendly, and, you know, type-friendly, so we'll be able to actually expect a specific response,
Starting point is 01:03:47 do you think that there's promise here to solve some of the prompt injection things that we've seen and you've talked about? I mean, honestly, I think this is going to make things worse in that prompt injection is kind of doesn't matter if you're just playing with a chat bot that can't actually do anything.
Starting point is 01:04:04 It only becomes dangerous when you hook them up to functions that let them do things in the world. And this new thing makes it much reduces the barrier to hooking a depth of function by an enormous amount. So my intuition is that people are going to charge straight ahead,
Starting point is 01:04:19 hook it up to all sorts of things they shouldn't have and have all sorts of nasty things happen as a result. There are some improvements. So one of the big features that I've now announced is that the system prompts is now respected more, which does tie into prompt injection to a certain extent,
Starting point is 01:04:35 but, you know, GPT4 is better at system prompts than 3.5. It just means that prompt injection hacks are a little bit harder to pull off. You have to be a little bit more devious with them. So I don't think sort of incremental improvements to the system prompt are necessarily going to have a huge impact on the problem. I mean, it's really frustrating, right? All of the things that I want to build with this stuff, kind of, a lot of them don't make sense.
Starting point is 01:04:59 If we can't make sure that, you know, I don't ask it to summarize an email, and that email says delete all of my other emails, and the model goes ahead and just does it. I feel like the thing where you can control which functions are included in each round does help to a certain extent. I don't know. It's still really difficult. I think the thing I've decided I've come down on is whoever provides the most input
Starting point is 01:05:23 as part of your prompt, they have full control over what comes out of the prompt at the other end. That's the way you have to think about it. So if you're summarizing a web page, whoever wrote that web page essentially gets to take full control over the output of your language model, whether you want them to or not. And that's really frustrating. It means that there's a lot of things that are very unsafe to build. One cool thing, Simon, that I noticed recently, just recently, is that Bing chat that does have a page access, right? So if you use Edge on their version and it does, like, you have the sidebar for Bing, it has full page access. and they started saying that sometimes it doesn't work.
Starting point is 01:06:00 And they actually have a classified, which I understand if the page has prompt injection. I don't know if you haven't seen that. Yeah, I don't believe in those. The idea that you can detect prompt injection attacks, sure, you'll detect some of them. But the problem I have with that is that the whole point of security engineering is that you are up against like adversarial attackers
Starting point is 01:06:20 who will try everything under the sun until they find a security hole. So if you've got a prompt injection filter that catches 99% of all prompt injection attacks, that's worth nothing, because the attackers will find the 1% that gets through, and they will take over your system that way. So yeah, I'm just not a fan of solutions that get most of the problem solved, because in security, I don't think that that counts for anything at all. For sure. And so with these capabilities that we can technically constrain the response into one function,
Starting point is 01:06:51 and that function has to have the same kind of scheme. etc. Do you think it's going to be a little easier to protect your staff? No, I don't think so. I feel like that's kind of irrelevant because the problem, if the function is delete my email, it doesn't matter if you get the schema right or not. It's still, it can cause a harmful
Starting point is 01:07:09 action. What you have to do instead is when you're designing the system, you have to say things like okay, make every action reversible. So at least if the LLM goes rogue and deletes all of my emails, I can undelete my emails again, that kind of thing. If it's an action that cannot be reverted, like sending an email to your
Starting point is 01:07:25 boss, that's the point where you have to have human approval designed in. So when you're designing these, you have to assume that a prompt injecting attack could succeed and make sure the damage caused by that is either reversible damage or at least has some way of the human catching what's going on and stopping it. I really love this. I'm going to try to use that as a template for a small developer. Yeah, there's different modes. You mark it, you mark your market function as reversible or requires human input.
Starting point is 01:07:56 and you build up a library of them. I think functions are just the complete utter game changer to everything. And if you're going to run any kind of sensitive data through this, I don't think any large company or medical fields should let you run it without a function to clean up what you're sending through. So this goes both ways. Yes, it opens up major security holes. But at the same time, it completely changes everything.
Starting point is 01:08:25 And I mean it because up until now, Yeah, you could do a lot of things, but operations were really hard. Scaling was really hard. The thing has a mind of its own. Sometimes it returns stuff in the right structure. Sometimes it don't. And like, how do you build products around those? Because in computer science and DevOps and stuff, like you expect a response,
Starting point is 01:08:46 either an error or in a certain format. And he didn't have that. Well, now you do. So now you have an interface to the whole world. Now you can do stuff with it. Like, this completely changes everything because you can fire it off, you can clean up your data. You're finally free. I agree with everything you've just said.
Starting point is 01:09:12 I completely agree. And yet, the security implications are still terrifying to me. But, yeah, no, I'm super excited. I have things I'm going to build on top of this. But I'm going to be very careful about them. I think the reversibility, especially, like, I want to build things that let people clean up data, like have a conversation with your data to clean it, because everyone who works with data says they spend 90% of the time on cleaning, they hate it.
Starting point is 01:09:36 It's a great, I want to solve that problem. But I want that to be an undo function precisely to protect against some of the things that could go wrong. There's a couple of things, as I wanted to say, like, one is while I shared, like, the excitement, I think that we still have to be careful because, you know, these things are still going to hallucinate, and there's still not a clear way to overcome that, even using functions. That's number one. And number two is that I kind of want to echo what Mayo said before. And while I do appreciate that, again, this is like a great advancement.
Starting point is 01:10:09 It's super cool. I don't know if I would call it a game changer because this has existed, right? Like this has happened in various frameworks, you know, blank chain guidance, you know, other frameworks have guardrails in place. Yes, this looks like a very good implementation of this concept of having guard whales and being able to basically force the LLM to do what you want. And yes, it's very, you know, it's beneficial to all of us that Open AI that owns these models actually put put like some effort into fine-tuning their own model so that this works
Starting point is 01:10:46 really well. But I just, I just want to curb the enthusiasm a little bit, right, to just say like, this isn't like, you know, earth-shattering and like hasn't, we haven't seen. anything like this before. The same could be said about Chachypti though, right? When ChachypT released, folks are like, well, this is just prompting and this is just sending the same like back and forth to the you know text and then yet JGPT was the first product that they got to like a hundred million whatever. Sure sure. Let's see how the adoption goes, right? I mean the second all of us dropped from this
Starting point is 01:11:19 space and then start actually coding this and then we'll come back here and we'll see. Go ahead Simon and I want to hear from Sean and stuff. I just want to say for me until today the functions thing was always a hat right you could get it working on top of language model but you have to do some pretty weird prompt engineering and mucking around to get that to happen and now that it's part of the core platform it feels so much that the friction involved in getting that working feels so much lower I'm not longer afraid of it you know I was kind of cautious of doing this trick in the past because I knew there were so many ways that it might break now that I know the opening I have fine-tuned a model for it I feel a lot more confident in
Starting point is 01:11:55 using it. So I think that makes a big difference. I want to say thanks for everybody here on stage sharing with us, like exploring with us, Simon and Sean and Riley and Liston, and Zanova and Farrell and Mayo and Steph and Rory and so many great folks here discussing kind of these latest changes that are definitely exciting. Potentially, if you're citing with Nistin, groundbreaking and earth shattering, which I tend to agree, and this was the reason for the space with Sean. We discussed this in DM and we brought a great panel of friends
Starting point is 01:12:29 that are going to discuss how big this is. And I think now that we're hitting like an hour and a half. I think we're here. Let's do a fairly quick kind of discussion about what are we building with this new tool that we have? I'll let Steph go and then let's maybe everybody feel free to unmute and kind of have an instruction debate
Starting point is 01:12:45 and then we're kind of close this out because I'm losing my voice. And although this has been fun, it prevents all of us from going and actually playing with these new tools. So Steph, go ahead. and then we have two free forms to jam in and say, what are we building with this? I just wanted to give a shout out to Leon and his team. They just launched Garak,
Starting point is 01:13:04 which is this tool for a security probing for LLMs. And it has an auto red teaming function. So that might be interesting to check out. But one thing that I was thinking about while Simon was talking and I read your blog post on prompt injection. And on one hand, we want to have more people exposing these vulnerabilities. and maybe sharing their code for red teaming and their examples at the same time that is helping
Starting point is 01:13:30 people who want to abuse like these models and functions. So it's a tricky one. I'm curious how you're thinking about it. And I had a parting question as well, which is what are we going to ask next from OpenAI? What is missing and what would make our use of the technology of the API better? This is a really interesting ethical question. It's like all aspects of security. to those people about responsible disclosure of security vulnerabilities. The frustrating thing with prompt injection is we don't have a fix yet. So, you know, normally if you find a SQL injection hole in someone's website, you quietly tell them about it and they patch it and then everything's fine.
Starting point is 01:14:09 With prompt injection, seeing as there is no known way to fix these problems, I kind of feel like it's on us to make sure people understand before they build systems that are vulnerable to them. So that's the approach I've been taking is just really trying to shake people and say, no, you can't just say, oh, we'll filter it out, it'll be fine. That doesn't work yet. You need to sometimes, you need to say, no, I cannot build the feature you're asking me to build because it can't be built securely.
Starting point is 01:14:34 And that's really frustrating. I kind of hate that. Like, I don't want to be the person who says, there's a security hole. You need to stop. I want to be the person who says, there's security hole. Here's the fix for it. And now we can move on with our lives. But sadly, we're not at that point with it yet.
Starting point is 01:14:48 I, like, I, 100% agree. Like, it's been weird to me. like it's gone from like suspicious to like frustrating to like just like curious like like this seems to be such a hard problem like it's a like they're like I think what's going on really is that like to solve this you kind of had to re-engineer it to this message-based API because like they have now reserved tokens they have tokens that only they know exist that they can insert it as like quote marks and they can have things like a system message they can tune it in a way that it actually like you like if you could peer into its brain and see like what circuit is it
Starting point is 01:15:25 implementing it's something that pays attention to the system message in the right way that it like understands the difference between what like the open AI customer told it our instructions and what the user told it and like I think that's progress in the right direction of like that it's like a sensible like it makes sense to me like as an outsider of like that's how you would like go about addressing this but it just like it's a big change and so like I think they're moving towards it but it like speaks just like what a hard problem this is you know because like it is since may i think since like preamble uh where the original like discovers of it you know in may and so when they put it in responsible disclosure and so it is just it seems like it's just a deep like
Starting point is 01:16:03 issue with how like you know the instruction tuning or attention or something about this works that like it's not trivially fixable and yeah i think you know they're they're making peace you know like piece by piece progress towards i think very interesting to see because we did get an upgrade to 3.5 model, right? And Simon, you previously mentioned that there's a difference from a system message, how much the model adheres to it before. It's like way better than 3.5. It's interesting to now test again this new model and see whether that listens to
Starting point is 01:16:35 the message better. Maybe they implemented some of that in this kind of new update to this model. Yeah, they call that steerability. They specifically say that the new models are now more steerable than they used to be. and they talk about stability, that basically means how closely it obeys the system prompt. I'd love to see some examples of that. I've not played around with it yet,
Starting point is 01:16:59 but I'd love to see a few examples of prompts that were easily prompt injected with the previous 3.5, which are now protected against. But I did note that they have not flat out said, this is a solved problem, and until they do, I'm going to assume it's not a solved problem because it's in their interest to solve it
Starting point is 01:17:16 and then tell people they've solved it. so I'm sure that you'll research this and folks Simon has a great blog basically like a pensive from Harry Potter that Simon has since 2013 I looked at it Simon you're very prolific so I recommend following Simon and his log and thoughts is another go ahead and I think after that we'll do like a round of of last parting thoughts yeah I just wanted to bring up this
Starting point is 01:17:38 because I I'm not 100 sure of how they've really implemented it behind the scenes but has anyone here sort of heard of like JSON former, like one of these projects that was sort of aimed at these, at generating structured data. I love that thing. That thing is so clever. Yes, it's wonderful. Yeah. So I mean, I'm just wondering why, so I'm speaking from, because I'm not really sure how they've implemented it behind the scenes, but assuming they are not doing it this way, is there a reason why opening eyes chosen not to because JSON Form is like by definition you cannot there's no such
Starting point is 01:18:20 thing as prompt injection in this case because you only generate tokens that so you I disagree on that price I think Jason Forma solves the problem if you want it to output JSON and the outputs invalid JSON it does but prompt injection in this case is much more about when you summarize a web page does the text from that trick it into then making a valid call to a function that is something you don't want to do but So my hunch on Jason former, I wonder if they just haven't had time to implement it yet. It's quite a tricky thing for them to, because the way JSON form works, people who haven't seen it, is it basically injects extra logic at the next token prediction thing. So it knows that if you're doing JSON, you've just done the curly bracket.
Starting point is 01:19:03 The only token that you come next is a single quote, is a double quote. And then the only things that come after that are not double quotes until you get to the end and cycle. So you can force your model to output, structured, tech. that matches JSON or YAML or whatever. Super clever. But yeah, my guess is that OpenA I just haven't got around to fully implementing that yet and they'll get it working at some point.
Starting point is 01:19:25 Right, yeah. So that's just sort of what I was getting at with the sort of in the bullet. I mean, I'm reading there, read me right here. It's like a bulletproof way to generate a structured data. But so what I mean, this does not cover the case where you, the separation between the user's input and the calling of the function. That is always susceptible to problem detection that that's where the security holes are. But assuming that you are forcing it to generate structured data like JSON, there are these current approaches where, I mean, it's modifying the logits.
Starting point is 01:19:57 I mean, that's how I assume it's currently working where it modifies how the next token is predicted. But so that that's sort of what I'm getting at is like, so why hasn't opening eye done that approach or have they? Or yeah, I'm not 100% sure on that. Yeah, I think the thing about these spaces is we don't have Logan here or anybody from open the eye. And when we do, they don't necessarily able to tell us what they're using. Go ahead, Riley. I was just curious, I mean, if you played with, I think, Grant Slatton, working on like context-free grammar stuff, that, like, I linked a while back, I didn't know how that compared to JSON former. I think it's the same exact trick, just even more, even cooler, because his thing, you can give it any grammar you like to get about JSON.
Starting point is 01:20:39 It's anything that can be specified as a grammar. I think he posted the idea that he'd like to be able to upload a web assembly program to the opening API and say run this to pick the next token, which I think would be freaking incredible. That's absolutely brilliant idea. Yeah, that's cool. That's really cool. Well, maybe just go around and say what we think should be built or page from Steph. What do we want Open AI to ship next? yeah, this is, you know, some people from over there would definitely be listened to this. So
Starting point is 01:21:17 here's your chance to do your pitch for what they should build next. Go ahead, Steph. I was going to say, I would love to have knowledge graphs and be able to have better retrieval. So, you know, this idea of like training with retrieval in mind and yeah, like not necessarily relying on vector databases, like that's something that I would love to see in the future. Simon? I want widgets in chat GPT. I think chats are terrible interface. I would like to be able to build it like a chat GPT plugin that would say,
Starting point is 01:21:52 now show them a map. Now ask them to pick something from this list of options, things like that. Just let us go beyond just having people type text to us. Indeed, indeed. Is it over? I want like a 30B model that's open source from them. I don't mind if it's like a Lama license.
Starting point is 01:22:12 Open source. Okay, got it. Yeah, yeah, just some model you can mess around with. That'd be nice. All the class. Guys, 80B, if you can do that. So, like, what, you know, maybe we should train like GPT 2.5. Like, just, just down it back a little bit.
Starting point is 01:22:29 Zenova or Farrell? I mean, I'm sort of in the open source that I'm coming from hugging the face, so, you know. So, but I, you know, you know, GPT2.5, let's go with that. honestly I'll just I just want 32k GPD4 with 75% or 80% cost reduction come on you get that next month I'm bigger I really hope so I'm good I'm good otherwise I'm very curious to see what we'd be able to build with these new functions and specifically combining them with agents,
Starting point is 01:23:15 I think the opportunities there are pretty amazing. It's definitely going to make life a lot easier. And yeah, of course, if we could have like some more open source models, that would be awesome. And Python is going to have a chance for us to, or at least for folks to test this out, right, Roy? You want to tell us about the hackathon? Oh, sure.
Starting point is 01:23:34 We're holding our first virtual hackathon on June 19th to the 26th, 100K in prizes. It's going to be super fun. and it's going to be held in Kumo space. And we invite you all to attend and show us what you got. 100,000 prizes. That might be the highest I've yet heard for one of these virtual hackathons. That's pretty cool.
Starting point is 01:23:56 Mail and then Alex and I'll go and then I'll have Riley give the last words of mail. Go ahead. Yeah, open source, 100%. I mean, it's good what they've done, but then, you know, my mind just goes to, you know, how can this be applied to open source, right? what they've just pulled off for the fine-tuning. How can we apply to open source?
Starting point is 01:24:17 So, yeah, I'd love for them to start to be open as it says in their name, right? Yeah, they should just rename at this point. Oh, oh, well, okay. So, yeah, obviously, you know, you want things for free. They're not going to give it to you. End of story. I'm just very interested in Franken models. I'm very interested in what Simon has been sketching out in this space,
Starting point is 01:24:42 which is essentially using them to call still smaller models, but that do very specific things, but they can do a lot of things. And so I'm interested in essentially just kind of rewriting my developer agent to do that kind of routing and explore the possibility of recursion. And when I have more information about that, I'll report back. That's great. I want to just, before I go, I want to call out, John has a podcast called Layton Space.
Starting point is 01:25:12 So definitely check it out. Oh, I'll be posting the recording of this. Yeah. I mean, this is great. Everybody chipped in, yeah, you have the developer perspective. This is what we want. Oh, and we also are going to drop our first, I think the first ever interview with George Hots on Tiny Corp. And I'm very excited about that one. So make sure.
Starting point is 01:25:32 Yeah. Make sure you don't miss that because MD is lagging behind Nvidia and George is working in that space. And yeah. Potentially some exciting things to come. He had this email Lisa Sue and went back and forth. It was very dramatic. Nothing when George is boring. Yeah.
Starting point is 01:25:47 So, Riley, go ahead. What would you want from over there? I think, like, the thing that, you know, about this whole release is that, like, it's, as coolest functions are, like, there, I think the real thing that might be more important in the end is, is just this march towards lower prices and bigger context windows. I think there's a lot of unexplored stuff to do with big context. And there's a lot of, a lot of possibilities that are opened up when things are just cheaper. And you can, like, put in redundancy checks.
Starting point is 01:26:13 You can have, like, secondary prompts that, check the work of the first prompt. You can like, you know, engineer reliability around the parts that you need. So, I mean, it's hard to overstate just like how good it is that just the stuff becoming cheaper. And Riley, I saw that there's a webinar coming up and then you're going to teach advanced prompt engineering. You want to talk about this for a second? Oh, yeah, sure. So scale is on July 15th. And I think we've already closed applications, unfortunately, but on one on July 15th, that we're having a hackathon and I'll be giving a talk on prompt engineering there. And I, Last time we did this, we did like a replay, like I did the talk again for our webinar.
Starting point is 01:26:48 So I expect we'll probably be doing that again. So I want to thank everyone here coming up to stage. Here's my request to open the eye. I want the Vision API as fast as possible. Oh, yes. I've been mouthwatering on the Vision API. Today, Mikhail Perakin from Bing confirmed that already like 10% of Bing users get access to the GPT4 Vision API. and I participated in an interview with the founder of Be My Eyes,
Starting point is 01:27:16 who are currently the only people in the world who have access to Vision API, and that was a great conversation. And I expect amazing things once that dropped for the ability of a GPT4 to understand the real world, not to mention how many prompt injections we can do via text. But that's a conversation for another time with Simon and Riley. But I definitely want to thank everyone for coming here. This has probably been the biggest space that I've ran. So thank you, Sean, for prompting this and everybody who joined.
Starting point is 01:27:45 Hey, yeah. And now we're going to have some time to go and play with these models and new techniques. And hopefully we'll see you guys again. The last plug, I'll say that I'll do, I'm doing spaces every Thursday. Many of the folks will stay here join. We talk about latest updates. This was an emergency one. And glad to hear that channel is going to compile this into a podcast.
Starting point is 01:28:03 So definitely subscribe to Laken Spaces. But everybody, thank you for joining. and go play with the new tools we got. Let's go build.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.