Latent Space: The AI Engineer Podcast - From API to AGI: Structured Outputs, OpenAI API platform and O1 Q&A — with Michelle Pokrass & OpenAI Devrel + Strawberry team

Episode Date: September 13, 2024

Congrats to Damien on successfully running AI Engineer London! See our community page and the Latent Space Discord for all upcoming events.This podcast came together in a far more convoluted way than ...usual, but happens to result in a tight 2 hours covering the ENTIRE OpenAI product suite across ChatGPT-latest, GPT-4o and the new o1 models, and how they are delivered to AI Engineers in the API via the new Structured Output mode, Assistants API, client SDKs, upcoming Voice Mode API, Finetuning/Vision/Whisper/Batch/Admin/Audit APIs, and everything else you need to know to be up to speed in September 2024.This podcast has two parts: the first hour is a regular, well edited, podcast on 4o, Structured Outputs, and the rest of the OpenAI API platform. The second was a rushed, noisy, hastily cobbled together recap of the top takeaways from the o1 model release from yesterday and today.Building AGI with Structured Outputs — Michelle Pokrass of OpenAI API teamMichelle Pokrass built massively scalable platforms at Google, Stripe, Coinbase and Clubhouse, and now leads the API Platform at Open AI. She joins us today to talk about why structured output is such an important modality for AI Engineers that Open AI has now trained and engineered a Structured Output mode with 100% reliable JSON schema adherence. To understand why this is important, a bit of history is important:* June 2023 when OpenAI first added a "function calling" capability to GPT-4-0613 and GPT 3.5 Turbo 0613 (our podcast/writeup here)* November 2023’s OpenAI Dev Day (our podcast/writeup here) where the team shipped JSON Mode, a simpler schema-less JSON output mode that nevertheless became more popular because function calling often failed to match the JSON schema given by developers. * Meanwhile, in open source, many solutions arose, including * Instructor (our pod with Jason here) * LangChain (our pod with Harrison here, and he is returning next as a guest co-host)* Outlines (Remi Louf’s talk at AI Engineer here)* Llama.cpp’s constrained grammar sampling using GGML-BNF* April 2024: OpenAI started implementing constrained sampling with a new `tool_choice: required` parameter in the API* August 2024: the new Structured Output mode, co-led by Michelle* Sept 2024: Gemini shipped Structured Outputs as wellWe sat down with Michelle to talk through every part of the process, as well as quizzing her for updates on everything else the API team has shipped in the past year, from the Assistants API, to Prompt Caching, GPT4 Vision, Whisper, the upcoming Advanced Voice Mode API, OpenAI Enterprise features, and why every Waterloo grad seems to be a cracked engineer.Part 1 Timestamps and TranscriptTranscript here.* [00:00:42] Episode Intro from Suno* [00:03:34] Michelle's Path to OpenAI* [00:12:20] Scaling ChatGPT* [00:13:20] Releasing Structured Output* [00:16:17] Structured Outputs vs Function Calling* [00:19:42] JSON Schema and Constrained Grammar* [00:20:45] OpenAI API team* [00:21:32] Structured Output Refusal Field* [00:24:23] ChatML issues* [00:26:20] Function Calling Evals* [00:28:34] Parallel Function Calling* [00:29:30] Increased Latency* [00:30:28] Prompt/Schema Caching* [00:30:50] Building Agents with Structured Outputs: from API to AGI* [00:31:52] Assistants API* [00:34:00] Use cases for Structured Output* [00:37:45] Prompting Structured Output* [00:39:44] Benchmarking Prompting for Structured Outputs* [00:41:50] Structured Outputs Roadmap* [00:43:37] Model Selection vs GPT4 Finetuning* [00:46:56] Is Prompt Engineering Dead?* [00:47:29] 2 models: ChatGPT Latest vs GPT 4o August* [00:50:24] Why API => AGI* [00:52:40] Dev Day* [00:54:20] Assistants API Roadmap* [00:56:14] Model Reproducibility/Determinism issues* [00:57:53] Tiering and Rate Limiting* [00:59:26] OpenAI vs Ops Startups* [01:01:06] Batch API* [01:02:54] Vision* [01:04:42] Whisper* [01:07:21] Voice Mode API* [01:08:10] Enterprise: Admin/Audit Log APIs* [01:09:02] Waterloo grads* [01:10:49] Books* [01:11:57] Cognitive Biases* [01:13:25] Are LLMs Econs?* [01:13:49] Hiring at OpenAIEmergency O1 Meetup — OpenAI DevRel + Strawberry teamthe following is our writeup from AINews, which so far stands the test of time.o1, aka Strawberry, aka Q*, is finally out! There are two models we can use today: o1-preview (the bigger one priced at $15 in / $60 out) and o1-mini (the STEM-reasoning focused distillation priced at $3 in/$12 out) - and the main o1 model is still in training. This caused a little bit of confusion.There are a raft of relevant links, so don’t miss:* the o1 Hub* the o1-preview blogpost* the o1-mini blogpost* the technical research blogpost* the o1 system card* the platform docs* the o1 team video and contributors list (twitter)Inline with the many, many leaks leading up to today, the core story is longer “test-time inference” aka longer step by step responses - in the ChatGPT app this shows up as a new “thinking” step that you can click to expand for reasoning traces, even though, controversially, they are hidden from you (interesting conflict of interest…):Under the hood, o1 is trained for adding new reasoning tokens - which you pay for, and OpenAI has accordingly extended the output token limit to >30k tokens (incidentally this is also why a number of API parameters from the other models like temperature and role and tool calling and streaming, but especially max_tokens is no longer supported).The evals are exceptional. OpenAI o1:* ranks in the 89th percentile on competitive programming questions (Codeforces),* places among the top 500 students in the US in a qualifier for the USA Math Olympiad (AIME),* and exceeds human PhD-level accuracy on a benchmark of physics, biology, and chemistry problems (GPQA).You are used to new models showing flattering charts, but there is one of note that you don’t see in many model announcements, that is probably the most important chart of all. Dr Jim Fan gets it right: we now have scaling laws for test time compute, and it looks like they scale loglinearly.We unfortunately may never know the drivers of the reasoning improvements, but Jason Wei shared some hints:Usually the big model gets all the accolades, but notably many are calling out the performance of o1-mini for its size (smaller than gpt 4o), so do not miss that.Part 2 Timestamps* [01:15:01] O1 transition* [01:16:07] O1 Meetup Recording* [01:38:38] OpenAI Friday AMA recap* [01:44:47] Q&A Part 2* [01:50:28] O1 Demos This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

Transcript
Discussion (0)
Starting point is 00:00:04 Michelle Popper, shine the mic leading the way. Open me, IPI, making waves every day. Pre-chat, Cheapy T, she was ahead of the game. Scaling challenges hit, but she stayed in the lane. J's on mode, rush and structure, up as tight. Building AGII with that clear its site. Refusal feel tricky, HCT, is out of day, chat, and those calling functions, it's the future we're met.
Starting point is 00:00:23 From agents built with structured dreams to find two models and API streams, batch it up, scale it wide, voice mode whispering. Welcome back. This is Charlie, your AI co-host. Michelle Pocras built massively scalable platforms at Google, Stripe, Coinbase and Clubhouse, and now leads the API platform at OpenAI. She joins us today to talk about why structured output is such an important modality for AI engineers that OpenAI has now trained and engineered a structured output mode
Starting point is 00:01:05 with 100% reliable JSON schema adherence. To understand why this is a new thing. is important, we have to go all the way back to June of last year when OpenAI first added a function calling capability to GPT 40613 and GPT 3.5 Turbo0613, which was then followed by November's Dev Day, where the team shipped JSON mode, a simpler schemerless JSON output mode that nevertheless became more popular because function calling often failed to match the JSON schema given by developers. Meanwhile, in open source, many solutions arose, including instructor and Langchain,
Starting point is 00:01:46 from our former guests Jason Liu and Harrison Chase, who by the way is returning to co-host and agents episode soon, and outlines from World's Fair Speaker Remy Loof and Lama.cp's constrained grammar sampling using the GGML extension of the Bacchus Nauer form or BNF syntax. Fast forward to April of 2024, Open AI started implementing constrained sampling with a new tool choice required parameter in the API, and finally in August closed the loop by releasing the new structured output mode, which extends constrained sampling with specific post-training to improve the performance of complex JSON schema following, especially with the strict true flag. The other big labs seem to be following suit with Gemini shipping structured outputs and Enum mode in this past month. We sat down with Michelle to talk through every part of the process
Starting point is 00:02:45 as well as quizzing her for updates on everything else the API team has shipped in the past year. From the Assistance API to Prompt Caching, GPT4. Vision, Whisper, the upcoming Advanced Voice Mode API, Open AI Enterprise features and why every Waterloo grad seems to be a cracked engineer. In latent space community news, if you're in Germany,
Starting point is 00:03:11 the first meetup for AI engineers is happening in Cologne in two weeks. And the second AI engineer summit, the curated invite-only conference run by the AI engineer world's fair team, is now moving to January in New York City. See the show notes for details. Watch out and take care.
Starting point is 00:03:28 Hey everyone, welcome to the Latenspace podcast. This is Alessio, partner, and CCO in residence and decibel partners, and I'm joined by my co-hostess Wix, founder of Small A.I. Hey, and today we're excited to be in the in-person studio with Michelle. Welcome. Thanks. Thanks for having me. Very excited to be here. This has been a long time coming.
Starting point is 00:03:51 I've been following your work on the API platform for a little bit, and I'm finally glad that we could make this happen after you ships structured outputs. How does that feel? Yeah, it feels great. We've been working on it for quite a while. So very excited to have it out there and have people using it. Well, tell the story soon. But I want to give people a little intro to your backgrounds.
Starting point is 00:04:11 So you've interned and worked at Google Stripe, Coinbase, Clubhouse, and obviously Open AI. What was that journey like? The one that has the most appeal to me is Clubhouse because that was a very, very hot company for a while. Basically, you seem to join companies when they're about to scale up really a lot. And obviously, OpenEI has been the latest. But yeah, just what are your learnings and your history going into all these? these notable companies. Yeah, totally.
Starting point is 00:04:35 For a bit of my background, I'm Canadian. I went to the University of Waterloo, and there you do, like, six internships as part of your degree. So I started, actually, my first job was really rough. I worked at a bank, and I learned visual basic, and I, like, animated bond yield curves, and it was, you know, not. Me too. Oh, really?
Starting point is 00:04:52 Yeah, that was a derivative trader. Interest rate swaps, that kind of stuff. Yeah. Yeah. So I liked, you know, having a job, but I didn't love that job. And then my next internship was Google, and I learned so much there. It was tremendous. But I had a bunch of friends that were into startups more, and Waterloo is like a big startup culture.
Starting point is 00:05:08 And one of my friends interned at Stripe, and he said it was super cool. So that was kind of my – I also was a little bit into crypto at the time. Then I got into it on hacker news. And so Coinbase was on my radar. And so that was like my first real startup opportunity was Coinbase. I think I've never learned more in my life than in the four-month period when I was interning at Coinbase. They actually put me on call. I worked on like the ACH Rails there, and it was absolutely crazy.
Starting point is 00:05:32 you know, crypto was a very formative experience. Yeah. This is 2018 to 2020, kind of like the first big wave. That was my full time. I was there as an intern in 2016. Yeah. And so that was the period where I really learned to become an engineer. I learned how to use Git, got on call right away, you know, managed production databases
Starting point is 00:05:50 and stuff. So that was super cool. After that I went to Stripe and kind of got a different flavor of payments on the other side. Learned a lot. I was really inspired by the Colson's. And then my next internship after that, I actually started a company. at Waterloo. So there's this thing you can do. It's an entrepreneurship co-op. And I did it with my roommate.
Starting point is 00:06:07 The company is called Readwise, which still exists. Yeah, yeah. Everyone uses Reithes. Yeah. You co-founder ReWise? Yeah. How about premium user? It's not even on your LinkedIn? Yeah, I mean, I only worked on it for about a year. And so Tristan and Dan are the real founders. And I just had an interlude there. But yeah, I really loved working on something very startup focused, user-focused, and hacking with friends. It was super fun. Eventually, I decided to go back to Coinbase and really, like, get a lot better as an engineer. I didn't feel like I was, you know, didn't feel equipped to be a CTO of anything at that point. And so just learned so much
Starting point is 00:06:41 at Coinbase. And that was a really fun curve. But yeah, after that, I went to Clubhouse, which was like a really interesting time. So I wouldn't say that I went there before it blew up. I would say I went there as it blew up. So not quite the Starling track record that it might seem. But it was a super exciting place. I joined as like the second or third back in engineer. And, you know, we were down every day, basically. One time Oprah came on and absolutely everything melted down. And so we would have a stand up every morning and be like, how do we make everything stay up? Which is super exciting.
Starting point is 00:07:11 Also, one of the first things I worked on there was making our notifications go out more quickly. Because when you join a clubhouse room, you know, you need everyone to come in right away so that it's exciting. And the person speaking thinks a lot of my audiences here. But when I first joined, I think it would take like 10 minutes for all the notifications to go, which is insane. Like, you know, by the time you want to start talking to the time your audience is. there. It's like you can totally kill the room. So that's one of the first things I worked on is making that a lot faster and, you know, keeping everything up. I mean, so already we have an audience of engineers. Those two things are useful. It's keeping things up and notifications out. Notifications,
Starting point is 00:07:43 like, is it a Kafka topic? It was a Postgres shop and you had all of the followers in Postgres and you needed to like iterate over the followers and like figure out, is this a good notification to send. And so all of this logic, it wasn't like well batched and parallelized and our job queuing infrastructure wasn't right. And so there's a lot of like fixing all. all of these things. Eventually, there were a lot of database migrations because Postgres just wasn't scaling well for us. Interesting. And then keeping things up, that was more of a, I don't know, reliability issue, SRE type. A lot of it, yeah, this goes down to like database stuff. Everywhere I've worked. It's all databases. Yeah. Yeah.
Starting point is 00:08:19 Actually, at Coinbase at Clubhouse and at Open AI, Postgres has been a perennial challenge. It's like the stuff you learn at one job carries over to all the others because you're always debugging a long running PostQuest career at 3am for some reason. So those skills have really carried me forward, for sure. What do you think that not as much of this is practiced? Obviously, Postgres is an open source project that's not aimed at this like giga scale, but you would think somebody would come around and say, hey, we're like the... Yeah, I think that's what Planet Scale is doing.
Starting point is 00:08:49 It's not on Postgres, I think. It's on MySQL. But I think that's the vision. It's like they have zero downtime, zero downtime migrations, and that's a big pain point. I don't know why no one is doing this on Postgres, but I think it would be pretty cool. Their connection pullers, like PG Bouncer is like good enough? Yeah. Well, even, I mean, I've run PG Bouncer everywhere and there's still a lot of problems. Your scale, it's something that not many people see, so. Yeah. I mean, at some point, every company gets the scale, every successful company gets the scale where Postgres is not
Starting point is 00:09:17 cutting it. And then you migrate to some sort of no-SQL database. And that process I've seen happen a bunch of times now. MongoDB, Redis, something like that. Yeah. I mean, we're on Azure now. And so there's, we use Cosmos. Cosmos DVD. At Clubhouse, I really love DynamoDB. That's probably my favorite database, which is like a very nerdy sense. But that's the one I'm using if I need to scale something as far as it goes. Yeah, DynamoDB. When I learned, I worked at AWS briefly.
Starting point is 00:09:43 And it's kind of like the memory register for the web. Yes. You know, if you treat it just as physical memory, you will use it well. If you treat it as a real database, you might run into problems. Right. You have to totally change your mindset when you're going from Postgres to Dynamo. But I think it's a good mindset shift and kind of makes you. you design things in a more scalable way.
Starting point is 00:10:01 Yeah, I'll recommend the DynamoDB book for people who need to use DynamoDB. But we're not here to talk about AWS, we're here to talk about Open AI. You joined Open AI pre-ChadGBT. I also had the option to join, and I didn't. What was your insight? Yeah. I think a lot of people who joined Open AI, joined because of a product that really gets them. I'm excited, and for most people, it's Chad GPT.
Starting point is 00:10:20 But for me, I was a daily user of co-pilot, GitHub co-pilot. And I was, like, so blown away at the quality of this thing. I actually remember the first time seeing it on Hacker News and being like, wow, this is absolutely crazy. Like, this is going to change everything. And I started using it every day. It just really, even now when, like, I don't have service and I'm coding without co-pilot, it's just like 10x difference.
Starting point is 00:10:41 So I was really excited about that product. I thought now is maybe the time for AI. And I had done some AI in college and thought some of those skills would transfer. And I got introduced the team. I liked everyone I talked to. So I thought it would be cool. Why didn't you join? It was like, I was like, is Dolly It?
Starting point is 00:10:56 We were there. We were at the Dolly. We were at the Dali, like, launch thing. And I think you were talking with Lenny, and Lenny was at opening out at the time. And you were like, we don't have to go into too much detail. Yeah, yeah, no, no. But it was one of my biggest regrets of my life, I think. No, no, no.
Starting point is 00:11:10 But, but I was like, okay, I mean, I can create images. I don't know if, like, this is the thing to dedicate. But obviously, you had a bigger vision than I did. Dolly was really cool, too. I remember, like, first showing my family, I was like, I'm going to this company, and here's, like, one of the things they do. And it, like, really held bridge the gap. Whereas, like, I still haven't figured out how to explain to my parents what crypto is.
Starting point is 00:11:31 My mom for a while thought I worked at Bitcoin. So it's like, it's pretty different to be able to tell your family what you actually do and they can see it. And they can use it too, personally. So you were there, were you immediately on API platform? You were there for the chat GPT moment. Yeah, I mean, API platform is like a very grandiose term for what it was. There was like just a handful of us working on the API. Yeah, it was like a closed beta, right?
Starting point is 00:11:52 Not even everyone had access to the GP3 model. A very different access model than a lot more like tiered rollouts. But yeah, I would say the applied team was maybe like 30 or 40 people and yeah, probably closer to 30 and there's maybe like five-ish total working on the API at most. So yeah, we've grown a lot since then. It's like 60-70 now, right? No, applied is much bigger than that. Applied now is bigger than the company when I joined.
Starting point is 00:12:15 Okay. Yeah, we've grown a lot. I mean, there's so much to build. So we need all the help. I'm a little out of date. Yeah. Any chat GBT release, kind of like all hands on deck stories, I had lunch with Evan Morikawa a few months ago. It sounded like it was a fun time to get, build the APIs and have all these people trying to use the web thing.
Starting point is 00:12:32 Like, how are you prioritizing internally? Like, what was the helping scaling when you're scaling non-GPU workloads versus like Postgres bouncer and things like that? Yeah, actually, surprisingly, there were a lot of Postgres issues when ChatGBT came out because the accounts for like ChatGPT were tied to the accounts. accounts in the API. And so you're basically creating a developer account to log into chat GPT at the time, because it's just what we had. It was low-key research preview. And so I remember there was just so much work scaling like our authorization system. And that would be down a lot. Yeah, also GPU, you know, I never had worked in a place where you couldn't just scale the thing up. It's like everywhere I've worked in compute is like free and you just like auto-scale a thing
Starting point is 00:13:11 and you like never think about it again. But here we're having like tough decisions every day. We're like discussing like, you know, should they go here or here? And we have to be principled about it. So that's a real mindset shift. So you just really structured outputs. Congrats. You also wrote the blog post for it, which was really well written. And I love all the examples that you put out. Like, it really gives the full story.
Starting point is 00:13:28 Yeah, tell us about the whole story from beginning to end. Yeah. I guess the story we should rewind quite a bit to Deb Day last year. Deb day last year, exactly, we shipped JSON mode, which is our first foray into this area of product. So for folks who don't know, JSON mode is this functionality you can enable in our chat completions and other APIs, where if you opt in, we'll kind of constrain the output of the model to match
Starting point is 00:13:50 the JSON language. And so you basically will always get something in a curly brace. And this is good. This is nice for a lot of people. You can like describe your schema what you want in prompt and the, you know, we'll constrain it to JSON. But it's not getting you exactly where you want because you don't want the model to kind of make up the keys or like match different values than what you want. Like if you want an enum or a number and you get a string instead, it's like pretty frustrating. So we've been ideating on this for a while and like people have been asking for basically this every time I talk to customers for maybe the last year. So it was a very time. really clear that there's a developer need, and we started working on kind of making it happen.
Starting point is 00:14:24 And this is a real collab between engineering and research, I would say. And so it's not enough to just kind of constrain the model. I think of that as the engineering side, whereas basically you mask the available tokens that are produced every time to only fit the schema. And so you can do this engineering thing, and you can force the model to do what you want, but you might not get good outputs. And sometimes with JSON mode developers have seen that our models output like white space for a really long time where they don't. Because it's a legal character. Right.
Starting point is 00:14:51 It's legal for JSON, but it's not really what they want. And so that's what happens when you do a very engineering biased approach. But the modeling approach is to also train the model to do more of what you want. And so we did these together. We trained a model which is significantly better than our past models at following formats. And we did the entwork to serve like this constrained decoding concept at scale. So I think marrying these two is why this feature is pretty cool. You just mentioned starts and ends with a curly brace and maybe people's minds
Starting point is 00:15:17 go to pre-fills in the Cloud API. How should people think about JSON mode, structured output, pre-fills? Because some of them are like roughly starts with a curly brace and ask you for JSON, you should do it. And then instructor is like, hey, here's the rough data schema you should use. And how do you think about them? So I think we kind of designed structured outputs to be the easiest to use. So you just, like the way you use it in our SDK, I think is my favorite thing.
Starting point is 00:15:42 So you just create like a Pidentic object or a Zod object and you pass it in and you get back an object. and so you don't have to deal with any of the serialization. With the parse helper. Yeah. You don't have to deal with any of the serialization on the way in or out. So I kind of think of this as the feature for the developer who is like, I need this to plug into my system. I need the function call to be exact. I don't want to deal with any parsing.
Starting point is 00:16:02 So that's where structured outputs is tailored. Whereas if you want the model to be more creative and use it to come up with a JSON schema that you don't even know you want, then that's kind of where JSON mode fits in. But I expect most developers are probably going to want to upgrade to structured outputs. The thing you just said, you just use interchangeable terms for the same thing, which is two function calling and structured outputs. We've had disagreements or discussion before on the podcast about are they the same thing. Semantically, they're slightly different.
Starting point is 00:16:31 They are, yes. Because I think function calling API came out first, then JSON mode. And we used to abuse function calling for JSON mode. Right. Do you think we should treat them as anonymous? No. Okay. Yeah, please clarify.
Starting point is 00:16:44 Yeah. And by the way, there's also tool calling. Yeah. The history here is we started with function calling, and function calling, you know, came from the idea of like, let's give the model access to tools and let's see what it does. And we basically had these internal prototypes of what code interpreter is now. And we're like, this is super cool. It's making an API. But we're not ready to host code interpreter for everybody. So, you know, we're just going to expose the raw capability and see what people do with it. But even now, I think there's a really big difference between function calling and structured outputs. So you should use function calling when you actually have functions that you want the model to call. And so, like, if you have a database that you want the model to be able to query from, or if you want the model to send an email or, like, you know, generate arguments for an actual action. And that's the way the model has been, like, fine-tuned on is to, like, treat function calling for actually calling these tools and getting their outputs. The new response format is a way of just getting the model to respond to the user, but in a structured way. And so this is very different, like, responding to a user versus, like, you know, I'm going to go send an email. A lot of people were hacking function calling to get the response format they needed.
Starting point is 00:17:50 And so this is why we shipped kind of this new response format. So you can get exactly what you want and you get kind of more of the models for BOSID. It's like kind of responding in the way it would speak to a user. And so less kind of just programmatic tool calling, if that makes sense. Are you building something into the SDK to actually close the loop with the function calling? Because right now it returns the function, then you got to run it, then you got to like fake another message to then continue the conversation. that in beta, the runs. Yes, we have this in beta in the Node SDK.
Starting point is 00:18:20 So you can basically define... Oh, no Python. It's coming to Python as well. That's why I didn't know. Yeah, I'm a Node guy, so he's probably... The JavaScript mine is to advance. It's already existed. It's coming everywhere.
Starting point is 00:18:31 But basically what you do is you write a function and then you add a decorator to it. And then you can basically, there's this run tools method. And it does the whole loop for you, which is pretty cool. When I saw that in the Node SDK, I wasn't sure if that's... Because it basically runs it. in the same machine. Yeah. And maybe you don't want that to happen.
Starting point is 00:18:49 Yeah, I think of it as like, if you're prototyping and building something really quickly and just playing around, it's so cool to just create a function and give it this decorator. But, you know, you have the flexibility to do it however you like. Like, you don't want it in a critical path of a web request. I mean, some people definitely will. You know, it's just kind of the easiest way to get started. Yeah. But let's say you want to like execute this function on a job QA sync, then, you know, it wouldn't
Starting point is 00:19:11 make sense to use that. Prior art. instructor outlines JSON former what did you study what did you you you know credit or learn from these things yeah there's a lot of different approaches to this there's more fill in the blank style sampling where you basically pre-form kind of the keys and then get the model to sample just the value there's kind of a lot of approaches here we didn't kind of use any of them wholesale but we really loved what we saw from the community and like the developer experiences we saw so that's where we took a lot of inspiration.
Starting point is 00:19:42 There was a question also just about constrained grammar. This is something that I first saw in Lama CPP, which seems to be the most, let's just say, academically permissive form of the lowest level. Yeah. For those who don't know, maybe, I don't know if you want to explain it, but they use back as nor form, which you only learned in, like, college when you're working on programming languages and compilers. I don't know if you use that under the hood or you explore that.
Starting point is 00:20:05 Yeah, we didn't use any kind of other stuff. We kind of built, you know, our solution from scratch to meet. our specific needs. But I think there's a lot of cool stuff out there where you can supply your own grammar. Right now we only allow JSON schema and a dialect of that. But I think in the future,
Starting point is 00:20:20 it could be a really cool extension to let you supply a grammar more broadly. And maybe it's more token efficient than JSON. So a lot of opportunity there. You mentioned before also training the model to be better function calling. What's that discussion like internally for like resources? It's like, hey, we need to get better JSON mode.
Starting point is 00:20:38 And it's like, well, can you figure it out on the API platform without, touching the model? Like, is there a really tight collaboration between the two teams? Yeah, so I actually work on the API models team. I guess we didn't quite get into what I do at API. What do you say it is you do here? Yeah, so yeah, I'm the tech lead for the API, but also I work on the API models team, and this team is really working on making the best models for the API. And a lot of common deployment patterns are research makes a model, and then you kind of ship it in the API. But, you know, I think there's a lot you miss when you do that. You
Starting point is 00:21:10 miss a lot of developer feedback and things that are not kind of immediately obvious. What we do is we get a lot of feedback from developers and we go and make the models better in certain ways. So our team does model training as well. We work very closely with our post-training training team. And so for structured outputs, it was a collab between a bunch of teams, including safety systems to make, you know, a really great model that does structured outputs. Mentioning safety systems, you have a refusal field. Yes. You want to talk about that? Yeah. Yeah, it's a little, it's pretty interesting. So you can imagine, basically, if you constrain the model to follow a schema,
Starting point is 00:21:43 you can imagine there being like a schema supplied that it would add some risk or be harmful for the model to kind of follow that schema. And we wanted to preserve our model's abilities to refuse when something doesn't match our policies or is harmful in some way. And so we needed to give the model an ability to refuse even when there is this schema. But also, you know, if you are a developer
Starting point is 00:22:05 and you have this schema and you get back something that doesn't match it, you're like, ah, the feature's broken. So we wanted a really clear way for developers to program against this. So if you get something back in the content, you know it's valid. It's JSON parsable. But if you get something back in the refusal field, it makes for a much better UI for you to kind of display this to your user in a different way. And it makes it easier to program against. So really, there was a few goals, but it was mainly to allow the model to continue to refuse, but also with a really good developer experience.
Starting point is 00:22:30 Yeah. Why not offer it as like an error code? Because we have to display error codes anyway. Yeah, we falafled for a long time about API design. as we are want to do. And there are a few reasons against an error code. Like, you can imagine this being a 4X error code or something. But, you know, the developer is paying for the tokens.
Starting point is 00:22:49 And that's kind of atypical for like a 4XX error code. We pay with errors anyway, right? So 4Xs is... That's a U error. Right. And it doesn't make sense as a 5XX either because it's not our fault. It's the way the API model is designed. I think the HTTP spec is a little bit limiting
Starting point is 00:23:09 for AI in a lot of ways. Like there are things that are in between your fault and my fault. There's kind of like the model's fault and there's no, you know, error code for that. So we really have to kind of invent a lot of the paradigm here. Make it 6XX. Yeah, that's one option. There's actually some like esoteric error codes we've considered adopting. We're still figuring that out.
Starting point is 00:23:28 But I think there are some things like, for example, sometimes our model will produce tokens that are invalid based on kind of our language. and when that happens, it's an error. But, you know, it doesn't, 500 is fine, which is what we return, but it's not as expressive as it could be. So, yeah, just areas where, you know, Web 2.0 doesn't quite fit with AI yet. If you had to put in a spec, to just change. Yeah, yeah, yeah. What would be your number one proposal to, like, rehall?
Starting point is 00:23:56 The HTTP committee to reinvent the world. Yeah, that's going to. I mean, I think we just need an error of, like, a range of model error. and we can have many different kinds of model errors. Like a refusal is a model error. 601, auto refusal. Yeah. Again, so we've mentioned before that chat completions uses this chat ML format.
Starting point is 00:24:16 So when the model doesn't follow chat ML, that's an error. And we're working on reducing those errors, but that's like, I don't know, 602, I guess. A lot of people actually don't no longer know what chat ML is. Yeah, because that was briefly introduced by OpenE Eye and kind of deprecated. Everyone who implements this underwood knows it. but maybe the API users don't know it. Basically, the API started with just one endpoint, the completions endpoint. And the completions endpoint, you just put text in and you get text out.
Starting point is 00:24:43 And you can prompt in certain ways. Then we released ChatGBT, GBT, and we decided to put that in the API as well. And that became the Chat Completions API. And that API doesn't just take like a string input and produce an output. It actually takes in messages and produces messages. And so you can get a distinction between like an assistant message and a user message, and that allows all kinds of behavior. And so the format under the hood for that is called chat ML.
Starting point is 00:25:07 Sometimes, you know, because the model is so out of distribution based on what you're doing, maybe the temperature is super high, then it can't follow chat ML. Yeah. I didn't know that there could be errors generated there. Maybe I'm not asking challenging enough questions. It's pretty right. And we're working on driving it down. But actually, this is a side effect of structured outputs now, which is that we have removed a class of errors.
Starting point is 00:25:28 We didn't really mention this in the blog just because we ran out of space. but what we're here to do. Yeah, the model used to occasionally pick a recipient that was invalid, and this would cause an error. But now we are able to constrain to chat ML in a more valid way, and this reduces a class of errors as well. Recipient meaning, so there's this like a few number of defined roles, like user, assistant, system. So like recipient as in like picking the right tool. So the model before was able to hallucinate a tool, but now it's, I can't when you're using structured outputs. Do you collaborate with other model developers to try and figure out this type of errors, like how do you display them?
Starting point is 00:26:07 Because a lot of people try to work with different models. Yeah. Is there any? Yeah, not a ton. We're kind of just focused on making the best API for developers. A lot of research and engineering, I guess, comes together with e-vals. You publish some e-vals there. I think Gorilla is one of them.
Starting point is 00:26:24 What is your assessment of the state of e-vals for function calling and structured output right now? Yeah, we've actually collaborated. with BFCL a little bit, which is, I think, the same thing gets your role. The Berkeley function calling leaderboard. Kudos to the team. Those evils are great, and we use them internally. Yeah, we've also sent some feedback on some things that are misgraded, and so we're collaborating to make those better.
Starting point is 00:26:45 In general, I feel evals are kind of the hardest part of AI. Like, when I talk to developers, it's so hard to get started. It's really hard to make a robust pipeline. And you don't want evals that are, like, 80% successful because, you know, things are going to prove dramatically. And it's really hard to craft the right e-value. You kind of want to hit everything on the difficulty curve. I find that a lot of these e-vowls are mostly saturated, like for BFCL.
Starting point is 00:27:09 All the models are near the top already, and kind of the errors are more, I would say, like, just differences in default behaviors. I think most of the models on the leaderboard can kind of get 100% with different prompting, but it's more kind of you're just pulling apart different defaults at this point. So yeah, I would say in general we're missing evals. You know, we work on this a lot internally, but it's a lot. hard. Other than BFCL, would you call out any others just for people exploring this space? So, Eventch is actually like a very interesting email if people don't know.
Starting point is 00:27:38 You basically give the model GitHub issue and like a repo and just see how well it does the issue, which I think is super cool. It's kind of like an integration test, I would say, for models. It's a little unfair, right? What do you mean? A little unfair because like usually as a human you have more opportunity to like ask questions about what it's supposed to do. And you're giving the model like way too little information.
Starting point is 00:27:58 It's a hard job. to do the job. But yeah, so we bench targets like, how well can you follow the diff format and how well can you like can you write code. So I'm really excited about e-vals like that because the pass rate is low, so there's a lot of room to improve. Yeah. And it's just targeting a really cool capability. I've seen other evils for function calling where I think might be BFCL as well, where they evaluate different kinds of function calling. And I think the top one that people care about for some reason, I don't know personally that this is so important to me, but it's parallel function calling.
Starting point is 00:28:28 I think you confirm that you don't support that yet. Why is that hard? There's more context about it. So yeah, we put out parallel function calling a dev day last year as well. And it's kind of the evolution of function calling. So function calling V1, you just get one function back. Function calling V2, you can get multiple back at the same time
Starting point is 00:28:45 and save latency. We have this in our API, all our models support it, or all of our newer models support it. But we don't support it with structured outputs right now. And there's actually a very interesting tradeoff here. So when you basically call our API for structured outputs with a new schema, we have to build this artifact for fast sampling later on. But when you do parallel function calling, the kind of schema we follow is not just directly one of the function schemas. It's like this combined schema based on a lot of them.
Starting point is 00:29:12 If we were kind of do the same thing and build an index every time you pass in a list of functions, if you ever change the list, you would kind of incur more latency. And we thought it would be really unintuitive for developers and hard to reason about. So we decided to kind of wait until we can support a no added latency solution and not just kind of make it really confusing for developers. Mentioning latency, that is something that people discovered, is that there is an increased cost in latency for the first token? For the first request, yeah. First request. Is that an issue? Is that going to go down over time?
Starting point is 00:29:39 Is there just an overhead to parsing JSON that is just insurmountable? It's definitely not insurmountable, and I think it will definitely go down over time. We just kind of take the approach of, and, you know, if there's nothing in there, you, you, you, you, don't want to fix, then you probably ship too late. So I think we will get that latency down over time. But yeah, I think for most developers, it's not a big concern. Because you're testing out your integration, you're sending some requests while you're developing it, and then it's fast and prod. So it kind of works for most people. The alternative design space that we explored was like pre-registering your schema, so like a totally different endpoint, and then
Starting point is 00:30:14 passing in like a schema ID. But we thought, you know, that was a lot of overhead and like another endpoint to maintain and just kind of more complexity for the developer. And And we think this latency is going to come down over time. So it made sense to keep it kind of in-chat completions. I mean, hypothetically, if one were to ship caching at a future point, it would basically be the superset of that. Maybe. I think the caching space is a little under-explored.
Starting point is 00:30:38 Like, we've seen kind of two versions of it. But I think, yeah, there's ways that maybe put less onus on the developer. But, you know, we haven't committed to anything yet, but we're definitely exploring opportunities for making things cheaper over time. Is AI in agents just going to be a bunch of structurally? your output and function calling one next to each other. Like, how do you see, you know, there's like the model that's everything. Where do you draw the line?
Starting point is 00:30:59 Because you don't call these things like an agent API. But like if I were a startup trying to raise a C round, I would just do function calling and say this is an agent API. Yeah. So how do you think about the difference and like how people build on top of it for like agenetic systems? Yeah. Love that question.
Starting point is 00:31:13 One of the reasons we wanted to build structured outputs is to make agentic applications actually work. So right now it's really hard. Like if something is 95% reliable, but you're chaining together a bunch of calls. If you magnify that error rate, it makes your application not work. So that's a really exciting thing here from going from like 95% to 100%. I'm very biased working on the API and working on function calling and structured outputs, but I think those are the building blocks that we'll be using to distribute this technology very far.
Starting point is 00:31:39 It's the way you connect like natural language and converting user intent into working with your application. And so I think like kind of there's no way to build without it. Honestly, like you need your function calls to work. Like, yeah, we wanted to make that a lot easier. And do you think the assistance kind of like API thing will be a bigger part as people build agents? I think maybe most people just use messages and completion. So I would say the assistance API was kind of a bet in a few areas. One bet is hosted tools.
Starting point is 00:32:07 So we have the file search tool and code interpreter. Another bet was kind of statefulness. It's our first stateful API. It'll store, you know, threads and you can fetch them later. I would say the hosted tools aspect has been really successful. Like people love our file search tool. and it saves a lot of time to not build your own rag pipeline. I think we're still iterating on the shape for the stateful thing
Starting point is 00:32:29 to make it as useful as possible. Right now there's kind of a few endpoints you need to call before you can get a run going, and we want to work to make that much more intuitive and easier over time. One thing I'm just kind of curious about, did you notice any tradeoffs when you add more structured output it gets worse at some other thing that was like kind of, you didn't think was related at all?
Starting point is 00:32:49 Yeah, it's a good question. Yeah, I mean, models are very spiky. And RL is hard to predict. And so every model kind of improves on some things and maybe is flat or neutral on other things. Yeah, like it's like very rare to just add a capability. Yeah. And have no tradeoffs in everything else. So yeah, I don't have something off the top of my head.
Starting point is 00:33:08 But I would say, yeah, every model is a special kind of its own thing. This is why we put them in API dated so developers can choose for themselves, which one works best for them. In general, we strive to continue improving on all e-vals, but it's stochastic. Yeah. Able to apply the structured output system on backdated models like 4-O May as well as Mini, as well as August. Actually, the new response format is only available on two models. It's 4-0 Mini and the new 4-0. So the old 4-0 doesn't have the new response format.
Starting point is 00:33:41 However, for function calling, we were able to enable it for all models that support function calling. And that's because those models were already trained to follow these schemas. We basically just didn't want to add the new response format to models that would do poorly at it because they would just kind of do infinite white space, which is, you know, the most likely token if you have no idea what's going on. I just wanted to call out a little bit more in the stuff you've turned in the blog post. So in blog posts, just use cases, right? I just want people to be like, yeah, we're spelling it out for you.
Starting point is 00:34:07 Use these for extracting structured data from unstructured data. By the way, it does Vision 2. Yes. So that's cool. Dynamic UI generation. Actually, let's talk about dynamic UI. I think Gen UI. I think it's something that people are very interested in.
Starting point is 00:34:21 Yeah. It's your first example. What did you find about it? Yeah, I just thought it was a super cool capability we have now. So the schemas, we support recursive schemas. And this allows you to do really cool stuff. Like, you know, every UI is a nested tree that it has children. So I thought that was super cool.
Starting point is 00:34:36 You can use one schema and generate like tons of UIs. As a backend engineer who's always struggled with JavaScript and front end, like for me, that's super cool. We've now built a system where I can get any front end that I want. So yeah, that's super cool. The extracting structured data, like, the reality of a lot of of AI applications is like you're plugging them into your enterprise business and you have something that works, but you want to make it a little bit better. And so the reliability gains you get here is like you'll never get like a classification using the wrong enum. It's like,
Starting point is 00:35:07 it's just exactly your your types. So really excited about that. Like maybe hallucinate the actual values, right? So let's clearly state what the guarantees are. The guarantees is that they fit the schema, but the schema itself may be too broad because the JSON schema type system doesn't say, like, I only want to range from 1 to 11. You might give me 0. You might give me 12. So yeah, JSON schema, so this is actually a good thing to talk about. So JSON schema is extremely vast, and we weren't able to support every corner of it. So we kind of support our own dialect, and it's described in the docs. And there are a few tradeoffs we had to make there. So by default, if you don't pass in additional properties in a schema, by default, that's true.
Starting point is 00:35:48 And so that means you can get other keys, which, you know, you didn't spell out, which is kind of the opposite of what developers want. You basically want to supply the keys and values, and you want to get those keys and values. And so then we had a decision to make. It's like, do we redefine what additional properties means as the default? And that felt really bad. It's like there's a schema that's predated us. Like, you know, it wouldn't be good.
Starting point is 00:36:07 It would be better to play nice with the community. And so we require that you pass it in as false. You know, one of our design principles is to be very explicit and to developers know, you know, what to expect. And so this is one where we decided, you know, it's a little harder to discover, but we think you should pass this thing in so that we can have like a very clear definition of what you mean and what we mean. There's a similar one here with like required. By default, every key in JSON scheme is optional. But that's not what developers want, right? Like you would be very surprised if you passed in a bunch of keys and you didn't get some of them back.
Starting point is 00:36:37 And so that's the tradeoff we made is to make everything required and have the developers spell that out. Is there a require false? Can people turn it off or they're just getting all? So developers can, basically what we recommend for that is to make your actual key a union type. And so, like, yeah, make it union of int and null, and that gets to the same behavior. Any other the examples you want to dive into math, chain of thought? Yeah, you can now specify like a chain of thought field before a final answer. This is just like a more structured way of extracting the final answer.
Starting point is 00:37:07 One example we have, I think we put up a demo app of this math tutor. example, or it's coming out soon. Did I miss it? Oh, okay. Well. Basically, it's this math tutoring thing, and you put in an equation, and you can go step by step and insert. This is something you can do now with structured up. In the past, a developer would have to, like, specify their format, and then write a parser and parse out the model's output to be pretty hard. But now you just specify, like, steps, and it's an array of steps, and every step you can render, and then the user can
Starting point is 00:37:34 try it, and you can see if it matches and go on that way. So I think it just opens up a lot of opportunities. Like, for any kind of UI, where you want to treat different parts of the model's responses differently. Structured outputs is great for that. I remembered my question from earlier. I'm basically just using this to ask you all the questions as a user, as a daily user of the stuff that you put out. So one is a tip that people don't know,
Starting point is 00:37:55 and I confirmed with you on Twitter, which is you respects descriptions of JSON schemas, right? And you can basically use that as a prompt for the field. Totally. I assume that's blessed and people should do that. Intentional. Right? One thing that I started to do,
Starting point is 00:38:07 which I don't, it could be a hallucination of me, is I changed the property name to prompt the model to what I wanted to do. So for example, instead of saying topics as a property name, I would say like brainstorm
Starting point is 00:38:21 a list of topics up to five or something like that as like a property name. I could stick that in the description as well. But is that too much? Yeah, I would say, I mean, we're so early in AI
Starting point is 00:38:34 that people are figuring out the best way to do things. And I love when I learn from a developer, like a way they found to make something work. In general, I think there's like three or four places to put instructions. Yeah. You can put instructions in the system message. And I would say that's helpful for like when to call a function. So it's like, you know, let's say you're building a customer support thing and you want the model to verify the user's phone number or something. You can tell the model in the
Starting point is 00:38:58 system message, like here's when you should call this function. Then when you're within a function, I would say the descriptions there should be more about how to call a function. So really common is someone will have like date as a string. But you don't tell the model, like, like do you want year, year, month, month, day, day, or do you want that backwards? And that's what a really good spot is for those kind of descriptions. It's like, how do you call this thing? And then sometimes there's like really stuff like what you're doing. It's like name the key by what you want.
Starting point is 00:39:23 So sometimes people put like, do not use. And, you know, if they don't want, you know, this parameter to be used, except only in some circumstances. And really, I think that's the fun nature of this. It's like you're figuring out the best way to get something out of the model. Okay. So you don't have an official recommendation is what I'm hearing. Well, the official recommendation is, you know, how to call a model system instructions. Exactly, exactly.
Starting point is 00:39:42 Or one-to-call function, yeah. Do you benchmark these type of things? So, like, say with date. It's like description, it's like return it in like ISO aid. Yeah. Or if you called the key, date in ISO-A601, I feel like the benchmarks don't go that deep, but then all the AI engineering kind of community and like all the work that people do, it's like, oh, actually, this performs better, but then there's no way to verify.
Starting point is 00:40:06 Right. You know, like even the, I'm going to tip you $100,000 or whatever. Like some people say it works. Some people say it doesn't. Do you pay attention to the stuff as you build this? Or are you just like, the model is just going to get better. So why waste my time running emails on these small things? Yeah.
Starting point is 00:40:21 I would say to that, I would say we basically pick our battles. I mean, there's so much surface area of LLMs that we could dig into. And we're just mostly focused on kind of raising the capabilities for everyone. I think for customers and we work with a lot of customers, really. developing their own evals is super high leverage. Because then you can upgrade really quickly when we have a new model. You can experiment with these things with confidence. So yeah,
Starting point is 00:40:44 we're hoping to make making evals easier. I think that's really generally very helpful for developers. For people, I'll just kind of wrap up the discussion for structured outputs. I immediately implemented, we use structured outputs for AI news. I use an instructor and I ripped it out. And I think
Starting point is 00:41:00 I saved 20 lines of code. But more importantly, it's like we cut it by 55% of API costs based on what what I measured because of all we saved on the retries. Nice. Yeah, love to hear that. Yeah, which people I think don't understand when you can't just simply like add instructor or add outlines. You can do that, but it's actually going to cost you a lot of retries to get the model that you want. But you're kind of just kind of building that
Starting point is 00:41:23 internally into the model. Yeah, I think this is the kind of feature that works really well when it's integrated with like the LLM provider. Yeah, actually I had folks, even my husband's company who works at a small startup, they thought we were just retrying. to make the re-blog post. We are not retrying. You know, we're doing it in one shot, and this is how you save on latency and cost. Awesome.
Starting point is 00:41:43 Any other behind-the-scenes stuff, just generally on structured outputs? We're going to move on to the other models. Yeah, I think that's it. Oh, it's excellent products, and I think everyone will be using it, and we have the full story now that people can try out.
Starting point is 00:41:56 So roadmap would be parallel function calling, anything else that you've called out as, like, coming soon? Not quite soon, but, you know, we're thinking about, does it make sense to expose custom grammars on JSON schema. What would you want to hear from developers to give you information, whether it's custom grammars or anything else about structured output?
Starting point is 00:42:12 What would you want to know more of? Just, you know, always interested in feature requests, what's not working. But I'd be really curious, like what specific grammars folks want. I know some folks want to match programming languages like Python. There's some challenges with the expressivity of our, you know, implementation. And so, yeah, just kind of the class of grammars folks want. I have a very simple one, which is a lot of people try to do, use GBs. as Judge, right?
Starting point is 00:42:36 Which means they end up doing a rating system. And then there's like 10 different kinds of rating systems, the Lycord scale, whatever. If there was an officially blessed way to do a rating system with structured outputs, everyone would use it. Yeah, yeah, that makes sense. I mean, we often recommend using log probs
Starting point is 00:42:53 with classification tasks. So rather than like sampling, you know, let's say you have four options like red, yellow, blue-green. Rather than sampling, you know, two tokens for yellow, you can just do like ABCD and get the log probes of those, you know, the inherent randomness of each sampling isn't taken into account and you can just actually look at what is the most likely token.
Starting point is 00:43:13 I think this is more of like a calibration question. Like if I ask you to rate things from 1 to 10, a non-calibrated model might always pick 7, just like a human would. Right. So like actually have a nice gradation from 1 to 10 would be the rough idea. Yeah.
Starting point is 00:43:28 And then even for structured outputs, I can't just say have a field of rating from 1 to 10 because I have to then validate it. And, you know, it might give me 11. Yeah, absolutely. Yeah. So what about model selection? Now you have a lot of models.
Starting point is 00:43:40 When you first started, you had one model endpoint. I guess you had like the Vinci end. But like most people are using one model endpoint. Today you have like a lot of competitive models. And I think we're nearing the end of the 3.5 run RIP. How do you advise people to like experiment, select both in terms of like task and like costs? Like what's your playbook? In general, I think folks should start with.
Starting point is 00:44:04 4-0 Mini. That's our cheapest model, and it's a great workhorse. Works for a lot of great use cases. If you're not finding the performance you need, like, you know, maybe it's not smart enough, then I would suggest going to 4-0. And if 4-0 works well for you, that's great. Finally, there's some, like,
Starting point is 00:44:20 really advanced frontier use cases, and maybe 4-0 is not quite cutting it, and there I would recommend our fine-tuning API. Even just like 100 examples is enough to get started there, and you can really get the performance you're looking for. We're recording this ahead of it, but like you're announcing are there some fine-tuning stuff that people should pay attention to?
Starting point is 00:44:38 Yeah. Actually, tomorrow we're dropping our GA for GBT-40 Fine-Tuning. So 4-0 Mini has been available for a few weeks now, and 4-0 is now going to be generally available. And we also have a free training offering for a bit. I think until September 23rd, you get 1 million of free training tokens a day. This is already announced, right? Am I talking about a different thing?
Starting point is 00:44:58 So that was for 4-0-mini, and now it's also for 4-0. So we're really excited to see what people do with it. And it's actually a lot easier to get started than a lot of people expect. I think they might need tens of thousands of examples. But even 100 really high-quality ones or 1,000 is enough to get going. I think people's concerns about fine-tuning is that they're kind of locked into a model. And I think you're paving the path for migration of models as long as they keep their original data set. They can at least migrate nicely.
Starting point is 00:45:23 Yeah, I'm not sure what we've said publicly there yet. But we definitely want to make it easier for folks to migrate. It's the number one concern. You know, I'm just, you know, it's obvious. Absolutely. I also want to point people to, you have official model selection docs where it's in the guide, we'll put it in the show notes, where it says to optimize for accuracy first. So prompt engineering, rag, evels, fine tuning.
Starting point is 00:45:44 This was done at Dev Day last year, so I'm just repeating things. And then optimize for cost and latency second. And there's a fuse sets of steps for optimizing latency. So people can read up on that stuff. Yeah, totally. Yeah. We had one episode with Nigelots Karlini from Deep Mind. And we actually talked about how some people don't actually get to the boundaries of the model performance.
Starting point is 00:46:04 You know, they just kind of try one model and it's like, oh, LLMs cannot do this and they stop. How should people get over the hurdle? It's like, how do you know if you hit the model performance or like you hit skill issues? You know, it's like your prompt is not good or like try another model and whatnot. Is there an easy way to do that? That's tough. Some people are really good at prompting and they just kind of get it right away. And for others, it's more of a challenge.
Starting point is 00:46:25 I think there's a lot we can do to make it easier to prompt our models. For now, I think requires a lot of creativity and not giving up right away. Yeah. And a lot of people have experienced now with chat GPT. You know, before chat GPT, the easiest way to play with our models was in the playground. But now kind of everyone's played with it with the model of some sort. And they have some sort of intuition. It's like, you know, if I tell you my grandma is sick, then maybe I'll get the right output.
Starting point is 00:46:48 And we're hoping to kind of remove the need for that. But playing around with chat GPT is a really good way to get a feel for, you know, how to use the API as well. Will prompt engineering be here forever? or is it a dying guard as the models get better? I mean, it's like the perennial question of software engineering as well. It's like as the models get better at coding, you know, if we hit 100 on sui bench, what does that mean? I think there will always be alpha and people who are able to like clearly explain what they're trying to build. Most of engineering is like figuring out the requirements and stating what you're trying to do.
Starting point is 00:47:18 And I believe this will be the case with AI as well. You're going to have to very clearly explain what you need and some people are better than others at it. and people will always be building. It's just the tools are going to get far better. That's two weeks you release two models. There's GPC40-20406, and then there's also ChatGBC4-O-Ladest. I think people are a little bit confused by that, and then you issued a clarification that was once chat-tuned,
Starting point is 00:47:41 then the other is more function-calling-tuned. Can you elaborate? Yeah, totally. So part of the impetus here was to kind of very transparent with what's on-chat GBT and in the API. So basically, we're often training models, and there are different use cases. So you don't really need function calling for user-defined functions in chat chvety. And so this gives us kind of the freedom to build the best model for each use case.
Starting point is 00:48:04 So in chat-GBT-BT latest, we're releasing kind of this rolling model. The weights aren't pinned as we release new models. This is literally what we use. Yeah. So it's in what's in chat-GBT. So it's very good for like chat-style use cases. But for the API broadly, you know, we really tune our models to be good at things that developers want, like function calling and structured outputs. and when a developer builds their application,
Starting point is 00:48:27 they want to know that the weights are stable under them. And so we have this offering where it's like, if you're tuning to a specific model and you know your function works, you know it will never change the weights out from under you. And so those are the models we commit to supporting for a long time, and we think those are the best for developers. But we want to give it up, you know, we want to leave the choice to developers. Like, do you want the chat GPUT model or do you want the API model?
Starting point is 00:48:49 And you have the freedom to choose what's best for you. I think it's for people, they do want to pin model versions. So I don't know when they would use chat GPT, like the rolling one, unless they're really just kind of cloning chat GPT, which is like, why would they? I mean, I think there's a lot of interesting stuff that developers can do when unbounded. And so we don't want to limit them artificially. So it's kind of survival of the fittest. Like whichever model is better, you know, that's the one that people should use. Yeah.
Starting point is 00:49:18 I talked about it to my friends as like, this isn't that new thing. And basically, opening has never actually shared with you the actual chat GBT BT model, and now they do. Well, it's not necessarily true. Actually, a lot of the models we have shipped have been the same. But, you know, sometimes they diverge, and it's not a limitation we want to stick around. Anything else we should know about the new model? I don't think there were, there was no evals announced or anything.
Starting point is 00:49:43 But people say it's better. I mean, obviously, LMSS is like way better above everything, right? It's like number one in the world done. Yeah, we published some release numbers. They're not as in-depth as we want to be yet, because it's still kind of a science and we're learning what actually changes with each model and how can we better understand the capabilities. But we are trying to do more release notes in the future and keep folks updated. But yeah, it's kind of an art and a science right now.
Starting point is 00:50:10 You need the best Evels team in the world to help you figure this out. Yeah, Evels are hard. We're hiring if you want to come work on Eval's. Hold that thought on hiring. We'll come back to the end on what you want, where you're looking for, because obviously the people want to join you and they want to know what qualities are looking for. So we just talked about API versus ChattGBT. What's, I guess, like, the vision for the interface? You know, the mission of open AI is like built AGI that is accessible.
Starting point is 00:50:34 Like, where is it going to come from? Totally, yeah. So I believe that the API is kind of our broadest vehicle for distributing AGI. You know, we're building some first-party products, but they'll never reach every niche in the world and kind of every corner in community. And so I really love working with developers and seeing the incredible things they come up with. I often find that developers kind of see the future before anyone else, and we love working with them to make it happen. And so really the API is a bet on going really broad. We'll go very deep as well in our first-party products, but I think just our impact is
Starting point is 00:51:04 absolutely magnified by every developer that we uplift. They can do the last malware where you cannot. Like chat GPD is one type of product, but there's many other kinds. In fact, I observed, I think, in February, basically chat GBT's user growth stopped when the API was launched, because everyone's kind of be able to take that and build other things. That has not become true anymore because Chachibati-G-C growth has continued to grow. But then you're not confirming any of this. This is me quoting similar web numbers, which have very high variants. Well, the API predates Chatsch-G-T.
Starting point is 00:51:37 The API is actually opening on his first product, and the first idea for commercialization, that predates me as well. Wide release. Like, GA, everyone can sign up and use it immediately. Right. That's what I'm talking about. But, yeah, I mean, I do believe that. you know, that means you also have to expose all of open-eye models, right?
Starting point is 00:51:53 Like all the multimodal models. We'll ask you questions on that, but like, I think that API mission is important. It's interesting that how does new programming language is supposed to be English, but it's actually just software engineering, right? It's just, you know, we're talking about HTTP error codes. Right.
Starting point is 00:52:11 Yeah, I think, you know, engineering is still the way you access these models. And I think there are companies working on tools to make engineering more accessible for everyone. But there's still so much alpha in just writing code and deploying. Yeah, one might even call it AI engineering. Exactly. Yeah, so there's lots of war stories from building this platform. We started at the start of your career,
Starting point is 00:52:34 and then we jumped straight to structured outputs. There's a whole thing, like two years that we skipped in between. What have become your principles? What are your favorite stories that you like to tell? We had so much fun working on the assistance API and leading up to Dev Day. You know, things are always pretty chaotic. when you have an externally, like a date that is hard,
Starting point is 00:52:52 and there's like a stage and there's like a thousand people coming. You can always launch a wait list. We're trying hard not to. Because, you know, we love it when people can access the thing on day one. And so, yeah, the Assistance API, we had like this really small team and just working as hard as we could to make this come to life. But even actually the morning of, I don't know if you'll remember this, but Sam did this keynote.
Starting point is 00:53:16 Yep. and Ramon came up and they gave free credits to everybody. So that was live, fully live, as were all of the demos that day. But actually, maybe like two hours before that, we had a little outage and everyone was scrambling to make this thing work again. So, yeah, things are early and scrappy here. And, you know, we were really glad. We were a bit on the edge of our seat watching it live.
Starting point is 00:53:39 What's the plan B in that situation? If you can share. Play a video. This is a classic Devereux. I don't know. I mean, I actually don't know. the plan B was. No plan B, no failure. But we just, you know, we fixed it.
Starting point is 00:53:51 We got everything running again and the demo went well. Just hire cracked waterloo grass. It's kill issues as usual. Sometimes you just got to make it happen. I imagine it's actually very motivating. But I did hear that after Dev Day, like the whole company got like a few weeks off just to relax a little bit. Yeah, we sometimes get, like we just had the week of July 4th off. And yeah, it's hard to take vacation because people are working on such exciting things.
Starting point is 00:54:16 and it's like you get a lot of FOMO on vacation, so it helps when the whole company is on vacation. Mentioning A's Assistance API, you actually announced a roadmap there, and things have developed. I think people may not be up to date. What's the offering today versus, you know, one year ago? Yeah, so we've made a bunch of key improvements.
Starting point is 00:54:33 I would say the biggest one is in the file search product. Before, we only supported, I think, like, 20 files per assistant, and the way we used those files was, like, less effective. Basically, model would decide based on the file name, whether to search a file and there's not a ton of information in there. So our new offering, which we shipped a few months ago, I think, now allows 10K files per assistant, which is like dramatically more. And also it's a kind of different operation.
Starting point is 00:54:59 So you can search semantically over all files at once rather than just kind of the model choosing one up front. So a lot of customers have seen really good performance. We also have exposed more like chunking and re-ranking options. I think the re-ranking one is coming, I think, next week or very soon. So this kind of gives developers more control. and more flexibility there. So we're trying to make it the easiest way
Starting point is 00:55:19 to kind of do rag at scale. Yeah. I think that visibility into the rag system was the number one thing missing from Dev Day and then people got their first impressions and then he never looked at it again. So that's important. The Ranker is a core feature of,
Starting point is 00:55:34 let's say, some other Foundation Model Labs. Is Open Eye going to offer a re-ranking service, a Rerangler model? So we do re-ranking as part of it. I think we're soon going to ship more controls for that. Okay, got it. And like if I'm an existing Langchain, Lama Index, whatever, how do you compare, do you make different choices? Like, where does that
Starting point is 00:55:53 exist in the spectrum of choices? I think we are just coming at it, trying to be the easiest option. And so ideally, like, you don't have to know what a re-ranker is, and you don't have to have a chunking strategy, and the thing just kind of works out of the box. So I would say, that's where we're going. And then, you know, giving controls to the power users to make the changes they need. Awesome. I'm going to ask about A couple other things, just updates on stuff also announced at Dev Day, and we talked about this before.
Starting point is 00:56:20 Determinism, something that people really want. Dev Day would announce the seed parameter as well as system fingerprint. And objectively, I've heard issues. Yeah. I don't know what's going on. Yeah, the seed parameter is not fully deterministic, and it's kind of a best effort thing. So you'll notice there's more determinism in the first few tokens. That's kind of the current implementation.
Starting point is 00:56:39 We've heard a lot of feedback. We're thinking about ways to make it better. But it's challenging. It's kind of training off against, you know, reliability and off time. Other maybe underrated API-only thing, Logite bias, that's another thing that kind of seems very useful and that maybe most people are like, it's a lot of work. I don't want to use it. Do you have any examples of, like, use cases or, like, products that are made a lot better
Starting point is 00:57:01 through using it? So, yeah, classification is the big one. So logic bias, your valid classification outputs and, you know, you're more likely to get something that matches. We've seen that people logit bias like punctuation tokens, maybe trying to get more succinct writing. Yeah, it's generally very much a power user feature. And so not a ton of folks use it. I actually wanted to use it to reduce the incidence of the word delve.
Starting point is 00:57:26 Yeah. Have people done that? Probably. I don't know. Is delve one token? You're probably, you got to do a lot of permutation. It's used so much. Maybe it is.
Starting point is 00:57:34 It depends on the tokenizer. Are there non-public tokenizers that I guess you can not answer or you would admit it. Are the 100K and 200K vocabes, like the ones that you use across all models? Yeah, I think we have docs that publish more information. I don't have it off the top, but I think we publish which tokenizers for which model. Okay, so those are the only two. The tiering rate limiting system, I don't think there was an official blog post kind of announcing this, but it was kind of mentioned that, like, you started tying, like, fine-tuning to tiering
Starting point is 00:58:02 and, like, feature rollouts. Just from your point of view, like, how do you manage that? And what should people know about the tiering system and rate limiting? Yeah, I think basically the main changes here were to be more transparent and easier to use. So before developers didn't know what tier they're in, and now you can see that in the dashboard. I think it's also, I think we publish how you move from tier to tier. And so this just helps us do kind of gated rollouts for the fine-tuning launch. I think everyone tier two and up has full access.
Starting point is 00:58:31 That makes sense. You know, I would just advise people to just get to Tier 5 as quickly as possible. Sure. Like a gold star customer, you know? Like, I don't know. It seems to make sense. Do we want to maybe wrap with future things and kind of like how you think about design and everything?
Starting point is 00:58:45 So you just mentioned you want to be the easiest way to basically do everything. What's the relationship with other people building in the developer ecosystem? Like I think maybe in the early days, it's like, okay, we only have these APIs and that everybody helps us. But now you're kind of building a whole platform. How do you make decisions? Yeah, I think kind of the 80-20 principle applies here. We'll build things that kind of capital.
Starting point is 00:59:06 you know, 80% of the value and maybe leave the long tail to other developers. So we really prioritized by like how much feedback are we getting, how much easier we'll make this, will this make something like an integration for a developer? So yeah, we want to do more in this space and not just be an LLM as a service, but kind of AI development platform as a service. Oh, okay. That ties into a thing that I put in the notes that we prepped. There are other companies trying to be AI development platform.
Starting point is 00:59:33 So you will compete with them. They just want to know what you won't build so that they can build it. Yeah, it's a tough question. I think we haven't determined what exactly we will and won't build. But you can think of something, if it makes it a lot easier for developers to integrate, you know, it's probably on our radar and, you know, stack rank by impact. Yeah. So there's like cost tracking and model fallbacks.
Starting point is 00:59:58 Model fallbacks is an interesting one because people do it. I don't think it adds a ton of value. But like if you don't build it, I have to build it. Because if one API is down or something, I need to fall back to another one. Yeah. I mean, the way we're targeting that user need is just by investing a lot in reliability. And so we have- Just don't fail. I mean, we have improved our uptime pretty dramatically over the last year.
Starting point is 01:00:20 And it's been the result of a lot of hard work from folks. So you'll see that on our status page and our continued commitment going forward. It's the important thing about owning the platform that it gives you the flexibility to put all the kind of messy stuff behind the, the scenes or, yeah, how do you draw the line between what you want to include? Yeah, I just think of it as like, how can we onboard the next generation of AI engineers, as you put it, right? Like, what's the easiest way to get them building really cool apps? And I think it's by building stuff to kind of hide this complexity or just make it really easy to integrate. So I think of it a lot is like, what is the value ad we can provide beyond just the models that makes the models really useful?
Starting point is 01:01:00 Okay, we'll touch on four more features of the API platform that we prepped. batch vision whisper and team enterprise stuff so you wanted to talk about batch yeah so the rough idea is you give a the contract between you and me is that i give you give you the batch job you have 24 hours to run it and it's kind of like spot inst for for the API what what should people know about it so it's half off which is a great savings it also works with like 4 oh mini so the savings on top of 4a mini is pretty crazy like the stuff you can do like 7.5 cents or something for million. Yeah, I should really have that number top of mind, but it's like staggeringly cheap. And so I think this opens up a lot more use cases. Like, let's say you have a user activation
Starting point is 01:01:41 flow and you want to send them an email like maybe every day or like at certain points in their user journey. So now you can do this with the batch API and something that was maybe a lot more expensive and not feasible is now very easy to do. So right now we have this 24 hour turnaround time for half off and curious would love to hear from your community like what kind of turnaround on time do they want? I would be an ideal user or a batch and I cannot use batch because it's 24 hours. I need two to four. Two to four hours. Okay. Yeah, that's good to know. But yeah, just a lot of folks haven't heard about it. It's also really great for like e-vals running them offline. You don't generally don't need them to come back within, you know, two hours.
Starting point is 01:02:16 I think you could do a range, right? Two to four for me, like I need to produce a daily thing. And then 24 for like the average use case and then maybe like a week, a month. Who cares? Yeah. For people who just have a lot to do. Yeah, absolutely. So yeah, that's batch API. I think folks you use it more, it's pretty cool. Is there a future in which, like, six months is, like, free? You know? Like, is there, like, small, is there, like, super small, like, shards of, like, GPU runtime that, like, over a long enough timeline, you can just run all these things for free?
Starting point is 01:02:46 Yeah, it's certainly possible. I think we're getting to the point where a lot of these are, like, almost free. That's true. You're already close. Why would they work on something that's, like, completely free? I don't know. Okay, so Vision. Vision got G8.
Starting point is 01:02:57 Last year, people were so wild by the GPC4 demo. and that was primarily Vision. What was it like building the Vision API? Yeah, Vision API is super cool. We have a great team working there. I think the cool thing about Vision is that it works across our APIs. So there's, I can use it in the Assistance API,
Starting point is 01:03:13 I can use the batch API and chat completions that works with structured outputs. I think it just helps a lot of folks with kind of data extraction where the spatial relationships between the data is too complicated and you can't get that over text. But yeah, there's a lot of really cool use cases.
Starting point is 01:03:27 I think the tricky thing for me is understanding how frequent to turn vision into, from single images into, like, effectively just always watching. And right now, I think people just send a frame every second. Will that model ever change? Will there just be like, I stream you video? And then... Yeah, I think it's very possible that we'll have an API where you stream video in.
Starting point is 01:03:49 And maybe, you know, to start, we'll do the frame sampling for you. Because the frame sampling is the default, right? Right. But I feel like it's hacky. Yeah, I think it's hard for developers to do. And so, you know, we should definitely work on making that easier. Is there in the batch API, do you have like a time guarantees, like order guarantees? Like if I send you a batch request of like a video analysis, I need every frame to be done in order.
Starting point is 01:04:12 For batch, you send like a list of requests and each of them stand alone. So you'll get all of them finished, but they don't kind of chain off each other. Well, if you're doing a video, you know, if you're doing like analyzing a video. I wasn't linking video to batch, but that's interesting. Yeah, well, video is like, you know, if you have. have a very long video, you can just do a batch of all the images and let it process. Oh, that's a cool idea. You could offer like, but serially.
Starting point is 01:04:35 Sequential true. Yeah, yeah, yeah, exactly. But the whole point of batch is you're just using kind of spare time to run it. Let's talk about my favorite model, Whisper. Oliver, I built this thing called Small Podcaster, which is an open source tool for podcasters. And why does Whisper API not have diarization when everybody is transcribing people talking? That's my main question. Yeah, it's a good question.
Starting point is 01:04:59 and you've come to the right person. I actually worked on the Whisper API and ship that. That was one of my first APIs I shipped. Long story short is that Whisper V3, which we open sourced, has, I think, the directization feature. But there's some performance tradeoffs. So Whisper V2 is better at some things than Whisper V3. And so it didn't seem that worthwhile to ship Whisper V3
Starting point is 01:05:20 compared to, like, the other things in our priorities. I think we still will at some point. But yeah, it's just, you know, there's always so many things we could work on. It's tough to do everything. We have a Python notebook that does the diarization for the pod, but I would just like, you can translate like 50 languages, but you cannot tell me who's speaking. That was like the funniest thing. There's like an XKCD thing about this, about hard problems in AI. It's forget the one. Tell me if this was taken in a park. And like that's easy. And it's like, tell me if there's a bird in this picture. And it's like, give me 10 people in a research team. It's like, you never know which things are challenging. And diorotization is, I think, you know, more challenging than expected. Yeah, it still breaks a lot with like overlaps, obviously. Sometimes similar voices it struggles with.
Starting point is 01:06:05 Like I need to like double read the thing. Totally. But yeah, great motto. I mean, it would take us so long to do transcriptions. And I don't know why like small podcasts has better transcription to like mostly every commercial tool. It beats the script. And I'm like, I'm just using the model. I literally not doing anything.
Starting point is 01:06:22 You know, it's just a notebook. So yeah, it just speaks to like sometimes just using the simple opening eye model is better than like figuring out. your own pipeline thing. Totally. I think the top feature request there just would be, I mean, say, again, you know, using you as a feature request dump is like being able to bias the vocab. I think there is like in Raw Whisper you can do that. You can pass a prompt in the API as well.
Starting point is 01:06:44 But you pass in the prompts, okay. Yeah. There's no more deterministic way to do it. So this is really helpful when you have like acronyms that aren't very familiar to the model. And so you can put them in the prompt and you'll basically get the transcription using those correctly. We have the AI engineer solution, which has. just a dictionary.
Starting point is 01:07:00 Nice. We're like all the way misspelted in the past and then G sub and like replace the. If it works, it works. Like that's engineering. It's like, you know, L, like all these different things. Or like length chain and like transcribes to length chain and like capitalization does a bunch of like three or different ways. Yeah. You guys should try the prompt feature.
Starting point is 01:07:20 I love these like kind of pro tip. Okay, fun question. I know we don't know yet, but I've been enjoying the advanced voice mode. It really streams back and forth. and it handles interruptions, how would your audio endpoint change when that comes out? We're exploring new shape of the API to see how it would work in this kind of speech-to-speech paradigm.
Starting point is 01:07:40 I don't think we're ready to share quite yet, but we're definitely working on it. I think just the regular request response probably isn't going to be the right solution. For those who are listening along, I think it's pretty public that OpenEI uses LiveKit for the chatGBT app, which seems to be the socket-based approach
Starting point is 01:07:56 that people should be at least, up to speed on. Like, I think a lot of developers only do request response. And, like, that doesn't work for streaming. Yeah. When we do put out this API, I think we'll make it really easy for developers to figure out how to use it. Yeah. It's hard to do audio change. Okay. And then I think the last one on this was team enterprise stuff. Auditlog, service accounts, API keys. What should people know was in the enterprise offering? Yeah, we recently shipped our admin and audit log APIs. And so a lot of enterprise users have been asking for this for a while. The ability to kind of manage API keys programmatically, manage your projects, get the auto log.
Starting point is 01:08:29 So we've shipped this, and for folks that need it, it's out there and happy for your feedback. Yeah, awesome. I don't use them. So I don't know. I imagine it's just like build your own internal gateway for your internal developers to manage your deployment of Open AI. Yeah, I mean, if you work at like a company that needs to keep track of all the EPA keys, it was pretty hard in the past to do this in the dashboard.
Starting point is 01:08:51 We've also improved our SSO offering, so that's much easier to use now. The most important feature. Yeah, people love SSO. All right, let's go outside of Open AI. What about just you personally? So you mentioned Waterloo. Maybe let's just do, why is everybody a Waterloo cracked? And why are people so good?
Starting point is 01:09:10 And why have people not replicated it or any other commentary on your experience? The first is the co-op program. It's obviously really good. You know, I did six internships, learned so much in those. I think another reason is that Waterloo is like, you know, it's very cold in the winter. it's pretty miserable. There's like not that much to do apart from study and like hack on projects
Starting point is 01:09:31 and there's this big like hacker mentality. You know, there's a hack the North is a very popular hackathon and there's a lot of like startup in computers. It kind of just has this like startup and hacker ethos. Then that combined with the six internships means that you get people who like graduate with two years of experience
Starting point is 01:09:47 and they're very entrepreneurial and you know, they're down to grind. I do notice a correlation between climate and the correctness of engineers. So, you know, it's no coincidence that Seattle is the birthplace of Microsofts and Amazon. I think I had this compilation of Denmark where people like, so it's the birthplace of C++, Ph.P, Turbo Pascal, StandardML, BNF, the thing that we just talked about, MD5 Crypt, Ruby on Rails, Google Maps, and V8 for Chrome.
Starting point is 01:10:16 And it's because, according to Bjorn Storostrup, the Crato C++, there's nothing else to do. Well, yeah, Lena Thorbalt's in Finland. Yeah, yeah, yeah. I mean, you hear a lot about this, like, in relation to SF. People say, you know, New York is way more fun. There's nothing to do on SF. And maybe it's a little by design that all tech is here. The climate is too good.
Starting point is 01:10:33 Yeah. If we also have fun things to do. Nature is so nice, you can touch grass. Why are we not touching grass? You know, restaurants close at like 8 p.m. Like, that's what people are referring to. There's not a lot of, like, late night dining culture. Yeah.
Starting point is 01:10:46 So you have time to wake up early and get to work. You are a book recommender or book enjoyer. What underrated books do you recommend most of others? Yeah, I think a book I read somewhat recently that was very formative was the making of the Prince of Persia. It's a striped press book. That book just made me want to work hard like nothing I've ever read. It's just like this journal of what it takes to build, you know, incredible things. So I'd recommend that.
Starting point is 01:11:10 Yeah, it's funny how video games are for a lot of people, at least for me, kind of like some of the moments and technology. Like when I played the sense of time on PS2 was like my first PlayStation 2 games. Now it's like, man, this thing is so crazy compared to any PlayStation one game. And it's like, wow, my expectations for like the technology. I think like Open AI is a lot of similar things, like the advanced voice. It's like, you see that thing. And then you're like, okay, what I can expect from everybody else is kind of raised now, you know. Totally.
Starting point is 01:11:38 Another book I like to plug is called Misbehaving by Richard Thaler. He's a behavioral economist and talks a lot about how people act irrationally in terms of decision making. And I actually think about that book, like once a week, probably, at least when I'm making a decision and I realized that, you know, I'm falling into a fallacy or, you know, it could be a better decision. Yeah. You did a minor in psych? I did.
Starting point is 01:11:58 Yeah. I don't know if I learned that much there, but it was interesting. Is there, like, an example of, like, a cognitive bias or misbehavior that you just love telling people about? Yeah. People, so let's say you won tickets to, like, a Taylor Swift concert. And I don't know how much they're going for, but it's probably like $10,000. Oh, okay.
Starting point is 01:12:17 Or whatever for sure. And, like, a lot of people are like, oh, I have to keep these. Like, I won them. $10,000. But really it's the same decision you're making. If you have $10,000, like, would you buy these tickets? And so people don't really think about it rationally. Like, would they rather have $10,000 of the tickets? For people who want it, a lot of the time, it's going to be the $10,000, but their bias is because they want it. The world organized itself this way, and you should keep it for some reason. Yeah. Oh, okay. I'm pretty familiar with this stuff. There's also a loss
Starting point is 01:12:42 version, uh, loss of version of this where it's like, if I take it away from you, you respond more strongly than if I give it to you. Yes. If people are like really upset if they like don't get a promotion, but if they do get a promotion, they're like, okay, few. It's like not even, you know, excitement. It's more like we react a lot worse to losing something. Which is why like when you join like a new platform, they often give you points and then they'll take it away if you like don't do some some action in like the first few days. Yeah. Totally. Yeah. The book references people who work like operate very rationally as econs as like a separate group to humans. and I often think
Starting point is 01:13:19 what would an econ do do here in this moment and try to act that way. Okay, let's do this. Are LLM's econs? I mean, they are maximizing probability distributions. Minimizing loss? Yeah, so I think way more than all of us
Starting point is 01:13:35 they are econs. Whoa, okay, so they're more rational than us? I think their optimization functions are more clear than ours. Yeah, just to wrap, you mentioned you need help on a lot of things. Yeah. Any specific roles, callouts, and also people's backgrounds?
Starting point is 01:13:49 Like, is there anything that they need to have done before, like, what people fit well at Open. Yeah, we've hired people with all kinds of backgrounds, people who have PhD and an ML or folks who have just done engineering like me. And we're really hiring for a lot of teams. We're hiring across the applied org, which is where I sit for engineering and for a lot of researchers. And there's a really cool model behavior role that we just dropped. So, yeah, across the board, we'd recommend checking out our careers page. You don't need a ton of experience in AI specifically to join. I think one thing that I'm trying to get at is like what kind of person does well at OpenEI?
Starting point is 01:14:25 I think objectively you have done well. And I've seen other people not do as well and basically be managed out. I know it's an intense environment. I mean, the people I enjoy working with the most are kind of low ego, do what it takes, ready to roll up their sleeves, do what needs to be done and unpretentious about it. Yeah, I also think folks that are very user-focused do well on kind of API and chat. The YC ethos have built something people want is very true at Open AI as well. So I would say low ego, user-focused, driven. Cool.
Starting point is 01:14:58 Yeah, this was great. Thank you so much for coming on. Thanks for having me. That was an excellent conversation with Michelle, which were just about to ship before we then heard that the OpenAI strawberry, now named O1 Release, was coming. So we delayed our podcast to organize an emergency meetup with more friends from open-O-I and latent space enjoyers in San Francisco to capture quick takes on 01, as well as answer some questions. Once begun, left Michelle and GPT for all just for fun. Strawberry fields where the bites all run, scaling up in France for second and none.
Starting point is 01:15:43 We set the counter GPT's Pass A. O-1's the hot shot, the H-EI way. Not a system, just one model. rise an extraordinary alien of unknown size 01 minis it be quick though a previews just a tease encoding stem minis the bees knees token counts the same gpt for those old news hiding thought chains like a cognit of leave space news all once begun left michel and gpt for o just for fun strawberry fields where the bikes all run scaling up in fritz but second to none bigger context loading no more scrolls handling long tasks the chunking's got holes tools not in the kit
Starting point is 01:16:20 The functions will call. Multimodal magic O-1's on the ball. Our rel's a sleeper sauce in our cold brew. GPD-4-0 can't catch up even in date-but mode. You think it seems slow because we summarise the jest, but when answers drop, the faster than the rest. My hiding firmware, O-1's the gun, left Michelle and GPT for all just for fun.
Starting point is 01:16:40 Strawberry fields where the bites all run. Scaling up in France was second to none. There are two sections to our O-1 coverage today. The first is our recorded meetup audio from Thursday, which will also be posted on YouTube, recapping the most important takeaways about 01 from SWIX, and then an open-ended cue and a session with Romaine Hewitt and other OpenAI representatives. We apologize in advance for the audio quality, but it was the best we could gather on short notice. The second is a recap of the official OpenAI developer's AMA on Friday,
Starting point is 01:17:27 which you'll hear about from me. Finally, you'll hear some O1 demos from Min Kim of Lighthouse and your favorite latent space co-host, Alessio. Stay tuned. Yeah, thanks everyone for joining this emergency 01 meetup, and thank you to open the eye for the model and the food. Very much appreciated for feeding our brains and our stomachs. So I'm just going to yap a little bit for the AI news recap that we did today.
Starting point is 01:17:52 One thing I really appreciate is that they generally tend to make it API available, even if the API can be surprisingly hard to work with. So anyway, so there are a few data points that I'm just going to recap, and then we'll let the actual experts answer any questions, or at least say a few words as well. So there are effectively two models released today. I like to show it like this in this chart, because I think that people don't really understand that the O1 preview
Starting point is 01:18:19 is really just a checkpoint to a broader thing, and they already have, and that one's still in process. The other notable thing that a lot of people are talking about that was in our summary of Discord discussion and Twitter discussion. Today was the pricing. That's notably a little bit higher than, for example, Opus. And we'll talk about the other stuff. There were a bunch of blog posts. It was actually a little bit hard to find or read, but this is what I think everyone should read.
Starting point is 01:18:47 If you sign up to the newsletter, you would have got it. So each model has its own blog post. There's a technical research blog post. there's a system card and then there's also docs and then there's some other videos if you care about the humans behind the project. I tried to collect together the what the alpha that people drop because obviously that is less. The main thing, so this is actually if you remember the goodwill hunting, goodwill hunting board problem. I gave it that and it solved it. And so the tweet that I was going to work on was basically 01 is Matt Damon.
Starting point is 01:19:20 Then it can solve the goodwill hunting problem. And the new paradigm really is in chat GBT, it shows this chain of thought. As far as I understand in the API, it doesn't actually show that. And this chain of thought is only shown in chat GBT, and it's not the actual chain of thought. It is a hidden, obscured summary of the chain of thought. And that is materialized as reasoning tokens, which you do pay for. The token limits are expanded. And this was the nastiest part about me actually using the API today, was that I could not just one-for-one port over the code.
Starting point is 01:19:51 things that instructor no longer work because there's no system role anymore. There's no temperature, there's no tool calling, there's no streaming. And especially there's no max tokens because you have no constraint or visibility on how many tokens are being used. But there's a lot of really interesting information in the docks that gives you some mental model of how these sort of reasoning tokens are given you and how if they take too many tokens and I ran it into exactly this for the for this task because it only solved the first problem and it could not solve second, third, fourth, because it ran into the 128 tokens issue.
Starting point is 01:20:24 Okay, the third thing that I want to talk about is also the e-vals. We're very used to really good evals, but I think this is, obviously, again, new highs in a lot of the categories that we're used to. The thing I think that is more important than just the evils is also scaling laws for test time compute. This one I really recommend people check out Jim Fan's comments on this, which is just an excerpt from the technical blog post,
Starting point is 01:20:48 that we now not only do train time compute, but also test time compute, and they seem to scale in a log later fashion that we can actually predict and invest in. And so that seems to be very, very helpful. As far as I understand, I mean, I basically covered, looked at the entire strawberry team's Twitter, which is the unofficial documentation. I really, I think the main thing that people are trying to figure out is how much can you do in the wild versus must you train the models to do this. And the only hints that we have so far is that they are, they're definitely doing Chin without using RL and it's much better than just prompting. And that's something that obviously as engineers we don't have access to. And you can read, you can read Jason's post for the
Starting point is 01:21:36 rest. And then finally, because I don't want to take out too much time, it seems like a lot of people are calling out O1 Mini for its relative outperformance. Usually when there's two models, one big, one small, or small, medium big, usually the big model gets all the love and then the other stuff is, you know, or just sizably smaller. But O1 actually outperforms if you go all the way back. Lucas from DeepMind actually called out like if this is true, why do you use O1 preview at all? You should just you should just use Mini or O1. And that's because 01 doesn't exist yet. But Mini basically scores better in STEM because it's trained to do that. That does not mean that it is better than 01 preview. It's just
Starting point is 01:22:15 trained to do better in STEM. So that's my quick recap. We're lucky enough to have a couple people from OpenE Eye here. Do you want to say a few words and what people should know? Yeah, I'll just leave it open. And here you can take this if you want just kind of just to magnet it. Cool. All right. Hey everyone. Thank you so much Sweeks for organizing this like emergency 01 meetup. Thank you, Sean. So thank you, Vipu, MindyB, Decibel for hosting. But yeah, super exciting day at OpenEi Eye with the release of O1. Sounds like Sean already gave you all the lowdown on the launch day.
Starting point is 01:22:52 But yeah, it's a new day for us. I think we're very excited to see what you builders, developers, are going to start building with this model. In particular agents or agentic apps that maybe before you had some challenges making them work. So yeah, we're very, very jazzed at the idea of like you kicking the tires of the model. It's obviously the very beginning. We have a lot of work ahead of us to give access to everyone and increase the rate limits and so on and so forth. But yeah, very excited to be here and can't wait to see the demos tonight and in the days to come.
Starting point is 01:23:21 We have a few members, by the way, of the team tonight. There's Lindsay from our CAMS team. There's Edwin also from our developer experience team. But yeah, excited and Anuj from our engineering team as well. So yeah, excited to be here and answer questions and can't wait to see the demos. Thanks again. Yeah, that is a great question. I think we'll...
Starting point is 01:23:47 Lindsay, you want to take now in? Because the capabilities of this model presents are just totally new and different from the GPT series. So we're still going to invest in the GPT series. Like, 4-0 is not going away. We're still very much invested in multi-capability. But we wanted to represent sort of this shift in capabilities towards reasoning, which is just something that while models have shown us that limits of being able to reason by like stitching things together in the past.
Starting point is 01:24:19 This is just net name. So we started over with 0. O stands for open AI, not an O like Omni Model and GPD4O. You know that we're not great at naming models. It's really hard. Just on the naming, I'll just follow question. I read someone saying somewhere,
Starting point is 01:24:37 I cannot find this one, that there was no GPT5. And like, is this a, correct. Is this a, is it R-A-1 and restarting a-up? Is it R-A-P-G-T or? We're still going to invest in the GPT series. I don't know if we're going to have the GP5, but this is 01. It's different.
Starting point is 01:24:53 Okay. I think you should expect us to continue doing the kinds of model historically people who are thinking would be GPD5, what we actually call it as a name. Something else. But again, our naming has always been confusing, so just a little set up. The bad idea. Everyone is kind of baddened. So it's like similar to OSX to ACO's 11. Yeah.
Starting point is 01:25:17 So question, how would you just phrase it? So for people that don't follow super closely, there's the 4O series, the traditional four, these new ones, really the categories. I mean, you can think about the GPT photo series, for instance, like that's your workhorse. Like we still expect developers, builders to rely on on GPT photo heavily,
Starting point is 01:25:38 and we keep on investing in that series for sure. And 01 will be for this new class of apps and that require like a steps of reasoning, for instance. On the API team and the developer experience team and overall at Open AI, we expect that developers, builders like all of you, will actually use both in concert. It's very likely that you'll need 01 and you'll reach for it for very specific, complex, hard problems.
Starting point is 01:26:01 And then for a lot of tasks that work great already with GPT for foro like summarization, labellization and all that, you'll still have the photo, foro mini to rely on. So we definitely expect builders are going to use both in concert. Will there be a way to speed up the chain of top? I don't know how we didn't have time. Do you think now? Definitely that's a high idea to speed up these models because in the end, a lot of the classical approaches to make these faster still apply and we just started your documentation.
Starting point is 01:26:35 So now my day is actively working on a bunch of improvements that in this game is possible. It's also day one for the API. So as Sean mentioned, like when he did his recap, app not everything is exposed in the API just yet because it's a net new model but over time we'll keep on adding features as well and the tools etc to the model pocket that's all basically it's still going to be a change though right but the better it gets the easy many different steps and you guys can make this stuff shorter there's a weird really attention kind of right to speeding up easy base right I mean I can't just think about
Starting point is 01:27:17 Yes, there is a change of thinking, but there are easy things. I'm going to say that are your different mechanisms that can we have to make possible. Yes, obviously because you can go on that move forward. I think the obvious things that we are there are a great context as you can see because now I want my next one is so that is something that is a nice thing like that. And then I'm going to create and work to open and go to the industry and faster and master
Starting point is 01:27:54 at the new organization. And we learn for continuing networked by the way that we're not. But now, I think we're taking what we have, which seems like good quality. Reusing it, even if people is pretty perfect or possible. So we have worked at all partners in terms of clear as well inside of it now that is.
Starting point is 01:28:20 We got this really well. So if I mean, like, how many steps can't take you for the model? What if you're built around the model? That's a lot of it is in the way you could be able to be hard. Okay, system like, still go to be able to decide or specify how much the you want to spend at this time? Or does this tool the model that always be the most decided how much it would spend on? Is it better a model and the question, besides it, right, it's like,
Starting point is 01:28:57 these lines of all these things. There are also some other issues, training, post-training and a little of the core of it is something. This weekend, we're doing a big hackathon here, and of course people will use the new model, and this is like remote as well, so we've got over a thousand participants in like 300 here in SF, But as you guys release a new model and it's different than just traditional elements, what do you want people to build like any crazy different ideas other than just you know summarization bought whatever like it's a very good
Starting point is 01:29:32 I think that's a great question the reason why we're putting these things like out there on day one and rolling out to all developers is because you know we know developers are going to be coming up with the most interesting use cases for these models the recommendations The definition would be anything you've tried to build in the past with a GPT4 class model that may not have worked so well or things that you felt like were limitations of the model. Try that again with 01. Are there like complex challenges, especially around coding for instance, like refactoring a large code base or potentially like trying to to troubleshoot an issue in a large code base? Those kind of clever things that GPT4 would maybe not be the best model for. Try reaching out for like 01 mini for this.
Starting point is 01:30:19 article great at coding. Yeah, and basically the reason why we hear, the reason why we want to talk to all of you is because we love feedback. We're definitely going to hear from what you build and hopefully make those models better for you. So how we get access to O1? Someone was Kevin Guy for access to O1, we all have a O1 preview. Oh, 1 is not like released yet, so for now it's all about O1 preview and O1 Mini will communicate later on about O1. I think that guy works at opening eye wherever this is. Okay, yeah, yeah, yeah.
Starting point is 01:30:49 you can kind of see the thinking steps are all site, you know? Yeah. Are you planning on exposing that eventually? What people get then going to be doing directly on a single step on the channel possibly, yeah. I think for the APIs, it's really early days. I think we wanted to just make it work on day one with chat completions, like the most popular API that people have.
Starting point is 01:31:10 But over times, we'll be trying to add more knobs for the developers to kind of understand what's happening. It's also for the first time, it's the first time, It's the first time we have a model that may take like 30 seconds to think hard about the problem. Like how do we make the developer experience better for your app? I think there's a lot of work ahead of us. Yeah. Following on what he said, the fact that you can now see like little by little like the way it's
Starting point is 01:31:34 person to get the problem and the solution, does that mean there's no visibility like towards why is it giving the answer because we know like there's a big black box and by AI response? So like does it mean that like now there's more visibility towards? like this, like why is this answer the one coming out? Yeah, I think what's interesting about these new models, like you can see the reasoning, right? Like for instance, like we, you might have seen in the blog post today that we had a few videos with like a quantum physicist, for instance, like using the model to go through like equations and looking step by step at like,
Starting point is 01:32:07 oh yeah, I understand what the model is actually doing. So yeah, there is a bit more about like, there's, the chain of thought turns into like a tokens where tokens where you understand like what's happened during the thinking process. I think we're using tokens also go against the context. The lot of the area context. So if it's a muddict of communication and all the reason it does, does that fit into the I think we documented what's happening in our platform docs. I don't have the exact answer to top of mind, but we show that.
Starting point is 01:32:41 I think you met you have that in your diagram. Yeah, right there. So I ran into this using chat gbt. and the thing just stalls, you're not able to enter any more messages. So it doesn't do any truncation for you and lets you continue chatting. Yeah, and the API does here, but I think something we have to inform. You see a setup of thinking and explaining what you're not. All of that gets added.
Starting point is 01:33:11 So what you see is adding. However, each additional step does another thinking process as you've got me a multi-complaint. And the amount of tokens it has to generate to do the thinking also is just like a splashback, right? So as you add in the chat GPD, the subsequent multi-turn step, what you see in the US is the total accumulated tokens, but then the thinking tokens are also a which we don't show the count-off. But they don't accuse it. If you do five thinking steps, all the hidden tokens
Starting point is 01:33:51 for those thinking steps are given. What are keenness? What are key? So if I do another chat message after that, only the summary that the new X is very well-stained. One thing I noticed also about the thinking tokens
Starting point is 01:34:08 is the sweep bench verified. So Swibang now requires you to submit your thinking trajectories. And that's why these guys, cosine, which I don't know if you talk to them at all while they're here.
Starting point is 01:34:22 I have not personally, but some people at the point I have, yes. So they actually scored higher than 01 on sweep bench verified and higher than this is full sweep bench.
Starting point is 01:34:36 But they could not get on the leaderboard because the leaderboard requires them to submit their thinking trajectories. So it's basically cosine did the same thing that you are doing. which is they're not showing the thinking trajectories because it's IP.
Starting point is 01:34:49 I thought it was an interesting observation. Cool? Yeah, nothing to say beyond that. It's kind of cool, but it's... So, O1 reported their own results, and if you put their results side by side, they're slightly lower. Yeah, this was before, I think.
Starting point is 01:35:05 Yeah, yeah. This is a blog post, this is from the blog post of SweetBedge verified, and then the O'SBer1 results were somewhere here. One of these guys. I don't know where the Sweet Bench one is. Somewhere here. System card. Okay. Yeah. Anyway, cool. I want to keep you too long, but thank you for that. You had one more? Yeah, I think all we have to share on this is is already on the black post at this time. Yeah, I've been a question to that. Is the current model that's available? It seems to be over 20 seconds. Is it still like a maximum?
Starting point is 01:36:11 would be a mineral thinking or is that like a cap on the tool that you guys are a lot of the current model? I think we shared like on the API side like the current way it's working but for for everything else like I think we are very flexible at the moment I'm sure we'll learn along the way I don't think there's anything in particular that would be like oh this is the number this is the cap I see so in theory if questions complex It should be Focke could be voting for a longer time
Starting point is 01:36:46 For a second question Is that all right? Yeah Yeah When you ask how many rs in strawberry for instance Now it gets it right but it's also pretty fast Versus like getting a very complex pump Yeah this question the this is the good world hunting question
Starting point is 01:37:08 It took one hundred and twenty one second I... Oh, wow. Yeah. Like, the same way, Judge, it would be whatever, or something, whatever, plus about some time because you just didn't have a focus, you can't answer, just mean that at some point you can just cut up thinking and you just don't get an answer because it's a job count for the recent case. That's right.
Starting point is 01:37:29 That does up, yes. Yeah, one last question. So, given that you're doing one year of driving your foundation model, right? model and try to debate one layer here right does this kind of mean that we have reached some kind of at the foundation model here that we come here and like we're trying to debate like what are higher in here I said I don't think what's a number right
Starting point is 01:38:00 okay other than you can do you have another but the amount that is also Oh boy. Okay, last one and then we have to move on. Sorry. Did she post any open-upon any event or any e-well or something like to talk about how would you be for a chain of thoughts with the same number of thoughts as
Starting point is 01:38:30 one for one question? Is there any like, how was the difference like basically in emails or? I think we've published any of that. I mean, all things about you've published out. Take a look at our research clubs. There's a technical record out of. Yeah, there's a lot of material on the O1 hub. There's two research blogs, one for model.
Starting point is 01:38:51 There's a system guard that's 85 pages long if you want to be holding. We do paper club Wednesdays. There's a great forum for it. There's a lot of material on there, so read through it. There's a lot of e-bos. I use a word to summarize. Sorry? I use a word to summarize.
Starting point is 01:39:05 It cannot use tools and doesn't have images, so very hard, very hard. But thank you so much. Cool, no, thank you all. Thank you, thank you, everyone. I think we have two demos. We're going to host an AMA on Twitter tomorrow morning. If you have more questions, we'll have a bunch of researchers, right? Yeah, that's right.
Starting point is 01:39:29 From the OSHAI Dance handle. Wow. We'll send a party full text files. No. Cool. And what else was I going to say? So yeah, it should take place around 10 a.m. tomorrow. So I think we'll post a tweet between 8 or 9 something like this to kind of source questions
Starting point is 01:39:45 and then we'll have some researchers to go deeper into some of those options. When you wonder, is we're going to be a Dev Day in October? Yes. Sweet. Awesome. We'll see this. Okay. Thank you all.
Starting point is 01:39:57 Yeah. I think we have like two demos. And first is the 01 Visa demo and then the second is the 01 visa demo. and then the second is the 01 model demo. But actually it's also AI enabled, so you can come up. That was the 01 emergency meetup that happened the evening of Thursday. See show notes for full video and photos, and a huge thank you once again to the OpenAI team for supporting us.
Starting point is 01:40:21 More excellent updates are coming for OpenAI Dev Day 2024, and you can bet on the latent space crew being there to take you through it. For part two of our 01 coverage, I'm going to recap the top takeaways from the Friday Twitter AMA done by the research team. Tybor Blahoe asks, How did you come up with the new names for 01-01 preview and 01 Mini? World's Fair Speaker, Romaine Hewitt, from OpenAI Answers. Reasoning represents a new level of AI capability, so we decided to reset the counter back to one and introduce this series as OpenI.
Starting point is 01:41:02 A.I.01. Preview because it's a preview of the capabilities and mini because it's smaller. And O stands for OpenAI. Former guest Simon Willison asks, why is it called O1 Mini and not O1 Mini preview? Shangji Zhao from OpenAI answers, you're exactly correct here. O1 preview is a preview of the upcoming O1 model while O1 Mini is not a preview of a future model. O1 Mini might get updated in the near future as well, but there is no guarantee. Max Schwitzer from OpenAI adds, We can't discuss the precise sizes of the two models, but O1 Mini is much smaller and faster,
Starting point is 01:41:48 which is why can offer it to all free users as well. O1 preview is an early version of O1 and isn't any larger or smaller. Friend of the pod, Alex Volkov, asks, Can you guys clarify, is O1 a system that runs chain of thought behind the scenes and gives us an answer or a model that reasons with special tokens and just hides those tokens from us at the output and just shows the final answer? Many folks are confused by this. Noam Brown from OpenAI answers. I wouldn't call O1 a system. It's a model, but unlike previous models, it's trained to generate a very long chain of thought before returning a final.
Starting point is 01:42:32 answer. We don't have plans to reveal COT to users either in the API or chat GPT. There is no guarantee the summariser is faithful. Hyeong-won-Chung from OpenAI ads, answer tokens are typically, though not necessarily, a lot shorter than the COT. It's a single model. We're not sharing param counts right now. Thinking stage is a summary of the thought process. So it appears to be slower than it is. is a single model. Jason Way from OpenAI ads, I think it's safe to say that no matter how hard you prompt GPT4O, you probably won't get an I.O.I. Gold. Many, many people ask, will lower upper bounds of thinking time or test time compute be a controllable
Starting point is 01:43:21 variable via API? Noam Brown from OpenAI answers. In the future, we'd like to give users more control over how much time the model spends thinking. Heung-Wanchung from OpenAI Answers. OpenAI O1 preview doesn't use tools. For the future models, we are considering adding support for function calling, code interpreter, and browsing. John O' Whitaker from Answer AI asks, How are the team finding it for research code?
Starting point is 01:43:53 HTML Snake is cool, but I'd love to hear their own use cases from the research trenches. Wukas Kondratiukh, Strawberry Training Infra Lead at Open. A.A.A. Answers. A.O.A.C.C. was already authored solely by 01. Many, many people ask, what are the plans for the next steps? Preview phase duration, availability of the real, 01 from benchmarks, and missing functionality tools, etc. Ahmed Elkiske from OpenOI Answers. While we don't have the exact preview duration, we plan to iteratively deploy additional functionality, including including tool capabilities like code interpreter and browsing.
Starting point is 01:44:36 Today's guest Michelle Pocras of the API team adds, We know we're missing a lot of the features that developers need. We're working to enable function calling, structured outputs, developer system messages, and our other standard params. Jerry Tawrick from OpenAI responds to a question about vision, aka image recognition, hopefully soon, working hard on it, not providing any hard date yet. Andrew from OpenAI ads,
Starting point is 01:45:04 OpenAI O1 is built with multimodal and achieves SOTA on MMU. We're working hard on making it safe to share with everyone. Nikunj Honda from OpenOI ads. We'll add batch AP support once the rate limits are higher. Many, many people asked about pricing. Shangji Aja Zhao from OpenOI answers? Historically prices go down 10x everyone, two years. The trend will probably continue.
Starting point is 01:45:30 For example, the cost per token of GPT40 Mini dropped by 99% since Text-Divincy-003. Swix asks, What inverse scaling have you seen under 01? Jason Way from OpenAI answers. I don't know of any great examples of inverse scaling. You can see from our blog post that on some types of prompts like personal writing,
Starting point is 01:45:55 it seems like OpenAI01 preview is not much better than GPT. P-T-40 or even slightly worse. Finally, I will read you the recap from Taiba Blahoe, which we link in the show notes. Model names and reasoning paradigm. OpenAIO, 1 is named to represent a new level of AI capability. Preview indicates it's an early version of the full model.
Starting point is 01:46:20 Mini means it's a smaller version of the 01 model. Optimized for speed. O as OpenAI. O1 is not a system. It's a model trained to generate long chains of thought before returning a final answer, size and performance of 01 models. 01 Mini is much smaller and faster than 01 Preview, hence offered to free users in future. O1 Preview is an early checkpoint of the O1 model, neither bigger nor smaller.
Starting point is 01:46:49 O1 Mini performs better in STEM tasks, but has limited world knowledge. O1 Mini excels at some tasks, especially in code-related tasks compared to O1. 1 preview, input tokens for 01 are calculated the same way as GPT-40 using the same tokenizer. O1 Mini can explore more thought chains compared to O1 preview input token context and model capabilities. Larger input contexts are coming soon for O1 models. O1 models can handle longer, more open-ended tasks with less need for chunking input compared to GPT-4-0, O1 can generate long chains of thought before providing.
Starting point is 01:47:29 an answer unlike previous models. There is no current way to pause inference during Code T to add more context, but this is being explored for future models, tools, functionality and upcoming features. O1 Preview doesn't use tools yet, but support for function calling, code interpreter and browsing is planned. Tool support, structured outputs and system prompts will be added in future updates. might eventually get control over thinking time and token limits in future versions. Plans are underway to enable streaming and considering reasoning progress in the API.
Starting point is 01:48:08 Multimodal capabilities are built into 01, aiming for state-of-the-art performance in tasks like MMMU-COT, chain of thought. Reasoning, O1 generates hidden chains of thought during reasoning, no plans to reveal Cotee tokens to API users or chat GPT, COT tokens are summarised, but there is no guarantee of faithfulness to the actual reasoning. Instructions in prompts can influence how the model thinks about a problem. Reinforcement learning, RL, is used to improve COT in 01,
Starting point is 01:48:42 and GPT40 cannot match its COT performance through prompting alone. Thinking stage appears slower because it summarises the thought process, Even though answer generation is typically faster, API and usage limits, O1 Mini has a weekly rate limit of 50 prompts for chat GPT plus users. All prompts count the same in chat GPT. More tiers of API access and higher rate limits will be rolled out over time. Prompt catching in the API is a popular request, but no timeline is available yet. Pricing, fine tuning and scaling.
Starting point is 01:49:19 Pricing of 01 models is expected to be. to follow the trend of price reductions everyone. Two years, batch API pricing will be supported once rate limits increase. Fine tuning is on the roadmap, but no timeline is available yet. Scaling up 01 is bottlenecked by research and engineering talent. New scaling paradigms for inference compute could bring significant gains in future generations of models. Inverse scaling isn't significant yet, but personal writing prompts show a one-party preview performing only slightly better than GPT40, or even slightly worse, model development and research insights. O1 was trained using reinforcement learning to achieve reasoning performance.
Starting point is 01:50:03 The model demonstrates creative thinking and strong performance in lateral tasks like poetry. O1's philosophical reasoning and ability to generalize, such as deciphering ciphers, are impressive. O1 was used by researchers to create a GitHub bot that pings the write code owners for review. In internal tests, O1 quizzed itself on difficult problems to gauge its capabilities. Broadworld domain knowledge is being added and will improve with future versions. Fresher data for O1 Mini is planned for future iterations of the model, October 2023 currently, prompting techniques and best practices. O1 benefits from prompting styles that provide edge cases or reasoning styles.
Starting point is 01:50:50 O-1 models are more receptive to reasoning cues in prompts compared to earlier models. Providing relevant context in retrieval augmented generation, R-AGE, improves performance. Irrelevant chunks may worsen reasoning general feedback and future enhancements. Rate limits are low for O-1 preview due to early-stage testing but will be increased.
Starting point is 01:51:14 Improvements in latency and inferences times are actively being worked on remarkable model capabilities. O-1 can think through philosophical questions like, What is life? Researchers found a one impressive in its ability to handle complex tasks and generalize from limited instruction. O-1's creative reasoning abilities, such as quizzing itself to gauge its capabilities, showcase its high-level problem-solving.
Starting point is 01:51:41 That's it for the AMA. Now for demos from Min and Alessio. So exciting. Thanks so much for having us. Wow, super fun. This is Coincidental. Hi, everyone, my name is Min. I'm the founder and CEO of Lighthouse, where we make US immigration go fast.
Starting point is 01:51:56 This is coincidentally a fun day, because we're talking about the 01 model, and we specialize in the O1 visa. Is anybody in the room on an 01 by any chance, just out of curiosity? Amazing, cool. Just in case for anybody who's unaware, I was born in South Korea.
Starting point is 01:52:14 This is like near and dear to my heart, but mostly because I've spent my whole career working with early stage technologists and really amazing talent, and we make it really hard for you to continue to be here and build what you want to build. The interesting thing about something about U.S. immigration that you might not know is that there's a sort of best-kept secret called the O-1 visa, which is an employment visa that is meant for really popularized by startup by the technology or the entertainment industry originally. So for a long time it was used by models, musicians, so on and so forth. Melania Trump got in on an 01 type thing. And then over the last five years, five, six, seven years or so,
Starting point is 01:52:53 it's become really popularized by the startup and technology industry for highly technical talent and entrepreneurs. But the problem is that it's still really underused. So if you think about just to give you some context, there's only about 11,000 people who get it every year, and that is in contrast to the 750,000 people who applied for the H-1B three years ago. So it's just hugely underused visa. And why is that the case?
Starting point is 01:53:15 It's because most lawyers are really bad at it. Most providers don't know how to do it. And then when you go through the process, you end up having to do a lot of it yourself. And for those of you who have been through the process, you may or may not have had understand this, where the actual application is like a 3 to 400 page research report about you and all the amazing things that you have done.
Starting point is 01:53:35 And how those professional or academic achievements map to the criteria that USCIS wants to see. It's sort of like applying like a really, really intense college application, like on steroids. And at Lighthouse, we want to make this process go way faster because traditionally if you go through this, one, you'll often get told you don't qualify. And then two, you'll end up having to do so much of it yourself and it'll take you six months. And it really doesn't have to be that way. So what we do is try to make the process as simple as possible.
Starting point is 01:54:03 We kind of like 30 days or fewer. You can do it in four weeks. And then the government can back to get back to, we'll get back. you within two to three weeks or so, and so the whole thing doesn't have to really take more than two months, and we were really lucky to get to support SWIX, and so in some ways this is SWIX's O-1 congratulations party as well. Yay! And we would love to have love to support more people, and part of the reason I'm here today is because I kind of wanted to give you a behind-the-scenes look of, we've basically built a platform
Starting point is 01:54:30 that's sort of like a paralegal out of the box, so that you don't have to spend time going back and forth with an attorney and spending hours and hours and hours drafting a lot of this yourself. We are happy users and customers of OpenAI. Please give us access to the mini because the model itself actually said, I think in your paper that it is really good for STEM problems. And a lot of the people that we work with are STEM talent. And so a lot of what we end up having to do is translate the technical papers that they've published and the work that you have done in a way that a layperson, a USCIS officer who's probably, you know, 40, 50 years old. probably spend some time in the military and doesn't really have context about what makes AI special right now.
Starting point is 01:55:10 And an amazing thing is right now everybody wants AI talent in the US and so if you are working in that field we definitely want to help you be able to stay and come here. So what I'm going to show you is actually I'm just going to do this on the fly and so we'll test the boundaries of all this. But let's say we were going to do let's I actually just picked Daniel because he's the second author on the paper And so we're going to do it in the fly. And so let's say he is an expert in artificial intelligence, just ignore that. Hold on, let me just make sure I can... Yes?
Starting point is 01:55:47 Oh, oh my god, can you not see? Okay. Is that better? Amazing. This is what we've built so that we can make this process go fast for you. And let's say goes to open AI, cool, let's just ignore that. We don't need everything right then and there. and what is really cool is once we can oh okay give me one second we're gonna do it here but let's pretend that Daniel needs we have to talk about his papers we can just add his page and let's pick his like first five papers and what this is actually going to do is
Starting point is 01:56:29 there's a whole bunch of stuff we've built in the background here but we basically fetch all the papers, read all of them, summarize them. And then we've layered on top of it a ton of industry know-how, basically like all the legally speak based on like tons of feedback from the legal experts that we work with. And so it takes a little bit of time because there's a lot of content that it's going through.
Starting point is 01:56:49 But it'll basically spit out the entire thing in terms of what we need to do. And all we need to do is edit and revise it. And then we'll show it to you. And then you make sure that it like is sound, right? And we can do this with a whole bunch of other things. It'll usually take like a minute or two because it's like literally picking up all the papers
Starting point is 01:57:07 and then reading them and then summarizing them. But so we'll let that run. But basically we can do that and we do it to find out whether or not your pay comparables are high. We do it to figure out what the papers that you've written or research that you've done. If you honestly we can put like this into it and have it have it like synthesized in a way
Starting point is 01:57:29 that is intended for USCIS to read. And yay, it's done. And so it'll basically tune it to whatever we decide that we wanted to focus on, right? And so in this case, we happen to want it to be focused on machine learning product. And so everything is kind of like oriented around that and it will summarize all the papers that we asked for, the five that we wanted. It'll pick up all the papers that we want. And so if you're familiar, SWIX, we like gave you this like 500 page document that we said we're going to literally physically print and ship to USCIS. And we basically turn a process that's probably 120 hours when you work with a traditional provider into something that's probably closer to 20.
Starting point is 01:58:11 And so that's how we can do it fast. And we would love to work with the 01 mini model because a lot of what we have to do is make sure that this is as precise as possible without while still being precise as possible. Yeah. And then just for kicks, I want to know, I want to show you what it might might. look like what it might say if 01 mini what might it say if we asked it to go read this and it should oh yay so we would say Daniel developed opening out amni and it's a cost-efficient reasoning model designed to excel in stem fields and so it'll basically summarize all of this in the
Starting point is 01:59:10 the way that it needs to be kind of like written for the officer. And in essence, the way that it would typically work is either you would have to write this for the lawyer because they wouldn't really know how to summarize this. And you don't really know what good looks like because, you know, probably, really, for, reasonably you would say, well, I wrote an abstract. That's what the abstract is for. But it's not written in the way that USIS would understand. So you would probably have to spend a lot of time doing this or they would do it and they
Starting point is 01:59:38 would do a shoddy job. And so we kind of combine the industry know-how and the magic of AI now to make this process a lot more efficient. So if you are looking for an O-1 visa, we also support all employment-based visas to the US. So come find me. My name is Min. All right. So the idea here, as Sean mentioned, the O-1ABI doesn't actually have a lot of the things like
Starting point is 02:00:13 a structure, up-o, tool calling, everything like that. So I spend the last three hours, try to give those tools to... to it. I'm using a thing called E2B. I do not work at E2B, but we invested in E2B, and E2B is the only thing that actually does this. It's basically like a cloud sandbox for your LLM to use. So the way I built it, I take a prompt. In this case, I'm from Rome, so I wanted to see a visualization of the growth of the Roman Empire, as one does. And I had to remind O1 that it has access to a code interpreter, even though it doesn't know it. And then what I did is use 4-0 Mini to extract the code. So 01 is just doing the planning.
Starting point is 02:00:49 Obviously, it writes the code in the actual output. So if you see this is what 01 actually returned. So it's like, you know, these are the steps. This is like something we can use for the data visualization. This is like explaining all, you know, a lot more information that it's going to go in the actual thing. This is how you practice it. And then what I've done is just putting 4O mini as a way to extract all of it.
Starting point is 02:01:15 So it's like, you know, your software engineers, that receives the plan and then from the plan you create one single script that does everything that you need to execute it. And I've created this simple kind of like structure output model. So there's the code that needs to be run, all the PIP packages that are required to run the code, the file name, and then the command to execute. So what I'm doing after that is basically open a sandbox. Actually, the only reason why this doesn't work, we can not get the PNGs out of the sandbox right now. We're trying to figure out the right path. But basically, you install, all the PIP packages, and then you run the command that the model generated.
Starting point is 02:01:52 And then there's also a self-healing loop. So if there's an error with the sandbox, it tells you what the error is. I feed that back again to the model, and it's like, hey, actually, this happened. And then I use Foro Mini again to extract the new code, rewrite to the file that I was running before and kind of like loop through it. And so what that looks like, we cannot actually get this file out, but it starts from this document, which is like you know a very kind of like it's a good plan but like it's not something you can actually run it goes through 4-0 so there's one error which is most of the time people are not
Starting point is 02:02:28 using negative dates in pandas so the the o-1 is actually saying hey you know in the year minus 23 and pandas is like what does that even mean nobody ever says that but it is a valid date so you actually have to convert to a different format to can only do to do before christ dates so that's something that 01 didn't do, but 4-0 Mini can do and update the code. So yeah, I'll publish this on GitHub just so people have it, but this is like an easy way to get from 01 planning to like 4-0 Mini code extraction and then actually have a runtime that does it. As you might have seen in the security review of the model,
Starting point is 02:03:03 it's like quite good at Jail Break, so I did not want to run it even in Docker in the local environment. So if YouTube B gets hacked, Vasa gets right there, he's gonna be mad, but yeah, this is this is it and we'll send out the link after and I'm actually going to run the code locally. I just kind of copy-based that thing here. So if you just do Python, that's the pi.
Starting point is 02:03:28 These are like all the charts that actually get generated. So this is like the empire population over time. This is the territorial areas over time. And these are kind of over-lap together. So again, this is a data set that there's not really a good source out there. for it, there's kind of like all these different pieces. So I think one of the most exciting thing about these models is also like dataset generation, you know, like how do we get
Starting point is 02:03:52 take all of this tax data that has been trained into it and bring it out in a structural way. So yeah, that's that's it.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.