Latent Space: The AI Engineer Podcast - ⚡️GPT 4.1: The New OpenAI Workhorse
Episode Date: April 15, 2025We’ll keep this brief because we’re on a tight turnaround: GPT 4.1, previously known as the Quasar and Optimus models, is now live as the natural update for 4o/4o-mini (and the research preview of... GPT 4.5). Though it is a general purpose model family, the headline features are:Coding abilities (o1-level SWEBench and SWELancer, but ok Aider)Instruction Following (with a very notable prompting guide)Long Context up to 1m tokens (with new MRCR and Graphwalk benchmarks)Vision (simply o1 level)Cheaper Pricing (cheaper than 4o, greatly improved prompt caching savings)We caught up with returning guest Michelle Pokrass and Josh McGrath to get more detail on each!Full Video EpisodeTimestampsPart 100:00:00 Introduction and Guest Welcome00:00:57 GPT 4.1 Launch Overview00:01:54 Developer Feedback and Model Names00:02:53 Model Naming and Starry Themes00:03:49 Confusion Over GPT 4.1 vs 4.500:04:47 Distillation and Model Improvements00:05:45 Omnimodel Architecture and Future Plans00:06:43 Core Capabilities of GPT 4.100:07:40 Training Techniques and Long Context00:08:37 Challenges in Long Context Reasoning00:09:34 Context Utilization in ModelsPart 200:10:31 Graph Walks and Model Evaluation00:11:31 Real Life Applications of Graph Tasks00:12:30 Multi-Hop Reasoning Benchmarks00:13:30 Agentic Workflows and Backtracking00:14:28 Graph Traversals for Agent Planning00:15:24 Context Usage in API and Memory Systems00:16:21 Model Performance in Long Context Tasks00:17:17 Instruction Following and Real World Data00:18:12 Challenges in Grading Instructions00:19:09 Instruction Following Techniques00:20:09 Prompting Techniques and Model Responses00:21:05 Agentic Workflows and Model PersistencePart 300:22:01 Balancing Persistence and User Control00:22:56 Evaluations on Model Edits and Persistence00:23:55 XML vs JSON in Prompting00:24:50 Instruction Placement in Context00:25:49 Optimizing for Prompt Caching00:26:49 Chain of Thought and Reasoning Models00:27:46 Choosing the Right Model for Your Task00:28:46 Coding Capabilities of GPT 4.100:29:41 Model Performance in Coding Tasks00:30:39 Understanding Coding Model Differences00:31:36 Using Smaller Models for Coding00:32:33 Future of Coding in OpenAIPart 400:33:28 Internal Use and Success Stories00:34:26 Vision and Multi-Modal Capabilities00:35:25 Screen vs Embodied Vision00:36:22 Vision Benchmarks and Model Improvements00:37:19 Model Deprecation and GPU Usage00:38:13 Fine-Tuning and Preference Steering00:39:12 Upcoming Reasoning Models00:40:10 Creative Writing and Model Humor00:41:07 Feedback and Developer Community00:42:03 Pricing and Blended Model Costs00:44:02 Conclusion and Wrap-Up This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Transcript
Discussion (0)
Hey, everyone. Welcome to the Lidenspace podcast.
This is Alessio, partner, and TTO Addacible, and I'm joined my co-host, Swix, founder of Small A.I.
Hey, and today, and today we have a returning guest as well as a new friends. Welcome, Michelle and
I guess, Michelle, I think used to introduce you as manager on the API team. It seems like
you've changed your role since we last talked on the podcast. Yeah. Now I lead a team
on the research side, specifically in post-referential.
training. Yeah. And Josh, you are also on post-training. Yep. I'm a researcher on Michelle's team.
Yeah. And I just found an interesting commonality you guys have. You're also both from Waterloo,
continuing the tradition of extremely cracked engineers. Oh yeah, we talked about the last time.
That's right. Okay. So we're gathering to talk about GPC 4.1. You launched it. I mean,
we got a little preview and it was a little bit rumored, right? It was pre-released, I guess,
with Open Router as Quasar Alpha and then it was also an Optimus.
version and I think people are trying to figure out like why are we going back from 4.5
through 4.1. You know, there's a whole bunch of other things. But like what are the headline
facts I guess you guys want to emphasize about 4.1? Yeah, I'll just say we released three new models
today, GPD 4.1, GPD 4.1 mini and GPD 4.1 data. And the real focus on these were just
making models that were great for developers. So we improved instruction following, coding, and shipped
our first one million context models.
Josh, anything to add?
I don't know if there's anything else
that people should really,
that are like in the fine print.
No, I think the only thing that I would touch on maybe twice
is that there's actually a new model in the lineup,
Nano, which is even faster for developers
that are making low latency applications.
And cheaper.
What's the, any fun story behind the code names?
Or, you know, I got the strawberry hat
as another fun time in the lore of a binet.
Yeah, yeah, we really wanted to get as much developer feedback as possible on this model to make sure it worked well in the real world.
And so we tested it kind of through open router and it was super cool to see people latch on to the names and get the theories going.
But the feedback we got from there was super helpful.
Yeah, yeah. It's not even like the name. It's more about just like the API shape.
Once we saw like chat cumple, it was like very obviously open AI.
Yeah, it's a good note.
Yeah.
But like, I mean, okay, is there like an emphasis on stars?
Like, what inference were we supposed to draw from, you know,
quote-unquote super massive black holes?
I don't think there's anything really to draw from there.
Okay, they're just cool.
They're just cool.
You know, they make you think of cool concepts.
The vibes are good.
The vibes are good.
You know, yeah.
Yeah.
The other thing about the examples, we're just mining for lore here, right?
The interesting animal comes up a few times.
on the live stream and on the blog posts,
what's up to Tapirs?
Who likes Tapirs here?
Yeah, our team is just a super big fan of Tepirs.
So they just happen to work their way into a lot of our content.
Okay, cool.
Awesome.
Yeah, go ahead.
Yeah, go ahead.
I think like the first thing that, yeah,
we just want to run through is obviously the 4.1 to 4.5.
I think that's the first thing that everybody was maybe confused about.
So I don't know you're demigating 4.5.
It sounds like 4.1 is just like a kick-ass.
model and the 4.5 size, maybe it's not as good of a fit. That was just a research preview.
So yeah, I don't know. Whatever you want to say to address that, I think it's something we've seen
come up also in the Discord. Yeah, totally. Okay, naming is really hard, and we've tried to make
this as less confusing as we can, but, you know, nothing's perfect. Basically, the way we got here
is that GPD 4.1 is like a pretty big improvement over the 4-0 line, and we really wanted to
signify that.
However, it's a model that's like much smaller and cheaper than GPT 4.5 and as a result, you know, doesn't achieve the same like Amy or other intelligence e-vows.
So it doesn't beat 4.5 on all of the evils and so we didn't think it made sense to increment beyond 4.5.
But we do think for most developers, they can kind of replace a lot of their 4.5 usage with 4.1.
And then the mini is strictly better than 4a mini.
Yeah.
Yeah.
with the nano.
But like, we don't know if 4.1 is a distillation of 4.5 or there's no relationship there.
Like, what can we say about like the shared lineage?
Yeah, what I'll say there is we're always using various research techniques to improve our models.
And distillation is something we talked about before.
It's really meaningful, especially for the small models.
And we've kind of pulled out some of the things that made 4.5 really good.
Like it has a lot of the instruction following greatness and also rolled that into 4.1.
Awesome.
I think one of the, because I strongly remember on the 4-0 launch that their communication was that we're kind of moving to a new model architecture that is Omni-Model, right?
That's what that's the O in the 40.
And then 4.1 is part of this subsequent trend of trying to merge everything, like the reasoning model, the Omni-Model, everything.
And I think there's just doubt about whether 4.1, I think it's basically trying to be sold as a strict replacement for 4-0.
But I don't know, is it going to be fully omnipal?
Is it roughly the same architecture that we think 4-0 has?
So we already have different slugs on the real-time API and like responses API.
So they're already, you know, somewhat different checkpoints.
We don't have any current plans to release 4.1 in the real-time API.
But, you know, things may change.
Yeah, and then there's image gen and all that, right?
Like, so as far as we know, no plans may be about nothing announced.
Not right now.
The focus for 4.1 was kind of these three core capabilities for developers.
Yeah.
Our Discord actually also did a launch watch party for the recent 4.5 podcast that Sam Altman did,
where I think for the first time, it was basically kind of confirmed that, like,
something that people already knew, like, Andre Carpathie was already talking about this.
that 4.5 was like 10x the size of 4.
And I think there's a question about like,
do we do the linear interpolation of 4.1 is like, you know,
0.0. It's like, I don't know, 2x size or something.
That's not really how we think about naming the models.
There's a whole lot of different parts that go into the recipe.
And so, you know, it doesn't really reflect on just the pre-training recipe,
our version numbering.
But I think the 4.1 is just because of the large jump that we have in like coding capabilities,
long contexts,
and so on. It's more so what it's like for the end user, more so than anything about the training
recipe. We can go a little under the hood on training, though, and we'll say that, you know,
nano is obviously a new pre-trained. We also have a new pre-trained for mini, and then the larger
version is a new mid-train. But we find that actually a significant amount of the gains
come from new post-training techniques. And so I think in the past, the narrative is that
you need to pre-train these larger and larger models to get better performance.
performance and we're finding that we're able to squeeze a lot more out of post-trading now.
Talking about how big a model is, the other side of it is the context window. You have a 1 million
context. I know that Sam at that day last year. You said that 1 million was like months away,
so right on right on time. Can you talk about how hard that was to get to 1 million and then
maybe where the end game is in your mind? Is it 10 million, 100 million infinite? What really
matters as you start to scale this. Yeah, Josh worked a lot on long context, so he's the right
person to ask. Definitely. So I think the first thing that I thought was really interesting when we
were going to long context is actually some of the evils that you see as like headlines on maybe
other blogs where it's needle in a haystack. Actually, most of the models do really well right out of
the box. But then we had to actually first get a lot of measurement on the longer context for long
context reasoning. So, you know, we actually just open-sourced two new evaluations that are about
using the context in a more complex way. So, you know, one of them you have to reason a lot about
ordering and the other is actually walking through graphs. So there's a lot of reasoning that you
have to do in those data sets. And that's where doing long context is actually much harder.
But single needle in a haystack, we were able to saturate pretty easily. And then the most
of the work came at these harder tasks. Yeah, I was going to say how, no, just
just how you think about length of context in terms of consuming documents versus like active,
kind of like thinking and planning.
I think there's obviously a whole part on the prompting side around billing a gentic workflows.
Do you think that people maybe like still think too much of it about, yeah, neither and a
haystack kind of document retrieval versus like traversing very long plans and kind of like iterations
in context?
Yeah, I think the mental model that I have is maybe actually has some more variables in it.
there's the needle in a haystack where you have like some amount of distractors and some, you know,
needles that you're trying to find. And I think that it's more so about how dense of the context
do you need to use. So like summarization, you're actually just using the entirety of the context,
whereas, you know, needle in a haystack, it's very sparse. And then I also generally think about
orderness. If you're going to make some sort of inference on this, do you, are you just looking,
you know, sort of front to back or do you need to move around in the context in order to
to generate a good answer while the model's sampling.
Yeah, is that something that you worked on with graph walks?
Is that the thing?
Yeah, that was sort of the most synthetic and clean way to measure the model.
And then, you know, we worked on a lot of other training techniques, data to sort of test
the model's ability and train in the model's ability to reason throughout the context
in a sort of shuffled way.
Yeah, you know, actually I have the ability.
I like to give people a little bit of visual aid with these things.
So I actually went into your Hugging Face release and got an example of the graph task.
And so there's a few versions of this, right?
There's like the BFS and DFS version.
And also it's, I guess it's very character specific.
So I don't know, maybe could you tell us like, you know, design choices around this?
Like what was surprisingly hard?
You know, anything like that.
Yeah.
So the idea here is you take a graph and you encode it into the context by looking at the edge lists and just putting that into the context.
and then asking the model to do an operation.
And then, you know, under the hood,
we're actually just executing the real operation
and using that to then evaluate the model's ability to work.
One of the things that I found surprising at first
was what the model would do when it wasn't sure how to use its context.
You know, early versions of the model just sort of looping,
saying like, oh, no, I can't find this edge that I think should be there.
And, yeah, I think I was actually very surprised
how all models seem to have more difficulty than I would have expected on a task that we would find very simple or like, you know, maybe an undergrad could write a Python script to run in a couple of minutes.
Yeah, right.
Okay, so like, what is the, like, what is the real-life task that this is meant to model, I guess?
You know, I feel like the other one, MRCR, seems a little bit more intuitive where,
you know, you have like four different stories and you pick out the second one. And that's a real task that
people have. But people don't really traverse graphs. Like this is a bit more theoretical. But like,
you know, was there any sort of correlation study done? Yeah. This is actually meant to be sort of the
idealized version of like a multi-hop reasoning benchmark. So we have a lot of things where, you know,
you're putting hundreds of documents into the context. And then you might ask a question that you
actually have to traverse 10 documents for. But there, the edges are.
they're implicit, right? Like, there is some underlying graph that's connecting all of these
documents that you need to traverse in order to answer the question. But they're actually much
harder to traverse because the edge isn't actually given to you. And so the question there was like,
okay, if I actually just give you all of the IDs of these things that you need to traverse,
can the model even do that? Where it's like it's actually just a lower bound on how well the
model can do. And I think that's actually somewhat well reflected in some of the internal
benchmarks we have that are using more natural data.
Imagine something like a tax return, right, where you like upload the entire tax code.
Like to figure out what to put into this, you know, box, you'll need to reference all of these boxes.
And so this is like a similar level of multi-hop reasoning.
But again, like Josh said, all of the references are implicit.
Yeah, I think that like some kind of backtracking if it's needed is also super interesting, especially for agent work.
For listeners who've been listening to us for a while, we actually covered this paper in New Reps last year.
two years ago, Coggyvel, where they actually modeled graphs for graph traversals for agent planning.
And it reminds me closely of that.
It's just that they never came up with this exact format that you have here, which basically is the same thing.
I also like that you included blank answers because sometimes people do hallucinate,
or models do hallucinate answers.
And you have a fair amount of blank ones.
Thanks, the random sampling over graphs I did, I guess.
Yeah.
Is this tied also to the file search API that you release recently?
Like, how should people think about how everything kind of comes together in the API?
Yeah, I think oftentimes with retrieval, you might be using RAG to fill the context.
And a lot of this is like to get around a limitation of a short context window.
So we do expect a lot of developers to start, you know, uploading their full context more directly to the model.
So for smaller tasks, you maybe don't need the whole vector store.
But we do anticipate this to play well with that paradigm as well.
Like maybe you can just insert way more chunks into the context.
So we think it'll play nice.
Yeah.
Any relationship to the memory upgrades in chat GPT that we recently got,
is long context just directly usable for memory,
or should we just always have a separate memory system?
Yeah, it's a good question.
And so right now, the dreaming feature, we kind of have some of these memories embedded in the context.
But, you know, they are separate features.
So 4.1 is powering the API, whereas the enhanced memory is chat should be too only.
Yeah.
Awesome.
Yeah, I think that's interesting.
I guess the one last thing I'll call out on long context, which is kind of unintuitive or maybe there's an explanation, which was the, you had, you had, you had,
two needle for MRCR and then we had four and eight and everything kind of just regresses to some
kind of baseline of like let's say 30 percent or 20 percent as that but it's interesting to see
where the smaller models sometimes match or outperform the larger models i was wondering if there's
anything unusual there or do you think it was like a bad roll of the dice i think it's probably
just a bad roll of the dice i think i would probably look more so at the the larger amount of
These things regresses you and increase the number of needles because there's sort of more complex reasoning that has to do about the order of different things in its context.
Awesome.
Yeah, cool.
Happy to move on from there.
Yeah, we have a whole bunch of other evals that we can go over.
So I had in my notes that we could talk over, you know, anything that you want.
There was also, like, Collie from Shun You, who we have on a podcast for instruction following.
And I realized that, you know, he joined OpenEye and I wonder if he had a role to be.
play in that one. No, we did not collab on it. Honestly, I think it's best when e-vail authors and
model developers don't collab too much because you, you know, as objective as possible, not trying
to gain any e-laws. Yeah. And then I think there was also, like, for the first time, the announcement
of the, or shout-out of the internal instruction following benchmark from API data. People have
had the ability to opt-in to share data for a while. Actually, I, like, published a, I post-
a tweet out because I found it in the dashboard that you can just opt in and like there's a
there's basically 16 days left for this program where you can to get free inference and like so I'm
just kind of curious like what you found from that kind of IF eval that that might be different
from the normal IFEval that people have yeah totally a lot of the instruction following evils that are
open sourced you know or crafted in a way that are easy to craft so for example like
craft walks is is somewhat easy to craft
Like you can create this graph and verify it easily, but it is not exactly aligned with what the users are doing.
And this is true for some of the instruction following evals, where you ask the model to output exactly four words or, you know, three paragraphs or stuff like that,
things that you can verify easily in code.
And these are useful instructions, but we find that many of the really interesting instructions are actually challenging to grade.
And so the open source evals often don't have them.
And so getting this like real world diverse set of data actually helps us find like what are the commonalities and what developers are doing.
What is a really good example of like a negative instruction?
And then we can go from there and figure out how to how to evaluate it.
Yeah.
I think that's also an interesting question of like what domains do people use you on?
And I wonder if like there's a way to tell you.
Because sometimes it can be very confusing if I, for especially because maybe I'm building.
building an app and letting people use my key, but other people are building apps on top of me.
So you have just a lot of chaos of like multiple degrees of abstraction where you just have to parse through the prompts.
Yeah, it's true. Well, I will say we do use our own products internally where we can. And so we're not manually by hand reading every prompt.
After they're like anonymized, we scrub them of any identifying data, then we use our models to take passes to categorize them.
And so if we get feedback that, like, we're not doing well on ordered instructions,
then we can kind of do a pass over all of our data and find some good examples of those.
So there's an instruction following section in this great prompting GVT4-1 models.
I think maybe we can go through some of these examples.
The first one that caught my mind that it's not necessary to use all caps and other incentives
like bribes or tips, but developers can experiment with this for extra emphasis.
So I think that second part leaves me confused.
Are you saying that people should still try and do this?
And sometimes the model responds positively to it.
Do you feel like it's still just part of the lore?
I'm curious why I would have loved for you to say either, yes, it works.
Or like, no, you should stop.
It looks silly.
I guess the truth is somewhere in the middle.
The truth is always messy.
Reality is that our models have gotten a lot better at following instructions.
Just stated once and clearly.
But we find, honestly, developers often,
and become the best experts at prompting our models because, you know, you're building your
livelihood on this thing and get to know the details of it really intimately. So I will say stuff
like that won't hurt the performance of the model. We kind of always want to leave it open to people
to figure out what works best. Yeah. Yeah. And then you had to always start with a response rules
or instructions section. Are those keywords meant to be taken kind of like verbatim?
Like those are kind of like the tokens that work the best or is it just like an example?
War of an example.
Yeah.
Okay, cool.
Yeah, this is great.
I feel like until today we did an episode with like the prompting report on like all
these prompting techniques, but then it's also unclear for which model, which ones
work best.
So it's super useful.
And then you had a in the agentic workflows one, you have a persistence thing.
It's like, please keep going.
How much?
And I think I read.
that improves like the suite, the sweet bench like 20% just by having like the persistence.
I wouldn't.
It's not that this one prompt improves sweet bench 20%.
It's that we found this is the most effective harness for our model.
And combined with all the post-training improvements, it results in the big improvement.
But yeah, like the model is trying a lot to be helpful.
And often it wants to check back in with the user and be like, you know, should I keep doing
this?
Like, am I on the right track?
And so a prompt like this mixture it keeps going, doesn't bother you again, and just gets the task done.
Yeah.
Yeah, I think like there's this interesting tradeoff between persistence and yielding back to the user.
The more agentic a model wants to be, the more persistent it should be, but then sometimes it just goes off the rails.
And I wonder how you solve this tradeoff because sometimes it just goes too far.
there's been criticisms of Claude Sonnet
trying to rewrite too many files at once
when I just wanted to make one thing, for example.
And that's a form of bad persistence.
What are the axes here in which you think about it?
Yeah.
I think one of the interesting thing that comes to mind here
is that we had an extraneous edits evel
where you asked the model to make an edit
and classify like,
were all its changes related to what it was asked to do
or did it go off and do a little too much?
and we found that from 4-0, which got 9%
was pretty crazy 9% of the time
making an experience edit is a lot.
4.1 is at 2%.
So it's a pretty big improvement.
So yeah, I'll just say, like,
focusing on this, we've heard feedback about this,
we made an e-val and we made sure to track it
and improve it during training too.
Yeah, yeah.
I mean, everything comes down to e-vails as,
as is no surprise to anybody.
That's true.
There's another interesting e-vel that I think
is causing some noise.
For the first time, I think also that you,
being the master of structured outputs, should know,
that JSON is bad now, and we should all use XML?
I wouldn't say that.
I don't know which Eval you're talking about, but...
It's in the prompts guide, which maybe you guys didn't write,
so we're kind of springing this on you.
Yeah, Noah and Julian on our team wrote the prompt guide and did great job.
I do think XML is very helpful for structuring prompts,
whereas for parsing outputs, maybe the story is a bit different.
Like, sometimes it's really useful to get outputs in JSON, so you can plug them directly into your application.
But I do think the models work particularly well with XML as inputs.
But, Chris, you need that?
No, no.
Cool. I mean, I think people always just care a lot about tool calls and structured outputs, as you well know.
And so any updates to instructions over there is good.
people also are interested in this concept of that apparently putting the instructions and user query at the top and the bottom, so duplicating it at the top and the bottom in the context, it's much better than putting it top only and much better than putting it bottom only.
Again, this is from the prompt guide, so I don't know how aware you guys are on this.
Yeah, I think part of that was just like, you know, empirical.
We tried all three for when we were evaluating the model and having that redundancy is definitely the best.
but then using the instructions at the beginning,
the model is going to be able to then take that into account
as it does processing.
Yeah.
I think a lot of people would see this as running counter to prompt caching
because obviously you want to put the things that change a lot at the bottom.
Basically, is this fixable in post-training?
Can we just tell models to take instructions or user queries
only at the bottom because we want to optimize for prompt caching?
When we figure it out, we will do that.
I mean, it seems doable.
It seems like a post-training thing.
I don't know.
Maybe my mental model post-training is wrong.
So I think actually having things at the beginning of the prompt,
you would still get prompt caching there.
If you're putting in, for example, like a big needle on Hasek,
and you have the data changing each time, like per user,
there's still different ways that you can be putting the prompt at the beginning
and getting a lot of the cache hits.
It sort of just depends on your use case.
Yeah, awesome.
The other thing I noticed, I know you made a note of this, Sean, too, is that our chain of thought and reasoning and how people should think about this model versus a reasoning model?
Yeah, what's your, yeah, should I just use 4.1 and prompt it to do a chain of thought?
Should I use 01 and make a plan and then use 4.1 to implement the plan?
How should people think about composability?
Yeah, it's a great question.
We have found that 4.1 is a lot better at doing planning and things.
thinking through its steps in COT when prompted than our previous non-reasoning models.
But our reasoning models are designed to have kind of more coherent plans and be able to reason
over longer horizons than these non-reasoning models. And you can see that reflected in things like
intelligence benchmarks. So Amy, GPQA, stuff like that, you'll see the reasoning models
do much better. So in general, I would say, like the question you're really getting at is like
I'm a developer, which model should I be using?
And I think the answer is always going to be the fastest model that accomplishes your task, right?
So maybe you start prompting 4.1 as a starting point.
If it does your task super well, then maybe you can drop down a 4.1 mini and save latency or even nano.
Whereas if 4.1 is struggling a bit little, maybe needs more coherent reasoning over longer time horizons,
then maybe you upgrade to a reasoning model.
is there a quick way to get through this heuristics i know one thing that a lot of people do is like they use
a one for like a plan and then they put that plan in cursor and then have the plan apply to the code base
it sounds like there's maybe not a rule to when to do which it's just like task dependent yeah i would say
we're all kind of figuring out the best way to use these models together and so i do think
reasoning models for planning uh and using kind of more targeted models to
execute is definitely a good architecture.
Cool. If there's nothing else on that side, I'd love to go into the coding, which is something
that we're emphasizing a lot. It's doing super well. It's better than O-1 in Sweet Bench.
Was that expected?
Not really.
Yeah. Like, what's the story there? There's also Sui Lancer, which is a newer one, which
attaches a money value to things. And basically, like, what should people understand is going on
here? Like, it's a better coding.
base model or just a coding agent model.
And I think there's also a question about like, you know, how important to coding is it
if I'm not using a coding use case?
Yeah.
So I'll start by saying we just set out to make model that was great at coding, both in your
terminal or in your editor or wherever you want to use it.
And so we kind of broke that down into the problems that it, you know, encompasses.
So like developers want the model to produce better diffs, for example, or they want the model
to explore the code base correctly, or they want to produce code that compiles or produce code that
writes tests. And so our approach was kind of teaching the model all of these various facets.
There's kind of just a bunch of work streams that all coalesced around GBT 4.1.
Yeah, I think much improved post-training all over to make for a better coding model.
Yeah, I think there's like different kinds of coding, right?
Like it's interesting for me to observe that there, for example, so I'm just going to pull it up on the chart here because I always like to show people visuals.
You're 55 on sweet bench and 01 gets like a 41, but then on, oh, I don't think I have the others.
But Ader is it is less, it is not at O1 level.
And so I think I struggle to get some kind of intuition of when, like, what are the different elements of coding?
I guess there is like, you know, single file edits.
whether it's like a diff or a whole file,
and then there is entire project edits?
Is that a reasonable split?
Are there more to this?
Yeah, that's one way to think about it.
Basically, where GBT 4.1 can I kind of explore and go through a repo?
Yeah.
It's been trained to do that particularly well.
Whereas, you know, to just get some code and produce a change,
a reasoning model might do better because it can kind of reason over the entire file.
And so that's one good way to think about it.
Yeah, yeah, that's fair.
Any understanding of like the smaller ones, the smaller models,
like basically for coding I should only use 4.1 and forget the rest.
You might like want to use the smaller models.
Maybe if you have like, if you have an IDE where you need an auto-complete feature, for example.
Or if you want something super fast, if you're building like, I don't know, a text to SQL thing,
you might want the first version to populate instantly.
So you can see like 4.1 mini is actually quite significantly better than 4O mini,
but not that far away from the old 4O.
So I do think that model will find use case in a bunch of these coding niches.
And I know you might not be able to talk about this,
but the clip of an AI CFO talking about the agentics suite as being going viral,
I think today.
It seems like every lab is putting a lot of emphasis into coding.
So yeah, I'm just curious if there's anything.
you can share about how people should think about open AI encoding.
You know, obviously today you don't have, you know, clottis, clock code.
You don't have anything related to coding.
And I think the Winsarf partnership today, they're giving 4.1 for free for free for a couple weeks.
It's maybe like one of the first Open AI endorsement, I guess, on the live stream.
But yeah, just I know there might not be an answer that the PR team might approve,
but I'm curious if you have any takes and thoughts.
I think just stay tuned.
Yeah, I think coding is an important use case for our users.
And so that's why we focused it on it a lot for 4.1.
We also love to use our own products internally.
And so making 4.1 selfishly helps us move faster as a company.
And so that's where the real focus has been for this model.
Do you track what percentage of code is written by 4.1 internally now?
We do have some metrics like that.
I don't have it off the top.
But I was actually just talking to one of the researchers on the team who worked on something over the weekend.
And he said that this model GBT 4.1 was able to like get 49 out of 50 of his commits on this massive PR done.
So we were pretty happy to hear that.
I'm excited to use that.
Awesome.
Yeah.
I think on the, yeah, I think coding is a super exciting use case.
And I think like open AI has always been very developer first as you've been to Michelle.
So it's great to see the convergence.
Yeah. The other, I think the last capability that I kind of
vectored in on was vision or just multimodality in general. It is a lot better. Basically,
I think like, I really like these niche benchmarks like Math Vista and chart side.
Yeah, just any extra color on like the vision side that you wanted to talk about,
but maybe you couldn't fit into the blog post. Yeah. Yeah. Go ahead.
I was saying, I think one maybe small nugget there is actually, I think that 4.1 mini,
is really exciting on that front as we were talking about. It's a different pre-training base,
and I think that really shows up in some of the vision e-vals.
And yeah, we talked about like coding and structure following along context,
a lot of gains coming from post-training, but in particular multimodal, like basically
everything you're seeing, the gains are there from pre-training. So kudos to the pre-training
teams there. They've done incredible work on perception and multimodal.
Yeah, totally.
Something that we've been exploring on the podcast for a while, and I'm curious if there's any takes on your side, is, is there a strong split between, like, sort of what I call screen vision versus embodied vision, right?
Like, are you taking pictures of, are you training on snapshots of a computer for computer use or, you know, and anything with charts, anything on a PDF is very similar to that?
Or pictures from the real world, which is more embodied, right?
like where a robot might be able to use that.
People have argued back and forth.
I'm curious where the movement is or the emphasis is.
I think one of the, first off, I think that 4.1 is better at both of those things,
regardless of how it was actually trained.
I think I would probably somewhat defer to the pre-training team
when it comes to which one you should be using,
or using a mixture of both.
But we've improved our results across e-vals on both.
Awesome.
That's something that I think people should definitely do want to explore the more embodied stuff as well,
because the benchmarks tend to focus on the screen vision stuff, you know, more chat, more controllable.
It's always easy to make an e-val that is easy to grade.
Yeah, exactly.
Those are the things they get looked at the most for sure.
I think one of the things that was really funny with both the 4.1 mini and nano is we had some strange internal eval results.
And it turns out that actually these new vision capabilities, they were able to read like, you know,
signs in the background and stuff, which was actually changing, like, some of the validity of our results.
And so we were, you know, just running into different eval problems as you actually improve the models.
Is there a feature of a 4.1 image gen? Or is that, like, a completely different part of this vision?
Like, you know, in some sense, vision is image to text. And the other way around is image gen.
Is it that simple or is something else? It is not. No plans right now to get 4.1 image gen.
Well, you know, it's very, very popular.
You like it's like melting your GPUs.
I mean, talking about GPUs, right?
Like, you know, part of this whole deprecation of 4.5 and moving people to 4.1 is to get back your GPUs.
That's a message that both Shuki and Kevin Weil have mentioned.
But, like, you are running all these models concurrently for the next three months.
Like, I don't know if you get back at GPUs.
I think you just grow their usage even more.
Yeah, I do think, you know, people get the message on deputables.
and start moving over.
So as developers use this model a little less,
we can kind of reclaim that compute.
But you're right, it takes a while.
And the tradeoff there is really our commitment to developers.
Like if you have something in the API,
we won't take it away without sufficient notice.
That's the tradeoff that is right for us.
Okay, awesome.
Then a couple other smaller announcements,
fine-tuning available day one,
which is, I think, new for Open AI.
Usually you have to wait like a month or two for the fine-tuning capability.
For one, 4.1 only and mini only and nano and future.
Any specifically call-outs for fine-tuning?
I guess like this is general discipline that always applies.
But any wins that you guys can talk about?
So first of all, yeah, shout out to the fine-tuning team.
They've worked really hard to get this ready on day one.
One thing I will say is that I think people have slept on the preference fine-tuning offering
or the, I think that's what we call the product.
Yeah.
So, SFT is, people know it pretty well.
It's the original fine tuning we had,
whereas this preference fine tuning is super helpful
for steering in a particular style.
And so I think not enough people are using that.
Isn't that only for reasoning models,
or is that for everything?
No, that's reinforcement fine tuning is only for reasoning models.
Right.
Preference fine tuning is offer the pairs.
Yeah, exactly.
Yeah.
And I thought it was in alpha.
It's just why I haven't looked into it.
I think it's RFT that's still in alpha.
Okay.
Well, that's a lot of confusion that we just cleared up.
Yeah, I think we're going to, you know, I'm doing my conference again in June,
and I think we're going to do a workshop on just general, all the fine-tuning options,
and I think that will clear up a lot of things, which is good.
Okay, new models, I know that we can talk a lot about a lot of them.
Norm Brown from your reasoning team just said that there should be a follow-up on reasoning models soon.
What can we say about that?
Sounds like he's...
We're not the right people to ask, but stay tuned for...
Yeah, but like 4.1 is a good basis for whatever comes next, right?
Yeah, not all of our models kind of build on each other necessarily,
but we think 4.1 is a great standalone offering for developers,
and we also think, you know, reasoning models are a good tool and toolbox.
Yeah.
Like, more just generally, like, I always want to explore the relationship between non-reasoners and reasoners,
and then also like how we merge them.
Are we doing routing?
You know, anything on that sort.
Obviously, you have a lot of secret sauce.
Cool.
And I think the other thing that a lot of people are demanding or asking about is the creative
writing model.
Will that ever see the light of day?
We're working on incorporating kind of those improvements into the models more generally.
Not a separate place.
People love about 4.5 is like the humor, the green text, the nuance.
So we've heard that feedback.
And I know, yeah, there's lots of folks working on that
and trying to bring it into our next models.
Awesome.
Alessio, anything else?
No, this was great.
Any requests for the developer community?
Things that you want them to try out that maybe people are not doing
things you want them to build for you using the new one, the new APIs.
I feel like first off, send us feedback.
It was really useful to look at different partners and customers
who are using our models and to get this nice, wrapped feedback from them.
it allows us to iterate a lot faster.
And on that vein, you know, opt in to data sharing.
This just helps us make the model better for you.
And one kind of slept-on way to do this is the e-vals product.
So you can upload an e-vow and opt-in such that we'll pay for the inference costs
if we can also use the e-vail.
And this is just another great way.
Like, we'll use those e-vails to make sure our models are getting better for people over time.
Yeah, I think the e-vals,
is permanent. There's no end date announced, but the opt-in in the API is at least until April 30th.
I think a lot of people still don't know about it. We might want to extend that so that people can
do more. Yeah, it's a good flag. All raised with the team. Yeah. Awesome. And I think the last question
I had was on just on pricing. I think pricing, you know, it's basically just generally cheaper than
4-0, but like not a ton, but like cheaper. And then you're also introducing this concept of blended pricing
for the first time that I've seen it,
but maybe it's just been out there for a while
because you have caching and all that.
Just generally, what is the cash to non-cash ratio
that we should be thinking about
when thinking about workloads?
Like, is there a general rule of thumb?
So one clarification, which is that GPD 4.1 Mini
is not cheaper than GPD40.
So it's not just like a blanket decrease in all the models,
but however, 4.1 Mini is cheaper.
than 4.1. Also, not sure if this is widely reported, but we've increased our prompt
cashing discount from 50% to 75% on these models. Yeah, I saw that. So that's a big input,
you know, into figuring out what kind of application you built. And then your question was on,
like, what kind of... Like, yeah, blended pricing, right? Like, I think there's this question of
comparability of prices across models and across providers, because, like, I, you know,
Like some people are three to one in terms of context to output, and then some part of that is
cached.
I selfishly, I make a chart that just plots all the model labs versus all the prices, and I'm
sure you guys have seen it.
And I don't know what numbers to plug in there.
So what are people seeing in real life?
What's the median, you know, cashing rate?
I don't think we have that off the top.
The blended pricing is more to just make it easier to compare, like so you can say,
Something like GPT4.1 is 25% cheaper than GPT4?
Yeah, you want one number.
Yeah.
Yeah.
No.
All right.
We'll all have to figure it out.
But thank you so much.
That was fantastic.
Thanks for all the work.
I think people are very excited to get to work testing this out, giving you feedback.
And I'm sure we'll be back again for the next one, probably the reasoner.
Nice.
Thank you guys.
Thank you.
