How I AI - GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark

Episode Date: July 9, 2026

GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against GPT-5.6 Terra, Luna, Claude Fable 5, and Sonnet 5 across five categories: PRDs, prototypes, wireframes, debugging, and... agentic voice. Sol won by a meaningful margin on my Claire Weighted Index (70% my taste, 30% Terminal Bench 2.1), and I also tested two use cases I can't stop thinking about: building a gamified homework tracking app for my kids in one shot with Codex, and browser automation with Chrome that burned through 500 LinkedIn replies while I did literally nothing.What you’ll learn:How I scored five AI models (including GPT 5.6 Sol, Fable 5, and Sonnet 5) using my “Claire Weighted Index” benchmark across PRDs, prototypes, code, and agentic voiceThe difference between GPT-5.6 Sol (Terra) and Sol for PRD writingHow Fable’s precision and pedantry made it harder to collaborate with, and the exact moment Sol broke through where Fable got stuckWhy Sonnet 5 is still my go-to for agentic voice in OpenClaw, even after this whole benchmarkHow I used GPT-5.6 Sol in Codex to build a fully gamified homework tracking app for my kids in one shotThe video editing use case that saved me hours clipping a talk I gave at Cursor’s eventHow to use Codex plus GPT-5.6 and Chrome for browser automation, and why this is my single most-loved use case right now—In this episode, I cover:(00:00) Intro(01:10) The three GPT-5.6 models: Sol, Terra, Luna(02:17) Pricing: Sol vs. Fable API costs(03:24) The How I AI benchmark(05:03) Claire-weighted Index results(07:00) Per-task winners: prototypes, PRDs, agentic voice(11:59) What Claire actually rewards(13:20) Full-fidelity prototype side-by-sides (Sol vs. Fable)(17:45) Wireframes(18:19) Agentic voice(19:15) Where Sol is better than other models(23:56) Gamified kids’ homework app, built in one shot(28:02) Fable’s pedantry problem and how Sol broke through it(31:49) Two bonus use cases: video editing and browser use(35:08) Final summary and model recommendations—Tools referenced:• GPT 5.6 (Sol, Terra, Luna): https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna• Codex: https://openai.com/codex• ChatPRD: https://www.chatprd.ai/• CapCut: https://www.capcut.com/• Math Academy: https://www.mathacademy.com/—Other references:• Cursor event where Claire spoke on the future of PM: https://www.youtube.com/watch?v=4CAFK-rc26A• ChatPRD blog (where benchmark outputs will be published): https://www.chatprd.ai/—Where to find Claire Vo:ChatPRD: https://www.chatprd.ai/Website: https://clairevo.com/LinkedIn: https://www.linkedin.com/in/clairevo/X: https://x.com/clairevo—Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Transcript
Discussion (0)
Starting point is 00:00:00 I have been very, very, very sad the last week because for the last week, I have not had access to my true favorite, top-of-the-line model, GPT-56. But guess what, babes? It is back. And I am here to walk you through. GPT-56, soul. GPT-56, Luna. GPT-56 Terra. I'm going to tell you, what are these models?
Starting point is 00:00:28 How have I been using them? Why are they my heart's favorite and is Fable better than all of them or not? I have been testing this model for a couple weeks. There was a few days there where we didn't have access and I found myself desperate to get this workhorse model back. Now, we're not just relying on my own opinion. We are going to run the very famous, very new How IAI Vibe Review benchmark against common tasks from PRD writing to prototyping to whether or not a lot.
Starting point is 00:01:00 it's cute in my OpenClaw agent, and I'm going to tell you very scientifically if this is the model that you should be working with all the time now. Let's get to it. Okay, you all can read these blog posts, so I'm not going to go into too much depth about the models and the benchmarks. I'll just give you the hits. First, OpenAI is releasing three new versions of their GPT 5.6 model. Soul, which is the next generation frontier model, the brainiest of the brainiest, Terra, which is a balanced model for efficient everyday work, and Luna, which is sort of akin to their mini or nano models, which is cheap and affordable for high volume work. So you're going to have these three versions of the models. I don't know if these beautiful images are
Starting point is 00:01:46 exactly how we should think about the relative capabilities, this big sun, this medium earth, and this tiny moon. But I will say, My love letter that is this podcast today is written directly to GBT 56. Soul. This big model is the one I love. Now, I have tested Tara and Luna, so I will give you my input there. But really, this is going to be all about Soul versus Fable and which one I would use for the type of work that I'm doing every day. Okay, quick note on pricing, Sol is a lot more affordable than Fable.
Starting point is 00:02:21 So it's $5 per million input tokens, $30. per million output tokens. I believe Fable, at the time I'm recording this, is 10 on a million input tokens and 50 on a million output tokens. Now, again, you're going to get a little bit of subscription usage built into your Open AI subscription. So you are going to get a decent amount that you can test with and use. You know, there's been some challenges with the Fable rollout. They've limited when it's been included in the subscription. And so it was supposed to be available till early this week. I think they extended that a little bit at Anthropic so subscription clod users could use Fable under their subscription. So we have to see how much sole usage we get
Starting point is 00:03:06 and if like Anthropic they're going to take soul out of the subscription. I suspect not. I suspect this is a model they want people to use. I also suspect this might put pressure on Anthropic to put Fable back into the cloud subscription. But for now it's more affordable even at API pricing. Now I'm not going to read through all the benchmarks for you. You can go to this OpenAid blog and read them for yourselves. All I will say is it is the brand new state of the art model from Open AI. It is the highest performing when using the ultra mode on terminal bench 2.1. And then they've also evaled it against a couple cybersecurity benches.
Starting point is 00:03:44 So I do think as we get these smarter models, you're going to see a lot more evals and benchmarks around exploits and security. and then very similar to what we're seeing with Fable, there's a lot of conversation in this blog post about the safeguards and security frameworks around the release of this model. I do believe like Fable it's going to fail over in some tasks that are maybe a little bit riskier, but I have not run into that myself. Now let's get back to how I eval these models. If you missed my episode on Fable, I got kind of bored of the Vibe Vibe Vibe Check and I built an extremely scientific How IAI benchmark. Now this HowIAI benchmark tests basically a couple things. It tests the ability to generate good PRDs. It tests the ability for it to wireframe against a couple different app ideas,
Starting point is 00:04:40 develop fully designed, robust designed prototypes, debug code, and then talk to me like a human, which is the thing that I care about the most. And I'm just going to remind you how I did these benchmarks and then scan you through a couple of the outputs. And since I know what the models are now after I've done the grading, I can show you which ones map to Fable and GPT 56. Okay, so this is my vibe review. What I tested was Fable 5, Sonnet 5,
Starting point is 00:05:09 and then the three versions of GPT 5.6. I did it against my common use cases of PRD's, prototyping, coding, and chit-chatting with an agent. And then what I have the eval harness do is it runs all the evals against each of these models. And it does a LLM-based judge. The LLM that I've decided is the hardest judge is GPT 5.5. So that's the one that judges. But it also gives me this page where I can actually go through and give what's called the Clairvote Taste Test, which is I read all the assets, I look at all the designs, I score it and give it notes. And so you can see here I went through PRDs.
Starting point is 00:05:50 We went through sort of some complex prototypes here in terms of a dock scheduler. We did a consumer app. So lots of beautiful different habit tracker apps, different versions you can see here, a pretty complex dev tool, wireframe versions of those same prototypes, which I graded. I also give notes. This one great note says my fave, but not. not great. And then I let the code grader just evaluate the agentic multi-step debug because I wanted it to be really about accuracy there. And I didn't feel like I could eyeball that and give a strong
Starting point is 00:06:29 opinion. And then the last thing that it generates is an agentic voice. So basically how it would respond to me answering a couple questions. Very important on agentic voice. These models, somebody, please hire somebody to get rid of the m-dashes and slothes. talk. I cannot stand it. Now, one of the things that I will say as an observation for 5-6 is it's a great writer and I will show you some examples of that. But truly, a lot of my evals here were M-Dash Slop. I hate you. Okay, so let's go to what the Claire-Waided Index says. Now, this is my show. This is my podcast. And so I sort of strike the balance between what the LLM judge said about the performance of the models and what I said about the performance of the models and then I get to
Starting point is 00:07:21 strike the difference. And you know what? I've decided I like my own taste better. So I've decided it's going to be a 70 clairvo, 30 the machines split on evaluating these models. And so if you look at that 70-30 split, your girl loves 5-6 soul. She just does. It had be. highest taste score by a significant amount. So I just thought it output the best work. Again, I went through dozens of evals, looked at them, clicked through them, gave my own opinion, put notes. And I just have to say, I really like GPT 5-6 whole. I know. I spoiled at the beginning, but I did blind taste test these. And so I do really feel like it did a good job. And I will give you a couple examples of that. No, I don't hate Fable 5. So I'm not saying that Fable 5 is out of the game.
Starting point is 00:08:23 I will say, I did not have to talk to Fable 5 when running this benchmark. I hate talking to Fable 5 because it talks to me like an engineer that has never met a human before. It's like its first day on earth. But when I don't have to talk to Fable 5, it outputs pretty good work. And I would say had some good outcomes there. And then Tara Luna did fine work. Sonnet 5 at the bottom really haven't figured out how to get this one working, although there's a very specific use case that we think Sonnet 5 is good at, or actually two use cases.
Starting point is 00:09:00 Now this is heavily weighted on its front end prototyping, design, and app building capabilities. Since that is the chunk of the HowI AI Eval, it is heavily weighted there. But I do want to call out that. Her task, I do have a couple favorites. So for that prototype task, and we'll go to some examples in a minute, I just love 5-6 soul. I just really do. I think it was functional. The designs were the most interesting. I thought it was really good. For PRD, actually I liked Tara. Maybe it's down to earth. Maybe I like a basic, straightforward PRD. As I said, it was my favorite, streamlined, and to the point. And so if you want clean, crisp, direct business writing, maybe GPT-56 Terra is the, way to go. You know, the bug hunting eval, which I don't really feel like I've nailed exactly. So I'm not super confident in this one, but the LLM as a judge thought that Sonnet 5 did the most complete and accurate job. I will say, I only like talking to Sonnet models through my open claw. Really, I only like talking to them. I still really struggle with getting my open claw to work well
Starting point is 00:10:08 with the GPT models. I still did not like Fable in the agentic voice. Eval, which you should not be surprised at, but Sonnet 5 got a very good gold star for me because I said, aside from the M-Dash, you are a human. That is very, very high praise. And then I'm going to show some of these designs in a second, but you can see across the board on a full fidelity prototype. I just really preferred five-six soul three out of five times, five-six, four out of five times. And Sona did the best job at the editorial design. I will say Claude's design aesthetic tends to this sort of like editorial design. If you know, if you've seen it, you know it. It's like that beige background, that orange, burnt orange color, the italic serif fonts.
Starting point is 00:10:59 It's just very, very clawed. But I hated that design overall the most. So you can see here, I rated it still lower than almost anything else on this leaderboard. It just happened to be the best of the worst, I would say. Now, where GPT56 sold did a really good job and I'll show some of these examples is like complex, dense, technical, unique designed things. And so I will say I have been happy to extract myself out of Claude Slop, out of like blurple slop into more interestingly designed websites. And I'll even show an example, kind of like a meta example, which is this is the opus designed version of this page.
Starting point is 00:11:42 Like very slop-adjacent. We got the blurple. We got a gradient. I don't think the typography is particularly sophisticated. And I asked soul to redesign it. And I just think this is a lot cleaner, a lot nicer and easier to look at. Okay, let's talk about how Claire qualitatively evaluates models. Some of these quotes will just give give you a sense of what I value. And again, I gave 50 written reactions. There were some like unmistakable hits where I loved what the models came up with. 14 places where I was like, this is garbage. So let's see like kind of what I talk about when I review things. So I was definitely calling out uniqueness, creativity, and functionality in the design. And so in designs, I like non-slop, unique designs that are functional.
Starting point is 00:12:34 And so I'm definitely going to reward. This doesn't look like the generic prototype. And you've pulled the thread of functionality through the prototype. For writing, I just like Sistinct and to the point I cannot stand AI writing. It drives me nuts. I can see it a mile away. So I really like just direct, very frank, very crisp writing. I think 5-6 is good at that.
Starting point is 00:12:57 And then you can see the things that I hate. I hate slop. I hate slop. I hate slop. We all hate slop. It's the worst. It's the worst part of AI. If I hate one thing about AI, it is that I have to experience Slop. So you can see I like clawed design slop across this editorial page, typography, emojis and bad placeholders. Like I really held a high bar in terms of design quality. Okay, let's look at a couple of these and why I really liked Soul compared to other models, although where Fable did a perfectly serviceable job. Okay, so this dense operation dashboard, it's basically like an Eval for a doc scheduler app and it's full design. And what you can see here is both were pretty useful, Sol on the left and Fable on the right. I just think Sol was the most unique. All of the other ones
Starting point is 00:13:49 really just looked like this dark mode, monospace kind of layout. As you can see here, Sol actually has like a really clean kind of like neutral color layout with great visual hierarchy. semantic color and this thing was functional. So like everything I expected to be able to click and work and assign and do, all of it actually worked. And this was just my experience across a bunch of the different prototypes is the sole ones were just a lot more functional. And that made a big difference on how I'm evaluating things. Now let's look at the fable design again. It's pretty good. It's actually a lot harder to read though and the design I would say is not as unique and even some layout issues like this white space here at the bottom now it did do a lot of functionality but I would say
Starting point is 00:14:44 like the colors weren't semantically assigned the typography needed some work and I just really preferred this unique design of soul even though it wasn't crazy it was just opinionated which I think is nice now here is another design it was this creative pack website. Again, both of these got fives from me. I just really preferred that Sol went ahead and had like a personality. Look at these placeholder images versus what Fable came up with, which I will say is beautiful and clean and worked really well. And like I have no complaints about it. It's a good one, especially foresight of sort of a wireframe style prototype, it's great. I would just say, it's not this. This is pretty interesting. It's got a better
Starting point is 00:15:36 point of view. And it's got like nice little design affordances that I just didn't see in these other designs. And so I just really preferred, or at least I rewarded the fact that Soul, you know, use its brains to be a little bit more unique and give me some inspiration. Then on this dev tools page this is again where soul went really well and it's sort of the same as the dock scheduler it just does the job of this is a incident triage site it just does the job a lot better than i would say the fable five did fable five is fine it's just not that unique and again the thoughts around the design are not exactly what i would want. And so again, this like functionality point of view design, I really preferred soul.
Starting point is 00:16:36 And then last side by side comparison. And again, I think this is a good one to think about. If you see here, we did these habit tracker apps and just looking at the comparison side by side design. Like this is good old classic Claude stuff. You've seen this design a million times, especially if you used Claude Co-work. And if you look at this, it's just again a little bit more opinionated. There are some slot pieces to this design. Some things that I did not love. The one thing I will say I noticed about soul, which you will notice,
Starting point is 00:17:16 which I have told the delightful and lovely opening eye team, and maybe it's because they love me. It loves a forest green. It loves a forest green. In fact, I think this forest green is like in its system prompt called like, woodland some woodland elegance or something like that. I mean look I love a forest green look at my office it is forest green but you will see a lot of green and I think this is one of the gpt five six tells that you will start to notice and get really frustrated with now on wireframes
Starting point is 00:17:48 again let's just look at these side by side sole very functional very easy to read like as a person trying to convey a complex application. I think this does a really quite excellent job and just a better job of this. It's just a little harder to read. I'm not quite sure what I'm supposed to do here. It's not as functional. There are some interesting things here, but you know what Fable came up with was not my favorite. Now, final thing is its voice. I just want to call out. I do. I do. I do love Sonnet for agentic voice, so I cannot knock Sonnet for not sounding ridiculous. So I asked it in sort of a EA personal assistant, open clause style, a couple questions. Can you move my meeting?
Starting point is 00:18:41 Deploys Red again. Why did I start this company? Let's just show a little straight to prod. And how Sonnet replied and how Soul replied, you're missing the line break, so it read a little bit better in the eval. But if you read them, like Sonnet's still super cringe, but Soul was worst. I mean, Sol said this deploy is a bug, not a referendum. Like, please don't do with this, not that.
Starting point is 00:19:04 To me, do not do M-Dashes. So I could not get rid of M-Dashes. But I thought Sonnet 5 had the best voice. I tend to use Sonnet for whatever for my open clause. So I'm not surprised about that. Okay. So that is the Clarevote eval, but I want to go into a couple other things I really love about this model. So let's switch over to Codex. Okay, I'm going to zip through a couple examples of things
Starting point is 00:19:28 that I think Sol does a lot better than other models, and in particular, a lot better than Fable. Number one, it writes like a normal person. I cannot cope. I love Fable your brainy. As I showed, the Eval show you do a pretty good job. I cannot talk to Fable anymore. Fable. No. makes up, it seems like Faber's unfamiliar with the English language and communication with humans. Fable is very much like a four agents by agents communication mechanism. I can barely make out what it's talking about. It is incredibly inscrutable writing. And that makes it very hard to collaborate with your bottle.
Starting point is 00:20:15 And so what I would say is my experience using Fable has been, it is like incredible, technical, incredibly pedantic. And while it is super intelligent, hardworking, will like definitely fan out and solve very complex problems. Its ability to collaborate is low and it left me with a lot of frustration as an end user using Fable. Now Fable did knock off some like pretty complex work and I'm very happy to go through what that is. It helped me build a full prototype tool inside chat PRD, so like a V0 lovable, et cetera version prototype tool. It's helping me build this like synthesis, product, brain product that I'm working on. But I found it incredibly hard to break it out of its own sort of frameworks, its own
Starting point is 00:21:11 limitations, its own structured way of approaching problems. And what I really feel like the difference, if you would take away like one highlight, difference between Fable and Soul, is, like, Fable is theoretically hyper-intelligent and soul is practically effective. And so, like, I've been an executive long time. I've been a manager a long time. Like, I really struggle working with theoretically intelligent colleagues who can't get anything done.
Starting point is 00:21:43 Like, can't actually see the forest for the trees, get too much in their head. And so, like, when I want to ship stuff to customers, I need practical, get the job done, understand the end user goal, understand the end user, and like, willing to loosen constraints appropriately to get things done. And that has just so much more been my experience with Soul versus Fable. The writing is straightforward. The communication is clear and it's less pedantic. I'll just give you a quick example of this, which is I had Soul, look at my chat period repo, and like, Greenfield totally rebuild it. Just my idea was like completely rebuild your idea of what CHAPRD should be in
Starting point is 00:22:30 2026 and wouldn't do a much research. And it came, came back to this. And again, love me and executive recommendation started uses, you know, tables. What exists today is very straightforward and easy, easy to understand. This is a very long document. I did read a lot of it. And it's just easier to parse than anything. thing I've seen come out of fable. So writing communication, definitely plus in Soul's Corner.
Starting point is 00:22:59 The second thing is like full zero to one prototypes as we've seen in the Eval benchmark I just really like. So again, for this like rewrite chat here D from the ground up, it came up with this idea of like taking a problem space or a decision, validating it with external insights and then pulling it all the way through coding handoff and built this. pretty complex prototype. Now, do I love everything about this idea? No, are we doing some of things about this idea, including Insights Generation? For sure. But this was actually very nice from a prototyping perspective. And I thought it did a good job of giving me a robust thing to experiment with and gave me some good ideas about what I could do with the product next. So
Starting point is 00:23:52 So I was pretty happy with the like zero to one prototype. Now a little bit more fun example is I asked Seoul to make a fully gamified homework tracking system for my kids. Look, my kids are coin operated. I have a middle child who's basically going to be an enterprise sales rep. If he does his homework, I need to like give him a skittal or let him trade skittles for nerve guns and he will like learn calculus by the time he's in fifth grade. But I'm a vibe code lady.
Starting point is 00:24:19 And so I want to build a app. just sneak peek into our household. My husband sent me a XP system proposal via open claw this morning. So I'm taking an open claw generated PRD, dropping it into Codex and GPT 56 soul and generating something. Now, what it came up with was pretty ambitious. Now, do I love the design? Is it a little, like does it have some AI tells? It's like gradients, you know, fonts, all this kind of stuff.
Starting point is 00:24:51 But it's like cute in a way. Look at this, you know, it's using this emoji really well with the texture. It's doing some animated things here. And basically it's giving my oldest child and my youngest child two different summer quests they can do. They can enter focus mode. I think this is really good again from a design perspective. They can enter focus mode. What does this listen do?
Starting point is 00:25:15 Math Academy. Finish one focused math academy mission. Hero check. Hey. Let's stop. So it built in some voice to it. It even built things like focus mode where it could start a timer and start to track the time that it's spending, that my kids are spending on particular homework items. Yes, we are very fun here.
Starting point is 00:25:36 How many lessons reward them about how they pursued their task. Inching the quest, you get some nice little confetti here. They then get to get available rewards. My oldest child is earning a one-on-one basketball coach because he likes coaching. So we say if you practice your piano, you get a coach. So they put that front and center. And then came up with different sort of like prizes they can win, including picking family dinner, a movie and staying up late,
Starting point is 00:26:09 or buying like new basketball shoes, which man, the way these kids grow their shoe size, they buy a lot of basketball shoes. And the same with my middle. He's focusing on a couple different things. including playing piano. It's actually really short what he has to do. And so it built that. And then what I love is it gamified them together.
Starting point is 00:26:28 And so if they can work together, they can earn more XP. They also can earn like companion, I don't know, avatars like beat bot and comet fox. They can get power oras. They can like figure out which different kinds of subjects they're learning. So it really went ham on some. gamification. And then again to the sort of like full-fledged functionality, it even gave me a parent HQ. Now we got a little slop here with a border on the side, but I can review exactly what they've done. I can turn on and off quests. I can edit how many points they get per
Starting point is 00:27:09 quests. I can add things. So if I want them to start doing stuff, I can add it in here. I can change what rewards they get. Again, it really listened to me. My oldest is motivated by basketball and my youngest is motivated by Minecraft. And it gives me a history and other settings that we can set. And so get a very robust app. It built it basically one shot and put a lot of effort into the design of it. And this is something that I've seen from Seoul. Now, like, is this consumer grade exactly what I would ship? No. it's a lot better than what I've seen kind of one shot out of other models. And I do just like the polish that it's put in in terms of efforts. So again, writing good, one shot sort of prototypes good. We've seen that in the benchmark. Let me talk about another thing where I think
Starting point is 00:28:04 GPt 56 Seoul and its family does a lot better than Fable. And I understand, I'm going to preface this by saying I understand why Fable is a great cybersecurity researcher in that it is like, incredibly precise, incredibly detailed. We'll like look at every corner and every edge and score every risk and like try to be incredibly precise. The problem is when you're building products, exact precision is neither helpful nor possible. Like you literally, especially when working with AI, cannot be precisely deterministic when building a great product. And like understanding what a user would like is not a exercise in technical precision. It is a, is an exercise in intuition, design, all these things, and boldness and creativity and strategy
Starting point is 00:28:51 and all this stuff. And I was working on two projects, deeply with Fable and then with Soul. And I just had a very much better experience unlocking with Soul. Let me just talk to you through what those are. One was this chat PRD kind of like integrated prototyping tool where like VZ lovable, all these things. You could take your PRD and make, make a prototype and building like a good effective coding harness there and then trying to figure out what the right model was. The second thing is basically like an Insights Ingest product where you can like hook up intercom and linear and all these GitHub and all these signals and suck them in and like basically build a product brain.
Starting point is 00:29:29 It's going to be rad. And when I was having a fable working on this, it did a lot of the like technical heavy lifting. It got the like big, meaty pieces into place. but it was like a brutal scorer and it hardened these um the architecture of both of these products that it actually broke itself so my example is it like had this very hardened tool calling loop in my prototyping tool and only gpt 5.5 would run like i could not get any other model to run and i ran eval after eval after eval open weight sonit opus all of these could not get anything but GP 5.5 to run it. And I was insistent that this was an us problem, not the model problem.
Starting point is 00:30:15 These models can definitely create front and prototypes. And Fable was like, no, bro, that's, it's, it's totally these models, models fault. And as soon as I switched it to Codex and said, like, look, I'm just not convinced we can't get Sonnet 5 to work. This is ridiculous. Just do what you think is correct. It, it fixed it and it got it actually working. Now, did it get it working perfectly? No, Do I think this is a great design? No, I'm trying to figure out what the problem is. But in one shot, it got out of its own mind and fixed things. And again, this was like such an unlock.
Starting point is 00:30:50 Very similar to my insights generating engine. Fable really wanted to like score and lint this effort. And wanted to like be able to deterministically figure out if generating pros could be like reproducible, always verifiable, always citations, all these things. And at the end of the day, that wasn't what it was going to make a great product. It was just what was going to make like a code evaluation verification loop exit. But once I've told GBT56 and Codex like stop being pedantic, I ended up getting these really useful and helpful wiki pages generated out of this, all this structured and unstructured data. It was actually really really,
Starting point is 00:31:37 good and it just, I don't know, I don't know what Fables deal was. I could not get it unlocked, but 5-6 was very willing to reconsider its own kind of limitations and build something. I'm going to do two more quick use cases where I think GPD 5-6 is really good. I will get you out of here. Go start coding. I'm basically out of a model capacity anyway, so I'm going to have to take a break. Two use cases that I think are amazing. First one is video editing, video editing. I have to do a lot of social clipping. And it's really tedious to go through and clip videos. So taking something really long and shortening it.
Starting point is 00:32:19 So recently I spoke at Cursor's event and gave this talk on the future of PM and got the recording from the cursor team. Thank you very much. And I really wanted to make it a hype video. So all you have to do is literally drag the file in here. And I said, can you cut this video into five clips for social? and I gave some feedback. I said I want them horizontal. I want them height video cuts from various parts. I need them to be faster. I need them to be tighter. And then I got these like sharp and funny height videos. Let's see if it opens up. This one's for my talk. We're going to figure out what it means to be a product manager in the age where anybody can build anything. We have been coming up with creative ways to avoid. building things forever. Yes, PRDs, like these complicated documents where you had to describe. So like that would have taken me so much time to like find the right cute parts, clip it, cut it.
Starting point is 00:33:19 I was able to drop it into Capcut, put some music, ship it on social. It's like a really cute hype video. But this is one of my favorite use cases. I'm pretty sure it can do even more color grading sound, all this kind of stuff. But even just dropping videos in here and fixing things are great. Finally, the last and best use case. of 5,6, and I cannot believe I waited to the end to show this. It is a beast, beast, when it comes to browser use. I am deeply obsessed with letting Codex plus GPD 56 and Chrome. And at Chrome in Codex, if you didn't know how to do that, you do it like this, at Chrome, on a logged in page and just say, go with the stars and do some stuff.
Starting point is 00:34:06 And like, I'm sorry, LinkedIn. I know I'm not supposed to do this. But I opened up LinkedIn and I said, can you use Chrome to reply to messages that are a very high value to chat peer D or the How IA podcast? Keep the bar very high. Again, I love you all. I cannot deal with all the LinkedIn requests. So like only accept them if they're executives of tier one companies. I don't want random sets of connections.
Starting point is 00:34:28 It went through and burned through probably 500 messages. It replied to people that I, um, needed a reply to it said thank you to people who said nice things about the podcast. Thank you to those people. I do mean it. But it just rocked through browser use. I have used it to test web apps. I have used it to fill out annoying forms. Browser use and 5-6. And when I got rolled back to 5-5, my life was worse. So please, please, please. Learn to use at Chrome, at browser, and at computer. and just let Codex rip and let GPD
Starting point is 00:35:07 56 rip. Okay, that's it. That is the very scientific how I AI model benchmark, the love letter to Claire Vaux's favorite favorite mom. GPD 56, a honorable mention to our pal fable
Starting point is 00:35:23 who, if I don't have to talk to you, I'm actually pretty happy with your code and a broad set of use cases I think it's really good at. Excellent at writing web apps. The best of the AI writers, unless you want it to have a personality, then that sonnet. Great at unlocking sort of technical work that has gotten too complex for its own good and breaking through to the real user value, cutting videos, which I really love to do,
Starting point is 00:35:51 really love to do with GPD 56, and using the browser. Those are the things that I would try. I would love to hear what you think about these models. I would love to hear your feedback if I am totally off my rocker, what I should add to the Haua A.I.A.A.Benchmark, we will publish all this work to the chat parity blog. And I look forward to talking to you about the next model soon. Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app.
Starting point is 00:36:26 Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at how IAIIPOD.com. See you next time.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.