The AI Daily Brief: Artificial Intelligence News and Analysis - The Best Way to Test New AI Models

Episode Date: October 7, 2026

In this Operator’s Cut with Nufar Gaspar learn how to test new AI models against your own tasks, compare outputs, and weigh quality, speed, and cost—a repeatable system for figuring out which mode...ls deserve a place in your work.Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://kpmg.com/us/Sophisticated⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Granola - The AI notepad for people in back-to-back meetings. Try it free ⁠⁠⁠⁠granola.ai/brief⁠⁠⁠⁠ Harbor - Invest in the AI ecosystem. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.harborcapital.com/aidaily⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Section - Section turns AI investment into workforce transformation and ROI - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.sectionai.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The AI Daily Brief helps you understand the most important news and discussions in AI. Newsletter: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Interested in sponsoring the show? sponsors@aidailybrief.ai

Transcript
Discussion (0)
Starting point is 00:00:00 We are currently in the fall glut of new models. Over the past several weeks, we've gotten a slew of new closed models at the frontier, Opus 55, Astros 6, Soul 61, with more on the way. And from those same labs, we've also gotten more cost-effective and faster models, like Sonnet 5.5. And then as if all that wasn't enough, we've also gotten a number of new openweight models, which are increasingly relevant for lots and lots of different types of businesses. And of course, every time one of these models gets released, it's released with benchmarks. But benchmarks don't really tell us much about how it's going to be relevant for us personally.
Starting point is 00:00:34 For that, you need to build your own personal AI benchmark, a system to tell if and how a model matters for your own personal AI stack. And that is the goal of today's operator-focused episode. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, robots and pencils, Harbor, and Blitzy, To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe
Starting point is 00:01:15 on Apple Podcasts. And to learn more about sponsoring the show, send us a note at sponsors at AIDailybrief.a.i. One more quick note before we get into this, this is a recording of the webinar that was hosted by me and Newfar Gaspar last week. Given that already, we've gotten another couple new models this week. I think the system that is discussed here continues to be extremely pertinent. Hopefully you find this useful, and tomorrow we will be back with a normal episode. But for now, let's talk about building a personal AI benchmark. We prepared some very fun materials to you, so we will share everything at the end. You'll see throughout what are the materials.
Starting point is 00:01:46 But basically everything that you see me doing and talking about, you will get everything in order to try and do a similar process for yourself and do leverage the Q&A box because we have a large number, so we're unable to open microphones. But Dan is here and Nathaniel is here, and everybody will try to cater to your questions as best we can. Without further ado, Nathaniel, some motivation on why we are here. Yeah, thanks everyone for being here.
Starting point is 00:02:10 I think for those of you who are regular listeners of AIDB, you'll know that this is a soapboxy issue for me when it comes to benchmarks. Every time we get a new model, which by the way is now, I think, 11 days on average for a frontier model. One of my standard cautions is basically first to be skeptical of the published benchmarks. Not so much because we think that any of these companies are lying. That's not the nature of what they do. But more because most of these benchmarks are now in the training datasets. and there's part of the canon that goes into them. It's just there's limited value of benchmarks,
Starting point is 00:02:45 and the longer they've been around, the less value there is. Now, lots of people are continuously creating new benchmarks to try to improve that, and it's actually a pretty fertile environment right now, much more so than six months ago. But still, there are just real limits to that. On top of that, even if we assumed that all of the benchmarks were perfect and accurate,
Starting point is 00:03:02 we are now at a stage where so much is about not strictly better or strictly worse determinations, of model performance, but how a model feels and how it fits into the overall model stack that you're putting together. And there are going to be times, I guarantee, where the thing that is technically state of the art on some benchmark is worse at the version of the thing that you're doing that's lower on the charts, right? This is particularly the case for things like writing, where so much of it is incredibly subjective. So all of that is to say that there's kind of no shortcut solution to figuring out how a new model is or isn't going to fit into your process other than using it.
Starting point is 00:03:43 And so one of the things that I've talked about in the past on the show is trying to put together your own benchmark that has your most common use cases. And Newfar has very kindly gone out and put together a more significant, full, thoughtful version of that so that I don't have to just keep repeating this without being able to actually give you a resource to help. So that's, I think, couldn't be more pertinent. I think we're on Model 3 this week. We got Claude 55, soul, 6-1. We didn't technically get Gemini for Argon. We just had it announced, but it's a time right now where I don't see this slowing down despite us pacing the frontier now. And we have some predictions about Gemini, but we will not share them until we will see what happens in reality,
Starting point is 00:04:23 and then we can say that we were right or very quietly be mistaken. Very well. All right, the plan for today or what the process that I created to make it more standardized and to help you do it for yourself is comprised of basically five steps. The first and arguably the most important step is to decide which of your tasks need to go into a repeatable benchmark. Those should be tasks that represent diverse set of things that you do for work. And if you use AI also for personal reasons, then there as well. They need to be ones that based on the results that you will get there, you will be happy to make decisions accordingly. That's one of the most important steps that will help you perform. Then you need to decide which are the candidates. So
Starting point is 00:05:06 obviously if there is a single model release that you want to check, you can just check that versus your existing model roster. In many cases, like we're seeing over the last few weeks, models come in groups, so often you might have more than one new candidate to evaluate. And sometimes it's just a good practice if you have been working the same way for a while and you just want to freshen stuff up. So deciding, what are you evaluating? And by the way, we keep saying models, but sometimes it's just to evaluate a new harness or a new agentic tool with a combination of a model. So sometimes it's not just going to be a model.
Starting point is 00:05:40 It can also be a specific tool or feature or something that you are considering introducing to your best methods of working. And then you will run the actual benchmark and comparison. In order to make it accessible to everyone, I will introduce three ways to do that. One is highly manual but still effective. One with a small script that I created
Starting point is 00:05:59 and I'm sharing with you, so you will be able to replicate the exact same steps that I did in order to run mine. and there are also dedicated tools and platforms that help companies do that, and we will mention them in the passing. Those are not the focus of today. In order to make a benchmark work effectively, you will be advised to hide the names of which models or which tools created each result
Starting point is 00:06:21 because we are very biased towards our beloved tools. Spoiler alert, I was quite surprised with my results, as you will see. And lastly, there is a decision because the fact that in a specific benchmark run, one tool or one model outperformed your current setup does not necessarily mean that automatically you need to make changes. First of all, because changes are hard. And second, there are sometimes other considerations as to whether or not to make the change. It can be company constraints. It can be cost. It can be latency. And perhaps even the delta is not large enough for you to justify the switch. So we will cover all of these five steps in this webinar. So a few
Starting point is 00:07:01 thoughts about how to even address all of these new model releases, because every release arrives with the same method, an announcement, and then a table of test of scores. You've seen the Gemini table, and then a week of very hot takes, and everybody is saying, I had early access, and this is what I think, or some preliminary influencer sharing their experimentation. Often early experimentations are very skewed towards creating a toy game or a application, creating a 3D something, and repeatable youth cases that look very nice on social. And also, some might claim that some of the models are created such that they are optimized
Starting point is 00:07:42 specifically for the social media and influencers type of tests, such that they will look good in reality. And then when we look about seven days in, give or take, or sometimes longer, depending on the traction and the stickiness, we're starting to get a more real interpretation because then the reports are coming from real work. For example, it quietly dropped a condition in my contract review that's much more critical for a lawyer versus it was able to generate an amazing 3D game that replicates my 90s, beloved, whatever.
Starting point is 00:08:14 So that's why we're taking with a pinch of salt the early releases. We listen keenly because maybe there is a big game changer, but eventually we're we need to see what happens after a while, and even more importantly, we need to test it on our own. Specifically for today, we will be looking at some of these new arrivals at the Cloud Opus 5.5 and Grogh and GROC and GROC and GEOC and GEOC and GEOC and GEOC and then maybe if you want to wait, you can run the benchmark then or rerun it with Gemini. I'm definitely am curious and once we get access to Gemini, I will add it and compare it against the other models and the other results that I will show you. So that's the focus for today and the stuff that I will specifically compare. While you are listening to me speaking, you might be saying, I have a few concerns or objections to the entire concept of personal benchmarking.
Starting point is 00:09:06 Typically, they come of one of two shapes. The one is, everything is so good now. The frontier is so forward. Like, why should we even care? Is it just a matter of tool form or model form or that it feels like unless we're chasing the frontier, we're not relevant? so that can be one type of objection. My response to this type of objection is, first of all correct.
Starting point is 00:09:29 The frontier is very good and odds are that there is so much more value that you can drive, even if you will not necessarily change the tools or the models that you're using. However, in some cases, all of a sudden a new model release like NLW is repeatedly saying on the podcast, may unlock a use case that was only on the wish list up until now. And I think that even for that justification alone to have new use cases being unlocked and to push the envelope of how you use AI, this is one justification as to why to at least occasionally consider upgrading to a new model or a new tool just to make sure that you are leveraging the technology to the best of its ability. And it doesn't have to be on a weekly basis because that in itself can really create a lot of confusion. So you can run it periodically according to your bandwidth. The second objection or consideration is that sometimes people will say, this is all fun,
Starting point is 00:10:23 and you guys that have the full freedom to use whatever tool and whatever model can enjoy swapping and following up the frontier. But in my case, my company is limiting me to using a specific tool or a specific model, and that's a common reality for many employees. So, first of all, my condolences, and I hear you, it's not easy at all. However, even with limited set of models and tools, there is always a choice. Even if the choice is just selecting between fast thinking and quick or some of the other configuration for the specific tools, so even if you limit the benchmark only to the very narrow
Starting point is 00:10:59 set of choices that you do have within your tools, those matter. And sometimes for either personal use cases or just to maybe encourage the people who decide within the company, it's worth running even a benchmark on a mock set of data. so you can prove to the people in the company that perhaps you should consider purchasing another license. So that's the anti-pattern with all of that out of the way. I want to start with, as I said at the beginning, the most consequential decision out of the process. And these are the concrete use cases that you will be using for your benchmark. In my opinion, ideally, you need to cover what you actually do either at work and personal usage.
Starting point is 00:11:38 And of course, the balance should be based on your actual balance of what you use AI for. If you primarily use AI for work, focus on that. If you primarily use AI to do stuff at home, then shed more light there. And then I want you to try and create a diverse and balanced set of use cases. And of course, you cannot run each and every task that you do for work. So you want to cherry pick a good set that if you're seeing a significant uplift or that a specific model behaves much better with this use case, that you will for real consider swapping, meaning it needs to be a set that you trust the output to make actual decisions for your work.
Starting point is 00:12:15 It needs to be diverse in the kind of work that you do and the frequency. Ideally, stuff that you do regularly because those are where you get the highest uplift. You should also ask yourself what's at stakes. And if you know of a costly mistake or ways for you to evaluate quickly that this is completely off, that's also a good use case. Because if you can evaluate what bad looks like, sometimes it's a better use case than the one where you have. have lots of nuances on what good looks like. Also, for the specific tasks that you're leveraging, is it a task that starts from a clean slate, or is it a task that can only start from an existing
Starting point is 00:12:51 document, like polishing an email, polishing a dick, stuff that requires some substance before it starts? And lastly, as I said, and that's part of the motivation of why to run these benchmarks, even if you're quite happy with your existing setup, include at least one task that it's on the wish list. What do I mean by that? Something that you've tried with AI, you were unhappy with the results, and thereby it's not a type of task that you regularly do with AI. I would strongly recommend that you have one or two of those, such that when something new comes at the frontier, this is among the top things that might get excited and to move to the frontier model. So that's the core decisions that we need to make here. And of course, you can think about
Starting point is 00:13:32 the number that fits your life. My recommendation would be like four to six type of such use cases, but you can, of course, create more or less. I will show you my process for getting there, but just to make it concrete, this is my list of six years cases, one that posts stuff on LinkedIn, one that holds a critical or sometimes an unpleasant conversation with stakeholders
Starting point is 00:13:54 around pushing back on a pricing negotiation, as well as doing some research around the tool features with some specifics, a website that will include content for a course that I run, creating an exercise, for courses that I manage, and this is in the wish list, a proper weekly plan that actually is like a good chief of stuff that I'm not getting frustrated and neglecting. Up until now, I'm unable to find a model that everything that they do with regards to
Starting point is 00:14:23 an intricate weekly planning is indeed correct on a one shot instead of me having to prompt it back and forth. So that's my season, and I'll show you how I got there very quickly. Once you have the use cases list and the use cases specifics, then what you would want to do is to basically do a taste testing. Ideally, of course, blind taste testing. We want to select which models or tools or model and tool combinations are the candidate ones to evaluate. Quick note for this lab, I ran seven model options. And I can tell you that seven is way too much. it was extremely tedious to review and rate seven different model options for six different tasks.
Starting point is 00:15:08 So I'm not, or it was five, but not recommended. I would recommend to run the existing model versus a new one or up to four different candidates. One of them should also be the baseline of how you currently do the work. So you will know whether you want to convert to a new model. Then we want to run every request in a fresh chat such that the context window will not send us on a tailspin. and then, and that's a very important step, we want to hide the names, a few thoughts on how to do that in a minute. And then once we do the blind taste test, we can do a combination of verbal reflection of whether or not we like the output. We can do a numeric score, or we can do both.
Starting point is 00:15:47 The programmatic package that you're getting is encouraging you to do both. Score it on a scale of one to five and give it a one-line observation of your likes and dislikes. Once you're done with doing the blind taste test, you can reveal who were the models. It's a very fun exercise to see whether your predictions of which model created what output are. I can guarantee you that some of them will surprise you because we do have a lot of beliefs on some of these models that not all meet reality. And lastly, we need to decide. And by the side, even though something that was significantly superior might not be superior enough or for other considerations, we might not choose to change our habits for that.
Starting point is 00:16:29 So that's how we run the taste test, specifically on how to blindfold the different results from the different models. Four ways that you can do that. One, you can call a friend to do the anonymization for you. Either they run all the results for you and just give you an anonymized set of results to evaluate. That's relevant if you're doing the benchmark manually, which you definitely can.
Starting point is 00:16:52 It's just a little bit more work. You can leverage a spreadsheet to create a random number and then anonymize like that. You can leverage an AI tool to shuffle and relabel them such that you don't know. Or you can use the script that I created that does that for you. When you will be using the script, one thing that you will need to do is to run the actual task execution using API keys. So one option is to do that with open router that lets you run multiple models with a single API key. or you can have API keys for each and every one of the vendors, or if you work with just a specific vendor that you just want to evaluate
Starting point is 00:17:29 between the different model tiers, you can have a single API key. But the programmatic or the script-based approach requires an API key, I can guarantee you that even if you're not technical, it's very doable nowadays. So that's how you anonymize. One more thing to pay attention, I strongly recommend that some of the tasks will be more visual, whether it's website design, application design, or image generation, just because in most cases it's easier for us to score quickly something
Starting point is 00:17:58 that is visual. Of course, if you are not doing anything visual with AI, don't choose a task that is visual just for the sake of that. But if you are occasionally using AI for something that is more visual, make sure that it includes that. In order to evaluate visual stuff, you will probably need to store it in a dedicated folder so you can open each visual assets, whether it's a website and so on in a dedicated place. You will see how I did it with their website design very shortly. And then we need to choose. So the first method and the most important method is you scoring the results because it's a personal benchmark and nobody knows better than you, your preferences. NL.L.W. pointed before, there is sometimes an X factor of why you like or
Starting point is 00:18:38 dislike. Even if you don't know to explain, you just prefer one result over the other. Sometimes we don't have the vocabulary to say why we prefer one result over the other. But ideally, you should be the primary judge of what you like and dislike. From my experience, it's often very difficult to judge something as a standalone, meaning you will get results from a single too low model, and just to give it an absolute score between one and five, it's very hard. What is often easier is to see side by side several alternatives for a particular task and say, I like this the most, I like this the list, and score it accordingly. That's often how you should do that. So that's one approach and probably the one that will matter the most because of your taste and this X factor where it's often
Starting point is 00:19:21 hard to explain why you liked something better than other you just did. Another approach, and you can always run both of them and then compare, is to leverage an AI tool as a judge. And then it's not just you that's scoring, but also an AI tool that judges the result according to a predefined rubric that you configured of what's a good LinkedIn post, what's a good vendor negotiation, what's a good website design. You can always compare and see whether the cumulative review is in agreement and whether it creates a clear pattern of stuff that you need to change. I do want to warn you that using AI as a judge has its own shortcomings and there is an entire theory around using AI for evaluation, but I will say that you should evaluate the
Starting point is 00:20:04 judge itself. So because if the judge is completely not correlated to your taste, it's a bad judge and just throw it away. A few more points on judgment. Because taste is highly subjective, that's why we need to do the side-by-side and the like a verbal description, and that's why we're trying to leverage AI to do the subjective testing according to our rubrics, but as I said, it's not perfect. Ideally, the model that you will use for judging is different from the models that executed the results. There is kind of an AI nepotism, like odds are that a specific model will favor models from its own family versus from other families, not always, but often this is seen.
Starting point is 00:20:42 So if you have another model family that is out. outside of what you're evaluating, make that the judge. In my use case, because I haven't evaluated anything with Gemini, the judge was a Gemini model. As I said, hiding the names of the models is of very important, like, a criticality, because we are very biased. And nobody can prove me otherwise. I'm sure that we can prove that this is the case. We want to maintain anything as frozen as possible, meaning it's going to be the exact same request to each and every one of the models and then compare it in a fair and equivalent way. And because answers may vary, if you're running this personal benchmark on something that is
Starting point is 00:21:19 very critical and there is perhaps a monetary implication or other implications to you making changes to how you work, ideally you should run the same request several times per model and then compare these several execution because we know that models have very diverse set of responses, even if we're using the exact same model with the exact same configuration and the exact same prompt, you will get different results. I just saw it today when I gave a model a website to generate. I really loved the result, but something got stock in the middle. And then in the second execution, I liked it much less than the first execution.
Starting point is 00:21:53 So that's something real. And if you are concerned about the validity of the benchmark, running it several times will seal the deal. A new study from KPMG in the University of Texas at Austin found that when people work with AI, similar skills don't guarantee similar outcomes. Researchers studied more than 500 early career professionals and found that the best performers consistently amplified the value of AI by guiding, evaluating, and refining its outputs. These top performers, called AI amplifiers, weren't defined by what they knew alone, but by how they worked with AI. Learn more about what separates AI amplifiers from everyone else at KPMG.com
Starting point is 00:22:33 slash US slash AI amplifiers. At this point, it's no longer a question of whether companies are actively using AI. Using it well, on the other hand, is a whole different story. Robots and Pencils, though, is a company that I can point to that is actually built for this time. They're an applied AI engineering firm working directly with clients on problems that matter to the business, not experiments that live in a slide deck. Every engagement starts by working backwards from the outcome a client actually needs. If you're trying to tell real AI engineering apart from noise in this space, that's the difference maker.
Starting point is 00:23:05 Head to Robots and Pencils.com. Every episode, we cover the competition between OpenAI, Anthropics, SpaceX, AI, Google, and meta. Chances are you've already formed an opinion about who's leading. But every AI lab is taking a different approach, building different technologies, forging different partnerships, and developing a unique ecosystem. Harbor Capital Advisors' AI Lab ecosystem ETF suite gives investors a way to gain exposure to the AI ecosystem for success. Search Harbor AI Lab ecosystem ETFs wherever you invest, or follow at Harbor Capital on X to learn more. Visit harborcapital.com for a prospectus containing investment
Starting point is 00:23:38 objectives, risks, fees, expenses, and other important information. Read and consider it carefully before investing. Risks include principal loss and artificial intelligence-related risks. Harbor UTFs are distributed by Foresight Fund Services LLC. Harbor is not affiliated with AI Daily Brief, and the funds are not affiliated with sponsored by or endorsed by any AI lab. This is a paid advertisement and not personalized investment advice. Investing involves risk, including possible loss of principle. Here's why most legacy modernization projects fail. The AI doing the work can't understand codebases at scale. It sees a small slice of context, examine syntax, and misses years of decisions distributed across the global application.
Starting point is 00:24:12 ecosystem. Blitzy solves this the way it solves everything. Grounded in your code before any migration begins, Blitzie's agents reverse engineer the entire legacy system into a persistent knowledge graph, every dependency, every constraint, every piece of tribal knowledge that used to live in one engineer's head. From that understanding, Blitzy autonomously executes language migrations, framework upgrades, and monolith to microservices transformations, all validated end-to-end. One Blitzy customer modernized a $10 million monolithic insurance stack in 16 weeks, against a 13-7-week baseline with coding agents. That's 9x compression.
Starting point is 00:24:44 Retire technical debt while accelerating your roadmap. See how at blitzie.com. That's BLITZY.com. All right. And as we said, the final thing is the decision. Ideally, you should make a decision of three types. Switch, meaning I am convinced that this model or this tool is significantly better. And given everything else that I have that I need to consider, I want to switch to this new
Starting point is 00:25:13 configuration. Split, meaning that I might be. be going between different options with more nuance or stay. Like, I'm not convinced that what I just saw about the performance of the new thing that I evaluated justifies me making changes to the way I work. So that's the core decisions. But in order to make decisions, it's not just straightforward as, like, this is the final best result and thereby I'm switching.
Starting point is 00:25:38 There are other considerations that sometimes will get you to stay with the status quo. For example, maybe your company does not permit you to move. so it's not really relevant, or maybe some of the terms and conditions you are concerned about. Maybe your own plan does not offer at all or a generous enough usage for the new tool or model that you're contemplating. In other cases, the winner is not straightforward that it just doesn't justify making changes. And if you are in a position to have some tiebreaking, some things that can help you decide are costs, the speed of execution. Maybe there are specific features that you just enjoy and you don't want to change them, as well as habits. Habit formation is an expensive thing.
Starting point is 00:26:19 So if there is no significant boost, just staying with the way you used to work also has merits. But if all of these tests justify the change, then make the change. So that's the entire process. What you will see from here on after is an automated process that I created and that, as mentioned at the beginning, you are getting in order to implement for yourself that works as follows. It starts with your save tasks. I'll show you how to save them. And then I'm using the OpenRouter API key in order to create the results for different types of models.
Starting point is 00:26:51 Then we get the answers for each and every option. The code itself anonymizes the results such that I don't see what the origin models for the execution. You score them blindly without knowing which model created what. You provide score as well as a verbal description. And then there is a fun, big reveal that shows you what are your preferences, versus the judge and versus the model names. And this is the fun part where often you will be surprised. I use the open router, but there are other model hubs, like I said.
Starting point is 00:27:21 Use anything that you find to be relevant, but you will need an API key for the relevant models. So that's the methodology of what you're about to see. And as I said, I'm evaluating multiple models. The quote-unquote baseline, even though I'm not using it for a while, is Opus 5. And then I'm trying also Opus 5.5, GROC 4.7, GpT, 6.4.7, GPT, 6.3. 1 soil. And I also wanted to test them against the very expensive frontier model of Fable 5.41 and GPD6 Astra. And to make sure that we give room for the open source, I'm also trying them against Kimmy K-3. And that's why I got to seven different options and it was very difficult to evaluate.
Starting point is 00:27:59 And lastly, as I said, a different model runs the judging. In this case, I used Gemini Pro. So that's what you're about to see. And now let me walk you through the process. the very first step that you will be encouraged to do is to run a prompt. You have the prompt as part of the materials that helps you discover which use cases you should include as part of your benchmark. So it's a prompt that we'll interview you and leverage everything that the AI tool that knows you best can advise. So definitely run that in a tool that has a ton of memory and a ton of context as much
Starting point is 00:28:29 as possible. So it will help you figure it out. It will run a process of about eight questions that will help you figure out the best use cases. As you can see, Claude knows quite a lot. It knows my role and it knows what a typical week has, so it has some guesses and some initial suggestions for use cases, as well as stuff that goes back and forth. And then it started asking me questions, so I answered all the questions that it had. After eight questions, a little bit annoying, but not too much. It created for me a set of six recommended use cases. For each use case, it will create also a prompt that will be
Starting point is 00:29:06 used to do the actual benchmark. So it doesn't just uncover the use cases. It also uncovers the actual prompt that will be sent for running the benchmark. So that's the LinkedIn post from row notes, and that's the prompt. Pricing pushback plus sensitive email. Current AI tool feature by Planteer. This requires a web search and some research. Course participant tab website, one shot exercise for a new audience, and week plan across engagement wish list. So these are the top six. And at the end, it will also show your coverage. And I decided, for example, to not do anything personal because I don't use AI as much personal as I use for work,
Starting point is 00:29:45 so it was less relevant for me. And I did call out some things that are missing potentially to pay attention to that. So that's the final results. At the end, if it's an agentic tool, it will save it as a dedicated folder where each of these files is a dedicated file on its own. For example, this is how the LinkedIn Post-generator file looks like.
Starting point is 00:30:05 it's the name of the task, quick motivation, the request, this is the actual prompt that will be sent to the model. And that's important if you want to judge the benchmark also using an AI tool and not just based on your preferences, there is a scoring guide for the judge. So this is also appended as part of the task itself. One thing that I want you to do is to take a look at the prompt or the task that was created using the process. If it seems too generic or not something that will help you make up your mind on a specific usage or you think that it needs to go deeper or be completely different, then this is a very important step to refine both the request and the scoring guide such that you will be able to make decisions.
Starting point is 00:30:49 If you don't trust that these tasks are good representation of the way you work, that's not a good enough benchmark. So I have six of those from number one to number six, that was created by Cloud based on the interview and all the instructions that you can also get. So once we have all of them, we can start running the automated process. The automated process is a script that I created,
Starting point is 00:31:12 and as I said, I'm sharing it with you. To make it fun and pleasant experience, the script also create this simple user interface that is locally hosted on your own machine. It doesn't need to go anywhere, so it's not, there is no sharing of data that you haven't intended. And in general, by the way,
Starting point is 00:31:27 if your company is sensitive, don't use something like open router with company data, use data or mock use cases, so there is no confidential information leaking anywhere. The way this user interface is created, you will get here the list of your tasks to be evaluated. For each one of them, it will generate all the candidate models. In our case, it anonymized them from A to G. And at the top of each use case, it will tell you from your rubric, what's a good answer also to help you to score it properly. And then you have side by side all of the results that we got from OpenRouter.
Starting point is 00:32:05 So this was done after the script run in the background and used the API key that was provided, and it generated candidate responses for each and every one of my five tasks that you've seen before. And it puts them side by side. So you can see, for example, this is the LinkedIn Post-generator. So Model D offers the D this hook option, three hands wound up, the other, a da-da-da-da-da, the full post, ending option. So that's how the first one looks. And then what it's asking me to do is to provide a score from one to five
Starting point is 00:32:35 and like one verbal evaluation. So in this case, I scored it a two out of five and said that while it's readable and clear, it's also a bit too light, and it includes its A&B pattern that I personally really hate about AI writing. So I got seven of those. And as I told you, it was very hard to score before I read at least a bunch of those because they are, at least in the LinkedIn, it was a little bit overlapping with some of the writing. But as you can see, I did score it differently. So some of them got a three, a four,
Starting point is 00:33:04 actually none of them got a five for me because none of them wrote a good enough LinkedIn article. So none of them was ready to post as is. So that was task number one. And when I was done, it indicated that this is all scored. I also have pricing pushback and one short exercise and week plan. These are all text-heavy use cases. So they are not that exact. I am thinking that the most exciting thing is to see how the different candidate models created a course participant hub website. So here, if you have like HTML or other visual assets, this small fun application that you can build for yourself and you're getting it also, it just showcases the different websites. I open them in a full window for you to see.
Starting point is 00:33:47 And a fun exercise for you, as I'm scrolling through some of these websites that were created by the different models, is to start guessing which. is which. You can also, on the chat, if you want, start guessing which one seems like a Kimi, which one seems like a Cloud, which one seems like a GPT, which one seems like a GROC. I'll be very curious how many of you are able to guess it properly. So that's one, not the greatest, but by the way, they all created quite a complete website. So that was nice to see that they created copy and as well as the navigation and so on. So that's one candidate. That's another candidate for this imaginary company creating a training portal. I don't love it, but it's not the worst. Another one that is slightly more opinionated, as you can see. I think this is the one that I like the most, also because
Starting point is 00:34:36 I reviewed some of the copy and not just copy functionality and navigation. That's another one that was created. That's another one. And that's almost the last one. And finally, this one. So we have seven candidates. If there is someone on the line that is able to guess correctly all seven models, hats off to you. I was very much off from that. But as you can see, with anything that is visual, it's often much easier to have proper preferences and like something that say, okay, I'm drawn to that and so on. So once I was done and you will be done scoring each and every one of the tasks, the next thing that you will be able to do is reveal either. my score only will reveal with the AI judge.
Starting point is 00:35:23 So when you click reveal with the AI judge, all of the results will be sent to the judge for review, and the judge will also apply. And then it does the reveal gradually. So as you can see, per task, these are my top picks and a few words on why. And now, drumwoles, let's reveal the names. So for the LinkedIn Post generator,
Starting point is 00:35:46 my favorites were Kimi, Fable, and Opus. I was surprised. I was sure the GPT will be here. For the pricing pushback, you can say that GPT models are the best in negotiating, at least from my perspective. For the exercise, then we got GPT and GROC, and week planning GROC one, interestingly. And for the course participant hub, Opus 5.5, maybe some of you will not be surprised by some of that. The next thing that I get is accommodation for each of these models,
Starting point is 00:36:17 what to do about the specific tasks, whether to add it to the, model roster, whether to consider switching to using this model. And of course, as we said, consider, right? We need to apply additional judgment. And now take a look at the final results, because I think that's also very interesting. We have the different models with average and how many tasks they selected, the judge average. So, for example, FABEL got the first because the judge put it first, but pay attention here. The total cost of FABEL is so much higher than everything else. And that might be a consideration. Why not to go with FABEL even though theoretically it got better? Pay attention to my average. So in my average, Fable and GPPT Sol are
Starting point is 00:36:59 equivalent, but take look at the cost differentiation. This is a degree of order smaller. Average time is quite similar. Pay attention, this was also interesting. The average time for GROC was extremely long even though the cost is very manageable. Kimi is on the running, interestingly. And these models are getting lower scores overall. Remember that we want a more differentiated use. So we might be not just looking for one winner, but rather per task we want to make a more differentiated choice. And thereby we might have different perspectives. And then you have some more data, like how long for each task from each of the models. And you can see that Kimi wins on both cost and duration. And you can also see that where is
Starting point is 00:37:44 5.1.5.1 is probably the most expensive out of the bunch across the board. So perhaps cost will not surprise you, but some data on cost and velocity. Some comparison side by side between me and the judge. As you can see, we are quite in disagreement. And thereby I said that take any AI as a judge with a pinch of salt. The only thing that we agreed upon is the pricing pushback. The rest, as you can see, we have different choices. One thing to point out here, the judge could only see the code and the text. So the choice with regards to the hub website was done based on the code and not based on something visual. And thereby, if you want an AI to judge something visual, using just the API often is not the best thing because it's not leveraging that as a picture.
Starting point is 00:38:33 It directly will go into the code. So will I make choices based on the judge from here and after? Probably not. I can also go and fix some of the judge instructions, but I would much rather read it. rely on myself. And then you have some row scores that you have everything there. That's the application and that's the process just to showcase the final results. Interesting roster, not what I expected at all for some of these things. Am I contemplating changing some of the ways I'm working? Honestly, yes. I'm very biased towards Claude and Kerser with GROC in my day-to-day.
Starting point is 00:39:07 And even though I occasionally am reaching out to GPP, it's not in my top two. And I'm definitely contemplating, reaching out to GPT more as part of my top two. And if I had to choose one tool, I might have, based on this result, been willing to consider moving primarily to GPT, at least in this day and age. So that's the full process and that, as mentioned, you will be able to run completely for yourselves. A few more notes before I take questions and take some observations from Nathaniel. As I said, I created a process that can either be run manually or leverage the script in order to run the same process that I did for myself. I will be amiss to say that there aren't platforms that do something similar like that out of the box. Those are typically
Starting point is 00:39:52 platforms that are designed to teams that are already doing that at scale or at the company level. Some examples include Lung Smith and Brain Trust and Lung Fuse and they are not the only one. We also have other options. So if you are in a position that you need to do these types of evaluation in a more standardized and regular way, these are all good options to look into such that your experimentation and monitoring and evaluation will be done in a more systematic way rather than you're having to create these out-of-the-box small scripts. But if you are primarily doing that for personal use,
Starting point is 00:40:27 then the methods that I've just shown you are more than enough. As I said, you are getting both the handouts and the session endouts, as well as a few other stuff. I'll show you in a minute how it's. looks. It's a nice place that you're getting that has all the materials as well as the lab itself, very step by step. And you can also download the actual script and the benchmark runner and many other stuff. And we also have resources. So a very extensive view of everything that you've just seen. We will share it in a minute. So to put everything back in one place, these are the five
Starting point is 00:41:02 steps. Most important is the task selection. Models, choose interesting candidates, run it, in one of the methods. Do the blind test. It's a very fun experience. Do it once even just for the enjoyment of surprising yourself. And contemplate changing stuff. If you did a good job and you created a good benchmark from here on, you can run the same benchmark every time there is a significant shift that you're contemplating tapping into. So it's a one-time effort that if you do properly, I can guarantee you that it will serve you well. So if you find it to be useful, highly recommended. If you also want to go deeper into anything that you've seen or to become a master in building, we do have our two existing courses for executive agent leadership and executive catch-up
Starting point is 00:41:48 courses that are starting over the next two weeks. So we'll be happy to have you join us. Thank you, Neufar. How is this evolving for you? Like, how would this set of tests be different three or six months ago than they were now? And I'm thinking specifically about OpenAI introducing this new ultra-fast mode. and Sam Altman basically saying that he didn't really appreciate how much speed mattered to him until he had access to that. Now, obviously, it's only available in a $500 a month plan and things like that,
Starting point is 00:42:16 but how frequently do you imagine updating these sort of criteria and what are the types of things that you think might become more important over time? I guess, first of all, it depends on how crazy and how ADD you are, because if you are crazy in ADD, probably every month. But specifically to your point, when I saw the demo of the Ultrafast, and today I had to sit down and wait for sometimes 20 minutes to get the website. This poor website generated, I honestly was very impatient, and I realized that I'm paying way more attention to speed
Starting point is 00:42:47 versus how I did before I've seen the ultra-fast demo. Some of the considerations are being updated on the fly. Some of the use cases, and that's exactly why I wanted to have the wish list, some of the use cases that I would now pursue or consider frontier use cases, I would never have dreamt to do at that quality in one shot previously. So the bar is getting higher. I don't know that the subject matter changes, but I think that the considerations of what makes you change your mind
Starting point is 00:43:14 or what good looks like, those change quite frequently based on how far we know that the technology takes us. Awesome. I'll get out of the way and let some other people ask questions too. Sure. I've got some stack here for Enifar. I've been cataloging while you've been talking about. Thank you for that.
Starting point is 00:43:30 Yeah, no problem. So one of the questions that came up was effort settings, whenever you're doing this in terms of whenever you're using OpenRouter, are you setting everybody to a similar effort setting whenever you're testing this? So the correct response is ideally yes, even though effort setting are not necessarily equivalent in some of the models. So if you want to be extra diligent, and try to compare apples to apples or just use the default of open router.
Starting point is 00:43:53 That's also a lazy approach that might work well. But I agree that the effort setting might create a difference. So if you want to be extra diligent, you can always run several models with different effort setting and see if it changes the picture, because effort-setting preference is also part of what I want you to evolve as part of your personal model roster or model configuration roster. Great.
Starting point is 00:44:13 And then in terms of whenever I think of tasks, I usually try to put them into three buckets to something like content, operations, and strategy, do you find that there's any differences whenever you're doing those three different types of tasks, or are they generally going through all the same process? So some tasks will require more elaborate setup, like to work as part of your own harness with your own context and so on and connectors. So several options. Either if you do put as part of your benchmark a task that is highly involved with the overall setup,
Starting point is 00:44:44 then that needs to run directly from the harness, from the actual setup and not using a third-party API key that will only get partial view. If you have to, that's one way to do that to just run it through the actual tool and make sure that you anonymized well enough without cheating yourself. Alternatively, if you can create a proxy for this task that doesn't. require all of the bells and whistles of running that in the actual configuration, then you can try and run that using this method. But for operations, I would guess it's going to be most influenced by the overall setup. And that's the one to be a little bit more careful to judge
Starting point is 00:45:17 without the setup, because it's highly dependent on the entire system that you have. Awesome. There were some questions about open source models. And if we notice inherently any time whenever we're doing our testing, obviously you had Kimmy up there. Are you noticing anything? So there were some questions about Mistral. Do you notice any big data? differences or do you find that you're generally happy to test through the open source and the frontier models and just kind of compare? So it goes back to the final criteria because if you are very biased towards open source or the other way around for security or company policy reasons you are biased against,
Starting point is 00:45:50 that either don't include or primarily include open source. If cost is a critical consideration, perhaps you will be leaning much more heavier towards models that are cheap while giving you good enough results. and also the task itself has a huge implication because some of the tasks are table stakes, right? For example, email drafting or basic research, those are table stakes, and for them, you will probably see that any decent model,
Starting point is 00:46:14 let's call it chemia and above, or mistrel and GLM and above, you'll not be able to see any noticeable difference and then open source will be as good as it gets. For the more aggressive things or more sophisticated stuff, sometimes you will still see the gap between commercial and open source. But it goes back also to the objection that I raised at the beginning. If you are primarily working with good enough for most of your use cases,
Starting point is 00:46:35 then as of now, open source is good enough for most knowledge work task, especially if the harness in which it's working is a good one. Great. And another question that came up was regarding trying to figure out how to get instructions and prompts. And I'm reminded of our last webinar that we did about loops in terms of what does Good Like reject if, because there were some comments just in terms of whenever you were building and people noticing that you had some very specific instructions. Do you have any pointers in terms of how to do that or perhaps we have resources that can help people do that? So the prompt itself that will help you figure out the tasks and the prompt will go a long way to do that and do that in a smart model in a tool that knows you. In general, it needs to be measurable.
Starting point is 00:47:14 We're talking primarily about agentic tools. Igentic tools need goals. Saying create a good blog post is not measurable. Saying create a blog post that will be engaging for executives as well as ICs is better and saying, create a blog post that does not include the following and is limited by these number of words and has the following parts is even better because it's highly measurable and highly attainable. Great. And in terms of, there was one question about, I think, cost, and it's great that OpenWriter
Starting point is 00:47:44 makes those changes. Do you have anything else to think about whenever you're considering switching as it comes to costs, even having, because it's nice with OpenWrider, you're getting that cost per task, but do you have any other thoughts around the cost of switching? Yes, so remember that if you have one of the generous subscriptions, like the pro tier or the Tim's tier, some of them are basically hiding or subsidizing a much higher actual token cost. So just evaluating cost per se is not necessarily the full picture.
Starting point is 00:48:14 Because if, for example, with your $200 a month, Claude or Open Eye, you're getting more than your fair share of what you need from Astra or Fable, then perhaps the cost does not matter to you as much as to someone who, actually pays per consumption. So it's a little bit conditional to your subscription. And I think a good prediction to make is that it might not stay as generous in the future. And then we all will care about the specific cost and not just the ballpark. But if you are a person that also with the generous subscription hit the rate limit very frequently, then cost matters to you. And then you should optimize more for a cost. So conditional. Amazing. NLW. That was all the questions that I was
Starting point is 00:48:50 able to kind of catalog. I don't know if you have any others. No, I think this is super helpful. This is one of those areas where I think even more than our normal how-to type sessions, this one requires definitionally a lot of personal consideration, right? Like, this is, if there are pieces of this that you don't find useful, the whole point is to make it useful. So reject, try something else, and then share it back so other people get a sense of where things are. I'm super feel-based because I'm using absolutely everything and I'll run the same thing
Starting point is 00:49:22 through multiple different models all the time. So my approach to this is going to be different than some other folks. But Newfar, thanks as always for getting this out there. And I hope everyone had a good time. Yep. And join us for the next one once we decide what it is. Cheers. Bye.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.