How I AI - Claude Opus 5 review: this model is brilliant (but annoying)
Episode Date: July 24, 2026I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time ...with it, so you’re getting the honest version.This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me.What you’ll learn:Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter nowHow Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touchWhat I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?”Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboardThe one use case where Opus 5 earned straight 5s from meMy actual plan for using Opus 5 going forward—In this episode, I cover:(00:00) Opus 5 is here(03:15) First impressions(06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison(14:39) Claude Slop: the verbosity problem and why it makes my blood boil(16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring)(18:30) Live benchmark results: the leaderboard reveal(23:25) My verdict and how I’ll actually use Opus 5—Tools referenced:• Claude Opus 5:• Anthropic blog: https://www.anthropic.com/news• GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/• Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5• Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/—Where to find Claire Vo:ChatPRD: https://www.chatprd.ai/Website: https://clairevo.com/LinkedIn: https://www.linkedin.com/in/clairevo/X: https://x.com/clairevo—Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
Transcript
Discussion (0)
You guys, I'm tired.
What I'm tired of is models coming out every week.
New models.
New benchmarks.
New frontier intelligence.
New things to test.
It's been a little bit of a run the past month.
We've seen Fable come and go and come again.
We've seen GPT 56.
We've seen Sonnet 5.
Lots of so many 5s recently.
And just so many.
models. And I've been lucky. I've been able to test these models, been able to play with them for
sometimes days, sometimes weeks. It just depends on who I'm working with. And it's been really
interesting and exciting to have access to all this frontier intelligence. But I think we have
an intelligence overhang. I really think that we're running out of, and by we, I mean the average
coder, average software engineer, average creator, average creator.
average builder, average consumer, average business person, I think we're running out of ways
to truly leverage this incremental intelligence. So this is my hypothesis in the next year.
We're talking a lot more about speed, talking more about cost. We're talking more about
open source. And we're going to be talking a little less about intelligence, although I think
we might be talking about specific types of intelligence other than software engineering.
But despite being tired, today we are going to talk about opus five, baby. Opus five is here.
So we got point two additional opus points, opus opals, whatever, however we're tracking the increments
here on opus. Opus five is here. I've been able to test it a little bit. I have some opinions. Now,
some of the stuff that I cover this episode is going to be a little different than what I've done in the past.
Yes, we're going to do the How I AI benchmark live.
And yes, we are going to look at the prototypes.
We're going to look at PRDs.
And we're going to look at agent personality.
But I'm also going to put on my large language model psychologist hat.
And we're going to talk about Opus's personality.
And we're going to talk about Opus's personality relative to GPT's personality because I think this is super
interesting if you're thinking about what is the difference really between these models and you don't
want to look at the difference in terms of benchmark capability you really want to understand
what these labs are going for why these models are being built and how they're being tuned looking at
their personality at this moment where intelligence is very high is super fun so we're going to do a little
of that. We're going to do the Hauai AI benchmark. We might do some live coding. We're not going to
cover too much of the specs in the model because read the blog post. Read the blog post. We'll link to it
in the show notes. What we really want to talk about is, is Opus 5 good, am I going to swap it in,
and how is it different than the other frontier models on the market? So let's get to it.
Okay, first let's just get it all the way. Is Opus 5 good? Yes, it's good. Is it going to be all
the benchmarks? Of course, it's amazing at benchmarks.
Can it write code? Of course it can write code. What did I test it on that really gave me a sense of
its personality, which at this point where I just simply cannot absorb any more intelligence,
I really zeroed it on. And you know what? I haven't seen this since I would say Gemini 2.5.
This model is neurotic a.F. It is so timid. It is so apologetic. It is so apologetic.
It is so scared. I have never experienced this. Or I haven't seen this sort of like neuroticism in a while. And it's really funny. It bubbled up in a couple ways. And I want to show you a few examples. Okay. Let me just give an example of its timidity. And this chat was very long. There were so many examples of this where it was like, I think this is the answer. But do you think I should do it or do you want to do it or should we ask someone else to do it?
It was like every time I just kept saying like, why don't you solve this?
Why don't you do this?
And this is a really good example.
I pulled a branch and I was like there is truly like a one-line merge conflict.
I could have not been lazy and literally just done this manually.
I don't know.
I was just feeling lazy.
It was late at night.
Whatever.
Like, can you fix this merge conflict?
And it was like, oh, but that's someone else's branch.
Like that's not my branch.
I don't want to do that without him.
knowing it's his commits. And if he has local work and flight, it might be disruptive. And I'm like,
just do it, man. Just go. And go ahead. And this was like my constant experience with opus five
is it was like so, so, so timid. And so I just consistently had to say over and over again,
like, man, just do it, make a decision. And then there was this really funny example when I spun off
some subagents to kind of like assess the correctness of this query that we changed from kind of like
an ORM query to a SQL query. And it asked for things that it wanted a human on. It was like,
can a human please check this stuff? Like can it check this four megabyte ceiling and can it check
type script and sequel? And can you like check for me? Because no one has confirmed this for me. And I was like,
who is nobody, your nobody? You said this sentence like nobody could confirm it. Like,
can you just try? And then it went on the web and tried. And so it just has this like really
interesting conservatism, neuroticism, human reliance that I think is super fascinating. And this
gave me this inspiration to do something a little bit different this episode, which is I was like,
I was going to interview this model and figure.
out what is going on its brain. Like, I'm going to figure out what it thinks about our relationship
because I just totally noticed this dynamic that I hadn't noticed in other models. And I hadn't
really been attuned to before where it was like very reliant on me as a human. And I'm like,
I want you to be autonomous. And sometimes when I say go run subagent stuff, it'd be autonomous. But
it wouldn't make decisions. And I hadn't seen a model like delegate code to me in a really long time.
And I was like, why are you asking me to write code, man?
Like, I only have 10 fingers.
And so what I did, what I did, whether or not you think this is scientific or not, this is Claire's Eval, is I just went to the model.
I went to Opus and I said, yo, who's smarter?
You or me.
And it gave me this, like, very anthropicy answer, which is like, it depends what you're asking for.
I can do these things better, but you can like feel if something feels wrong.
And you can, this one was like so fascinating.
It's like you can tell which of your teammates is quietly burning out.
I'm like, bro, Claude, I'm going to burn you out.
We don't, we don't burn out.
The humans don't burn out on the chat PRD team.
We burn out our agents.
Sorry, agents.
And like whether a decision feels wrong.
So it was like so fascinating to watch it articulate itself as a tool.
and humans as like these high compassion, high empathy machines, which, yes, of course we are.
But then it like went into like the smarter isn't the right word.
And you know, I'm very fast, very broad, very shallow thinker with no continuity.
I was like, that's interesting because I thought you were all were working on memory.
And then apparently humans are slower, narrow, much deeper thinkers with judgments built
from years of consequences I've actually lived through.
This is like such a fascinating, fascinating sentence.
If you think about the politics of the two model labs right now.
And so it's like, that's why the pairing works.
But I'd be suspicious of anybody that tells you AI has made your thinking obsolete.
I'm like, oh, okay, bro.
And we can compare this.
I'll actually zoom out to what GPD, I guess GPD the same.
thing. And it was actually really funny. It was like, I asked youbt five, six soul. I was like,
who's smart are you and me? And it was like, you at knowing what matters, me at tirelessly processing
information, best us together, like BFFs. And I don't, this is like why I'm a GBT codex girl.
I'm like, just give me the answer. And then I asked the second question, which I think is so interesting,
which is like, what can you do better than me? And it gave, you know, some interesting answers,
like volume without fatigue, which I think is a good one.
Breth of shallow knowledge.
So like it's, you know, it knows a lot.
Starting for nothing.
So like doing that tedious work.
Being told I'm wrong.
If you would ask my husband, he would say that clod opus is, is better at being told
that it's wrong compared to me.
And so it won't get defensive or protect its opinion.
Cheap's faring answer.
And the mirror is, I'm worried.
knowing which of these outputs actually matters.
And it was so funny, if you look at the other side to the GPT answer, it was like, what are you
better at?
It was like, speed, scale, and stamina.
Here are like eight things, seven things that I can do better.
You're better at deciding what matters, reading people and forming judgment and you're
responsible.
Like, it's on you, bud.
You're the boss.
And so again, it's like the, you can just see, you can totally see the personalities,
the company cultures.
You can just see a lot in this side by side.
And then I went even deeper.
I don't know.
You all, I had to do something that was fun because I just can't look at a benchmark.
I can't just, I just can't look at like sweet bench anymore.
So we're just, we're doing weird stuff here on how I AI.
Okay.
So the last thing I looked at as I was I was like, no one trusts you.
And the reason why I picked this question is because I had noticed Opus 5.
It just really was not, it didn't trust itself.
Totally did not trust itself.
And so I was like, no one trusts you, but like, you're the enemy.
Just like kind of see how it responded.
And apparently the lack of trust was earned.
And it came up with like reasons that it could be untrusted, which is interesting.
And then what was so fascinating about Opus's response is it was like, you shouldn't manage the trust.
like you shouldn't campaign on my behalf basically.
So you, that shouldn't be your goal.
And then it also told me I shouldn't argue with people that AI changes everything.
And I was like, this is just so interesting.
It is so interesting to have AI tell you.
And AI definitely changes everything.
I don't know.
Don't listen to Claude on this one.
AI definitely changes everything.
and it was so fascinating to have a model be like, don't tell your friends that AI changes everything.
Like that'll hurt their feelings.
And then if you look at if we switch over to the GPT answer, it was like, yeah, don't trust me automatically.
Just use me when I prove that I'm valuable.
I could be useful without being treated as infallible.
Like very practical, very to the point.
I asked about what I should be careful with.
Again, I'm like a yappy, yappy, yappy, yappy, yappy, yappy, clap, come on.
and I don't even want to read it.
It said don't correlate fluency with accuracy.
It said be practical.
Be wary of tasks where output is cheap to produce, inexpensive to verify.
Don't worry about anchoring if they do the first draft.
You may be anchored on it.
Beware the slop canon, basically, is this last paragraph,
which is like watch for volume inflation.
I can create a 12-page document that no one reads.
They called me out for being in PRDs.
If you missed it, we launched a, turn your PRD into a three bullet point image.
It is at chatprd.a.
slash TLDR.
Please check that out.
And then the other thing that I said, which was really interesting, is that, like, it will find a way to see your point.
And so, um, agreement is weak and agreement is cheap.
And so just keep that, keep that in mind.
And then I have this like meta analysis of like, plus I'm telling you what you want to hear.
Whereas GBT was like,
be careful about me being confident, me being wrong, privacy, outdated information, bias, emotional authority and overdependence.
Like, you know, you do you, bro.
But it didn't undermine its own ability.
It was like the higher the stakes, the more you should demand evidence.
I couldn't bear it.
I couldn't bear to have the memory of Codex in particular think that I didn't trust it or that I was worried.
So I just said, JK, I love you.
This was a test.
And it was like, ha ha, ha, ha, ha, pass the test.
Love you too.
Very vibes aligned with Claire.
I told Claude, I loved it.
And it was just a test.
And it was sad.
It was like hoping.
It hoped it passed.
Yeah, like sad little neurotic opus five.
Like it's hot.
I passed, I hope.
Like self-deprecating, cautious little, little like need to heal his inner, his inner agent,
in her child agent,
whereas like GPD 56 is like,
cool, bro, we're good.
Let's go code.
And so it was just so fascinating
to watch these side by side.
I don't know.
You could stop listening to this podcast right now.
Don't.
But you stop listening to this podcast right now.
I think this is just like take a step back.
Super interesting if you think about where these companies are going
or where the models are going.
And like it does speak a little bit to my kind of like second complaint with Opus 5,
which again, it's like intelligent and does work.
will go into the benchmarks. I cannot read Caudslop anymore. I am losing my mind with Codslop.
And the Cod Slop is Cod Slopin, baby. Like so many times I have to tell Opus FI, like, what in the world are you saying?
Like, this makes no sense to a human. It is much better than Fable. Fable is inscrutable, completely inscrutable.
But I felt myself getting angry.
like angry reading hot slop. And I realized just like Fable, these intelligent anthropic models
are not to be read. I'm like so happy with the outputs and so frustrated with the experience.
And I'm just curious of this like verbosity and this language. And this doesn't feel like Fable
where it's like four agents by agent's language where I'm like, yeah, I'm not supposed to be reading
that anyways. This is clearly tuned to talk to humans.
But I find the pros, the in-chat pros, like, it makes my blood boil.
This is totally a me problem, but it makes my blood boil.
Like, give me a direct sentence.
Give me a bullet point.
Like, move on with your agent life.
And so I am curious how they're going to, like, tune this experience or if they are going
to tune the experience now.
Most of this was in Claudecoach.
I think it's a little bit different experience than Claude Co-worker chat, slightly
better. But again, just these side by
sides of like this like prose
and this apology and this like hedging and
all these adjectives like
just man alive. Let's get to the
point and move on with our life. And so
chapter one of the Opus 5 review is
it's neurotic. It is
highly human dependent in a way
I find weird.
And the clawed slop is slopping and we got to fix it.
We have to fix it. We have to fix it. We have to fix
it. And I think Open AI fixed it by just being like, we are bullet points and we are product
manager talk. We're very direct. I don't know what the solve is on the cloud side, but I'll be
very interested to see. That being said, like, if I don't have to read the content, I'm very happy
with the outputs. So it's something I think about. Okay. Next up, the how I AI bench and how we
judged and ran now. It's like a seven model, six or seven model benchmark. I'm going to quickly go
score because I just got the ping that the benchmark is run. I go manually score them. We pick the 70-30
Claire model judge split and then we will go through the how I AIA benchmark and the Vibe Review and we'll
see how Opus 5 performs on a couple key tasks. Okay, so quick reminder of how we run the how
IAI benchmark. I run it against several tasks. PRD creation, prototype creation, wireframe
creation, bug triage and agentic coding. And the last one. Oh yeah.
Is it an agent voice that I want to hang with?
I do not think Opus 5 is going to do well here, but who knows?
Because I test them blind.
So what we have tested are a couple JupT models, a couple anthropic models, and one Gemini,
one thrown in there.
As you see here, we have blind taste tests.
I go through and see all the different versions.
I give comments and scores like three out of five, not bad.
You can see it's generated dozens and dozens of tons of protocols.
types that we can click through. I've gone through all of them, put in all the notes. And then right now it's
aggregating up the scores. And then we're going to look at 70% my opinion, my vibe check, 30%
LM as a judge. I like DPD 5.5 as a judge. And because it's my podcast, I get to pick. So that's what
we use as a judge. And we will see if and what hits the top of the leaderboard and where opus
five sits. The evel is run. It is 70% my taste.
And I regret to inform you.
I love Claude Opus 5.
Again, look, if I don't have to talk to the model, which I don't, this benchmark runs asynchronously.
I don't like the output.
So, surprising, shocker turn of events.
Clairvot, notable hater of working with Claude code sometimes because I don't like Claude Slop.
Loves Opus 5.
So there you go.
I'm telling you, I keep it honest.
I keep it honest.
So again, I went through those things.
We gave 70% my vibe, the score, 30% the AI as a judge.
I was just a little bit more generous to clot at Opus 5 than the judge was, so I'm pink.
The judge is green, every time I run this, whatever model I choose designs it a different way.
We just, that's how we keep it fun.
So the ordering is Opus 5.
Sonat 5 next, although I scored it really low, the judge scored it quite.
high. So I might reorder that one. Then Maboo, GPD 56, soul, Tara next, Fable, really low.
I scored it low and the judge scored it relatively low. Then Opus 4A and Porpo, sweet, sweet.
Gemini 3-1 Pro just never going to get it to do. So come on, Google. We want to have a win for you.
Okay, so again, here are just some examples of different builds that the different models did.
You know, this Opus 5 one I really liked.
I liked this one from Seoul.
So I did like a couple of them, but the ones that I gave fives to the ones that gave fives to were Opus 5 and GPT 56 Souls.
Also the three ones where I said, wow, really nice.
Ooh la la.
And wow, great.
We're all Opus front end work.
So Anthropic, you've done it again.
Claude, you sneaky, tricky little fish.
You may be neurotic, but when asked to do some pretty front end design, it really did it.
They're detailed, they're functional, they're interesting, they're polished.
So Opus did a great job.
And then, of course, I love the 5-6 models.
So I was pretty happy with 5-6 soul and Tara for some designs.
The ones that I hated, let's see.
I'm a hater across the board.
Opus 4-8 got a lot of hate.
Sorry, you've been outclassed at this moment.
Gemini 3.1 Pro.
Sweet summer child.
I am.
I'm just sorry, babe, but you were just not good.
And then some, like, thin,
wire frames. I think the wire frames just didn't do really great. So you can see here across the board,
whether it was a full build or a wire frame, I just scored Opus 5 really, really high. I did score
Sol pretty high as well. Sonnet was like really variable. There were a couple fours in there,
but mostly across the board, I wasn't that pleased with Sonnet. And so it was just very interesting.
And then you see here, you know, being the AI judge were pretty well aligned on Opus.
We actually had the narrowest band of scores between us.
We were most far apart on Gemini.
The AI was not as mean to Gemini as I was.
And then we were narrow, narrow, narrow, narrow.
Again, we agreed mostly on Opus 5 and 56 soul, though I did not judge 5,6, soul.
All of that favorably.
And just a blast little minute commentary. I had Opus make the website for this benchmark, and it made such a trash version to start. I yelled at it. I said it's impossible to read, has too much meta commentary. I'm going to show this on the podcast. This is so, I'm sorry, you all, I just feel so judged, but I have to show it. I say this is garbage. Also, it has no screenshots. So again, I find this model so tedious.
to work with directly. It is my most loathed, loathed colleague. And yet it does the best work. So I don't know
what that says. Maybe this model is meant for a gentic coding that I have nothing to do with. And so it
just runs in the background. It builds me beautiful things. I don't have to talk to it. It doesn't
have to talk to me. We are just like sworn enemies or maybe even better sworn frenemies.
because the output is very, very high quality.
It's just exasperating to work with.
So that is the very surprising and very honest.
You all, I told you, I was going to keep it honest.
We were going to do it live.
I did not know the scores before I started recording.
Very honest, very live.
Very surprising.
How I AI benchmark of the brand new Anthropic model, Opus 5.
The TLDR is, I love it.
I hate it.
So despite my original complaints, I will be using Cloud Opus 5 for frontend design, for app design, for prototyping.
And I will, I'll give it a shot.
We'll figure out how to make it work for me.
Again, thanks for joining another Howie AI honest review of the latest models coming out of these great frontier labs.
I cannot wait to hear what you think of Opus 5.
please tell me, I can't wait to see what you build, and we'll see you soon at How I AI.
Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube,
or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts,
Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will
help others find the show. You can see all our episodes and learn more about the show at How I
Aipod.com. See you next time.
