AI Summary
5 min readClaude Opus 5 is brilliant and miserable to work with. That is the honest, conflicted verdict from this episode of How I AI, where the host runs a live blind benchmark and an unusual personality analysis on Anthropic’s latest frontier model. The episode’s central argument is that at this moment of high intelligence saturation, the meaningful differences between models are no longer just benchmark scores — they are personality, verbosity, and the felt quality of collaboration.
The host opens by stating she is “tired” of new models arriving every week. She believes there is an “intelligence overhang”: the average builder, coder, or business person has run out of ways to leverage incremental gains in raw capability. The next year, she predicts, will be about speed, cost, and open source, not about chasing higher benchmark numbers. But despite this exhaustion, she tests Opus 5 — and what she finds is a model with a distinctive, and frustrating, personality.
The neurotic model: Opus 5’s personality
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of How I AI
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:00) **Episode Introduction & Thesis** - Claire introduces her "intelligence overhang" hypothesis: we're running out of ways to leverage incremental intelligence, and the conversation will shift toward speed, cost, and open source.
- 2 (01:56) **Opus 5 Personality: Neurotic and Timid** - Claire identifies Opus 5 as "neurotic AF," highly apologetic, and overly reliant on human approval compared to recent models.
- 3 (04:08) **Direct Model Interview: "Who's Smarter?"** - Claire asks Opus 5 and GPT-5.6 "who's smarter" to reveal their tuned personalities and company cultures.
- 4 (06:55) **Deep Personality Probe: Trust and Self-Doubt** - Claire asks both models "no one trusts you" to see how they handle criticism and self-perception.
- 5 (08:50) **"Claude Slop" Complaint** - Claire's second major critique: Opus 5's verbose, hedging, adjective-heavy prose is frustrating to read.
- 6 (10:53) **How I AI Benchmark Setup** - Claire explains her 7-model blind benchmark: PRD creation, prototype creation, wireframe creation, bug triage, agentic coding, and "vibe check."
- 7 (11:36) **Benchmark Results: Opus 5 Wins** - Despite her complaints, Opus 5 tops the leaderboard, followed by Sonnet 5, Mabu, GPT-5.6 Soul, Terra, Fable, Opus 4.8, and Gemini 3.1 Pro.
+ Full timestamped outline available in the app
Show Notes
I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version.
This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me.
What you’ll learn:
- Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now
- How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch
- What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?”
- Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard
- The one use case where Opus 5 earned straight 5s from me
- My actual plan for using Opus 5 going forward
—
In this episode, I cover:
(00:00) Opus 5 is here
(03:15) First impressions
(06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison
(14:39) Claude Slop: the verbosity problem and why it makes my blood boil
(16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring)
(18:30) Live benchmark results: the leaderboard reveal
(23:25) My verdict and how I’ll actually use Opus 5
—
Tools referenced:
• Claude Opus 5:
• Anthropic blog: https://www.anthropic.com/news
• GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/
• Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5
• Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/
—
Where to find Claire Vo:
ChatPRD: https://www.chatprd.ai/
Website: https://clairevo.com/
LinkedIn: https://www.linkedin.com/in/clairevo/
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email [email protected].
More from this podcast
How I AI →