AI Summary
5 min readClaude Sonnet 5 arrived with a clear pitch from Anthropic: near-Opus performance at Sonnet pricing, making it the everyday agentic model the company thinks most users should be reaching for instead of abusing Opus. But when the host of How I AI ran 64 blind generations across five frontier models — including Sonnet 5, Opus 4.8, GPT 5.5, Sonnet 4.6, and Gemini 3 Pro — the results were anything but straightforward. The automated benchmarks and the host's own taste diverged sharply, and the model that topped the leaderboard was a surprise.
The problem with vibe checks
The host has grown tired of one-off "vibe checks" — dropping a new model into Cursor or Claude Code, one-shotting a landing page, and declaring an opinion. The feedback feels soft, unrepeatable, and untested over time. So he built a structured benchmark called the How I AI Bench, designed to be run every time a new model drops. The goal: consistently score models on tasks that actually matter to builders — writing product requirement documents (PRDs), one-shotting prototypes and wireframes, multi-step agentic codebase searches, and an "agentic voice" test that checks whether the model's conversational personality is one you'd want to work with daily.
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of How I AI
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:00) **Introducing Claude Sonnet 5 and the "How I AI Bench"** - Anthropic's new model claims Opus-level tasks at Sonnet-level prices; host is tired of vibe checks and introduces a repeatable, clairvoyant benchmark.
- 2 (02:46) **Sonnet 5's Claimed Strengths: Agentic Tool Use & Computer Use** - Anthropic says Sonnet 5 excels at long-running tool sessions and browser use, nearly matching Opus at lower cost.
- 3 (04:03) **Why Build the "How I AI Bench"?** - Host is tired of non-repeatable vibe checks; wants a clairvoyant, repeatable benchmark with human taste.
- 4 (05:32) **Building the Benchmark with Claude Code** - Host shows how he used Claude Code to brainstorm and build a custom eval set for builders.
- 5 (08:39) **The Benchmark Tasks: PRDs, Prototypes, Agentic Bug Hunting, Voice** - Detailed look at the four task categories and how they were scored.
- 6 (14:14) **Live Leaderboard Reveal: The How I AI Index** - Surprise results: Gemini 3 Pro ties with Sonnet 5 and GPT 5.5 at the top; Opus and Sonnet 4-6 at the bottom.
- 7 (16:20) **Why Human and LLM Judges Disagree** - Models are "easy judges" that miss taste, uniqueness, and visual appeal; host's vibe checks caught things rubrics missed.
+ Full timestamped outline available in the app
Show Notes
I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected.
What you’ll learn:
- What Anthropic claims Sonnet 5 improves over Sonnet 4.6, and where the benchmark data actually backs that up
- How I built the How I AI Bench in under 45 minutes using Claude Code, starting from my own stored session history
- Why I combined human vibe scoring (70%) with LLM as judge scoring (30%) instead of trusting either alone
- How to set up a local HTML scoring page so you can rate AI outputs on gut feel and export those scores as JSON
- Which model I recommend for PRDs, which for complex prototypes, and which for chatting with an agent daily
—
Brought to you by:
Runway—The creative AI platform for images, video and more
Hyperagent—Deploy fleets of agents that handle real work
—
In this episode, we cover:
(00:00) Sonnet 5 is out
(01:55) What Anthropic claims
(04:02) Why I’m done with one-off vibe checks
(05:05) Building the How I AI Bench live with Claude Code
(07:42) The scoring system
(10:43) Agent voice eval
(11:57) Quick recap
(13:58) Results: The How I AI index leaderboard
(21:21) What I’m improving for the next run
(22:16) Generating a Claire-weighted index
(23:53) Model-by-task recommendations
—
Tools referenced:
• Claude Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5
• Claude Opus 4.8: https://www.anthropic.com/news/claude-opus-4-8
• GPT-5.5 (OpenAI): https://openai.com/index/introducing-gpt-5-5/
• Gemini 3 Pro (Google DeepMind): https://deepmind.google/models/gemini/pro/
• Cursor: https://www.cursor.com/
—
Other references:
• SWE-bench Pro (agentic coding benchmark referenced): https://www.swebench.com/
—
Where to find Claire Vo: