AI Summary
5 min readOpus 5.5 vs. GPT-6 Sol: a blind taste test
Both Anthropic and OpenAI dropped new models on the same morning, forcing an unplanned live comparison. The host had been ready to record a dedicated Opus 5.5 review, but with GPT-6 Sol and GPT-6 Luna also landing, she pivoted to a blind taste test using an expanded version of her personal "How I AI" benchmark. The central question: which of these daily-driver models—not the frontier systems like Astra or Fable, but the workhorses—would actually win her day-to-day use?
The models and what changed
Three models entered the ring: Opus 5.5 from Anthropic, and GPT-6 Sol and GPT-6 Luna from OpenAI. Opus 5.5 is positioned as the first Opus-level model shipped with the same cyber and bio guardrails that Fable had, meaning it will refuse certain security and biology tasks. It is also cheaper than Opus 5 and Fable 1, and Anthropic claims it was built with external evaluators after "pacing the frontier." On the OpenAI side, Sol is the host's long-standing favorite—cheap, fast, and enjoyable to use. Luna is a variant. The pricing difference is notable: Opus 5.5 costs about twice as much as GPT-6 Sol.
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of How I AI
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:05) **Three New Models Drop Simultaneously** - The host explains that Opus 5.5, GPT-6 Sol, and GPT-6 Luna all launched the same morning, forcing a live, unscripted review.
- 2 (02:31) **Model Positioning and Pricing Landscape** - The host distinguishes these "daily driver" models from frontier models like Astra and Fable, and highlights the aggressive price competition.
- 3 (04:11) **Anthropic's Safety Guardrails and Personality Shift** - Opus 5.5 ships with Fable-level cyber and bio guardrails for the first time, and the host notes how safety philosophy leaks into model personality.
- 4 (07:50) **Speed and Communication Style Differences** - A comparison of perceived vs. actual speed between Opus 5.5 and Sol, driven by how each model narrates its work.
- 5 (09:04) **The Expanded How I AI Vibe Benchmark** - The host unveils her new, broader benchmark covering knowledge work, personal productivity, front-end/back-end coding, agentic tasks, and creative work.
- 6 (11:47) **Personal Productivity: Email Triage Results** - Blind evaluation of how models handled fake emails, feedback, and calendar invites in the host's voice.
- 7 (13:51) **Front-End Coding: Complex Dashboard Prototyping** - The host evaluates 10 different models on front-end tasks including editorial, dark mode, incident management, and B2B renewals.
+ Full timestamped outline available in the app
Show Notes
I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived.
What you’ll learn:
- How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it
- Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend
- Where Sol still wins me over on clear writing, readable PRDs, and price
- The character SVG results that completely overturned my prediction about Anthropic
- What happened when I asked these models to edit video, and why I think skills explain part of the disappointment
- Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t
—
In this episode, we cover:
(00:00) LIVE setup and new model launches
(01:30) What’s new in Opus 5.5, Sol, and Luna
(04:11) Guardrails, personality, and speed
(09:00) The How I AI bench and blind evaluation process
(11:31) Email and personal-productivity results
(13:50) Frontend prototype vibe checks
(24:10) Backend, agent personality, and long-running tasks
(28:25) SVG illustration test
(29:48) AI video-editing results
(30:43) Predictions before the reveal
(31:20) Barbie Bench: the 3D fashion-game test
(34:17) Results: Astra, Sol, and Opus 5.5
(35:04) Writing clarity and creative surprises
(36:51) Why the LLM judge disagreed with me
(37:24) What each model is actually best for
—
Tools referenced:
• Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5
• GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/
• Codex (OpenAI): https://openai.com/codex
—
Where to find Claire Vo:
ChatPRD: https://www.chatprd.ai/
Website: https://clairevo.com/