AI Summary
5 min readClaire, the host of How I AI, spent a week without access to GPT-5.6 Sol, her favorite model, and was miserable. Now it is back, and she has run her own "How I AI Vibe Review Benchmark" to compare GPT-5.6 Sol, Terra, and Luna against Anthropic's Fable and Sonnet 5. The central argument of the episode is that while Fable is theoretically hyperintelligent, GPT-5.6 Sol is practically more useful for the kind of product-building work she does every day. The distinction is not about raw intelligence but about collaboration, design taste, and getting things shipped.
The Benchmark: Taste Over Raw Scores
Claire built a custom benchmark because she got bored of standard vibe checks. It tests models on generating Product Requirement Documents (PRDs), wireframing and prototyping apps, debugging code, and "talking like a human" in an agentic voice. The evaluation uses a 70/30 split: 70% of the score comes from Claire's own qualitative "taste test" (the "Clairvo Taste Test"), and 30% comes from an LLM judge (GPT-5.5). She explains, "I've decided I like my own taste better."
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of How I AI
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:00) **Introduction & Model Lineup** - Claire introduces GPT-5.6 Sol, Luna, and Terra after a week without access, framing the episode as a love letter to Sol.
- 2 (02:16) **Pricing Comparison: Sol vs. Fable** - Sol is significantly cheaper than Fable at API pricing, and Claire speculates on subscription implications.
- 3 (04:12) **The "How I AI" Benchmark Methodology** - Claire explains her custom evaluation framework, which blends an LLM judge with her own "Clairvo Taste Test."
- 4 (07:44) **Overall Benchmark Results: Sol Wins** - GPT-5.6 Sol scores highest in the weighted index, especially on front-end prototyping and design.
- 5 (11:59) **Qualitative Evaluation Criteria** - Claire explains what she rewards and penalizes in her "Clairvo Taste Test."
- 6 (13:26) **Side-by-Side Design Comparisons** - Sol consistently produces more unique, functional, and opinionated prototypes than Fable.
- 7 (18:22) **Agentic Voice: Sonnet Wins, Sol is Cringe** - For conversational AI, Sonnet 5 has the best "human" voice, while Sol is criticized for M-dash slop.
+ Full timestamped outline available in the app
Show Notes
GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against GPT-5.6 Terra, Luna, Claude Fable 5, and Sonnet 5 across five categories: PRDs, prototypes, wireframes, debugging, and agentic voice. Sol won by a meaningful margin on my Claire Weighted Index (70% my taste, 30% Terminal Bench 2.1), and I also tested two use cases I can't stop thinking about: building a gamified homework tracking app for my kids in one shot with Codex, and browser automation with Chrome that burned through 500 LinkedIn replies while I did literally nothing.
What you’ll learn:
- How I scored five AI models (including GPT 5.6 Sol, Fable 5, and Sonnet 5) using my “Claire Weighted Index” benchmark across PRDs, prototypes, code, and agentic voice
- The difference between GPT-5.6 Sol (Terra) and Sol for PRD writing
- How Fable’s precision and pedantry made it harder to collaborate with, and the exact moment Sol broke through where Fable got stuck
- Why Sonnet 5 is still my go-to for agentic voice in OpenClaw, even after this whole benchmark
- How I used GPT-5.6 Sol in Codex to build a fully gamified homework tracking app for my kids in one shot
- The video editing use case that saved me hours clipping a talk I gave at Cursor’s event
- How to use Codex plus GPT-5.6 and Chrome for browser automation, and why this is my single most-loved use case right now
—
In this episode, I cover:
(00:00) Intro
(01:10) The three GPT-5.6 models: Sol, Terra, Luna
(02:17) Pricing: Sol vs. Fable API costs
(03:24) The How I AI benchmark
(05:03) Claire-weighted Index results
(07:00) Per-task winners: prototypes, PRDs, agentic voice
(11:59) What Claire actually rewards
(13:20) Full-fidelity prototype side-by-sides (Sol vs. Fable)
(17:45) Wireframes
(18:19) Agentic voice
(19:15) Where Sol is better than other models
(23:56) Gamified kids’ homework app, built in one shot
(28:02) Fable’s pedantry problem and how Sol broke through it
(31:49) Two bonus use cases: video editing and browser use
(35:08) Final summary and model recommendations
—
Tools referenced:
• GPT 5.6 (Sol, Terra, Luna): https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna
• Codex: https://openai.com/codex
• ChatPRD: https://www.chatprd.ai/
• CapCut: https://www.capcut.com/