How I AI
How I AI

GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark

July 9, 2026

AI Summary

5 min read

Claire, the host of How I AI, spent a week without access to GPT-5.6 Sol, her favorite model, and was miserable. Now it is back, and she has run her own "How I AI Vibe Review Benchmark" to compare GPT-5.6 Sol, Terra, and Luna against Anthropic's Fable and Sonnet 5. The central argument of the episode is that while Fable is theoretically hyperintelligent, GPT-5.6 Sol is practically more useful for the kind of product-building work she does every day. The distinction is not about raw intelligence but about collaboration, design taste, and getting things shipped.

The Benchmark: Taste Over Raw Scores

Claire built a custom benchmark because she got bored of standard vibe checks. It tests models on generating Product Requirement Documents (PRDs), wireframing and prototyping apps, debugging code, and "talking like a human" in an agentic voice. The evaluation uses a 70/30 split: 70% of the score comes from Claire's own qualitative "taste test" (the "Clairvo Taste Test"), and 30% comes from an LLM judge (GPT-5.5). She explains, "I've decided I like my own taste better."

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (00:00) **Introduction & Model Lineup** - Claire introduces GPT-5.6 Sol, Luna, and Terra after a week without access, framing the episode as a love letter to Sol.
  • 2 (02:16) **Pricing Comparison: Sol vs. Fable** - Sol is significantly cheaper than Fable at API pricing, and Claire speculates on subscription implications.
  • 3 (04:12) **The "How I AI" Benchmark Methodology** - Claire explains her custom evaluation framework, which blends an LLM judge with her own "Clairvo Taste Test."
  • 4 (07:44) **Overall Benchmark Results: Sol Wins** - GPT-5.6 Sol scores highest in the weighted index, especially on front-end prototyping and design.
  • 5 (11:59) **Qualitative Evaluation Criteria** - Claire explains what she rewards and penalizes in her "Clairvo Taste Test."
  • 6 (13:26) **Side-by-Side Design Comparisons** - Sol consistently produces more unique, functional, and opinionated prototypes than Fable.
  • 7 (18:22) **Agentic Voice: Sonnet Wins, Sol is Cringe** - For conversational AI, Sonnet 5 has the best "human" voice, while Sol is criticized for M-dash slop.

+ Full timestamped outline available in the app

Show Notes

GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against GPT-5.6 Terra, Luna, Claude Fable 5, and Sonnet 5 across five categories: PRDs, prototypes, wireframes, debugging, and agentic voice. Sol won by a meaningful margin on my Claire Weighted Index (70% my taste, 30% Terminal Bench 2.1), and I also tested two use cases I can't stop thinking about: building a gamified homework tracking app for my kids in one shot with Codex, and browser automation with Chrome that burned through 500 LinkedIn replies while I did literally nothing.


What you’ll learn:

  1. How I scored five AI models (including GPT 5.6 Sol, Fable 5, and Sonnet 5) using my “Claire Weighted Index” benchmark across PRDs, prototypes, code, and agentic voice
  2. The difference between GPT-5.6 Sol (Terra) and Sol for PRD writing
  3. How Fable’s precision and pedantry made it harder to collaborate with, and the exact moment Sol broke through where Fable got stuck
  4. Why Sonnet 5 is still my go-to for agentic voice in OpenClaw, even after this whole benchmark
  5. How I used GPT-5.6 Sol in Codex to build a fully gamified homework tracking app for my kids in one shot
  6. The video editing use case that saved me hours clipping a talk I gave at Cursor’s event
  7. How to use Codex plus GPT-5.6 and Chrome for browser automation, and why this is my single most-loved use case right now

In this episode, I cover:

(00:00) Intro

(01:10) The three GPT-5.6 models: Sol, Terra, Luna

(02:17) Pricing: Sol vs. Fable API costs

(03:24) The How I AI benchmark

(05:03) Claire-weighted Index results

(07:00) Per-task winners: prototypes, PRDs, agentic voice

(11:59) What Claire actually rewards

(13:20) Full-fidelity prototype side-by-sides (Sol vs. Fable)

(17:45) Wireframes

(18:19) Agentic voice

(19:15) Where Sol is better than other models

(23:56) Gamified kids’ homework app, built in one shot

(28:02) Fable’s pedantry problem and how Sol broke through it

(31:49) Two bonus use cases: video editing and browser use

(35:08) Final summary and model recommendations

Tools referenced:

• GPT 5.6 (Sol, Terra, Luna): https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna

• Codex: https://openai.com/codex

• ChatPRD: https://www.chatprd.ai/

• CapCut: https://www.capcut.com/

• Math Academy:

How I AI