TBPN
TBPN

Model Mayhem: OpenAI’s 5.6 and Meta’s Muse Spark 1.1 | Diet TBPN

July 9, 2026

AI Summary

5 min read

OpenAI’s GPT-5.6 Sol scored 7.78% on the Arc AGI v3 benchmark, up from 1.5% on the previous generation. That number is still tiny—humans can get 100% on Arc AGI—but the jump is the real story, and it signals that the frontier of what AI can generalize to is shifting.

The episode covers a burst of model launches that turned a slow summer into a competitive scramble. XAI released Grok 4.5, built specifically for coding and agentic collaboration with Cursor. Meta announced Muse Spark 1.1, its first serious paid API model, with Mark Zuckerberg returning to X for the first time in years to promote it. And OpenAI dropped GPT-5.6 Sol, a general-purpose model with expanded coding and agent capabilities, alongside GPT Live, a new real-time interactive voice experience. The hosts note that the pace of releases has accelerated so quickly that a "slow summer" remark from 24 hours earlier was already outdated.

The Arc AGI v3 Score and What It Means

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (00:05) **Model Mayhem: A Wave of New AI Models** - The hosts introduce a flurry of new model releases, including xAI's Grok 4.5, Meta's Muse Spark, and OpenAI's GPT-5.6, framing the "slow summer" as anything but for the AI race.
  • 2 (02:58) **GPT-5.6's Arc AGI Score: A Massive Leap** - The hosts analyze GPT-5.6's performance on the Arc AGI V3 benchmark, scoring 7.78%, a huge jump from GPT-4.8's 1.5%, signaling progress in generalization and spatial reasoning.
  • 3 (04:26) **GPT-5.6 Launch Games: Fun and Functional** - The hosts discuss the interactive mini-games included in the GPT-5.6 blog post, such as a sailing game, which demonstrate the model's ability to create polished, vibe-coded software.
  • 4 (06:37) **The Future of Entertainment is Interactive** - The hosts explore how AI is enabling the creation of mini-games and simulators, shifting from static memes to interactive experiences that can be distributed on platforms like Steam.
  • 5 (08:05) **GPT-5.6 Cracks a Magic Trick: A Test of AGI** - The hosts share a story from Stanley Tang of DoorDash, who claimed GPT-5.6 was the first model to figure out a bulletproof magic trick, framing it as a milestone for first-principles reasoning.
  • 6 (09:06) **Pareto Frontier: Fable vs. GPT-5.6** - The hosts discuss a comparison between Fable and GPT-5.6, noting that while GPT-5.6 is faster and more affordable, Fable still excels on the hardest problems, illustrating a healthy competitive landscape.
  • 7 (10:38) **AI 2040 and the Meaning of Model Numbers** - The hosts briefly discuss the release of "AI 2040," a sequel to "AI 2027," which advocates for a slowdown in AI development, and then delve into the confusion around model numbering.

+ Full timestamped outline available in the app

Show Notes

Diet TBPN delivers the best of today’s TBPN episode in 30 minutes. TBPN is a live tech talk show hosted by John Coogan and Jordi Hays, streaming weekdays 11–2 PT on X and YouTube, with each episode posted to podcast platforms right after.


Described by The New York Times as “Silicon Valley’s newest obsession,” the show has recently featured Mark Zuckerberg, Sam Altman, Mark Cuban, and Satya Nadella.


TBPN is made possible by:

Ramp - https://ramp.com

Public - https://public.com

Cisco - https://www.cisco.com

Console - https://www.console.com

CrowdStrike - https://www.crowdstrike.com

Figma - https://www.figma.com

MongoDB - https://www.mongodb.com

NYSE - https://www.nyse.com

Railway - https://railway.com

Shopify - https://www.shopify.com/


Follow TBPN: 

https://TBPN.com

https://x.com/tbpn

https://open.spotify.com/show/2L6WMqY3GUPCGBD0dX6p00?si=674252d53acf4231

https://podcasts.apple.com/us/podcast/technology-brothers/id1772360235

https://www.youtube.com/@TBPNLive

TBPN

More from this podcast

TBPN →