AI Summary
5 min readOpenAI’s GPT-5.6 Sol scored 7.78% on the Arc AGI v3 benchmark, up from 1.5% on the previous generation. That number is still tiny—humans can get 100% on Arc AGI—but the jump is the real story, and it signals that the frontier of what AI can generalize to is shifting.
The episode covers a burst of model launches that turned a slow summer into a competitive scramble. XAI released Grok 4.5, built specifically for coding and agentic collaboration with Cursor. Meta announced Muse Spark 1.1, its first serious paid API model, with Mark Zuckerberg returning to X for the first time in years to promote it. And OpenAI dropped GPT-5.6 Sol, a general-purpose model with expanded coding and agent capabilities, alongside GPT Live, a new real-time interactive voice experience. The hosts note that the pace of releases has accelerated so quickly that a "slow summer" remark from 24 hours earlier was already outdated.
The Arc AGI v3 Score and What It Means
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of TBPN
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:05) **Model Mayhem: A Wave of New AI Models** - The hosts introduce a flurry of new model releases, including xAI's Grok 4.5, Meta's Muse Spark, and OpenAI's GPT-5.6, framing the "slow summer" as anything but for the AI race.
- 2 (02:58) **GPT-5.6's Arc AGI Score: A Massive Leap** - The hosts analyze GPT-5.6's performance on the Arc AGI V3 benchmark, scoring 7.78%, a huge jump from GPT-4.8's 1.5%, signaling progress in generalization and spatial reasoning.
- 3 (04:26) **GPT-5.6 Launch Games: Fun and Functional** - The hosts discuss the interactive mini-games included in the GPT-5.6 blog post, such as a sailing game, which demonstrate the model's ability to create polished, vibe-coded software.
- 4 (06:37) **The Future of Entertainment is Interactive** - The hosts explore how AI is enabling the creation of mini-games and simulators, shifting from static memes to interactive experiences that can be distributed on platforms like Steam.
- 5 (08:05) **GPT-5.6 Cracks a Magic Trick: A Test of AGI** - The hosts share a story from Stanley Tang of DoorDash, who claimed GPT-5.6 was the first model to figure out a bulletproof magic trick, framing it as a milestone for first-principles reasoning.
- 6 (09:06) **Pareto Frontier: Fable vs. GPT-5.6** - The hosts discuss a comparison between Fable and GPT-5.6, noting that while GPT-5.6 is faster and more affordable, Fable still excels on the hardest problems, illustrating a healthy competitive landscape.
- 7 (10:38) **AI 2040 and the Meaning of Model Numbers** - The hosts briefly discuss the release of "AI 2040," a sequel to "AI 2027," which advocates for a slowdown in AI development, and then delve into the confusion around model numbering.
+ Full timestamped outline available in the app
Show Notes
Diet TBPN delivers the best of today’s TBPN episode in 30 minutes. TBPN is a live tech talk show hosted by John Coogan and Jordi Hays, streaming weekdays 11–2 PT on X and YouTube, with each episode posted to podcast platforms right after.
Described by The New York Times as “Silicon Valley’s newest obsession,” the show has recently featured Mark Zuckerberg, Sam Altman, Mark Cuban, and Satya Nadella.
TBPN is made possible by:
Ramp - https://ramp.com
Public - https://public.com
Cisco - https://www.cisco.com
Console - https://www.console.com
CrowdStrike - https://www.crowdstrike.com
Figma - https://www.figma.com
MongoDB - https://www.mongodb.com
NYSE - https://www.nyse.com
Railway - https://railway.com
Shopify - https://www.shopify.com/
Follow TBPN:
https://TBPN.com
https://x.com/tbpn
https://open.spotify.com/show/2L6WMqY3GUPCGBD0dX6p00?si=674252d53acf4231
https://podcasts.apple.com/us/podcast/technology-brothers/id1772360235
https://www.youtube.com/@TBPNLive
More from this podcast
TBPN →