World Models and the Future of Spatial AI with Justin Johnson
September 1, 2026
AI Summary
5 min read“There isn't a clear definition of world models that everyone in the field agrees on,” says Justin Johnson, co-founder of World Labs and University of Michigan professor. This ambiguity sits at the heart of one of AI’s most active frontiers. While large language models dominate headlines, a growing community of researchers believes the next leap requires systems that don’t just process text but understand space, predict change, and act in the physical world. Johnson joined the TWIML AI Podcast to untangle what “world model” actually means, where the term comes from, and what it will take to build machines that don’t just generate pixels but grasp the worlds behind them.
Three Threads of a Confusing Term
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (02:23) **Why World Models Are the Next Frontier** - Justin Johnson explains the shared belief that language models alone won't give us embodied, world-interacting AI, motivating the push toward world models.
- 2 (04:15) **The Definition Problem: Three Competing Threads** - Johnson clarifies that there is no single agreed-upon definition, outlining three distinct meanings of "world model."
- 3 (07:04) **All Definitions Are Valid, But Confusing** - Johnson argues all three threads are interesting and worth building, but the field needs better terminology to avoid confusion.
- 4 (09:25) **The Holy Grail: World Model as Theory Builder** - Johnson introduces a fourth, aspirational concept: a model that builds deep, compact, explanatory theories about the world, not just processes observations.
- 5 (12:26) **The POMDP Foundation: Agent, State, and Observations** - Johnson explains the Partially Observable Markov Decision Process (POMDP) formalism, the original technical home of the term "world model."
- 6 (15:56) **POMDPs Beyond Reinforcement Learning** - Johnson shows how the POMDP abstraction applies outside of RL, such as in behavior cloning for robotics.
- 7 (19:43) **Ground Truth State vs. Learned State** - Johnson distinguishes between the true, abstract state of the world and a learned neural vector that behaves like a state.
+ Full timestamped outline available in the app
Guests on this episode
Show Notes
In this episode, Justin Johnson, co-founder of World Labs, joins us to discuss world models and the emerging field of spatial AI. We explore why many researchers see capabilities beyond language as an important frontier for AI, and what it means to build models that can understand, generate, and simulate the environments around them. Justin explains the different approaches to world modeling, including explicit 3D representations and generative models, and why there is still no established recipe for building these systems. We also discuss World Labs’ Marble system, which can generate navigable 3D worlds from images and other inputs, the challenges of evaluating world models, and the role of simulation, planning, and action. Finally, Justin shares his vision for models that bring these capabilities together, supporting everything from interactive virtual environments to agents and robots that can operate in the physical world. 🗒️ Full show notes: https://twimlai.com/go/775.
More from this podcast
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) →