a16z Show
a16z Show

Fei Fei Li: The Race to Build World Models For AI

September 4, 2026

AI Summary

5 min read

From Pixels to Places: Atlas and the New Primitive for Spatial AI

World Labs launched Atlas, a new kind of world model that combines generation, reconstruction, and simulation into a single system. The core insight is deceptively simple: instead of predicting the next token (like language models) or the next frame (like video models), Atlas predicts the next view of a scene from an arbitrary camera position. The team believes this "new view prediction" could become the foundational primitive for spatial intelligence, much as next-token prediction became the foundation for large language models.

What Makes Atlas Different

Atlas accepts text, images, video, and camera poses as native inputs. When you give it one or more images—each tagged with its precise 3D camera position—the model can generate what a scene looks like from any other viewpoint you specify. This is fundamentally different from standard video generation. Video models produce plausible-looking frames but have no internal understanding of 3D space; ask them for a different angle and the scene rearranges itself. Atlas, by contrast, produces frames that are spatially grounded—the geometry stays consistent because the model has to reason about where things actually are.

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (00:00) **Introduction: The New Paradigm of Spatial Intelligence** - Host Martin Casado introduces the episode's core question: what happens when AI learns to predict the next view of the world instead of the next token.
  • 2 (01:44) **Atlas: A New Frontier Model Launched** - Justin Johnson describes the three core capabilities of the newly launched Atlas model: generation, reconstruction, and simulation.
  • 3 (02:39) **The "Bullet Time" Demo Explained** - The team explains how Atlas achieves the iconic "bullet time" effect from *The Matrix* using just three cameras, a 50-100x reduction in equipment.
  • 4 (03:33) **The Core Primitive: New View Prediction** - The team defines the fundamental primitive of Atlas, distinguishing it from next-frame prediction used in video models.
  • 5 (04:18) **How Atlas Differs from Other Video Models** - Ben Mildenhall explains that Atlas's spatial grounding is the key differentiator from other generative video models.
  • 6 (05:48) **A New Architecture, Not Just a Scaled Video Model** - Justin Johnson explains that Atlas is a fundamentally new architecture that jointly performs generation and reconstruction in a single model.
  • 7 (07:32) **The Unification of Pixel Generation and Reconstruction** - Fei-Fei Li provides a historical perspective, noting that this unification is a major milestone after decades of separate tracks in computer vision.

+ Full timestamped outline available in the app

Show Notes

World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join a16z General Partner Martin Casado to discuss Atlas, their latest world model, and what it reveals about the pursuit of spatial intelligence.

At the center of Atlas is what the team calls “new view prediction”: given images or views of a scene, the model predicts what that environment should look like from a different position in space and time. This brings generation and 3D reconstruction into the same model, and raises a broader question about whether predicting views could become a useful primitive for understanding the physical world.

They discuss the technical bets behind the model, what it can and can’t yet capture, and the importance of dynamics, editability, and simulation as world models develop. The conversation also explores applications in creative work, architecture, and robotics, where Fei-Fei argues that one of today’s biggest constraints is access to real-world training data.

 

Resources:

Follow Fei-Fei Li on X: https://x.com/drfeifei

Follow Justin Johnson on X: https://x.com/jcjohnss

Follow Ben Mildenhall on X: https://x.com/BenMildenhall

Follow Martin Casado on X: https://x.com/martin_casado

Learn more about Atlas: https://www.worldlabs.ai/blog/atlas



 

Stay Updated:

Find a16z on YouTube: YouTube

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Show on Spotify

Listen to the a16z Show on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

a16z Show

More from this podcast

a16z Show →