Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again
August 18, 2026
AI Summary
5 min readThe Bitter Lesson's Next Chapter: Why AI Must Learn Again
Rich Sutton, the AI researcher who wrote the seminal "Bitter Lesson" essay and co-invented reinforcement learning, has a straightforward way of describing his worldview: "I'm thinking the ordinary way. It's just everyone else that's thinking a bit weird." He means this literally. Before the current AI boom, he argues, nobody would have needed to say "continual learning" because "it wouldn't make any sense to talk about learning that wasn't continual. All learning is continual."
Sutton and his former student and now co-founder Khurram Javed are building Oak Lab to pursue what they see as the missing piece in modern AI: systems that actually learn from ongoing experience rather than being frozen after training. The conversation reveals a deep fault line in how the field thinks about intelligence, and a concrete technical agenda for crossing it.
The Bitter Lesson and Its Limits
Sutton's 2019 essay "The Bitter Lesson" argued that AI progress comes not from embedding human knowledge into systems, but from methods that scale with computation—search and learning. Large language models are a perfect positive example: they drank in the internet and got dramatically more capable just by scaling.
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of Training Data
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 Timestamped Navigation Outline
- 2 (00:00) **Why Sutton Thinks He's Not the Radical One** - Opening framing: Sutton argues that continual learning isn't a special subfield—it's just learning, and the field is weird for treating it otherwise
- 3 (01:03) **Sutton's Origin Story: Alberta Against the Odds** - How Sutton ended up founding the reinforcement learning program at University of Alberta during an AI winter
- 4 (03:41) **The Bitter Lesson Defined in 26 Words** - Sutton distills his famous essay into its essence
- 5 (09:53) **LLMs: Both a Positive and Negative Example of the Bitter Lesson** - Sutton evaluates large language models against his framework
- 6 (11:16) **Why Synthetic Data Is a Dead End** - Sutton and Javed explain the "Big World Hypothesis" and why synthetic data doesn't escape human bottlenecks
- 7 (18:05) **Prior Knowledge vs. Learning: False Enemies** - Sutton resolves the apparent tension in the Bitter Lesson
+ Full timestamped outline available in the app
Show Notes
Rich Sutton, who helped pioneer reinforcement learning and wrote the seminal AI essay The Bitter Lesson, has now cofounded Oak Lab with his former student Khurram Javed. Their goal: to build agents that continuously learn from their own experience rather than from us. Rich doesn't think he holds a radical view: "I'm not weird. The field is weird." He says all learning is continual, and the field is the one that needed a new name for it. Rich and Khurram argue synthetic data is "a big mistake." Their "big world hypothesis" is that the world is massively more complex than any agent or simulator, so approximations have to be updated continuously rather than frozen at deployment. Rich calls LLMs an unanticipated scientific breakthrough, but says they represent roughly a quarter of intelligence. He says catastrophic forgetting is "totally curable" with the ideas behind their continual backprop algorithm. Khurram explains why the frontier labs can't follow: they sit in a local minimum where a new paradigm gets worse before it gets better. Their target, five to ten years out, is a trillion-parameter mind that keeps learning, stays coherent, and runs on 20 watts.
Hosted by Sonya Huang and Alfred Lin, Sequoia Capital
More from this podcast
Training Data →