The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? with Greg Burnham

September 29, 2026

AI Summary

5 min read

In May of this year, an AI system solved a problem about how many points can fit in a plane all one unit distance apart from each other—a question the prolific mathematician Paul Erdős had posed decades earlier, and which had resisted significant human effort ever since. That solution was a warm-up act. Soon after, an AI solved one of the seven Millennium Prize Problems, the Navier-Stokes equations, a feat that would have been a career-defining achievement for any human mathematician just a few years ago. Greg Burnham, who leads capabilities research at Epoch AI, studies this trajectory for a living, and his job has become harder precisely because AI has gotten so good so fast.

From School to Work: The Phase Shift in AI Evaluation

Burnham argues that AI capabilities have crossed a fundamental threshold. Up through roughly 2024, the standard way to test AI was with human exams—grade-school math word problems (GSM8K), high-school math competitions, medical licensing tests. AI systems aced those. Now, the frontier has moved from "exam" to "on the job." The problems AI is solving are ones that professional mathematicians have tried and failed to solve. This shift creates a measurement problem: how do you benchmark a system when you don't know the right answer yourself?

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (01:38) **From School to Work: Why Capabilities Measurement Matters Now** - Greg Burnham frames the shift from academic exam-style AI testing to measuring real-world impact.
  • 2 (02:45) **The "School to Work" Transition in AI** - A deeper look at how AI evaluation has evolved from human exams to on-the-job performance.
  • 3 (04:35) **Debating the "No Slowdown" Claim** - Sam pushes back on the idea of no slowdown, pointing to diminishing returns on individual benchmarks.
  • 4 (07:07) **The Counterintuitive Linearity of AI Progress** - Greg explains why the smooth, linear trend across benchmarks is both surprising and suspicious.
  • 5 (13:25) **What "Linear" Really Means** - Greg clarifies the statistical methodology behind the Epoch Capabilities Index.
  • 6 (16:12) **The Engine of Progress: Compute, Data, and the Risk of Discontinuity** - Greg explains the drivers of AI progress and the key risk of a "software-only intelligence explosion."
  • 7 (17:58) **Tripwires for Recursive Self-Improvement** - How Epoch AI tracks whether AI is approaching the ability to improve its own algorithms.

+ Full timestamped outline available in the app

Show Notes

AI systems have gone from struggling with grade-school math to helping solve research problems that have resisted mathematicians for decades, including Navier-Stokes. In this episode, Greg Burnham, who leads AI capabilities research at Epoch AI, joins us to examine what that progress says about where AI is going. We look at how these systems are solving hard math problems, how much they rely on persistence and prior human work, and whether they are starting to produce genuinely new ideas. We also discuss how to measure progress as traditional benchmarks become less useful, why capability gains appear surprisingly steady across model generations, and where models still struggle with open-ended work, learning from experience, and identifying promising new research directions. 🗒️  Full show notes: https://twimlai.com/go/778.

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)