Dwarkesh Podcast
Dwarkesh Podcast

Ryan Greenblatt – What happens once AI can automate AI research?

August 11, 2026

AI Summary

5 min read

Ryan Greenblatt, chief scientist at Redwood Research, argues that once AI systems become capable of automating AI research and development (R&D), a rapid feedback loop could compress years of progress into a single year. He expects this milestone around 2030–2031, with the resulting AI systems becoming superhuman across virtually all cognitive tasks shortly after, by roughly 2033. The conversation focuses on whether this scenario is plausible, what mechanisms could drive it, and why it poses serious risks of misalignment and takeover.

The case for rapid recursive self-improvement

Greenblatt’s argument rests on three claims. First, AI R&D is unusually verifiable compared to other domains. Researchers can create containerized, small-scale environments — training small models, optimizing hyperparameters, or implementing algorithms — and use reinforcement learning (RL) to improve AI performance on these tasks. Because success is measurable and feedback loops are tight, AIs can be aggressively trained to become better at doing AI research itself.

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (00:06) **Introduction and Thesis** - Ryan Greenblatt (Redwood Research) joins to discuss recursive self-improvement: the idea that once AI matches top human researchers, it could kick off a fast feedback loop.
  • 2 (04:29) **Why AI R&D Is Verifiable** - Greenblatt explains that small-scale, containerizable tasks let you train AIs directly on the core components of AI research.
  • 3 (07:28) **Math as an Intuition Pump, and ML vs. Math Depth** - Greenblatt uses AI progress in mathematics to illustrate how verifiable domains enable rapid breakthroughs, but he argues ML research is even more favorable.
  • 4 (10:47) **The "Low-Hanging Fruit" Concern and the Role of Intuition** - Even if low-hanging fruit is exhausted by 2030, Greenblatt expects progress to come from increasingly complicated infrastructure and "in-the-weeds" intuition.
  • 5 (16:52) **Concrete Example: Compressing Five Years of Progress into One** - The host asks what it would look like to go from GPT-3-level compute to a Mythos-class model in one year.
  • 6 (19:29) **The Role of Expert Human Data vs. Algorithms** - Greenblatt argues that scaling up expert human data labeling has not been the primary driver of recent AI progress.
  • 7 (24:35) **How to Get Superhuman Generalists: Transfer and In-Context Learning** - Greenblatt explains how AIs trained on diverse RL environments can transfer to radically new domains like running a chip fab or negotiating a deal.

+ Full timestamped outline available in the app

Show Notes

Ryan Greenblatt is the Chief Scientist at Redwood Research, where he works on technical AI safety research. He's also lead author on the "Alignment faking in Large Language Models", and is currently working on a third party investigation into the OpenAI/HuggingFace incident. In my opinion, he's one of the most interesting thinkers on the future of AI.

Had him on to discuss/debate recursive self-improvement. This might be the most important question in the world right now – whether within a year or so of achieving human-level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields.

I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today.

If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman.

We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031.

We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what’s happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels.

And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world.

The first piece of advice you get when you’re learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy!

Watch on YouTube; read the transcript.

Sponsors

* Antithesis is a software testing platform that finds the failures no human or AI could ever anticipate. It runs thousands of copies of your code inside a fully deterministic computer, injecting faults and steering each trajectory toward the most insidious bugs. This lets you find critical issues in minutes rather than waiting months for your users to uncover them. Learn more at antithesis.com/dwarkesh

*

Dwarkesh Podcast