Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Academia is for Ambition — Alex Zhang, MIT

October 2, 2026

AI Summary

5 min read

Alex Zhang, a PhD student at MIT and one of the creators of Recursive Language Models (RLMs), argues that the most powerful lever in AI research right now isn't a bigger model or more data—it's the design of the harness. A harness is the scaffolding of code, tool calls, and memory that wraps around a language model to let it solve complex, multi-step problems. Zhang believes most current harnesses (like Claude Code or Codex) are essentially the same, and that the real opportunity lies in building more opinionated, compositional systems that let models generalize far beyond what they were trained on.

The Harness as the Missing Ingredient

The core insight is that a language model's next-token prediction format is a terrible interface for solving long, multi-step tasks. You can't just ask a model to navigate a large codebase or write a complex GPU kernel in a single prompt. The harness is what bridges this gap—it's the program that orchestrates the model's actions, decides when to call tools, and manages context.

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 Academia is for Ambition — Alex Zhang, MIT
  • 2 Timestamped Outline
  • 3 (00:03) **Introduction and GPU Mode Origins** - Alex joins the show; discusses how he got involved with the GPU Mode (formerly Cuda Mode) Discord community
  • 4 (04:41) **GPU Programming: From Niche to Saturated** - Alex reflects on how GPU programming has evolved from a niche skill to a nearly saturated field
  • 5 (07:21) **Human Expertise vs. Compute Scaling** - The value of domain knowledge when working with AI-generated solutions
  • 6 (11:21) **Bootstrapping and the Limits of Automation** - Why generating kernel code autonomously is still difficult
  • 7 (13:18) **Research Taste and Taking Big Bets** - Alex's background at Princeton with the SWE-bench team and his philosophy on academic research

+ Full timestamped outline available in the app

Show Notes

Last call for regular tickets for AI Engineer NYC! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for new tickets only, no refunds! See you in 2 weeks!

While we tend to cover industry on the pod, every so often we celebrate a clearly emerging superstar PhD. In 2024 we featured Shunyu Yao, who went on to build Operator at OpenAI and is now Chief AI Scientist of Tencent. In 2025 we featured Jack Morris, who went on to cofound Engram at $600m and is now a leading voice on continual learning.

This year we are proud to feature the work of Alex Zhang of MIT.

From GPU kernels and KernelBench to Recursive Language Models, Mismanaged Geniuses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems.

RLMs took over the timeline early this year:

and an RLM based harness was the first to ~solve ARC-AGI-3 before OpenAI’s Astra:

and is even today, influencing new research that has more extreme implications than RLMs:

We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research

Latent Space: The AI Engineer Podcast