The Professor of Outputmaxxing — Anjney Midha, AMP
June 18, 2026
AI Summary
5 min readThe Professor of Outputmaxxing
Anjney Midha, CEO of AMP, has a memorable teacher from his boarding school in India who would tell him: "Luck favors the prepared mind." It's a phrase he now applies to Anthropic's breakthrough in coding. When Claude cracked coding last October, many called it luck. Midha sees it differently: Anthropic had been "the most prepared company for four years," forced into ruthless efficiency by scarce resources while competitors raised billions. That hardship, he argues, was a feature, not a bug—and it's a lesson he's applying across the AI infrastructure stack.
The Two Types of Utilization
Midha draws a sharp distinction between two metrics that get conflated in AI infrastructure. The first is node allocation—what percentage of GPUs in a data center are actually turned on and assigned to someone. At Google, where his co-founder Seb built the Borg scheduler, 95% allocation is considered an outage; 96% should be standard. Most single-tenant clusters don't come close.
The second metric is MFU (model flop utilization)—what fraction of those allocated GPUs' theoretical compute capacity is actually being used for useful work. Best-in-class today sits between 60% and 70%. The gap between these two numbers represents massive waste.
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of Latent Space: The AI Engineer Podcast
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:03) **Two Types of GPU Utilization** - Defines node allocation (95% target, considered an outage below that) and MFU (best-in-class 60-70%)
- 2 (02:13) **"Common Sense" in AI Infrastructure** - Argues that AI scaling demands more, not less, infrastructure discipline
- 3 (03:21) **Community Backlash as a Data Center Risk** - Up to 20% of US data centers face community opposition
- 4 (05:07) **The "Neo-Cloud" vs. Trusted Operators** - Skeptical of marketing hype; prefers established data center providers with 20+ year track records
- 5 (06:28) **AMP as an Independent System Operator (ISO)** - Horizontal pooling layer across clouds and silicon, modeled on the electric power grid
- 6 (08:56) **The "Oxygen" Program and Interruptible Demand** - Google's internal system for guaranteeing base load while allowing research spikes
- 7 (09:45) **The Tragedy of Hoarded Research** - DeepMind's best work often never sees production; papers with business potential are embargoed indefinitely
+ Full timestamped outline available in the app
Guests on this episode
Show Notes
Last 4 days before regular tickets sell out at AI Engineer World’s Fair - this is the single biggest gathering of AI Engineers, Founders, Leaders, and Researchers in the world. Attendees get >$5000 worth of sponsor credits and talk tracks are looking FANTASTIC. Join us!
The AI scaling debate always focuses on the question of “how do we get more GPUs?” but the better question may be: how do we make the most of ones we already have.
The fact that a frontier lab like xAI could be running at sub-10% MFU (Model FLOPs Utilization) is just a hint at what the real problem may be.
For context, older frontier-scale training runs were already much higher than 10%. GPT-3 was around 21% MFU. Gopher was around 32%. Megatron-Turing NLG was around 30%. PaLM reached around 46%. And our guest Anjney says best-in-class MFU today is closer to 60–70%.
It’s not necessarily that xAI is uniquely incompetent (it’s clear they have talented folks) but rather the priorities may be flipped in the GPU arms race.
While GPU access is a bottleneck, simply increasing CapEx won’t automatically translate to better models as frontier AI is increasingly a systems problem: scheduling, utilization, networking, kernels, frameworks, data pipelines, parallelism, cluster reliability, and the thousand small decisions that determine whether your theoretical FLOPs become real training progress.
From building Discord’s developer platform and backing frontier AI companies like Anthropic, Mistral, Black Forest Labs, and Periodic Labs to now building AMP’s independent compute grid, Anjney Midha has spent years close to the real bottlenecks of AI scaling. In this episode, Anjney joins swyx at Periodic Labs to unpack why the AI race is not just about buying more GPUs, why 95% utilization would have been considered an outage at Google, and why the next era of AI infrastructure has to be more aligned, more efficient, and more responsible.
We go deep on AMP’s vision for a compute grid that makes FLOPs flow like megawatts, the difference between full-stack AI labs and horizontal pooling, why AI data centers need community buy-in, and how compute markets could evolve into something closer to an independent system operator. Anjney also explains why DeepMind’s unpublished research points to a market failure, why end-of-life prediction remains one of the most important AI applications he has thought about for fourteen years, and why “output maxing” may become a new discipline for frontier systems.
We also discuss Anthropic’s culture, why “luck favors the prepar
More from this podcast
Latent Space: The AI Engineer Podcast →