Invest Like the Best with Patrick O'Shaughnessy
Invest Like the Best with Patrick O'Shaughnessy

Neil Movva - Making AI 10x Cheaper

August 25, 2026

AI Summary

5 min read

Neil Movva, founder of Sale Research, is building what he calls a "token factory"—an inference company designed for a future where AI agents run in the background for hours or days, not for real-time chat. His central argument is that the entire AI industry is misaligned: it optimizes for low latency (fast responses for human-in-the-loop chatbots), but the real growth will come from long-running, asynchronous agents where latency matters least and cost matters most. Movva’s entire strategy—from software kernels to chip procurement to power sourcing—is built around driving the cost per token as low as physically possible, even if that means accepting 95% uptime and chips no one else wants.

The Latency-Throughput Trade-off and the Shift to Background Agents

The foundational insight of Movva’s approach is the unbreakable trade-off between latency and throughput inside every GPU. A GPU is a throughput machine: it is most efficient when processing large batches of work simultaneously. But serving a chatbot requires spinning out tokens as fast as possible, which forces the GPU to operate in a low-latency, low-batch regime—the equivalent of taking a private car instead of a bus. “The best latency is no latency at all,” Movva says. “When you wake up in the morning, the work’s already been done overnight.”

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (03:19) **What Sale Research Builds** - Neil introduces his company as a "token factory" serving open-source LLMs via API at unbeatable prices, plus long-running agent VMs (Sale Boxes) designed for hours or days of background work.
  • 2 (05:31) **Why the Market Left an Opening for Background Inference** - Neil explains that existing inference providers optimized for low-latency chatbots (driven by Cursor), but the future is long-horizon agents where latency matters less and throughput matters more.
  • 3 (06:43) **Evidence That Background Agents Are the Future** - Neil points to the trend of average task length getting longer, with models now capable of running for an hour, and predicts the market will shift to 90% background workloads.
  • 4 (11:57) **The Long-Term Vision: Verifiable Problems Become Cheap** - Neil argues that any verifiable problem (most software, math proofs, scientific discovery) can be tackled with abundant tokens, limited only by the questions humans can ask.
  • 5 (15:31) **The Software Layer: Squeezing Peak GPU Efficiency** - Neil describes building the entire LLM stack around GPU throughput, starting with low-level kernel programming, a skill he honed at Nvidia.
  • 6 (20:57) **Why Throughput vs. Latency Is an Unbreakable Trade-Off** - Neil explains that batching many users' work together on a GPU increases total work but slows any individual token, analogous to a bus vs. a private car.
  • 7 (25:07) **Cerebras and the SRAM vs. DRAM Trade-Off** - Neil contrasts SRAM (fast, on-chip, low capacity) with DRAM (slower, off-chip, high capacity), and explains how Cerebras builds wafers with massive SRAM for ultra-low latency inference.

+ Full timestamped outline available in the app

Show Notes

My guest today is Neil Movva, founder of Sail. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background for hours or days at a time rather than answering a human in real time. 

In that world, latency matters less and cost matters more, and Neil has built the whole company around driving the cost of a token as low as it can possibly go.

What makes this conversation special is that it is one of the most detailed tours I have ever done through the full stack of intelligence, the software, the chips, and the power, and how all three connect. 

Along the way we cover the trade-off between speed and cost that lives inside every GPU, his scavenger strategy for buying the chips and power nobody else wants, his contrarian view on Nvidia, and why the premium the frontier labs charge for being three to six months ahead may not last. 

Please enjoy my conversation with Neil Movva.

For the full show notes, transcript, and links to mentioned content, check out the episode page ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠here⁠⁠⁠⁠⁠

-----

Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at ⁠colossus.com/subscribe⁠.

-----

⁠Ramp’s⁠ mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ ⁠ramp.com/invest⁠⁠ to sign up for free and get a $250 welcome bonus.

-----

Trusted by thousands of businesses, ⁠Vanta⁠ continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to ⁠vanta.com/invest⁠

-----

WorkOS⁠ is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthro

Invest Like the Best with Patrick O'Shaughnessy