Training Data
Training Data

Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

June 30, 2026

AI Summary

5 min read

Why Hardware-Software Co-Design Is AI's Real 100x

Dylan Patel started SemiAnalysis on his 24th birthday, posting two blog posts that were "the best stuff you could find on the internet about semiconductors" at the time. Five years later, the company has 90 people—roughly half technologists and engineers from across the supply chain, half former hedge fund analysts—and has become a premier research firm covering the intersection of AI, semiconductors, and economics. Patel's path there was anything but linear: he grew up in a family motel and gas station, started moderating hardware forums at age 12, worked as a quant, got screwed out of a bonus, lost his grandmother to dementia, and spent months living out of his truck visiting national parks while reading textbooks about semiconductors. That obsessive, self-directed education became the foundation for a firm that now shapes how institutional investors and tech executives understand the AI hardware landscape.

The Co-Design Thesis: Why Layer-by-Layer Optimization Misses the Point

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (00:00) **Background & Origin Story** - Dylan Patel recounts his early life, first job, and how a broken Xbox led to a lifelong obsession with hardware.
  • 2 (06:42) **Founding SemiAnalysis** - Patel explains the personal and professional catalysts that led him to start the research firm.
  • 3 (11:05) **The Value of Arcane Conferences** - Patel describes how deep, niche technical conferences provide the most valuable, non-public information.
  • 4 (14:47) **InferenceX: The Living Benchmark** - Patel explains the motivation and mechanics behind the industry-standard inference benchmarking project.
  • 5 (19:20) **The Inference Cost Curve vs. Space** - Patel discusses the economics of inference, including the surprising potential of space-based data centers.
  • 6 (23:34) **The Real 100x: Hardware-Software Co-Design** - Patel argues that the biggest gains come from co-optimizing hardware, software, and model architecture, not any single layer.
  • 7 (29:17) **Near-Term Bottlenecks: Memory & Power** - Patel identifies the most acute technological bottlenecks he is tracking.

+ Full timestamped outline available in the app

Guests on this episode

Show Notes

Dylan Patel, founder of SemiAnalysis, argues the biggest gains in AI don't come from faster chips, they come from software-hardware co-design. Optimizing the model, the kernels, and the silicon together turns a 2x here and a 2x there into 100x. He explains why DeepSeek's experts were shaped for Nvidia's Hopper (and why TPUs struggle to run it), why OpenAI's sparser models and Anthropic's denser ones pull them toward different hardware, and why the so-called CUDA moat was never really about CUDA. Dylan breaks down InferenceX, his living benchmark that runs the latest models on over $50M of donated hardware daily, tracking a roughly 60x annual drop in cost per unit of quality. He makes the case that inference will be a bigger market than oil, that the compute crunch persists because models expand the value of useful work faster than compute grows, and why Jensen Huang is bankrolling neoclouds to engineer a multipolar world.

Hosted by Shaun Maguire and Sonya Huang, Sequoia Capital


Training Data

More from this podcast

Training Data →