AI Summary
5 min readJonathan Ross, founder of Groq, spent a decade building a company whose core thesis—that fast inference would matter more than anyone believed—was dismissed by almost everyone, including his own team. The turning point came when a viral video of an LLM running on Groq’s LPU hardware forced the market to see what he had seen years earlier. This conversation covers how he built the company, the leadership lessons he learned the hard way, and why he believes the AI age will reward people who ask the right questions rather than those who can answer them.
The speed insight and the Nvidia deal
Ross’s fundamental insight was that inference speed—how fast an AI model generates tokens—would become the critical bottleneck, not compute power for training. He saw this years before the market agreed. At Google, he had worked on the TPU and watched AlphaGo’s Elo score jump dramatically when it ran on faster hardware, even though the model was identical. The same model, running faster, was smarter because it could search deeper into the game tree. Ross realized that speed and quality were linked: “Being able to think faster makes you think smarter.”
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of David Senra
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:10) **The $20B Nvidia Partnership** - How the deal came together in just three weeks, from a call to money in the bank
- 2 (01:53) **Why Speed Matters for AI-to-AI Communication** - AI thinks faster than humans, so inference speed becomes critical for agentic workflows
- 3 (03:30) **Hobby Projects as Innovation Engines** - Jonathan builds cutting-edge side projects before bringing ideas to work
- 4 (06:21) **The Shift from Answering to Asking Questions** - Success in the AI age is about asking the right questions, not having the answers
- 5 (08:23) **Leadership Philosophy: Followers First** - Leadership is defined by having followers, not by a specific style
- 6 (12:36) **The Cost of Learning to Lead** - Jonathan's transition from engineer to founder cost Groq 3-4 years
- 7 (16:39) **Lessons from Jensen Huang** - No one-on-ones, no politics, and always focus on what the customer needs
+ Full timestamped outline available in the app
Show Notes
Jonathan Ross is the founder of Groq and the inventor of the Google Tensor Processing Unit (TPU), now a senior executive at NVIDIA following the company's $20 billion partnership with Groq.
Before Groq, Ross built something that didn't exist: a custom AI chip at Google called the TPU, which became the backbone of DeepMind's AlphaGo — the system that defeated world Go champion Lee Sedol in 2016. After watching the TPU push AlphaGo's ELO score up by hundreds of points overnight, Ross grasped a principle that would define his next decade: faster inference produces more capable models. He left Google to act on it.
Groq's first decade was brutal. Early West Coast VCs passed — and would later watch as NVIDIA announced what Ross describes as the firm's largest deal by nearly 3x. Ross came within weeks of running out of money. Rather than lay off the engineers he needed to hit a critical product milestone, he created "Groq bonds" — war-bond–style instruments that exchanged salary for equity. About 80% of the team participated; nearly half took statutory minimum wage. They saved two months of runway and kept the company alive.
The core bet Ross made — that fast inference would matter — was widely dismissed, inside Groq and out. When the CEO of GitHub called needing chips to run LLMs, Ross's own engineers told him it couldn't be done. He eventually stopped asking and started declaring: "I intend to do this." He describes that shift — from inviting pessimism to announcing direction — as the most important leadership change he made.
Now at NVIDIA, Ross carries what he calls manufactured discontent: a deliberate refusal to rest, convinced that every day without sufficient compute is a day the world waits longer for cures for cancer and aging.
Show notes: https://www.davidsenra.com/episode/jonathan-ross
Made possible by
Ramp: https://ramp.com
AppLovin: https://applovin.com/senra
Deel: https://deel.com/senra
Chapters
(00:00:00) The $20 Billion NVIDIA Deal Closed In 3 Weeks
(00:00:25) Why GPUs And LPUs Are Better Together
(00:01:46) When AI Talks To AI, Speed Wins
(00:03:30) Always Start With A Hobby Project
(00:05:55) Ask The Right Questions, Not Answer Them
(00:08:23) There Are Infinite Ways To Be A Leader
(00:13:00) I Was One Of The World's Worst Leaders
(00:14:34) Fewer Constraints, More Room To Surprise You
(00:16:31) At NVIDIA There Is No Politics
(00:19:44) You Have To
More from this podcast
David Senra →