AI Summary
5 min readIn the days before ChatGPT, running a large language model meant scrambling to download a torrent link and hoping it worked. Simon Moe, co-founder and CEO of Infra, remembers that weekend vividly—when a model lab dropped a new release, his team spent the entire time behind the scenes trying to get the inference engine to run on it. That frantic, enthusiast-driven era has since given way to something far more systematic: open-source AI has become critical infrastructure, and the software that makes it run—inference engines like vLLM—now powers half a million GPUs at any given moment.
What Makes Serving an LLM Fundamentally Different
Serving a large language model is not like running traditional machine learning workloads. It requires accelerators like GPUs or TPUs, and the process is computationally intensive. The core challenge lies in handling differences in input distribution (how long each request is) and output distribution (which is non-deterministic), along with batching and scheduling. This is the job of an inference engine like vLLM, which Moe describes as "kind of like databases and operating system and other critical software to power AGI." Its purpose is to turn available GPUs into a running endpoint for intelligence, ensuring cost effectiveness, efficiency, and reliability.
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of a16z Show
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:00) **Open-Source AI as Critical Infrastructure** - The episode opens with a thought experiment about GPU prices dropping 99% and the premise that open-source AI has become a foundational force requiring a new infrastructure layer.
- 2 (02:22) **Why Serving an LLM is Fundamentally Different** - Simon explains that serving a large language model requires running it on GPUs or TPUs, with computationally intensive processes that must handle non-deterministic output and variable input lengths.
- 3 (03:29) **From Open-Source Project to Critical Infrastructure** - Matt Bornstein traces how open-source was the norm for early AI models (like BERT) but shifted as models grew larger and required specialized hardware and software.
- 4 (08:21) **Where VLLM Sits in the Stack** - Simon describes VLLM as an inference engine that turns GPUs into running endpoints for AI, supporting over 1,000 model architectures with "day zero" model releases.
- 5 (09:49) **The Drama Behind Model Releases** - Simon shares stories of the chaotic co-design process when model labs release new open-weight models, involving multiple parties like hardware vendors, Hugging Face, and inference clouds.
- 6 (13:03) **Why Open Weights Matter: Cost vs. Control** - Infract signed the Nvidia open weights letter to stand against a world controlled by proprietary APIs, emphasizing that open development must not be blocked or banned.
- 7 (16:24) **The Real Economics of Frontier Open-Weight Models** - Simon explains that while some open-weight models (like Kimi K3) bridge a 10x cost gap with proprietary models, they are not always cheaper—but the value lies in ownership and customization.
+ Full timestamped outline available in the app
Show Notes
Elena Burger and Matt Bornstein are joined by Simon Mo, co-founder and CEO of Inferact, the open-source inference engine powering many of today's most advanced AI applications. Together, they explore how open-source AI evolved from a research project into critical infrastructure, why inference has become one of the most important layers of the AI stack, and what it takes to bring frontier intelligence to developers around the world.
The conversation covers vLLM's origins, the rise of open-weight models, why companies increasingly want control over their AI infrastructure, and how open-source inference enables the next generation of AI applications. They also discuss model licensing, the economics of open-weight AI, Kimi K3, distillation, AI infrastructure, and why Simon believes the gap between open and closed models is rapidly disappearing.
Resources:
Follow Simon Mo on X: https://x.com/simon_mo_
Follow Matt Bornstein on X: https://x.com/BornsteinMatt
Follow Elena Burger on X: https://x.com/VirtualElena
Follow Inferact: https://x.com/inferact
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
More from this podcast
a16z Show →