AI Summary
5 min readIn 2022, Mati Staniszewski co-founded ElevenLabs, and by early 2025 the company had grown to a $450 million ARR business valued at $11 billion, building foundational audio models that capture the humanness of speech. But as Staniszewski explains, the path from mechanical speech machines to today’s emotionally inflected voice agents involved a series of breakthroughs—and the field still has a long way to go before it passes a true conversational Turing test.
How voice models work: from phonemes to emergent accents
To understand ElevenLabs’ approach, it helps to know how speech generation evolved. Early attempts tried to physically replicate the human vocal tract with analog machines. Later, Bell Labs created structured digital signals to represent speech, and researchers began stitching together phonemes—the smallest units of human speech sound—based on probabilistic predictions of the next word.
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of Cheeky Pint
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 Timestamped Outline
- 2 (00:01) **Introduction to ElevenLabs and Mati Staniszewski** - Brief background on the company's founding in 2022 and its position as a leader in AI audio
- 3 (00:21) **How Audio Models Work: From Analog to Neural** - Explains the evolution of speech synthesis from mechanical to modern neural approaches
- 4 (02:17) **The Spectrogram Space Explained** - Defines the mel spectrogram as a visual representation of speech across pitch and energy
- 5 (03:33) **Voice Models: Context and Voice Characteristics** - Explains the two-part nature of voice models: sound of intonation and voice characteristics
- 6 (04:48) **What Is a Token in Voice Models?** - Discusses the fundamental representation units in voice AI
- 7 (06:12) **How ElevenLabs Achieved Human-Sounding Voices** - Describes the combination of architecture, compute, and data innovations
+ Full timestamped outline available in the app
Show Notes
Description
Mati Staniszewski is the co-founder of ElevenLabs, the research company making audio accessible across languages and voices. He sits down with John to discuss the "voice Turing Test" and why AI has conquered text but still struggles with conversational speech. They discuss the future of human-computer interaction, including why we still can't get our phones to read a PDF properly and the massive potential for voice agents in everything from farming to healthcare. Mati also opens up about ElevenLabs’ rapid ascent to an $11 billion valuation and gives a behind-the-scenes look at how Ukraine is using their tech for digital government services.
Timestamps
(00:00:27) How audio models work
(00:08:52) ElevenLabs business model
(00:17:50) The conversational Turing Test
(00:21:01) Link by Stripe
(00:26:02) Cascaded vs speech-to-speech
(00:31:53) Universal translation
(00:51:41) Designing an AI-native org
More from this podcast
Cheeky Pint →