From Voice Agents to AI Avatars with Alexander Smola
September 17, 2026
AI Summary
5 min readFrom Voice Agents to AI Avatars
Voice AI has gotten remarkably good, but it still lives in an uncanny valley. A little too much latency, an interruption handled awkwardly, the wrong emotional tone — and the illusion shatters. And the bar only rises as systems move beyond text and voice toward multimodal avatars that can see and be seen. Alexander Smola, co-founder and CEO of Bosun AI, argues that voice is an intermediate stepping stone. "Eventually we will have avatars," he says. "You'll be talking to an AI agent that looks and feels like a human — and we're still very, very far away from making this really natural."
The Systems Challenge Behind Natural Conversation
Smola frames the problem as a convergence of biology, engineering, and science. Human audio-visual perception operates at roughly six to ten hertz — it takes about 150 milliseconds for a visual stimulus to reach the cortex, slightly less for sound. Bosun AI designed its models to be interruptible within a similar timeframe. But the deeper challenge is economic: audio requires far more tokens per second than text. Text runs at three to five tokens per second; audio can easily require ten or more, with each token covering about 100 milliseconds of temporal granularity. More tokens per second means the model cannot have too many parameters if it needs to run affordably in real time. "You can always make your mo
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:53) **The Uncanny Valley Problem in Voice AI** - Host Sam Charrington introduces the episode’s core challenge: voice AI still feels unnatural due to latency, interruption handling, and emotional tone, and the bar rises as systems add vision.
- 2 (01:41) **Voice is an Intermediate Stepping Stone to Avatars** - Guest Alexander Smola states that the ultimate goal is human-like avatars, but we are still very far from making that natural.
- 3 (02:29) **The Brittleness of Current Voice AI Systems** - Sam describes the limitations of advanced voice mode, noting it only works in perfect conditions and is brittle with interruptions.
- 4 (04:42) **The Hardware and Latency Challenge** - Alex explains that a single microphone is insufficient; proper noise cancellation requires microphone arrays, which is a solved hardware problem.
- 5 (07:19) **The Biology-Engineering-Science Trade-off** - Alex breaks down the core challenge: human perception operates at about 6-10 Hz, requiring the model to be interruptible within ~150 milliseconds.
- 6 (11:29) **The Challenge of Continuous Video for Avatars** - Alex describes the difficulty of generating a visually consistent, hour-long video feed, contrasting it with short, ten-second video segments.
- 7 (13:43) **Affordability as a Design Constraint** - The conversation shifts to the economic viability of voice AI, where using a full server GPU for a single conversation is not sustainable.
+ Full timestamped outline available in the app
Show Notes
Voice AI has gotten remarkably good, but natural conversation remains a high bar. Small delays, awkward interruptions, or the wrong tone can quickly break the illusion—and adding vision and visual presence only raises the stakes. In this episode, Alex Smola, co-founder and CEO of Boson AI, explores the path from today’s voice agents to audiovisual agents and AI avatars. We discuss the technical tradeoffs behind real-time voice, including audio tokenization, latency, model size, and inference cost, as well as what changes when these systems can both see and be seen. We also explore the role of emotional intelligence in AI, how agents can learn from human interactions, and what it will take to move beyond impressive demos toward interactions that actually feel natural. 🗒️ Full show notes: https://twimlai.com/go/777.
More from this podcast
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) →