EP 58: Every Millisecond Matters: Diffusion LLMs and the Future of Voice AI | Aditya Grover, Inception
Recorded live at the Ai4 Podcast Pavilion, Sam wraps Day One with Aditya Grover, Co-Founder & CTO of Inception, on why the next generation of LLMs won't look anything like the ones we use today.
What's Covered:
"Every Millisecond Matters" — Why latency, not intelligence, is the real bottleneck holding back voice agents and multi-step AI agents alike.
How Mercury Actually Generates Text — Instead of predicting one token at a time like every autoregressive model, Mercury generates a rough draft of the full response and refines it into coherence — diffusion, applied to language instead of images.
Solving Voice AI's Impossible Tradeoff — Fast-but-lower-quality, or high-quality-but-too-slow: Aditya explains how Mercury 2 finally delivers both.
A Term Coined Live at This Conference — From Aditya's own Ai4 keynote: "We're moving from token maxing to value maxing."
Advice for the Next Generation — Ten-plus years into AI research, Aditya's honest take on why this is still the best time to pursue a PhD, join a startup, or do both.
The Next 5-10 Years of Voice AI — A prediction for a future where voice becomes humans' predominant mode of interacting with AI, the same way it is with each other.
Key Quote:
"Sequential generation is not a law of nature... AI can have a different way of generation, one that's more parallelizable."
Connect with Aditya:
LinkedIn: https://www.linkedin.com/in/aditya-grover/
Inception: https://www.inceptionlabs.ai/
Subscribe: Spotify | Apple Podcasts | Amazon Music | iHeart Radio | YouTube | Substack
#Ai4Conference #InceptionLabs #DiffusionLLM #VoiceAI #AsembleAI
More description
Recorded live at the Ai4 Podcast Pavilion, Sam wraps Day One with Aditya Grover, Co-Founder & CTO of Inception, on why the next generation of LLMs won't look anything like the ones we use today.
What's Covered:
"Every Millisecond Matters" — Why latency, not intelligence, is the real bottleneck holding back voice agents and multi-step AI agents alike.
How Mercury Actually Generates Text — Instead of predicting one token at a time like every autoregressive model, Mercury generates a rough draft of the full response and refines it into coherence — diffusion, applied to language instead of images.
Solving Voice AI's Impossible Tradeoff — Fast-but-lower-quality, or high-quality-but-too-slow: Aditya explains how Mercury 2 finally delivers both.
A Term Coined Live at This Conference — From Aditya's own Ai4 keynote: "We're moving from token maxing to value maxing."
Advice for the Next Generation — Ten-plus years into AI research, Aditya's honest take on why this is still the best time to pursue a PhD, join a startup, or do both.
The Next 5-10 Years of Voice AI — A prediction for a future where voice becomes humans' predominant mode of interacting with AI, the same way it is with each other.
Key Quote:
"Sequential generation is not a law of nature... AI can have a different way of generation, one that's more parallelizable."
Connect with Aditya:
LinkedIn: https://www.linkedin.com/in/aditya-grover/
Inception: https://www.inceptionlabs.ai/
Subscribe: Spotify | Apple Podcasts | Amazon Music | iHeart Radio | YouTube | Substack
#Ai4Conference #InceptionLabs #DiffusionLLM #VoiceAI #AsembleAI
2026-08-23
15 min
Listen elsewhere
Available Results
Generated results are saved to your library for reuse and search.
No generated results are available for this episode yet.
Transcript
No transcript is available for this episode yet.
No audio file is available for transcript generation.
Chapters
No chapters available.