THREAD 04
Voice Systems
Latency budgets, turn-taking, interruption handling, and streaming inference for natural voice interaction.
Open question
How much latency can each stage consume before a conversation stops feeling fluid?
How we investigate
End-to-end latency decomposition, interaction studies, and streaming-system analysis.
2 min readThe 200ms Turn-Taking Problem in Streaming Voice AI
A practical latency budget for voice agents, and why endpointing—not only model speed—determines whether a conversation feels responsive.
Voice SystemsRead
5 min readWhy Voice Agents Feel Broken at 800ms
Total latency is the wrong number to optimise. What decides whether a voice agent feels alive is time to first audio, and where those milliseconds go.
Voice SystemsRead