Key ideas
Voice agents map naturally to perception, planning, and control
The system is organized using a conceptual model borrowed from autonomous driving. Transcription…Show moreShow less
The system is organized using a conceptual model borrowed from autonomous driving. Transcription serves as perception by converting real-world audio into structured text. The language model performs planning by interpreting that text and deciding what should happen next. Text-to-speech acts as the control layer by turning the planned response back into an externally perceptible signal.
Why it matters: This separation creates clear optimization boundaries and allows latency improvements in one layer without replacing the entire system.
Supporting evidence
For voice, it's going to be the transcription. Basically, effectively turning these like signals from the real world into elements of data that the language model or whatever brain you're working on can process.
Cascading preserves intelligence while targeting latency
The presentation rejects the assumption that responsiveness requires relying exclusively on smal…Show moreShow less