Overview
Voice AI often fails not because of one catastrophic model error, but because small breakdowns accumulate across recognition, pronunciation, timing, vocabulary, and context. Midam Kim proposes linguistics as a unifying framework for diagnosing these failures. Human conversation is a joint activity: participants continually listen, speak, interpret intentions, manage turns, and update shared mental models. Voice agents must support the same interconnected process across two channels—listening and speaking—and four layers: sounds, words, interaction, and mental models. These layers cannot be optimized independently. Accurate transcription is insufficient if the agent interrupts the user, mispronounces a name, repeats failed prompts without clarification, or forgets what has already happened. Voice also unfolds on a disappearing timeline: unlike chat history, spoken sounds and words vanish, while their effects on the user's understanding and frustration persist. Effective systems therefore require dynamic orchestration of ASR, TTS, vocabulary, turn detection, latency, emotion handling, and context retention throughout each call and across different users. The practical conclusion is that there is no universal component-level fix. Teams need an end-to-end linguistic framework, suitable benchmarks, and continuous adaptation to users and changing language.
Sections
Higher-Order Insights
Implications derived from the framework's treatment of conversation as an interconnected, time-dependent system.
- Conversational quality is path-dependent: an early error changes how the user interprets and responds to every later turn, so identical model behavior can produce different outcomes depending on the preceding interaction.
- A user's request for a human agent is often a delayed indicator of accumulated coordination failures rather than evidence that the final prompt alone was defective.
- The mental-model layer acts as the integration point between technical performance and user satisfaction because it retains the consequences of otherwise transient speech events.
- Scalability in voice AI is not achieved solely through serving more calls; it also requires orchestration that adapts across users, emotional states, speaking styles, and linguistic change.
Risks and Failure Modes
Key pitfalls that can degrade task completion and user trust.
- Optimizing ASR, TTS, or another component in isolation may leave the overall interaction broken because linguistic layers are interdependent.
- Premature turn detection can cause the agent to interrupt users who pause while locating or reading unfamiliar information.
- Repeating a failed prompt without targeted clarification can trap users in an error loop and trigger escalation or abandonment.
- English-centric pronunciation and vocabulary assumptions can misrepresent names and reduce trust for users with different linguistic backgrounds.
- Static conversational behavior may fail as the user's emotional state, speaking style, or expectations change during the call.
Contrasting Design Models
The main distinctions used to explain why conventional voice systems fail.
- A pipeline treats recognition, synthesis, and dialogue as successive technical stages; joint activity treats the bot and user as participants continually coordinating meaning and action.
- Chat preserves a visible textual history, while voice signals disappear after they are spoken and leave only their effects on the user's evolving mental model.
- Simple repetition asks the user to retry unchanged, whereas interactive clarification identifies uncertainty and collaboratively repairs the misunderstanding.
- Static orchestration applies fixed behavior throughout the call, while dynamic orchestration adapts to changing context, emotion, user characteristics, and conversational history.