From Voice Activity to Active Listening: A Semantic Turn-Taking Controller for Full-Duplex Voice Agents
Abstract
Cascaded voice agents often use voice-activity detection and a silence timer to decide whentoreply. These rules can confuse a thinking pause with a completed turn or fail to distinguish an interruption from noise. We introduce a small semantic turn-taking controller, separate from the language model that generates responses. At approximately 100 ms intervals, it reads dialogue text and a pause label and predicts one of five states. It supports full-duplex interaction and active listening: the agent can give a short acknowledgment, or backchannel, while the user continues speaking. A text-based iteration loop converts live failures into matched training examples without collecting and labeling new audio. A 270M controller reaches 0.849 macro-F1 on the internal evaluation, compared with 0.851 for a same-family Gemma-3-4B controller and 0.847 for Qwen3-4B with trained embeddings. Its small size supports on-device serving. External benchmarks, live traces, and pipeline latency measurements assess system behavior.