When Does a Speech Model Know to Hand Off? Reading Competence from Hidden States for Real-Time Escalation
Abstract
A local speech model can keep a conversation going, and a stronger expert can take over the hard questions. Calling the expert on every turn adds cost and waiting time, so the local model needs an early estimate of whether its own answer will fail. Hidden-state probes give such estimates for text models. It is not known whether this signal survives speech and duplex fine-tuning, audio input, and the timing of a native listen/speak protocol. We study these questions on frozen MiniCPM-o~4.5. We use lightweight probes to read failure signals from hidden states and separately calibrate a gate for expert handoff at answer onset. We find that failure signals can transfer from text to speech, and that readout design matters in native audio generation. On the full 240-query internal benchmark, selective expert handoff raises unweighted answer-content accuracy from 40.8\% to 62.9\%. The full real-time timing evaluation on the same query IDs gives mean server waiting time to first answer audio of 6.10\,s with aggressive handoff, versus 9.43\,s with always-escalate. P50 is 0.70\,s locally and 2.87\,s with aggressive handoff.