ECHO: Diagnosing Spatial Cue Access in Audio-Language Models
Abstract
Stereo localization requires an audio-language model—a system that accepts audio and answers in language—to recover spatial evidence from left and right signals and use it to infer direction. Traditional end-to-end accuracy reports only whether the final direction is correct, hiding whether a model failed to obtain the spatial cue or failed to use it. We introduce ECHO (Explicit Cue versus Heard Observation), a paired diagnostic framework, and StereoMusicQA, a benchmark built from real instrument recordings rendered at known positions. Together, they test whether a model can use a spatial cue that it fails to obtain through audio. Each of 1,045 test items has two views: cue-text provides the measured left-right level difference and conversion equation; audio-only provides labeled channels and requires the model to obtain a useful relationship from audio. Because cue-text also supplies the rule, this comparison measures the practical advantage of a written cue and rule, not the effect of one internal layer. For Qwen2-Audio, Gemini Flash-Lite, and GPT-audio, cue-text left/center/right accuracy is 71.9%, 93.6%, and 90.4% compared with 28.5%, 36.6%, and 42.0% audio-only. These 43.3-57.0 percentage-point gains remain positive when complete source tracks are resampled. Furthermore, a model can identify the correct side yet still predict an inaccurate angle: cue-text within 5 degree accuracy ranges from 0.4% to 80.7%, and written-calculation faithfulness ranges from 0.4% to 77.8%. ECHO therefore shows why spatial evaluation should separately test evidence access, cue use, and whether a stated calculation follows from the cue.