The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models
Abstract
Large audio-language models have made rapid progress in recognizing what is present in an audio clip, yet spatial audio-language understanding still lacks a clear task interface. A model must not only identify sound events, but also decide where they occur, which semantic and spatial attributes belong to the same auditory object, how multiple objects are arranged, and whether a scene-level answer is physically plausible. We formalize this missing capability as audio scene analysis (ASA), a three-level problem spanning atomic perception, relational integration, and cognitive reasoning. We propose The World is Not Mono (TWNM), a framework that instantiates this definition by equipping audio-language models with explicit spatial evidence. TWNM uses physically grounded First-Order Ambisonics (FOA) simulation to obtain controllable supervision, learns slot-regularized spatial representations from multichannel audio, and fuses these representations with semantic audio features before reasoning with a language model. We further train the model with a progressive curriculum, ending with preference optimization over metadata-derived correct answers and auxiliary format/evidence rewards. To operationalize the ASA definition, we build a controlled benchmark from scene metadata, covering localization, attribute binding, spatial comparison, scene abduction, and counterfactual reasoning. On this ASA benchmark, TWNM achieves 70.8% overall accuracy, 66.4% on spatial-family tasks, and 79.76% on mixed L3 scene-level question answering (QA) under exact multiple-choice question answering (MCQA) scoring. We also audit monaural and binaural reference systems as diagnostic references with explicit audit labels, because they differ in spatial input, training interface, and output format. The supported claim is that a clearly defined ASA task hierarchy, FOA-conditioned spatial representations, and metadata-grounded training together enable controlled, auditable spatial audio-language reasoning, with STARSS23 providing a limited real-recording diagnostic.