Hear, Localize, and Reason: Spatially Aware Scene Understanding for Audio-visual LLMs
Abstract
Audio-visual large language models (AV-LLMs) have made strong progress in multimodal understanding, but they typically treat audio as monaural semantic content and therefore struggle to reason about where sounds originate and how they relate to the visual scene. Recent spatial audio-visual studies address parts of this problem, but often focus on specific abilities, such as spatial correspondence or direction and distance reasoning. In this paper, we propose HLR-AVSceneQA, a comprehensive benchmark for spatial audio-visual scene understanding. HLR-AVSceneQA evaluates whether models can jointly hear, localize, and reason by recognizing what is heard, grounding where it comes from in egocentric and allocentric views, and inferring how sound sources relate to visible, hidden, or nearby objects. We further introduce HLR-LLM, which extends a strong audio-visual foundation model with a binaural spatial audio branch. HLR-LLM is trained with a three-stage curriculum, applying chain-of-thought supervision in the final stage to help the model learn structured multi-step spatial reasoning. Experiments show that existing AV-LLMs remain limited in spatial grounding and relational reasoning, while HLR-LLM substantially improves spatial audio-visual scene understanding without sacrificing semantic audio-visual perception.