OmniWhere: Agentic Audio-Visual Perception for Embodied Spatial Reasoning
Abstract
Embodied spatial reasoning requires locating objects that are not currently in view. A static object can be located by reconstructing the scene from the frames observed earlier, whereas a moving object may have changed position while out of view. Sound can fill this gap: hearing is not limited to the facing direction. Building on signal-processing estimates of direction and distance from a microphone array, SAVVY located a moving target out of view on real egocentric video with a training-free pipeline. In SAVVY, however, the acoustic and visual estimates are combined by a fixed rule. The rule does not consider the reliability of an acoustic estimate, and it measures only at the queried moment. In this paper, we introduce OmniWhere, a code-writing spatial agent that uses acoustic estimators of direction and distance as a tool, without any training. Each estimate is returned with the statistics of the measurement behind it, and the agent chooses the time window of each measurement. The agent thereby decides how much to trust an estimate and when to take it. On SAVVY-Bench, OmniWhere outperforms SAVVY by 7.4 points overall, with the largest gain on direction questions behind the camera. An ablation with the same acoustic estimates indicates that the gain behind the camera depends on exposing their reliability to the agent, not on the estimates alone.