Audible World Models: Spatially Aware Sound Generation for 3D Worlds
Abstract
Text- and image-conditioned world generators can now produce visually rich 3D environments, but these worlds are usually silent or paired with a soundtrack generated only from text or rendered video. Such audio can describe what should be heard, but it does not explicitly represent where sounds live in the world or how they should change as a listener moves. We propose Audible World Models, a training-free framework that treats sound as part of a generated world state. Given a text prompt, our system builds a panoramic 3D proxy, decomposes it into semantic layers, identifies audible foreground objects and ambient background regions, synthesizes dry audio for each sound label, attaches sources to reconstructed geometry, and renders listener-dependent spatial audio through geometric acoustic propagation. This explicit coupling of semantics, geometry, and propagation produces audio that remains tied to persistent source locations and responds to viewpoint and motion. Across 80 generated scenes, our method substantially improves spatial consistency over text-, video-, and panorama-conditioned baselines while maintaining competitive semantic alignment. VLM-based and human evaluations further show that the resulting soundtracks are preferred for audio--visual consistency, spatial plausibility, and motion-dependent behavior.