Latent Spatial Reasoning: Building Innate 3D Awareness via Latent-Space Distillation
Abstract
Humans effortlessly infer 3D structure, such as depth, occlusion, and spatial arrangement, from 2D images and reason about it fluidly. Multimodal large language models still struggle with such spatial reasoning. Current approaches attempt to bridge this gap by injecting explicit 3D priors at inference time or aligning features with 3D-aware teachers during training. While these strategies improve geometric perception, they typically treat 3D knowledge as static representations, rather than enabling the model to reason with 3D cues during inference. We argue that closing this gap requires internalizing 3D knowledge in a form that can participate in the model's reasoning dynamics. To this end, we introduce \textbf{Latent Spatial Reasoning (LSR)}, a framework that distills geometric knowledge into \emph{latent spatial tokens}: geometry‑aware representations that the model actively queries and conditions on to guide its spatial reasoning process. Through hierarchical distillation, LSR transfers 3D knowledge from foundation models, aligns it with the vision-language space, and trains the model to interleave spatial and linguistic reasoning, all without external 3D modules at inference time. Experiments on diverse spatial reasoning benchmarks demonstrate state-of-the-art performance.