Despa: Resolving Spatial Collapse in VLMs via Depth-Grounded Geometry
Abstract
Vision-Language Models (VLMs) excel at open-world 2D perception but struggle with precise metric spatial reasoning, a key requirement for applications like autonomous driving, navigation, and embodied intelligence. We identify this as a structural limitation: images are 2D projections of the 3D world, inherently suffering from projective ambiguity, while VLM components favor semantic understanding and rely on 2D positional bias, leading to Spatial Collapse. To address this, we propose Despa, a Depth-grounded spatial VLMs that injects depth-based geometric information via Geometric Positional Embedding (GPE) and Depth-aware Rotational Positional Embedding (DoPE). To expose masked limitations in current benchmarks, we construct SpaDataset (1.7M QA pairs) and SpaBench (2,051 QAs), to our knowledge the largest fine-grained training corpus and benchmark suite for metric spatial reasoning. They cover three difficulty levels and ten sub-tasks across diverse settings. The fine-tuned Despa-4B model consistently outperforms general-purpose, closed-source, and specialized VLMs on SpaBench by 27.9% over Qwen3-VL-235B-A22B and 31.1% over GPT-5.2, generalizes to MSMU (66.3%), RefSpatial-U (42.4%), and the multi-frame VSI-Bench (66.0%), and consistently lifts LLaVA-1.5, Qwen2.5-VL, and Qwen3-VL backbones, achieving the performance with only a 4B model. The code will be publicly released.