RelationVGGT : Visual Geometry Transformers for 3D Spatial Relation Segmentation
Abstract
Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit—yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We introduce \emph{3D spatial relation segmentation}, a task that requires models to identify a target object satisfying a given spatial relation with respect to a subject, consistently across multiple views. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction—requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.