Hierarchical Graph Alignment for Cross-Modal 3D Scene Grounding
Abstract
Text-guided point-cloud grounding enables robots to ground natural-language descriptions in 3D environments, making it important for embodied AI and human-robot interaction. Existing coarse-to-fine methods primarily rely on global descriptors for submap retrieval, but these descriptors often compress object semantics, spatial relations, and scene layouts into a single vector, thereby limiting discriminability in large-scale scenes. We propose HiGraLoc, a multi-level cross-modal alignment framework that improves the coarse retrieval stage through three complementary branches. The instance branch uses a Hyperbolic Instance Structure Encoder to model hierarchical object semantics, the relation branch aggregates reliability-weighted pairwise spatial relations, and the global branch employs a Global Spectral Graph Encoder to capture multi-frequency scene structure. A Text2Loc-style coordinate regression stage then refines the location estimate within the retrieved submap. Experiments on KITTI360Pose show that HiGraLoc achieves a 19\% improvement in Top-1 recall@10m over existing state-of-the-art methods. Our code and dataset are available at https://github.com/Anonymous09871745/HiGraLoc.