CrossWeave: Emergent Cross-Modal Scene and Instance Retrieval from Sparse 2D-3D Alignment
Aadith Warrier ⋅ Gnana Prakash Punnavajhala ⋅ Siddharth Tourani ⋅ Muhammad Haris Khan ⋅ Avinash Sharma ⋅ Madhava Krishna
Abstract
Cross-modal scene and instance retrieval aims to retrieve a corresponding 3D environment or object given query 2D images, serving as a foundational mechanism for spatial reasoning in robotics and AR/VR. Existing methods achieve this by compressing the entire 3D map into a single global feature vector, which discards the fine-grained local geometrical structure in indoor scenes. In this work, we show that scene- and instance-level retrieval capabilities emerge from supervision on sparse, local 2D-3D correspondences, without any explicit global labels. Specifically, we propose $\textbf{CrossWeave}$, a novel framework for cross-modal retrieval that learns sparse patch-point alignment, which alone outperforms global encoding methods. Additionally, an optimal transport-based aggregator is applied over the aligned features to attain further gain by clustering structurally redundant regions, freeing descriptor capacity for semantically distinct local features. CrossWeave achieves state-of-the-art performance on both scene- and instance-level retrieval on ScanNet and 3RScan with a $\textit{single model}$, whereas existing methods require separate models for each task. Our results suggest that whole-scene encoding is not necessary for cross-modal retrieval, and that sparse local correspondence supervision is a more effective and flexible alternative.
Chat is not available.
Successful Page Load