LaSA-Net: A Language-Guided Network for Outdoor Generalized 3D Referring Expression Segmentation
Abstract
3D Referring Expression Segmentation (3D-RES) aims to segment target objects in point clouds via language. However, extending 3D-RES from indoor scenes to large-scale outdoor environments introduces three major challenges, i.e., large-scale geospatial complexity lacking in indoor benchmarks, sparse-target perception overwhelmed by massive geometric clutter and distractors, and context-induced query dilution, where highly relation-intensive expressions cause globally-initialized queries to be absorbed by salient reference objects. To address these issues, we introduce UrbanRefer, the first outdoor 3D-GRES benchmark with 195 urban road point cloud scenes and 7,840 textual descriptions for complex geospatial reasoning. We further propose LaSA-Net, a Language-guided Semantic Segmentation Network. First, to enhance sparse-target perception, LaSA-Net incorporates pixel-level dense features from DINOv3 as 2D priors and fuses them with 3D features through depth-consistent verification. Then, to alleviate query dilution, we design a semantic-adaptive query generation module that applies text-weighted farthest point sampling to focus on text-relevant regions, and performs multi-granularity sampling to maintain both local grounding and global semantic consistency. Finally, a reliability-aware decoder adaptively suppresses noisy language cues through cross-modal reliability-gated injection, generating precise prompts for accurate mask prediction. Extensive experiments demonstrate that LaSA-Net achieves 52.7\%, 55.9\%, and 47.9\% mIoU on the ScanRefer, Multi3DRefer, and UrbanRefer datasets, surpassing prior state-of-the-art methods by 2.3\%, 4.2\%, and 2.8\% respectively.