RelGS: Relation-Aware Gaussian Splatting for Open-Vocabulary 3D Scene Understanding
Abstract
Open-vocabulary 3D scene understanding provides an important interface for querying and interacting with reconstructed scenes through natural language. Although recent 3DGS methods have enabled text-driven object selection and open-vocabulary segmentation, they still struggle with compositional queries involving attributes, spatial relations, and part-whole relations. Most existing approaches learn per-Gaussian or per-cluster language features, construct query-conditioned referring fields, or perform spatial reasoning only at inference time, but they do not organize the scene itself into a reusable relation-aware representation. As a result, object relations are not explicitly stored as typed and confidence-aware 3D structures, limiting compositional reasoning and multi-hop querying. To address this issue, we propose RelGS, a relation-aware 3D Gaussian Splatting framework for open-vocabulary scene understanding. RelGS learns semantic and instance embeddings for Gaussians under multi-view supervision, groups them into cross-view consistent 3D semantic nodes, and enriches each node with structured attributes verified by reverse CLIP consistency. It further constructs a confidence-weighted relation graph, where 3D spatial cues propose candidate relations and an LLM verifies their semantic plausibility. At query time, RelGS combines attribute matching, CLIP retrieval, and graph traversal to localize the target object. Extensive experiments demonstrate strong performance on word-level selection, sentence-level referring segmentation, multi-hop relational queries, and point-level semantic segmentation.