DIGS: Distribution-Informed Gaussian Splatting for Training-Free Open-Vocabulary 3D Segmentation
Abstract
Open-vocabulary 3D Gaussian segmentation typically binds semantic features to per-Gaussian state via per-scene gradient training (1–4 h) or the more recent closed-form distillation. Both commit each primitive to a single feature; therefore, boundary Gaussians spanning multiple 2D instances are forced into a hard decision that degrades retrieval accuracy and yields unstable mask boundaries at render time. We present DIGS, a training-free framework with two stages, class-agnostic map- ping and multi-modal encoding, that decouples instance discovery from linguistic naming. At the Gaussian level, our distribution-informed representation maintains an explicit top-M label distribution per primitive, updated by a tempered Dirichlet rule; preservation of multi-peaked posteriors at boundary primitives allows a late-argmax renderer to recover clean instance boundaries. A bidirectional state machine aggregates these distributions into a persistent class-agnostic instance ontology. The instance-level encoding stage then fuses the CLIP visual embedding of each instance’s isolated 3D footprint with the CLIP-text embedding of an LLM-generated attribute description, disambiguating visually similar instances. DIGS maps a scene in 1–3 minutes on a single A100 and serves 3D-native retrieval in 5.4 ms, reaching 71.95% mIoU on LERF-OVS, surpassing the strongest training-free baseline SFS and the strongest training-based baseline LaGa, and setting new state-of-the-art performance on LERF-Mask (91.71%) and 3D-OVS(96.45%). Codes are at: https://anonymous.4open.science/r/DIGS-C4EF.