Online Open-Vocabulary 3D Segmentation by Separating Recognition from Association
Abstract
Open-vocabulary 3D segmentation should label the surface seen so far and keep the same identity for each object as an RGB-D camera moves. Yet most methods wait for the complete sequence or add a visual encoder for object crops. We introduce, a training-free method that avoids both by separating recognition from association. Frozen SAM 3.1 processes each RGB frame once. Its pixel-wise scores recognize a region is, while overlap between its projected mask and the current 3D map links it to which object it belongs to. Reliability-weighted fusion stabilizes category evidence, and point assignment combines both cues only where objects compete for points. The result is a semantic-instance map after every frame, without a learned 3D network, extra encoder, or video memory. On all 312 ScanNet validation scenes, the final map with optional post-capture reconnection reaches 0.594 semantic mIoU and 0.544 joint AP at IoU 0.50. It outperforms online and complete-sequence baselines under one protocol. Matched-region analysis confirms the design: scores best recognize categories, while overlap best associates same-class objects among the tested cues. Together, these results show that image-model predictions and geometry can maintain a current 3D map without a learned association representation.