Online Open-Vocabulary 3D Segmentation by Separating Recognition from Association
Abstract
Useful 3D perception should update while a camera explores, so robots can act and users can inspect a scan before capture ends. Existing open-vocabulary methods often defer 3D decisions or repeatedly encode object proposals. We introduce, a training-free method that processes each RGB frame once with frozen SAM 3.1 and projects its masks and pixel-wise category scores through measured RGB-D geometry. Its key design separates two decisions that require different evidence: scores recognize what is present, whereas projected-mask overlap preserves which object it is. Reliability-weighted score fusion labels the observed surface and resolves competing instances without a trained 3D network, extra encoder, or tracking memory. On all 312 ScanNet validation scenes, the final result, including optional post-capture reconnection, reaches 0.594 semantic mIoU and 0.544 class-aware instance AP at IoU 0.50, outperforming online and full-sequence baselines under one protocol. Matched-region analysis supports the separation: category scores give the best tested recognition, while 3D mask overlap gives the best same-class association.