Geometry-Guided Semantic Reconstruction for 3D Instance Segmentation
Abstract
3D instance segmentation requires representations that are both geometrically structured and semantically discriminative. Pure 3D encoders model spatial structure well but often lack strong appearance semantics, while directly lifting 2D priors into 3D by concatenation injects rich semantics without ensuring that they align well with the underlying 3D structure. In this paper, we address this problem by formulating 2D-to-3D semantic transfer as a \textbf{masked semantic reconstruction} task. Specifically, we mask a subset of background lifted 2D priors and train the 3D encoder to recover the missing semantic priors, providing explicit supervision for distilling rich 2D semantics into the 3D backbone. We further propose \textbf{geometry-modulated semantic injection}, which uses raw 3D geometry to modulate lifted 2D priors before sparse encoding, so that geometry explicitly controls semantic injection and alleviates the distribution gap between 2D semantic features and 3D geometric features. Experiments on two baseline frameworks show that the proposed method achieves better or comparable results in both full-data and few-shot settings on ScanNetV2 and ScanNet200.