MonoChunk3D: Monocular Online 3D Instance Segmentation with Persistent Instance States
Abstract
Monocular online 3D instance segmentation enables embodied agents to build object-level 3D understanding while continuously exploring their environments. Unlike RGB-D settings, monocular RGB streams lack direct depth and camera pose information, making geometry reconstruction necessary for 3D segmentation. Recent reconstruction foundation models (RFMs) have made this task more feasible by recovering 3D geometry and cross-view cues from monocular images. However, reconstructed geometry remains noisy and incomplete, and streaming observations provide only partial views of object instances, making local observations insufficient for stable instance representations and temporally consistent segmentation. We propose MonoChunk3D, an online 3D instance segmentation framework that incrementally reconstructs, aligns, and segments incoming monocular RGB chunks. Rather than treating reconstructed geometry merely as input, the framework adaptively integrates complementary RFM-derived cross-view priors with 3D geometric features to initialize current-chunk instance queries. To overcome the limitations of local observations, we maintain compact persistent instance states composed of explicit prototypes, implicit query embeddings, and spatial bounds to preserve long-term instance context. These states are propagated as history-aware queries and jointly decoded with current-chunk queries, allowing the preserved instance context to directly guide current-chunk segmentation. The resulting predictions are associated with historical instances to update the persistent states, enabling temporally consistent segmentation over streaming observations. Experiments on ScanNet200, ScanNetV2, and SceneNN show that MonoChunk3D achieves consistent improvements over existing methods in the monocular online setting.