Revisiting Cross-View Completion: Self-Supervised Pre-Training via Reconstruction Error Comparison
Abstract
Self-supervised pre-training via cross-view completion learns strong features for 3D vision from co-visible regions of image pairs. However, the reference view provides little information for reconstructing non-co-visible patches, implicitly yielding a monocular training signal in these regions. We introduce Gekko, a self-supervised pre-training framework that turns this limitation into a useful signal. The relative improvement of the cross-view reconstruction error over a masked-autoencoder error serves as a self-supervised proxy for co-visibility: large improvements indicate co-visible regions, while negligible ones indicate non-co-visible areas. Gekko is a network, trained from scratch, that jointly performs cross-view completion, masked autoencoding, and per-pixel prediction of this relative improvement, providing an additional binocular signal for all masked regions without requiring any ground-truth 3D annotations. Under identical architectures and training data, Gekko consistently outperforms CroCo on zero-shot correspondence estimation, relative pose estimation, and pointmap regression. We further show that Gekko can be trained directly from raw videos using a simple stride-based curriculum, eliminating the cumbersome 3D preprocessing required by prior methods while matching the performance of models trained on curated data, thereby enabling fully self-supervised pre-training. Code and pre-trained models will be released.