Scaffold3D: SfM-Conditioned Pointmap Prediction for Multi-View 3D Reconstruction
Abstract
Multi-view 3D reconstruction is increasingly driven by feed-forward models, which are fast and robust but often imprecise or globally inconsistent across views, especially in sparse-view and low-overlap settings. Structure-from-motion (SfM) offers a complementary geometric signal through camera poses and sparse 3D points estimated by correspondence filtering and optimization. To propagate this globally consistent \emph{SfM scaffold} into dense geometry, we introduce \textbf{Scaffold3D}, an SfM-conditioned reconstruction framework that injects image patch-aligned SfM point tokens into a pairwise feed-forward pointmap predictor. Our approach preserves image-token reasoning while using explicit SfM structure to condition dense prediction. Since pairwise pointmaps are predicted in local frames, we fuse them with global alignment and an SfM anchoring term. Across established benchmarks (ScanNet++, ETH3D, and Tanks&Temples) and out-of-domain data (4D-DRESS and MV-dVRK), Scaffold3D achieves stronger overall performance than feed-forward reconstruction models, point-token propagation, depth-completion baselines, and recent geometry-conditioned models. The gains are especially clear in low- and no-overlap evaluations. Our two-view SfM-conditioned pointmap predictor also outperforms several multi-view geometry-conditioned baselines, indicating that how geometric inputs are fused with image-token reasoning is as important as the number of conditioned views. Together, these results establish SfM-scaffold conditioning as a practical interface between learned pointmap prediction and classical geometric optimization.