Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video
Abstract
Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet videos offer virtually unlimited raw data, utilizing them as a scaling source for geometric learning is challenging due to the absence of ground-truth geometry and the presence of observational noise. To address this, we propose SAGE, a framework for Scalable Adaptation of GEometric foundation models from raw video streams. SAGE leverages a hierarchical mining pipeline to transform videos into training trajectories and hybrid supervision: (1) Informative training trajectory selection; (2) Sparse Geometric Anchoring via SfM point clouds for global structural guidance; and (3) Dense Differentiable Consistency via 3D Gaussian rendering for multi-view constraints. To prevent catastrophic forgetting, we introduce a regularization strategy using anchor data. Experiments on applying SAGE to two architecturally distinct backbones (including both MV-DUSt3R and VGGT) show that Chamfer Distance is consistently reduced by 20-42% on unseen benchmarks (7Scenes, TUM-RGBD, Matterport3D). SAGE establishes Internet video as a viable and scalable adaptation source for geometric foundation models.