Extending 3D Reconstruction Models to Any Camera
Abstract
Feed-forward 3D reconstruction from unconstrained multi-view images has recently emerged as a dominant paradigm in 3D computer vision. Various foundational reconstruction models have been proposed with different architectures and reconstruction paradigms. Despite being trained on large-scale perspective datasets with substantial computational resources, these models generally fail when deployed on images whose camera models deviate from the training distribution. Applying such models to downstream applications with arbitrary, unknown, and potentially mixed camera types remains challenging. To maximally leverage the strengths of existing reconstruction models with minimal retraining efforts, we propose a low-rank adaptation framework that efficiently extends perspective-pretrained models to fisheye and panoramic imagery. Extensive experiments demonstrate that, with only 0.1% additional parameters and one day of training, our method can adapt a foundational 3D reconstruction model to non-pinhole cameras while fully preserving its original performance on perspective inputs.