Sequence-to-Sequence Modeling with Camera-Induced Priors for Multi-View Stereo
Abstract
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. In this work, we address the Multi-View Stereo (MVS) setting, where camera parameters are assumed to be known. Most existing learning-based methods rely on per-view cost volumes as a camera-induced prior, effectively casting the problem as a sequence-to-one mapping that predicts depth only for a designated reference view. We instead reformulate MVS as a sequence-to-sequence task, enabling the simultaneous prediction of depth maps and point maps for all input views. To this end, we propose a global transformer-based architecture with two key components that explicitly utilize camera-induced priors. First, we employ ray-map embeddings to inject camera parameters into image patch tokens, making the transformer architecture camera-aware for geometry prediction. Second, we replace conventional per-view cost volumes with a unified global cost-volume representation that jointly captures 3D structure across all views. Extensive experiments on multiple public benchmarks demonstrate that our approach achieves state-of-the-art performance, surpassing both multi-view stereo and feed-forward reconstruction baselines.