BridgeMVS: Bridging Multi-View Stereo and Monodepth via Bidirectional Dynamic Fusion
Abstract
Multi-View Stereo (MVS) and Monocular Depth Estimation (MDE) provide complementary cues for dense 3D reconstruction. MDE offers rich contextual and structural priors but lacks explicit geometric constraints, whereas MVS exploits multi-view geometry but often struggles in ambiguous regions, such as weakly textured or reflective surfaces. Existing MDE-assisted MVS methods mainly use monocular cues in a unidirectional manner, leaving the interaction between monocular representations and multi-view cost volumes insufficiently explored. To address this limitation, we introduce BridgeMVS, a unified cascaded MVS framework that bridges monocular depth features and multi-view cost-volume representations through Bidirectional Dynamic Fusion (BDF). At each cascade stage, BDF performs dynamic mono-to-volume and volume-to-mono updates: monocular structural features are injected into cost-volume regularization, while multi-view geometric evidence recalibrates intermediate monocular representations. The enhanced monocular features are propagated across stages together with MVS representations and supervised by stage-wise monocular losses, establishing an auxiliary structural guidance path for the MVS branch. Unlike stage-isolated designs, BridgeMVS preserves fused representations from coarse to fine stages, allowing low-resolution cross-branch interactions to guide subsequent high-resolution depth estimation. Moreover, BDF is designed as a plug-and-play module and can be conveniently integrated into existing cascaded MVS frameworks. Extensive experiments show that BridgeMVS achieves state-of-the-art performance on the DTU, Tanks and Temples, and ETH3D benchmarks.