U-MVP: Encode Locally, Decode Globally for Feed-Forward 3D Gaussian Splatting
Abstract
Multi-view transformers for feed-forward 3D reconstruction must jointly solve two tasks. They need to extract precise per-view features and reason about cross-view geometric correspondences. Existing architectures apply cross-view attention throughout the network, leaving each layer to handle both objectives together. We revisit this design through a layer-wise probing analysis of feed-forward geometry transformers and observe that trained models tend to divide these two roles across layers. Early layers tend to specialize in per-frame feature extraction, while spatial cognition and cross-view reasoning emerge predominantly in later layers. This suggests that separating the two roles across the network can allocate compute more effectively. Building on this observation, we propose U-MVP, a U-shaped multi-view transformer for feed-forward 3D Gaussian Splatting that reflects this structure in the architecture itself. The encoder applies frame-wise attention to preserve per-view detail, the decoder introduces multi-view attention for cross-view geometric reasoning, and skip connections fuse the two streams so that appearance and geometry are recombined at decoding time. To scale to dense view regimes, we replace several global interaction blocks at the bottleneck with a grouping and swapping scheme that preserves cross-view information flow at lower cost. U-MVP performs competitively with feed-forward and optimization-based baselines from 16 to 256 views, generalizes to unseen datasets, and achieves strong results on 3D reconstruction and novel-view synthesis in posed and unposed settings.