Geometry-Aware Flow Matching for Sparse-View 3D Gaussian Splatting
Abstract
Generalizable 3D Gaussian Splatting aims to predict renderable Gaussian scenes from sparse images without per-scene optimization at test time. Existing feed-forward methods improve this problem mainly by strengthening the evidence given to the predictor, such as multiview aggregation, geometric cues, auxiliary supervision, and surface priors. Yet the prediction itself is usually learned as a direct map to the final Gaussian scene. This endpoint-centered objective compresses scene layout, local geometry, visibility, opacity, orientation, and appearance into a single end-point target, leaving the construction path of the Gaussian scene implicit. We propose GRiF, a geometry-aware flow matching framework that augments endpoint target supervision with conditional transport in scene space. Starting from a controlled source scene, GRiF learns a view-conditioned velocity field that progressively evolves a hierarchical Gaussian state toward a renderable target. The hierarchy organizes updates from coarse layout to fine attributes, while geometry-aware paths respect the native structure of Gaussian variables with rotation-aware paths for orientation variables. Experiments on RealEstate10K and ACID, with zero-shot transfer evaluated on ScanNet, DL3DV, and DTU, show improved in-domain reconstruction and stronger cross-domain performance.