GAMMA: Scalable 4D Gaussian Reconstruction Model for Novel View Synthesis of Monocular Videos
Abstract
We present GAMMA, a scalable feed-forward model that reconstructs 4D Gaussian Splatting from monocular videos, enabling dynamic novel view synthesis in real-time. Most existing approaches for dynamic scene modeling require multi-view videos as input or costly per-scene optimization. In contrast, GAMMA learns a unified 4D scene representation from monocular videos and generatively predicts 4D Gaussian primitives within seconds. The key insight of GAMMA lies in its unified Gaussian reconstruction model, which jointly estimates temporally consistent geometry, appearance, and motion from the input video. We further explore its scalability through large-scale training on a comprehensive dataset encompassing both synthetic and real-world videos. The reconstructed GAMMA representation supports interactive scene exploration via real-time rendering across timesteps and local viewpoints. Extensive experiments demonstrate that GAMMA outperforms existing methods in both reconstruction fidelity and efficiency.