GGQR: Gaussian-Grounded Query Refinement for Feed-Forward 4D Gaussian Splatting
Abstract
Feed-forward 4D Gaussian Splatting (4DGS) offers an efficient paradigm for reconstructing dynamic scenes from monocular videos, bypassing the need for lengthy per-scene optimization. However, existing methods suffer from three fundamental bottlenecks: redundant frame-wise grid-aligned Gaussian representations, suboptimal single-step 2D-to-4D regression architectures, and restrictive dependencies on 3D/4D signals during training or inference. In response to these limitations, we introduce a novel Gaussian-Grounded Query Refinement (GGQR) framework. To reduce representation redundancy, GGQR learns a compact set of Gaussian queries equipped with a newly designed plateau-shaped temporal kernel, enabling single primitives to persistently model motion-consistent regions across extended frames. To alleviate regression ambiguity, these queries are iteratively refined through a Gaussian-grounded attention mechanism operating in 3D space. Specifically, GGQR leverages foundation models to lift 2D observations into a shared 3D space, explicitly grounds the 3D attention in intermediate Gaussian states, and residually updates those Gaussian states. Trained entirely self-supervised on casual videos, GGQR achieves a new state-of-the-art on dynamic reconstruction benchmarks among feed-forward 4DGS methods, delivering superior performance in novel view-time rendering alongside downstream motion modeling tasks.