Holo4D: Holistic 4D Reconstruction as Geometric Control for Video Diffusion
Abstract
Reconstructing a dynamic 4D scene along a novel camera trajectory requires both metric geometry from the source video and generative completion of disoccluded target-view regions. Feed-forward 4D reconstruction models recover increasingly comprehensive input-view geometry, including depth, point maps, and dense 3D tracks, but remain tied to the observed camera stream. Camera-controlled video diffusion models (VDMs) provide strong generative priors for novel-view synthesis, yet their camera control is typically not precise enough for metric novel-view reconstruction. We argue that comprehensive input-view 4D reconstruction is an effective geometric interface between these two capabilities: properly exposed to a VDM, it turns reconstruction outputs into camera-accurate novel-view RGB videos that remain geometrically consistent under downstream 4D reconstruction. We introduce HOLO4D, a geometry-aware video-to-video framework that controls a pretrained VDM with a hybrid 4D geometric cache built from the source video. A small set of source views covering the camera motion forms the static cache, while dynamic regions are rendered from the matching source frame. The resulting hybrid RGB-D representation is combined with explicit valid masks, target-camera ray embeddings, and dense 3D tracking tokens. Conditioned on this 4D scaffold, the diffusion backbone synthesizes the target-view RGB video. Under a common downstream 4D reconstructor, our generated videos yield target-view depth more consistent with the held-out reconstruction than that of the baselines, and an additional feature-level readout shows that intermediate VAE features already encode target-view depth. Experiments and ablations further show that progressively richer 3D/4D reconstruction signals improve camera adherence.