Harnessing Image Diffusion Prior for Photo-Realistic Video Restoration
Abstract
Recent image generation models produce photo-realistic results with rich details, but transferring their powerful priors to video restoration remains difficult due to temporal inconsistency from stochastic detail synthesis. Existing video restoration methods improve temporal coherence but often sacrifice visual fidelity, leaving a gap between image-level quality and video-level stability. We propose a generative video restoration method that bridges this gap by enabling temporally consistent detail synthesis from image generation priors. Our method combines noise energy rebalancing to suppress structure-disturbing low-frequency stochasticity, motion-aligned noise warping to propagate details along object motion, a self-supervised denoising encoder to mitigate encoder-induced temporal uncertainty, and decoupled temporal attention to model cross-frame dependencies. Together, these components preserve fidelity, enhance fine-grained realism, and maintain temporal consistency. Extensive experiments show that our method substantially outperforms existing video restoration approaches in visual quality, detail richness, and temporal stability.