Static-to-Dynamic: Animating Still Mattes via Generative Motion for Video Matting
Abstract
Video matting is essential for high-precision foreground extraction, yet its development is constrained by the scarcity of diverse annotated training data. To address this, we propose Static-to-Dynamic (S2D), a novel generative paradigm that scales video matting data by animating still labeled images into high-fidelity video-matte pairs. Starting from a static image-matte pair, S2D leverages an image-to-video generative model to synthesize realistic motion and scene dynamics, while simultaneously producing pseudo alpha labels from generative priors. Specifically, we introduce Latent Matte Propagation (LMP), which propagates the first-frame matte to subsequent generated frames via the attention mechanisms in the diffusion transformer. We further design a Coarse-to-Fine Matte Decoder to recover pixel-level alpha details from compressed latent representations using multi-scale generative features. Based on this paradigm, we build S2D-VM, a large-scale dataset containing 14K high-fidelity paired video clips. We also propose a Self-Corrected Progressive Training strategy to mitigate pseudo-label noise during optimization. Experiments show that S2D effectively scales video matting data and consistently improves the performance of existing models across multiple benchmarks.