Latent Motion Alignment for Video Diffusion
Abstract
Video diffusion models often fail in precisely the dimension that distinguishes videos from images: motion. Standard reconstruction and flow-matching objectives in pixel or VAE-latent space are dominated by appearance, so visually similar clips can be treated as close even when their underlying dynamics differ. We propose Latent Motion Alignment, a training-time objective that compares predicted and ground-truth videos in a learned motion space. This space is learned by training a compact encoder to describe how one latent frame changes into the next, while a decoder ensures that the resulting token captures information needed to model the transition. Once learned, the encoder is frozen and used as an auxiliary loss for fine-tuning video diffusion models. The motion module is used only during training and adds no inference-time cost. We validate that the learned space separates motion rather than appearance shortcuts, improves motion-sensitive generation metrics, and enables fine-grained Olympic-diving generation where jump type and score depend on subtle temporal structure.