Beyond Distribution Matching: Self-Supervised Representation Forcing for Few-Step Video Generation
Abstract
Few-step distillation is essential for deploying modern video diffusion models, and Distribution Matching Distillation (DMD) has emerged as a leading paradigm due to its strong output-space fidelity and its flexibility in supporting both bidirectional and causal student architectures. We find, however, that DMD has a structural blind spot: its objective lives entirely in the output space and provides no signal pushing the student to form a structured internal representation. As the step count shrinks, this output-only supervision becomes too coarse, and the distilled student degrades on both fine spatial structure and coherent temporal dynamics. To close this gap, we introduce Representation-Forcing, which adds a predictive representation loss on top of DMD without changing its output objective. By feeding the student and an EMA teacher with heterogeneous noise levels, we create an information asymmetry that forces the student to predict the teacher's representation from a more corrupted view --- explicitly bringing "representation compression" to distribution matching. Experiments show that Representation-Forcing consistently improves both spatial fidelity and temporal coherence. The mechanism is paradigm-agnostic across bidirectional and causal student architectures, requires no external feature extractor, and adds negligible cost.