Trajectory Forcing: Exploiting Diffusion Trajectories for Autoregressive Long Video Generation
Bohan Wang ⋅ Shuo Chen ⋅ Yunzhi Li ⋅ Zhaozheng CHEN ⋅ Junzhe Zhang ⋅ Hanwang Zhang
Abstract
Autoregressive (AR) modeling is arguably the only way to extend a fixed-length video diffusion model to longer or even streaming videos. However, existing AR diffusion methods use only the previously predicted chunk as context for the next, discarding the intermediate predictions that constitute the chunk's generation history. In contrast, conventional AR models such as LLMs condition on the entire history of preceding tokens. To this end, we propose Trajectory Forcing (TraF), an AR diffusion method that exploits the denoising trajectories of previous video chunks as context. TraF operates under two training regimes: (1) with ground-truth videos, where pseudo-trajectories are constructed from Gaussian interpolations and re-predicted by the model, and (2) with a teacher model, where real trajectories are obtained from the student's own denoising rollout. The first regime naturally provides a strong initialization for the second. On the 5-second VBench benchmark, TraF achieves a Total score of 85.15, outperforming Self Forcing (+0.84) and Causal Forcing (+1.11) under the same training configuration, with Quality and Semantic improving simultaneously. On 30-second generation — $6\times$ the training horizon — TraF achieves the highest overall quality among all compared methods.
Chat is not available.
Successful Page Load