TriTD: Tri-Partite Trajectory–Distribution Distillation for Real-Time Autoregressive Video Generation
Abstract
Autoregressive (AR) video diffusion has emerged as a leading paradigm for efficient video generation, yet its practical deployment remains bottlenecked by the trade-off between generation quality and inference efficiency. We attribute this bottleneck to three error sources in AR rollouts: the base model error, the train–inference distribution mismatch, and the upstream error propagation. Building on this decomposition, we derive closed-form upper bounds on the per-step and cumulative errors, providing a principled mechanism for jointly optimizing generation quality and inference efficiency. Guided by this analysis, we propose TriTD, a Tri-Partite Trajectory–Distribution co-optimized distillation framework with three complementary modules that each target a distinct error term in the bound. First, for ODE initialization, we adopt a Next-Node Trajectory Distillation (NTD) objective and prove a strictly tighter error bound than conventional consistency distillation, reducing the base model error at its source. Second, we introduce Chunk-Graded Noising (CGN), a monotonically non-decreasing per-chunk noising schedule that provably narrows the train–inference noise mismatch and removes redundant KV Cache updates, yielding substantial speedups at no quality cost. Third, we propose Orthogonal-Parallel DMD (OP-DMD), a distribution-matching loss that preserves the original DMD gradient direction while explicitly suppressing unsupervised off-axis drift in the orthogonal subspace, improving both training stability and final generation quality. Extensive experiments show that TriTD consistently surpasses state-of-the-art baselines in visual quality while reaching a real-time throughput of 26.8 FPS on a single H100 GPU, delivering a balanced quality-efficiency solution for interactive streaming video generation.