Dynamics-Aware Sparse Attention for Efficient Autoregressive Video Diffusion
Yuanyu He ⋅ Zhuokun Chen ⋅ Yefei He ⋅ Zhiwei Tang ⋅ Jiasheng Tang ⋅ Yinghao Yu ⋅ Jianfei Cai ⋅ Bohan Zhuang
Abstract
Autoregressive Diffusion Transformers have emerged as a powerful paradigm for long video generation. However, they are often constrained by prohibitive computational overhead and substantial memory footprints. This inefficiency arises from dense computation mechanisms that allocate uniform processing resources across all spatiotemporal regions, failing to leverage the inherent temporal redundancy of video where motion is typically confined to sparse areas. In this paper, we propose **Dynamics-Aware Sparse Attention (DASA)**, a unified framework that shifts from dense computation to sparse, motion-driven generation to eliminate both storage and computational redundancies. Inspired by video codec principles, our approach explicitly decouples video content into dynamic and static regions. To mitigate memory bottlenecks, we selectively compress the historical KV cache by retaining only dynamic features within the generated chunk. Furthermore, to reduce computational overhead, we employ a region-level computation allocation strategy based on localized dynamics. Additionally, a dynamics-guided distillation method is introduced to enhance performance via a lightweight distillation training process. Experiments on VBench demonstrate that our approach accelerates autoregressive video generation by 2.16$\times$, while preserving high visual fidelity.
Chat is not available.
Successful Page Load