DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion
Abstract
Leveraging the natural spatiotemporal energy decay in video diffusion offers a path to efficiency, yet relying solely on rigid static masks risks losing critical long-range information in complex dynamics. To address this issue, we propose DynamicRad, a sparse-attention framework that constrains adaptive selection using a radial locality prior. DynamicRad introduces a dual-mode strategy: static-ratio for speed-optimized execution and dynamic-threshold for quality-first filtering. To avoid online search over sparse indices, we integrate an offline Bayesian Optimization (BO) pipeline with a semantic motion router. The router maps prompt embeddings to BO-selected sparsity regimes with a single projection module. Unlike online profiling methods, our offline BO optimizes attention reconstruction error (MSE) on a proxy task and reuses the selected configurations during inference. Experiments on HunyuanVideo and Wan2.1-14B demonstrate that DynamicRad achieves a strong efficiency-quality trade-off among FlashAttention-compatible sparse-attention baselines, obtaining 1.7x-2.5x inference speedups with over 80% effective sparsity. In some long-sequence settings, the dynamic mode matches or improves over the dense baseline on automatic quality metrics, while optional Mask-Aware LoRA further improves long-horizon coherence. Code is available at https://anonymous.4open.science/r/DynamicRad-5D55/.