Bidirectional Sparse Attention for Faster Video Diffusion Training
Abstract
Video diffusion Transformer (DiT) models excel in generative quality but hit major computational bottlenecks when producing high-resolution, long-duration videos. Full attention requires computing dot products between all pairs of queries and keys, resulting in a quadratic computational complexity with respect to the sequence length ((O(L^2))), leading to high training and inference costs.To overcome this limitation, we propose a Bidirectional Sparse Attention (BSA) framework that sparsification from both the query and key–value directions to reduce the quadratic cost of full attention.Specifically, the sparsification of queries is achieved by pruning tokens that are locally redundant or semantically similar within the 3D spatiotemporal domain of a video. For the key–value pairs, only those that are highly correlated with the query at a global level are selected for attention computation. Furthermore, we design a dynamic threshold adjustment mechanism to adaptively regulate the selected number of key–value pairs, thereby maintaining quality without degradation. Extensive experiments demonstrate that BSA significantly accelerates DiT training across long sequences, reducing FLOPs by up to 20× and achieving 17.79× faster attention training, while preserving or even surpassing the generative quality of full attention.