AsymHP: Load-Balanced Sparse Attention for Video Diffusion Transformers
Xinwei Qiang ⋅ Yue Guan ⋅ Ruihan Zhu ⋅ Mihir Jagtap ⋅ Zaifeng Pan ⋅ Zhongkai Yu ⋅ Chang Chen ⋅ Zhengding Hu ⋅ Yufei Ding ⋅ Adnan Aziz
Abstract
Video diffusion transformers rely on self-attention over long spatio-temporal token sequences, making inference expensive even when sparse attention reduces the number of computed blocks. This paper studies a systems bottleneck that appears in dynamic sparse video attention: per-head sparsity can vary widely across denoising steps, layers, and prompts, so equal-head parallel execution leaves some GPUs waiting for dense heads assigned to other GPUs. We propose AsymHP, a runtime system that uses the previous denoising step's per-head sparse density to place a non-uniform number of heads on each GPU. AsymHP combines this lightweight online density estimate, a cost model for mask construction and sparse attention computation, and an asymmetric head redistribution primitive that avoids padding variable-size head shards into symmetric collectives. In our evaluation on H100 GPUs, AsymHP improves sparse attention latency by up to $1.54\times$ without changing the underlying sparse-attention algorithm.
Chat is not available.
Successful Page Load