FlashMask-3: Efficient and Expressive Mask-Aware Distributed Attention
Guoxia Wang ⋅ Qianyue He ⋅ Siming Wu ⋅ Haoyang Xie ⋅ Qiwen Bao ⋅ Jinle Zeng ⋅ Jiabin Yang ⋅ Dianhai Yu ⋅ TIAN WU
Abstract
Long-context training of large language and multimodal models couples expressive structured attention masks (document, prefix-LM, share-question, sliding-window, block-sparse) with multi-node context parallelism, exposing the mask as a dominant bottleneck. The column-wise sparse mask of FlashMask compresses such masks from $O(n^2)$ to $O(n)$ via four query-axis boundary vectors and covers a broad family of structured patterns. Implemented for Ampere GPUs, FlashMask does not exploit FlashAttention-3's TMA/WGMMA warp-specialized pipeline on Hopper GPUs, and offers no distributed support. We observe that the column-wise sparse mask is indexed by key positions but stores query-axis interval boundaries. Under any multi-chunk context-parallel assignment, this property reduces per-rank mask localization to a provably zero-communication \emph{clip--shift--max} on the boundary vectors, directly compatible with the single-GPU kernel. Building on this property, FlashMask-3 introduces a mask-aware Hopper kernel with a three-role warp specialization atop the FlashAttention-3 TMA/WGMMA pipeline, together with a mask-aware distributed extension: zero-communication mask localization, a sparsity-aware load balancer (HCC-IPO), and a mask-aware compute-communication overlap with mask-driven communication pruning, whose forward all-gather additionally adopts topology-aware hierarchical routing. On a single H100 GPU across 12 structured masks at context lengths from 4K to 128K, FlashMask-3 improves achieved TFLOPs over FlashMask, FlexAttention, and MagiAttention by 45\% to 141\%, 34\% to 71\%, and 2\% to 44\%, respectively. Under context-parallel degrees from 4 to 32 with 8K sequence per device, it attains 31\% to 75\% compute-communication overlap efficiency and improves throughput over Megatron-Core CP (ring), Megatron-Core CP (all-gather) and MagiAttention by $3.25\times$ to $5.11\times$, $1.97\times$ to $3.26\times$ and $1.06\times$ to $4.09\times$ respectively. The source code will be released on GitHub.
Chat is not available.
Successful Page Load