CHARM+: Cross-Hardware Attention with Re-Merge Multistream Mechanism
Xinhang Zhang ⋅ boning zhang ⋅ Chengchun Liu ⋅ Chunpu Li ⋅ Lei Sun ⋅ Weifeng Zhang ⋅ Limin Xiao
Abstract
As sequence lengths continue to grow, distributed attention based on sequence parallel (e.g., Ring-Attention) has become a mainstream solution for large language models (LLMs). Nevertheless, inherent communication dependencies limit computation–communication overlap, reducing GPU utilization. Moreover, in cross-hardware (NUMA or node) architectures, due to imbalanced communication bandwidth, frequent communication synchronization will limit overall performance. Based on these insights, we propose CHARM+, Cross-Hardware Attention with Re-Merge Multistream Mechanism. We reorganize sequential communications into several levels and achieve intra-level parallelism and computation overlap. Furthermore, we decouple cross-hardware communication from intra-hardware communication, eliminating synchronization overhead caused by bandwidth imbalance. Forward and backward algorithms were evaluated on multiple hardware architectures and consistently outperformed state-of-the-art methods. Compared to Ring-Attention, our forward and backward algorithms achieve average speedups of $1.3\times$ and $1.7\times$ on a dual-NUMA PCIe architecture, and $2.8\times $ and $3.9\times$ on a dual-node NVLink architecture, respectively.
Chat is not available.
Successful Page Load