Punctuation-aware Hybrid Trainable Sparse Attention for Large Language Models
Abstract
Attention serves as the fundamental mechanism for long-context modeling in large language models (LLMs), yet dense attention becomes structurally prohibitive for long sequences due to its quadratic complexity. Consequently, sparse attention has received increasing attention as a scalable alternative. However, existing sparse attention methods rely on coarse-grained semantic representations during block selection, which blur intra-block semantic boundaries and lead to the loss of critical information. To address this issue, we propose Punctuation-aware Hybrid Sparse Attention (PHSA), a natively trainable sparse attention framework that leverages punctuation tokens as semantic boundary anchors. Specifically, (1) we design a dual-branch aggregation mechanism that fuses global semantic representations with punctuation-enhanced boundary features, preserving the core semantic structure while introducing almost no additional computational overhead; (2) we introduce an extreme-sparsity-adaptive training and inference strategy that stabilizes model behavior under very low token activation ratios. Extensive experiments on general benchmarks and long-context evaluation tasks, together with ablation studies on punctuation types and the linguistic origin of punctuation tokens, demonstrate that PHSA consistently outperforms both dense attention and the state-of-the-art sparse attention baseline InfLLM v2. Specifically, for a model under the training-inference consistent setting with an input sequence length of 32k tokens, PHSA reduces information loss by 10.8\% at a sparsity ratio of 97.3\%.