BMAttn: Block-Aligned Mixed-Precision Attention Quantization for LLM Inference
Abstract
The deployment of Large Language Models (LLMs) with extended context windows is fundamentally constrained by the quadratic computational and memory costs of the self-attention mechanism. While acceleration techniques like sparse attention and quantization offer potential relief, they currently face a critical dilemma: uniform compression strategies fail to capture the non-uniform distribution of information importance, degrading performance on long-range dependencies; conversely, fine-grained, token-level adaptation often introduces irregular memory access patterns that negate theoretical efficiency gains on modern GPUs. To resolve this tension, we introduce BMAttn (Block-Aligned Mixed-Precision Attention), a framework that unifies importance-aware precision allocation with hardware-friendly execution. BMAttn partitions attention maps into block-aligned high-precision, low-precision, and sparse regions, using a novel affine windowing mechanism to dynamically adjust boundaries based on sequence length. We further propose a saliency-weighted calibration method and a layer-adaptive regularizer that adaptively aligns precision targets with perceptual importance and layer sensitivity. Extensive experiments demonstrate that BMAttn achieves a \textbf{3.3x} speedup on long-context tasks with negligible accuracy loss, and up to \textbf{5x} speedup with minimal degradation, effectively bridging the gap between algorithmic adaptivity and hardware efficiency.