AdaCal: Adaptive Calibration for Robust Sparse Attention in Long-Context LLMs
Abstract
Large language models face substantial computational challenges in long-sequence inference due to the quadratic cost of self-attention, with the prefill stage being particularly bottlenecked. Static sparse attention reduces compute via predefined masks but can miss task-critical evidence, while dynamic sparsification is more flexible yet can be brittle across tasks, leading to either insufficient coverage or distractor-induced noise. We propose AdaCal, an Adaptive Calibration framework for robust sparse attention that synergizes coarse, experience-driven task dispatch with head-wise physical profiling. By deriving entropy, concentration, and mid-zone coverage metrics from a lightweight probe pattern, AdaCal dynamically governs head activation and augmentation policies, allocating additional computation strictly on demand. This design is realized through task-conditional policies that effectively suppress unnecessary expansion for structure-dominant inputs while enhancing long-range coverage for retrieval-intensive tasks. Experiments show that AdaCal improves robustness over strong static and dynamic baselines, with clearer gains in long-context settings beyond 32K tokens.