RCSTAT: A Statistical Framework of Relative Contextualization in Transformers
Debabrata Mahapatra ⋅ Shubham Agarwal ⋅ Apoorv Saxena ⋅ Subrata Mitra
Abstract
Identifying which tokens and activations truly influence model predictions is critical for both the efficiency and interpretability of large auto-regressive language models. Yet in practice, token importance is typically inferred from attention distributions that entangle contextual relevance with softmax-normalization effects, obscuring relative influence across tokens and heads. We propose $\textbf{RCStat}$, a statistical framework for quantifying contextual influence in attention mechanisms. At its core $\textbf{Relative Contextualization (RC)}$ is a random variable that measures how strongly one subset of tokens contributes to another under the model’s internal scoring scheme. RCStat admits computationally efficient bounds on expected influence that can be estimated at inference time. No retraining is required. We apply RCStat to two tasks. For $\textbf{attribution}$, attention heads with high expected RC accurately identify the prompt spans that drive generation, enabling reliable span-level explanations. For $\textbf{key-value cache eviction}$, RC-based adaptive thresholding selectively evicts low-impact KV entries, substantially reducing cache size while preserving generation quality. Across question answering, summarization, and attribution benchmarks, RCStat achieves consistent gains, improving generation quality by 15–40\% with upto 36\% error reduction for KV eviction, and attribution accuracy by 2–16\%,. These results demonstrate that explicitly modeling contextual influence provides a principled and practical alternative to attention-based heuristics.
Chat is not available.
Successful Page Load