Position-aware eXplanation: A Model-Agnostic Framework for Positional Attributions
Abstract
Transformer predictions depend on positional embeddings, which are known to produce biases such as first- and last-token dominance. However, feature attribution methods implicitly entangle influence from positions with features into a single opaque importance score per feature. To uncover these positional effects, we propose Position-aware eXplanation (PaX), a model-agnostic framework that transforms standard attribution methods to jointly produce feature and position attributions. PaX additionally produces counterfactual positional explanations: actionable scores quantifying how the prediction changes when a feature is relocated. We use PaX to generalize perturbation-, gradient-, and boundary-based attribution methods across vision, language, and clinical time-series benchmarks with no architectural changes. We demonstrate that separating positional effects improves both feature and position attribution faithfulness by 17.5% / 55% and 16.8% / 43%, respectively (insertion / deletion). In a case study on sepsis forecasting, we find that PaX recovers clinically validated bedside signals that standard attributions miss.