Reading Attribution from Attention: Evidence Heads as Latent Attribution Mechanisms in LLMs
Abstract
Large language models (LLMs) show strong performance in multi-document question answering, but their practical deployment is limited by unreliable and unfaithful attribution to supporting evidence. Existing prompting and training-based methods often suffer from hallucinated citations and lack interpretability in how evidence is selected. In this work, we investigate whether attribution signals are inherently encoded within transformer attention mechanisms. We introduce a sensitivity-based diagnostic that identifies a small subset of attention heads, termed Evidence Heads, which are highly responsive to perturbations in supporting documents. Through causal interventions and semantic analysis, we show that these heads play a significant role in evidence identification and exhibit alignment with document-level entailment signals. Building on this finding, we propose a training-free Attention-based Attribution framework that extracts evidence signals directly from Evidence Head attention using global and local strategies. Extensive experiments show that our method consistently outperforms strong baselines while remaining lightweight and interpretable. Overall, our results suggest that structured attribution signals are implicitly encoded in LLM attention, and can be effectively leveraged for faithful multi-document reasoning.