Jacobian Scopes: A Unified Geometric Framework for Token-Level LLM Attributions
Toni Liu ⋅ Baran Zadeoğlu ⋅ Nicolas Boulle ⋅ Raphaël Sarfati ⋅ Gurbir Arora ⋅ Christopher Earls
Abstract
Gradient-based attribution methods for large language models (LLMs) have so far been limited to explaining scalar outputs - the logit of a single target token. Yet LLM predictions are inherently distributional, and many practically important questions concern the full predictive distribution: which tokens make the model uncertain? Which drive a broad range of plausible continuations? We propose **Jacobian Scopes**, a unified framework that fills this gap by projecting the input-to-output Jacobian onto different directions in output space via a single vector-Jacobian product, requiring only one backward pass. This yields three complementary methods: **Semantic Scope** attributes a specific target logit; **Fisher Scope**, grounded in information geometry, identifies tokens that most alter the overall shape of the predicted distribution; and **Temperature Scope** traces which tokens govern the model's predictive confidence. Fisher and Temperature Scopes are, to our knowledge, the first attribution methods to natively target distributional rather than pointwise features of LLM predictions. Through case studies spanning instruction following, translation, and in-context time-series forecasting, Jacobian Scopes reveal implicit political biases, uncover word- and phrase-level translation strategies, and illuminate the nearest-neighbor pattern-matching mechanisms underlying LLM forecasting. Quantitative evaluation on LAMBADA and IWSLT2017 across six leading LLMs (LLaMA-3.2, Qwen2.5, Gemma-3) confirms that Jacobian Scopes consistently match or outperform Input $\times$ Gradient and Integrated Gradients, at a fraction of the latter's computational cost.
Chat is not available.
Successful Page Load