The Fitting Corpus Is Part of the Measurement: Auditing a Jacobian Lens on Pythia
Abstract
A fitted interpretability lens that reads a language model’s hidden state at a given layer inherits the sampling choices made in selecting the corpus its map was estimated from. The Jacobian lens introduced by Gurnee et al. [2026] reads the residual-stream state at a chosen layer by projecting it through the average of the model’s own Jacobians over a corpus of prompts, and they find that the lens’s measured quality saturates with surprisingly few prompts. Their ablations vary the amount of fitting data, not which text distribution is averaged. We test source- corpus identity directly on EleutherAI’s Pythia model family, whose published training stream makes raw containment, the verbatim overlap between a corpus and the training data, measurable. An 8×8 factorial crosses eight fitting corpora with eight read-context corpora while holding the model, the evaluation battery, the cached activations, and the scorer fixed. Changing only the fitting corpus moves the Pythia 410M readout score from 22.8 to 26.0 points, and the effect reappears in controlled replications at Pythia 1B and 2.8B that meet pre-registered criteria without repeating the full factorial. In a two-source resampling control, variation between sources is much larger than variation across disjoint samples from either source. Raw verbatim containment does not account for the ordering: corpora with essentially no detectable 32-gram overlap with the published training stream produce readouts within the range of heavily contained ones. Over fitting sets of 25 to 800 prompts the score moves less within a corpus than across corpora, while a panel of empirical corpus and fitted-map statistics does not identify which property drives the effect. The fitting corpus is therefore part of the measurement specification. It should be reported alongside the model and layer, and sensitivity to fitting-corpus choice should be evaluated just as sensitivity to fitting-set size is. Code and results: https://github.com/AnonymousNeurIPSInterpScience/pythia-jlens-artifacts Artifacts: https://huggingface.co/AnonymousInterpScience/pythia-jlens-artifacts