Target-Token Asymmetry Can Masquerade as Depth Specialization
Abstract
Interpretability studies often compare how much early- and late-layer interventions hurt next-token prediction on different datasets. A larger drop is sometimes interpreted as evidence that a task is represented more strongly at one depth. We show that this conclusion can be misleading when the datasets contain different kinds of target tokens. In a GSM8K--Wikitext comparison, the overall effect repeats across models, but its positive sign comes from whitespace targets. Effects for content words, numbers, punctuation, and function words are all non-positive, and removing whitespace makes the estimate negative in all three fully decomposed models. We then evaluate 13 dataset pairs and 41 model--pair conditions. Cheap tokenizer statistics help identify comparisons that need this audit, including two comparisons motivated by prior work. The main finding remains under alternative target definitions, position controls, and a continuous prediction score. A separate, narrower content-token effect remains under a different intervention and prediction horizon. We therefore recommend logging individual targets, reporting results by target type, and checking whether a domain-level conclusion remains after removing or reweighting target types.