What Heads Are You Really Finding? Single-Token vs. Free-Form Localization
Abstract
Steering pipelines routinely find the components responsible for a behavior by constraining the model to emit one token and ranking components by their effect on that token's probability. The behaviors themselves, however, live in whole generations. This work investigates the transferability of interventions localized on single-token tasks to free-from generations. Across five open-weight instruction-tuned models and five behaviors, we build two versions of every localization contrast that are matched in every respect except the length of the response the model produces, run identical attribution patching and difference-in-means steering on both, and evaluate the resulting interventions on shared held-out prompts. Steering from free-form generation-derived heads produces the target behavior more often on generative evaluation with a lower head budget. The gain is concentrated on the behaviors whose constrained format asks for an option letter and vanishes on those whose constrained answer word already carries the trait. Structurally, the localized head sets diverge across the two methods, but the vectors they inject into the residual stream are strongly aligned and both place the behavioral trait on a leading principal component. The separation lies in what accompanies the trait: the generation-derived concept subspace is an order of magnitude lower in rank and puts twice the share of its variance on the axis it steers along, while the constrained concept subspace spends the remainder on the answer symbol.