Proxy measures and causal effects disagree in persona steering
Adhitya Rajendra kumar ⋅ Rohit Shenoy ⋅ Leilani Gilpin
Abstract
Studies of refusal locate a direction in a language model's residual stream and intervene on it, then report the fraction of harmful prompts the intervention makes the model answer. That fraction comes from an automated scorer validated once and applied unchanged across every intervention strength. We show this hides a bias that grows with the intervention: steering makes Gemma-2-2B-it verbose and list-prone, which a substring scorer reads as compliance, so two scorers disagree on $1.45\%$ of unsteered generations and $7.82\%$ of steered ones. The bias is largest exactly where an effect looks strongest: full directional ablation puts $64\%$ of its generations in the disputed cell, so its corrected rate spans $0.35$ to $0.68$ rather than resolving to one number, an effect that scorer validation performed once cannot catch. Using the same causal reference, we separately find that probe accuracy, used to pick an intervention layer, ties across seven layers and selects a different one than a direct causal test once the tie is broken. Applying the corrected scoring to the persona-vector decomposition that motivated this work, a vector's component parallel to the refusal direction reduces to $\pm\hat{r}$ once magnitude-matched and carries no persona information, while its perpendicular component raises the rate for all three harmful personas we test. The reference itself holds, since ablating twenty random directions leaves the corrected rate an order of magnitude below ablating the refusal direction. A scorer should therefore be validated at every intervention strength it is used to report, not once for the study as a whole.
Chat is not available.
Successful Page Load