Attributing Steering Outcomes: How Concept Selection Changes Method Comparisons
Abstract
Interpretability methods often select the behaviors used to evaluate them, so success may depend on both the method and the selected concepts. Natural Language Autoencoders (NLAs) describe model activations in text and reconstruct edited descriptions to steer output. We test these possibilities for NLA steering on Qwen2.5-7B and Gemma-3-12B. On concept sets selected using NLA results, NLA places concept words in 44.2\% of Qwen responses and 38.8\% of Gemma responses, more than the other activation methods. On a panel chosen before steering, DiffMean has the highest mean judge and word-based scores on both models. After correcting for multiple planned comparisons, no judge comparison involving NLA is statistically significant, so this panel does not establish a reliable winner. Editing an NLA description changes behavior for some concepts, but the ordering changes under some checks: one of three judges reverses the Gemma ordering, and changing steering strength or removing individual concepts changes the reported differences. The results do not show that differences between steering algorithms caused the observed ranking. Direct prompting outscores every activation method we test.