Stress-testing NLA causality through verbalizer steering
Abstract
Natural Language Autoencoders (NLAs) describe model activations in text and reconstruct edited descriptions in activation space. The difference between the two reconstructions can steer model outputs without additional training. We compare this procedure with established activation steering methods on Qwen2.5-7B and Gemma-3-12B. On concept sets selected using NLA steering results, NLA places concept words in 44.2\% of Qwen responses and 38.8\% of Gemma responses, more than the other activation methods, and NLA stays first after matching for output quality. On a third panel chosen before steering that lead disappears: DiffMean has the highest mean AxBench judge and lexical scores on both models, but no planned judge comparison involving NLA is significant after Holm correction, so this panel names no reliable winner. The panels also differ in their word lists and tuning objectives, so this experiment does not isolate concept selection from every other design choice. NLA can build useful directions for some concepts, but the results do not establish a general advantage, and direct prompting outscores every activation method we test. For interpretability-based discovery, explanations should be tested on cases chosen independently of the method.