When a Conceptor Complement Becomes Gain: Operator Geometry and Behavioral Attribution in Language Models
Shuhul Razdan ⋅ R Balamurugan
Abstract
Activation interventions are often described by the direction or matrix they contain, although generation receives the full composed transformation. We study a target-centered, scaled conceptor complement, $F_{C,t}(h)=m_t+\beta(I-C_t)(h-m_t)$, and isolate what the fitted matrix contributes beyond the same target-specific center and gain. At $\beta=2$, spectra fitted on three instruction-tuned models and seven assistant-behavior targets place every mode above unit gain. On question-held-out activations, the matrix-dependent correction is 0.093%--7.815% of the matched-control displacement in root-mean-square norm. We then remove only the fitted matrix in a 1,920-generation study over three models and four targets. Its removal produces modest aggregate changes in target expression and automated coherence despite frequent text-level differences. In contrast, dedicated erasure methods substantially reduce fresh linear accessibility. Together, these results give a practical attribution rule: analyze the complete map, ablate one component at a time, and measure representation and output endpoints separately.
Chat is not available.
Successful Page Load