Finite-Sample Feature Recovery and Gradient-Aligned Activation Steering of Personas
Abstract
Rank-1 activation interventions based on steering or persona vectors are an efficient but unreliable way to control the outputs of language models. Their efficacy exhibits substantial variance across concepts and intervention sites, which has been hypothesized to depend on how cleanly the empirically estimated vector is linearly extracted from the noisy activation space. We first make this claim more rigorous by theoretically characterizing the conditions under which a finite-sample difference of means steering vector can be \textit{recovered} (has positive cosine similarity with a target feature direction) under a linear feature-superposition model. Using the multiple choice logit gradients as empirical proxies for target feature directions, we show that positive alignment between empirical steering vector and these gradients can positively predict steerability across 26 persona datasets. Finally, we show that simply projecting the steering vectors to improve positive alignment with gradients can improve steering performance over standard difference-in-means vectors.