Minimally Invasive Steering of Language Models
Abstract
Test-time alignment aims to adapt large language models (LLMs) to runtime objectives without updating their parameters. This is especially useful when the reward signal is user-specific, time-varying, or available only through a black-box evaluator. We propose a minimally invasive pre-logit steering method that optimizes additive interventions to the hidden states of a frozen LLM in order to improve reward while preserving the base model's behavior. Rather than penalizing the Euclidean norm of the steering vector, we derive an effort penalty from the local KL geometry of the induced token distribution. Specifically, we show that the per-token KL divergence between the steered and reference policies admits a second-order expansion given by a Fisher-quadratic form in the steering vector. This regularizer has an analytic gradient computable through matrix-vector products with the frozen language-model head, making it as inexpensive as a standard quadratic penalty while retaining a principled KL interpretation. We further decompose the sequence-level KL gradient into an analytic Fisher component and a trajectory-dependent score-function component, and introduce a hierarchy of Fisher surrogates with provable first-order equivalence in the small-steering regime. The resulting algorithm provides a training-free, reward-driven test-time alignment procedure that improves reward while controlling deviation from the base policy.