Kernelized Activation Steering
Abstract
Activation steering provides a simple, training-free mechanism for controlling attributes of generative models (e.g., sentiment, style, helpfulness). However, standard approaches such as Difference-in-Means (DiM) apply a single input-independent steering vector across all activations, limiting expressivity and ignoring the local structure of the activation space. We propose Kernelized Activation Steering (KAS), a unifying framework that lifts activation steering into a reproducing kernel Hilbert space (RKHS). KAS formulates steering as an optimization problem expressed purely via kernel evaluations, yielding an implicit, activation-dependent steering score without constructing explicit feature maps. Unlike DiM, KAS induces locally adaptive steering: each activation is modified according to its relative position with respect to source and target reference sets, producing a nonlinear steering field over the representation space. Importantly, DiM is recovered as a special case under a linear kernel, offering a principled interpretation of steering, while richer kernels (e.g., RBF) enable geometry-aware interventions. Across standard activation steering tasks, including jailbreaking LLMs and image style control, KAS consistently outperforms existing methods.