The Sharp Directions Are Against You: Curvature Analysis of Activation Steering
Seyedarmin Azizi ⋅ Arya Fayyazi ⋅ Parsa Razmara ⋅ Erfan Baghaei Potraghloo ⋅ Souvik Kundu ⋅ Massoud Pedram
Abstract
Activation steering adds a vector to a language model's internal representations at inference time. It is a lightweight alternative to fine-tuning for behavioral control. The standard construction, _contrastive activation addition_ (CAA), is fragile, as on many inputs it shifts behavior in the wrong direction. To explain this fragility, in this paper we first study the geometry of the steering vector. Specifically, we introduce **steering Hessian**, _the matrix of second order derivative of the model's behavioral loss with respect to the steering vector_. It captures the output behavior response sensitivity of the model under small perturbations of the steering vector along each direction. Its top eigenvectors are _sharp directions_: tiny perturbations along them can cause large, erratic, input-dependent behavioral swings. On the other hand, bottom eigenvectors are _flat directions_: behavior is robust along them. We identify that removing the sharp component of a steering vector approximately preserves its direction and magnitude (cosine $> 0.974$, norm reduction $< 3\%$ for typical settings) while reducing the vector's directional curvature by up to 12$\times$ (median 4$\times$ across 36 conditions). We evaluate on three model families (Qwen2.5-7B, Llama-3.1-8B, Mistral-7B), with four behavioral tasks, and 36 token-level conditions. Steering along sharp directions reverses the intended behavior in every condition tested, with anti-steer rates of 53--100\%. Sharp directions account for only 1--5\% of the CAA vector's squared norm. However, they drive a disproportionate share of the gap between the steered and unsteered next-token distributions during autoregressive generation. We term this one-line correction as _sharp removal_; it outperforms standard CAA in open-ended generation across all 25 conditions tested. A jailbreaking experiment further shows that sharp directions act as _noise_ rather than signal: they disrupt coherent behavioral control regardless of steering polarity.
Chat is not available.
Successful Page Load