LLM Rheology: Auditing Refusal Geometry in Aligned Language Models
Abstract
Scaling improves language-model capability, but it does not necessarily strengthen internal safety mechanisms. We introduce LLM Rheology, a representation-level framework for auditing how aligned language models respond to adversarial perturbation inside activation space. Given a learned refusal direction, we define manifold sensitivity / compliance as the Fisher--Rao-normalized distributional response induced by controlled activation perturbation. This quantity measures how strongly refusal behavior remains coupled to adversarial task execution in representation space. Across six open model families and 22 checkpoints, including Qwen, DeepSeek-Distill, Mistral, Llama-3, Gemma, and Yi, we observe a recurring geometric scaling pattern. Small aligned models often exhibit elevated refusal response, intermediate-scale models frequently enter weakened-response regimes (compliance valleys), and larger checkpoints trend toward near-baseline response consistent with increasing task--refusal decoupling. Cosine measurements on continuous Qwen and DeepSeek scaling axes support this interpretation, showing monotonic decay between refusal and task-generation directions with scale. We further provide causal evidence through inference-time activation intervention on Qwen-72B. Injecting a learned refusal vector at a critical semantic layer restores refusal behavior on the evaluated jailbreak subset while largely preserving benign reasoning performance. Matched-norm placebo vectors fail to reproduce the effect, supporting the directional specificity of the intervention. Together, these results suggest that behavioral refusal can remain surface-level even when internal refusal geometry becomes weakly coupled to adversarial task execution. LLM Rheology provides a complementary representation-level perspective for auditing refusal safety beyond behavioral evaluation.