Emotion Probes Move Classification Without Reliable Behavioral Control: A Cross-Family Study of the Perception–Action Gap in LLMs
Abstract
Linear probes decode emotion from the hidden states of large language models (LLMs) with high accuracy. It does not follow that the decoded representation con- trols what the model does. We test two questions across ten open-weight decoder- only models from five families (1B–14B parameters). First, are appraisal-aligned directions the directions along which emotion classification can be moved? A 14-dimensional appraisal-probe space separates confusable within-cluster emotions in all ten models, but at matched dose the binary one-vs-rest emotion coefficient moves the probe-fused classifier about 49×more than the appraisal direction. The intervention vector is also a component of that classifier, so we use this result as a manipulation check and not as evidence about emotion content. Second, does a shifted classification reach behavior? In Llama-3.2-1B, all 301 benchmark prompts change their latent readout and none change refusal. Seven of the eight models with matched generation runs show no refusal change. Ministral-8B is the excep- tion, and only at α= 30, where a held-out wrong-emotion direction reproduces about 90% of the drop. An eight-dimension judge panel, sycophancy prompts, and TruthfulQA show the same effect is channel-specific. We call the dissociation the perception–action gap. Whether post-training produces it remains an open question that our data cannot settle.