From Representation to Intervention: Using Emotion Vectors to Monitor and Guide Language Models
Abstract
Recent work has shown that large language models form linear representations of emotion concepts that causally influence alignment-relevant behavior, including blackmail, reward hacking, and sycophancy. These findings raise a natural question: can such representations be used not only to characterize model behavior in controlled evaluations, but also as practical tools for monitoring and guiding models during deployment? We explore this possibility, investigating how emotion-vector activations can serve as interpretable signals of a model's internal state over the course of a conversation or agentic task. We discuss how these signals might support debugging, early detection of drift toward misaligned behavior, and lightweight steering interventions to nudge the model back toward desired behavior. Beyond the Assistant, the same representations apply to other characters in the model's context — including the user — suggesting a broader family of applications in which internal representations of emotional state inform how systems respond. We outline the opportunities and limitations of this approach and highlight open questions around validation, robustness, and responsible use.