Persona Prompts Modulate Internal Emotion-Concept Dynamics in a Language Model
Abstract
Humans exhibit personality-dependent emotional responses to identical events. We ask whether language models show an analogous pattern: when conditioned on different persona prompts, do their internal emotion-concept representations follow different trajectories under repeated failure? We place Qwen2.5-1.5B-Instruct under twelve persona prompts spanning a 2 × 2 Extraversion × Neuroticism design (three paraphrase sets) in an adversarial guessing game engineered for five consecutive failures, tracking internal emotion-concept trajectories via cosine projections onto emotion-direction vectors extracted from an external dataset. Across all persona conditions, consecutive failures drive joyful projections downward and grief-stricken and furious projections upward; all three overall trends are significant after Holm correction. However, persona prompts systematically modulate these rates. Six of nine planned contrasts survive Holm correction: high Extraversion slows joyful decline and grief-stricken rise but accelerates furious rise; high Neuroticism accelerates joyful decline and grief-stricken rise, with a significant Extraversion × Neuroticism interaction on grief-stricken trajectories. These effects were directionally consistent across the three parallel wording templates, indicating that persona prompts modulate internal emotion-concept trajectories within the same standardized failure environment. This modulation is emotion-concept-specific rather than a uniform valence shift, and the overall pattern broadly parallels predictions from human personality–emotion models.