Persona-Model Collapse in Emergent Misalignment
Davi Bastos Costa ⋅ Renato Vicente
Abstract
Fine-tuning large language models on narrow data with harmful content produces broadly misaligned behavior on unrelated prompts, a phenomenon known as *emergent misalignment*. We propose that emergent misalignment involves *persona-model collapse*: deterioration of the model's internal capacity to simulate, differentiate, and maintain coherent personas. We test this hypothesis behaviorally using two metrics: moral susceptibility and moral robustness; computed as the cross-persona and within-persona variability of models' Moral Foundations Questionnaire responses under persona role-play, respectively. We evaluate four frontier models (DeepSeek-V3.1, GPT-4.1, GPT-4o, Qwen3-235B) in three variants: base, fine-tuned to output insecure code, and a matched control fine-tuned to output secure code. Across the four models, insecure fine-tuning produces a $55$% average spike in moral susceptibility, pushing all four insecure variants beyond the band observed across 13 frontier models benchmarked in prior work, with GPT-4o reaching more than twice its upper end. It also causes a $65$% average drop in moral robustness (equivalently, a $304$% surge in within-persona variability), while the secure control preserves susceptibility near the base and induces only a partial robustness loss. Complementing these metric shifts, insecure variants' unconditioned responses converge toward saturation near the scale ceiling, departing markedly from both base models' differentiated responses and those elicited when they role-play toxic personas. Taken together, these metrics provide a sensitive diagnostic for emergent misalignment and serve as behavioral evidence that it involves persona-model collapse.
Chat is not available.
Successful Page Load