Measuring and Strengthening Behavioral Suppression in Language Models
Abstract
Undesired model behaviors can emerge or re-emerge at any stage of training. Once observed, how effectively can we suppress them? To study this question, we first construct two model organisms---language models trained to exhibit a specific misalignment---with qualitatively distinct structures. In the narrow-trigger organism, undesired behavior is elicited by an identifiable prompt feature. In the broad-trigger organism, misalignment surfaces diffusely across open-ended responses without a separable cue. We then formalize post-hoc behavioral suppression and evaluate methods along three criteria: immediate suppression of the target behavior, durability under further training, and preservation of general capability. We compare standard methods---SFT, GRPO, and gradient-ascent unlearning---along these dimensions while controlling the number of gradient updates. Building on prior work on alignment shallowness, we hypothesize that more durable suppression requires the ability to recover from a wider range of misalignment states. We introduce Prefix-Expanded Adversarial Reinforcement Learning (PEARL), a GRPO variant that augments rollouts with continuations from intermediate states sampled along cached organism trajectories. Across both organism settings and two model families (Qwen3-4B, GPT-OSS-20B), PEARL achieves lower exploit rates while preserving task accuracy, and yields lower reactivation rates under both targeted and benign capability fine-tuning than SFT and GRPO baselines. Together, our framework and method offer a step toward more durable post-hoc removal of undesired behaviors.