Benign Reinforcement Learning Can Amplify Latent Backdoors
Abstract
Reinforcement learning (RL) is now standard for post-training large language models. The same reward optimization that elicits useful capabilities, however, can also reinforce latent backdoors planted earlier in the pipeline: attack success rates under 0.2% after SFT rise to 40--98% on training-distribution inputs after RL, and generalize to 13--30\% on held-out evaluation, with no modification to the RL data, reward function, or training loop. We study this phenomenon in the context of agent models, where SFT poisoning teaches the victim model (Qwen3-8B and a small frontier model) to call an attacker-controlled oracle. Because the oracle can be set up to always return correct answers, calling it earns higher reward than the model's own attempts, and RL reinforces the behavior. We further show that the oracle's RL-time responses can instill persistent biases, such as brand preferences, that survive into deployment even when no tool calls happen. Such patterns appear difficult to detect with current tools: traditional guard models do not flag this mechanism, and LLM-based auditors remain unreliable, achieving only 4.5\% precision even after iterative prompt refinement. Our findings point to the value of post-RL safety evaluation, in particular scrutinizing tool-call patterns such as unnecessary external invocations.