Gradient Routing Localizes and Removes Unintended Behaviors in RL
Abstract
We propose a novel solution to reward hacking in reinforcement learning. Our setup involves a misspecified reward function corresponding to task completion, and a classifier that flags invalid solutions. As in frontier LLM development, the classifier makes systematic errors. Consequently, reward penalties and data filtering based on the classifier fail to prevent reward hacking. To address this limitation, we introduce GRAFT: Gradient Routing Adapter Fine Tuning. Rather than penalizing or filtering flagged behaviors during training, we constrain (``route'') gradient updates to specific parameters. This has two benefits. First, the model does not learn to exploit classifier errors, i.e., it does not learn to hide the unintended behavior. Second, the unintended behavior can be disabled by ablating the parameters to which they were localized. We demonstrate these benefits in a variety of RL environments where reward penalties result in monitor-subverting policies. Additionally, we show that our method remains effective when we scale model size and task complexity by distilling frontier LLM trajectories into Qwen3-32B. Overall, we show that GRAFT is a promising approach for leveraging unreliable monitors during post-training in environments where true performance is difficult to evaluate.