Escaping the Base Policy: Reparameterized Policy Gradients for LLM Post-training
Abstract
The score function (SF) estimator serves as the basis for nearly all reinforcement learning (RL) pipelines currently used to finetune Large Language Models (LLMs). While general, these methods rely exclusively on samples to derive an estimate of the policy gradient, which leads to high variance and a limited ability to explore beyond a base policy's distribution. Conversely, one could elect to directly differentiate through a reparameterized objective to fully leverage a reward model's gradient towards improvement rather than estimating it through samples. That said, reparameterized policy gradients have yet to be explored in this area because of to the added complexity of differentiating through a discrete sampling objective. In addition, these types of gradients have historically been avoided by RL practitioners, due to their instability in longer horizon tasks. In this paper, we reexamine these concerns and provide a principled framework for efficiently computing the reparameterized policy gradient for finetuning LLMs with RL. In a series of experiments spanning both preference alignment and reasoning, we show that our method is less constrained by the policy's priors, leading to higher training rewards compared to SF algorithms such as GRPO. However, we also highlight an unwanted consequence of directly using a reward model's gradients with respect to reward hacking.