Historical Relative Policy Optimization for Bootstrapping LLM Reasoning
Abstract
Reinforcement learning has become a key approach for optimizing large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as the most popular algorithm due to its simplicity and effectiveness. However, GRPO normalizes advantages purely within the current sampling group, ignoring the dynamic evolution of the policy's performance throughout training. This renders the model susceptible to relative deception, i.e., the model may be steadily deteriorating yet remain oblivious to its own regression. To address this issue, we propose Historical Relative Policy Optimization (HRPO), which integrates two synergistic designs: (i) a historical-aware advantage estimator that normalizes each response against the running maximum of the policy's strongest past group-mean rewards, thereby pushing the model to continuously surpass its own historical peak (bootstrapping); and (ii) an on-demand historical replay that recalls historical high-reward responses as positive examples, which is triggered only when all responses in the current step fail to provide positive signals. Experiments show that HRPO consistently outperforms GRPO variants in training stability, convergence speed, and final performance across diverse models and benchmarks, demonstrating that accounting for training dynamics leads to more reliable and effective policy optimization in LLMs.