Adaptive Robust Estimator for Policy Optimization in Reinforcement Learning
Abstract
Reinforcement learning (RL) has become a key ingredient in post-training large language models (LLMs), with Group Relative Policy Optimization (GRPO) and its variants widely adopted for improving reasoning performance. Despite their success, these methods rely on batch-level reward normalization, where the empirical mean can be severely distorted by noisy, skewed, heavy-tailed, or contaminated rewards. Such distortion directly affects advantage estimation and may lead to unstable or inefficient policy updates. We propose an \emph{Adaptive Robust Estimator} (ARE), a plug-and-play robustification module for GRPO-style policy optimization. ARE replaces the empirical batch mean with a two-level robust location estimator: adaptive loss minimization suppresses extreme rewards within each block, while median-of-means aggregation limits the influence of corrupted blocks. This design preserves the original policy optimization objective while improving the reliability of reward centering. We prove that ARE is consistent and achieves high-probability deviation guarantees under both finite-variance and heavy-tailed reward distributions. Experiments on mathematical reasoning benchmarks show that ARE improves GRPO-based training in both single-agent and multi-agent settings, with particularly clear gains under noisy and out-of-distribution evaluations. We further validate its generality on embodied vision-and-language navigation tasks, where ARE improves training stability and downstream navigation (VLN) performance.