Efficient Gradient-Aware Asynchronous Reinforcement Learning for LLM Post-Training
Abstract
Asynchronous reinforcement learning (RL) is an effective paradigm for improving the training efficiency of large language model (LLM) post-training by decoupling rollout generation from policy optimization. However, this decoupling introduces stale off-policy samples generated by earlier behavior policies, leading to distribution shift and unstable policy updates. Recent replay selection methods such as D-ARL mitigate this issue by selecting variance-aware samples, but require expensive current-policy evaluation over replay responses and mainly rely on variance reduction or current-policy matching as the selection criterion. To address these limitations, we propose GA-ARL, an efficient Gradient-Aware Asynchronous Reinforcement Learning framework for LLM post-training. GA-ARL derives an analytic target distribution from the KL-regularized RL objective and introduces an advantage-weighted variant to account for both policy preference and gradient strength. During training, GA-ARL maintains a replay buffer containing samples from recent behavior policies and selects low-discrepancy samples according to the advantage-weighted analytic target, avoiding additional current-policy evaluation during replay scoring. GA-ARL outperforms SOTA asynchronous methods across mathematical, logical, and code reasoning benchmarks, achieving the best average accuracy on both Qwen3-1.7B and Qwen3-4B, with improvements of up to 6.1%. Meanwhile, compared with the SOTA D-ARL, GA-ARL reduces wall-clock training time by 25.3% on average by avoiding additional current-policy evaluation during replay selection.