GCD: Correcting Hidden-State Bias in Off-Policy Agentic RL
Changyuan Chen ⋅ Jianyu Xiang ⋅ Jiasheng Luo ⋅ Ziye Wang ⋅ NALLAPPAN GUNASEKARAN
Abstract
Off-policy reinforcement learning (RL) for large language models bottlenecks at the rollout stage: every parameter update invalidates the key-value (KV) cache of in-flight agent trajectories, forcing either expensive recomputation or staleness that destabilizes training. Recent advances such as M2PO and asynchronous RLHF tolerate token-level staleness by reweighting the policy gradient with second-moment importance corrections, but they implicitly assume that a stale rollout is sampled from a well-defined behavior policy $\pi_{\theta_{\text{old}}}$. We show that this assumption is silently violated by every modern partial-rollout system: once a stale KV cache is reused under updated parameters $\theta_{\text{new}}$, the resulting token distribution is a hybrid $\pi_{\text{hyb}}$ whose attention projections consume stale keys and values while its MLP, output, and gating projections use the new weights. The token-level importance weight $w_t = \pi_{\theta_{\text{new}}}(y_t \mid x_{
Chat is not available.
Successful Page Load