Replay Buffer Distribution Dynamics Matter in Massively Parallel Off-Policy RL
Abstract
GPU-accelerated simulators enable off-policy reinforcement learning with thousands of environments in parallel. Such parallelism increases data throughput, but also changes how experience from successive policies is represented in the replay buffer. We distinguish the temporal composition of replay data over changing policies from the finite-sample resolution with which that history is represented. By preserving the temporal composition while increasing environment parallelism and replay-buffer size, we isolate how finite sampling affects learning. Experiments spanning a controlled diagnostic and massively parallel humanoid locomotion show that learning curves approach common dynamics as finite-buffer gradient variance decreases. These results are consistent with the eventual saturation of the benefits of additional parallelism and provide a data-distribution perspective on massively parallel, simulation-based robot learning.