Adaptive Stepsizes for Eligibility Traces in Deep Reinforcement Learning
Abstract
Modern deep reinforcement learning algorithms often store past sensory observations in a buffer for later replay and learning. However, this process is biologically implausible; animals do not store raw sensory observations. Instead, they store their imperfect understanding of the world and the events in it. An alternative to storing all past sensory observations is to learn directly from the observations as they come in and store the seemingly important parts of the incoming data. Eligibility trace algorithms in reinforcement learning (RL) sweep back updates for multiple steps, improving efficiency and providing update rules that do not rely on storing large data buffers. These algorithms were common and effectively used in the linear setting, but have not been widely adopted in the deep RL setting. Many questions remain open to make these algorithms more effective, even as basic as how to incorporate vector step-sizes. We first highlight how adaptive vector step-sizes need to be incorporated into the eligibility trace and then introduce a new vector step-size algorithm that is more amenable to the forward-backward view equivalence needed to derive an update with eligibility traces. We show that our new algorithms improve over other streaming algorithms across Mujoco and MinAtar environments, and are comparable to methods with replay buffers such as PPO.