Instability of Meta-Learning Intrinsic Rewards for Policy Gradient Reinforcement Learning
Abstract
Meta-learned intrinsic rewards are a powerful tool for shaping policy learning in reinforcement learning, particularly when extrinsic rewards are sparse, delayed, or unavailable. Learning Intrinsic Rewards for Policy Gradients (LIRPG) introduced this paradigm and serves as the foundation on which subsequent meta-learned intrinsic reward methods are built. While LIRPG and LIRPG-based methods have shown strong results in low-dimensional control, their behavior in high-dimensional domains, where the policy and value networks use a shared encoder, remains poorly understood. In this work, we analyze meta-learned intrinsic rewards in high-dimensional environments and uncover a consistent failure mode, particularly when training with intrinsic rewards alone, where performance collapses to near-random behavior. We identify the root cause as dense, non-stationary intrinsic rewards inducing large and high-variance value losses that dominate shared encoder updates, suppressing policy learning. We further demonstrate that decoupling policy and value optimization using phasic policy gradient methods is one simple and effective approach to addressing this issue.