Decoupling is the Key: Scaling Deep Value Networks in Reinforcement Leanring
Abstract
Scaling laws have driven remarkable performance breakthroughs in Computer Vision (CV) and Natural Language Processing (NLP) by increasing model depth. Compared to increasing width, deeper networks provide higher parameter efficiency and more expressive representations. However, increasing network depth in reinforcement learning (RL) still yields limited performance gains. In this paper, through theoretical and empirical analysis, we reveal that this performance degradation is primarily attributed to implicit low-rank bias and overfitting. While these two issues also exist in shallow networks, the detrimental effects caused by component coupling are significantly amplified as the network depth increases, leading to severe performance degradation. For the first issue, we theoretically prove the existence of bootstrapped spectral coupling, which causes high-frequency spectral components to couple with the low-frequency spectral components of the value network, thereby driving representations of deep value networks into severe rank collapse. For the second issue, we reveal that deep value networks severely overfit the noise induced by TD target couplings, which means the construction of TD targets is influenced by other components, such as the boundary of value distribution or current policy. Therefore, our insight is that decoupling is the key to successfully scaling deep value networks in RL. Motivated by this insight, we propose two simple yet effective solutions respectively: policy-independent spectral decoupling and boundary-target decoupling. Moreover, decoupling policy and value learning (as in AWR) is also a necessary target decoupling for deep value network training in offline RL. By integrating these decoupling components, we propose Spectrally Decoupled Distributional Learning (SEED). Empirically, to the best of our knowledge, SEED is the first approach to successfully unlock the potential of depth in both online and offline RL settings, enabling value networks to scale effectively from 3 to 32 layers with consistent performance gains.