A dynamical systems theory of reward-modulated learning in linear recurrent networks
Abstract
We derive closed-form ODEs for the learning dynamics of a linear recurrent policy network trained with REINFORCE on sparse-reward tasks in a high-dimensional teacher–student setting. Under alignment assumptions between student and teacher networks, the dynamics reduce to a finite system of coupled order-parameter ODEs whose reward terms are governed by trajectory-level orthant probabilities of a structured correlated Gaussian process. The theory predicts that recurrence can accelerate escape from sparse-reward plateaus by accumulating task-relevant memory, and that each episode length admits an optimal recurrent scale that minimizes learning time. Our results provide a quantitative theory of how recurrence, and sparse reward interact during learning.