Why Different Reset Sets Change Greedy Actions in Neuron Recycling: An Action-Boundary Analysis
Abstract
Neuron recycling can restore plasticity in continual reinforcement learning, but resetting several neurons together can change greedy actions and potentially disrupt learning. At a fixed reset-set size, these changes depend on which neurons are reset together. We call this the set effect. We define immediate disagreement as the proportion of states whose greedy action changes after a reset. With a linear action head under ReDo's reset, we can determine exactly whether removing the selected neurons' current Q-value contributions changes the greedy action. Across disjoint replay batches, directly evaluating each reset set produces more consistent rankings than scoring sets by the sum of individual-neuron disagreement scores. Controlled action-gap scaling further supports boundary crossing as the mechanism behind the set effect. As an initial application, the Action Boundary Balancer (ABB) uses this mechanism to select low-disagreement reset sets of adjustable size. In a continual Double DQN experiment, ABB substantially reduces ReDo's immediate disagreement on both held-out replay states and future states encountered after reset, while preserving much of ReDo's improvement in time-averaged return over no reset.