Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase
Abstract
Replay-based continual learning rehearses past data either in an offline phase, during which the agent stops acting, or interleaved with the live stream, where it perturbs the computation serving the current input. An agent that learns in deployment can afford neither. Motivated by local sleep, the use-dependent off periods of individual cortical circuits in awake animals, we show that replay can instead be written into the degrees of freedom the current input leaves unused. In a network with k-winner-take-all hidden layers, confining replay updates to synapses whose presynaptic unit is silent or whose postsynaptic unit is inactive leaves the hidden computation on the current batch invariant: exactly so for silent and suppressed units, and for all but 0.3% of samples in practice. A refractory rule under which units that have just fired sit out the next competition doubles the width of this channel and carries most of the accuracy. On class-incremental split-MNIST the resulting learner, with no offline phase, matches or exceeds the best offline rehearsal schedule and outperforms experience replay, ER-ACE and unmasked interleaved replay, each re-tuned under the same micro-batch schedule. Against DER++ the comparison splits by protocol: with five epochs per task DER++ leads by 1.5 points once it runs under that schedule, and in a single pass over the stream, the regime closest to the agent deployment setting, the system leads it by 1.6 while offline rehearsal falls 15 points behind. On split CIFAR-10 it again leads offline rehearsal and experience replay, but trails ER-ACE and DER++ by two to three points. The construction is not tied to the local learner: on a backprop network with k-winner-take-all hidden layers under the same schedule, refractory rotation adds half a point in five epochs and three in a single pass, and isolation again costs nothing on top of it.