Continual Learning as Bayesian Filtering
Abstract
Real data streams are rarely independent and identically distributed (iid): class frequencies, contexts, and label semantics can drift over time. In task-free online continual learning, the learner must adapt to this drift without task boundaries, replay schedules, or episode identifiers. We argue that this setting is best understood as approximate Bayesian filtering. Under a latent-state model of the stream, the Bayes-optimal causal predictor is the posterior predictive distribution, and catastrophic forgetting arises when the learner's state is too impoverished to represent the filtering posterior. This view reinterprets common continual-learning methods as restricted belief-state approximations: regularization preserves local precision, projection methods preserve constrained subspaces, replay approximates the posterior predictive in data space, and parameter-isolation methods avoid shared belief revision. From this filtering view, we derive a Mori--Zwanzig form of the Bayes filter and obtain a posterior-memory learning rule for retaining hypotheses that finite belief states would otherwise discard. In controlled one-hidden-layer neural-network streams, the resulting Mori–Zwanzig filter improves over the corresponding single-Gaussian closure and standard continual-learning baselines, with gains attributable to its ability to preserve competing posterior hypotheses.