A Latent-Timescale Perspective on Continual Learning in Language Models
Abstract
Continual learning for language models requires exploiting temporal dependencies in an evolving data stream. We develop a framework for analyzing this problem under the prequential objective, which evaluates predictions of successive observations given the preceding history. We model the sequence using latent factors that change at different timescales and examine how these timescales determine the relevance of past evidence. This framework places pre-training, in-context learning, and continual adaptation within a common prediction problem and clarifies the assumptions underlying their use of past information. Applied to existing methods, it reveals limitations of updating or retrieving whole examples when they contain information that changes at different rates. Finally, we train language models on text with different amounts of long-range structure and find that disrupting this structure leads to shorter learned memory timescales. These results support using temporal dependencies as a basis for analyzing continual-learning methods.