Conditioning Structure, Not Entropy Alone, Predicts Memorization Dynamics
Abstract
Autoregressive language models do not memorize every training sequence at the same rate; yet, commonly studied predictors, such as entropy, tokenization, context length, model capacity, and repetition, do not specify how a model uses its prefix to retrieve a continuation. We study a complementary axis: the complexity of the conditioning structure that disambiguates one training sequence from the others. In controlled GPT-2-style experiments, uniformly random prefixes with the same total information content are easier to use when their information is concentrated into fewer, more distinctive tokens. We construct synthetic datasets whose suffixes are deterministic under rules of increasing complexity: a local position-aware transition, a fixed prefix cue, or a position-dependent prefix cue. The resulting memorization curves preserve this ordering across one- and four-layer models and across dataset sizes from 500 to 24,000 sequences. In one-layer models, attention patterns visibly follow the planted cue locations. However, analogous interventions on WikiText yield no detectable separation, due to other interfering natural language cues. These results support disambiguation complexity as a useful controlled axis for studying memorization and, more broadly, training dynamics. Code and reproduction instructions are available here: https://anonymous.4open.science/r/memorization-llms-F1EF/README.md.