Tracking Memorization without Validation Data
Abstract
Large language models can learn to reproduce rare training sequences perfectly, creating privacy risks if those sequences contain sensitive or personally identifiable information. At the same time, a primary objective of learning is increasing the likelihood of training sequences. Detecting when a model shifts from learning generalizable patterns to memorizing training data traditionally requires continuously evaluating performance on a large holdout set. However, in data-constrained regimes such as supervised fine tuning, this incurs a severe sample-splitting penalty that reduces available training data and increases implementation complexity. We introduce a simulated holdout metric computed entirely from sequential training batches that circumvents the need for a separate holdout set and has minimal implementation and computation costs. Empirical evaluations show that our metric can identify the onset of memorization risk, enabling practitioners to halt training or take other action before dangerous behavior begins.