The Benchmark Blind Spot: Lag Profiles of Distributional Model Collapse and a Low-Cost Calibration-Gap Signal
Abstract
Recursive fine-tuning on purely synthetic predecessor output, the replace paradigm, degrades perplexity, diversity and distributional divergence, while standard likelihood-scored benchmarks stay flat or improve. We measure that gap across two scales (GPT-2 345M, SmolLM2 1.7B), two decoders (top-k, greedy), 8-13 generations per condition and 3-5 seeds. Perplexity, KL-unigram and Distinct-2 fire at generation 1 in every condition and survive Holm correction thereafter. At 345M under top-k, the only benchmark fire we count as a detection appears at generation 6 and does not survive correction, and at 1.7B under top-k, no benchmark signals within twelve generations. Benchmark rises are an early-collapse transient whose window runs long under top-k and shrinks to one or two generations under greedy, closing entirely in the most severe condition, where benchmarks instead fall sharply and fire at generation 1. We propose the calibration gap (log PPL minus predictive entropy) as a candidate monitor. It opens monotonically from ≈ 0 in all four conditions, is register-graded, and requires no forward pass beyond the perplexity computation itself.