Benchmark Shadows: How Data Regimes Shape Parameter Footprints and Generalization
Abstract
Large language models can achieve strong benchmark scores without proportional gains in broader capability, but diagnosing when this occurs remains difficult. We study this gap through benchmark shadows: support-concentrated data regimes that may concentrate learning around narrow evaluation-relevant patterns while limiting broader representational development. Under fixed architecture, tokenizer, optimizer family, and training budget, we compare a coverage-expanding baseline with two support-concentrated regimes: redundant repetition and frequency-concentrated rewriting. Spectral, rank-based, and layer-wise update diagnostics reveal distinct parameter-space signatures: repetition-concentrated training is largely recoverable after later diverse training, whereas frequency-concentrated support collapse leaves more persistent footprints. Correlational analyses of open-source multimodal model families reveal analogous structural patterns that co-occur with asymmetric benchmark profiles, while a prompt de-duplication case study shows that surface redundancy alone does not induce the same regime-level effects. Benchmark scores alone are therefore insufficient to characterize capability; parameter-space diagnostics provide complementary signals about training quality, coverage, and generalization-related regime effects.