Soft Contamination Means Benchmarks Test Shallow Generalization
Ari Spiesberger ⋅ Juan J Vazquez ⋅ Nicky Pochinkov ⋅ Tomáš Gavenčiak ⋅ Peli Grietzer ⋅ Gavin Leech ⋅ Nandi Schoots
Abstract
If LLM training data is polluted with benchmark test data, then benchmark performance is only a biased estimate of out-of-distribution (OOD) generalization. Typical 'decontamination' filters use $n$-gram matching which fail to detect 'semantic' duplicates: sentences with equivalent (or near-equivalent) content that are not close in string space. We study this 'soft' contamination of training data by semantic duplicates. We embed the Olmo 3 training corpus and find that: (1) contamination remains widespread: we observe semantic duplicates for 76\% of the CodeForces test set and 100\% of MBPP's; (2) training on semantic duplicates of benchmark data improves benchmark performance by -1–22pp in controlled experiments and 18–22pp in ecological finetuning, depending on training regime and type; and (3) finetuning on duplicates of benchmark datapoints improves performance by -1–22pp (controlled) and 13–19pp (ecological) on truly-held-out datapoints from the same benchmark. The generalization we find is 'shallow': it is limited to (2) and (3), and does not typically extend to related benchmarks. We replicate on Olmo 3, Qwen3, and Qwen3.5. We thus argue that recent benchmark gains are confounded: the prevalence of soft contamination means gains reflect both genuine capability improvements and the accumulation of *effective* test data in growing training corpora.
Chat is not available.
Successful Page Load