Towards Robust, Performance-Based Contamination Detection for LLMs
Allenna Tang ⋅ Jackie CK Cheung
Abstract
Data contamination---the exposure of benchmark items during training---inflates LLM benchmark performance, and being able to detect it is needed for trustworthy evaluation. An adversarial model provider has incentives to inflate scores while evading contamination audits, and the robustness of current detection methods to such an adversary remains poorly understood. In this work, we show that current detectors can be cheaply evaded and propose a performance-based detection paradigm that is invariant to the evasion by construction. Our evasion adapts confidence masking from membership inference defenses to LLMs: an argmax-preserving logit transform that makes signals on contaminated items resemble those on unseen items, while leaving every benchmark answer unchanged. On Llama-3.1-8B-Instruct fine-tuned on MMLU items, the transform reduces confidence-based detectors (loss, zlib, Min-K\%, Min-K\%++) from near-perfect detection (AUROC $> 0.93$) to chance ($\approx 0.5$). We then introduce a detector based on Item Response Theory (IRT) residuals, which calibrates item difficulty on a population of 102 LLMs and flags items where a model outperforms its estimated ability. Because it reads only binary responses, its AUROC is exactly unchanged under the evasion. We advocate for further exploration into performance-based contamination detection methods as a step toward more trustworthy evaluation tools. More broadly, we suggest evaluating detection robustness through targeted efforts to attack current detection methods, as is done in the membership inference attack literature.
Chat is not available.
Successful Page Load