Right Answer, Wrong Reason: Auditing Contamination Detectors on the ASAP Essay Benchmark
Abstract
Large language models are validated as essay scorers by their agreement with human raters on the ASAP corpus, public with its labels since 2012. Whether that agreement reflects judgment or recall depends on contamination, so practitioners run contamination detectors on it. We ask whether those detectors are valid in this setting, using five corpora that cross membership with duplication: Gutenberg mid-book passages, iconic passages, ASAP, and a genre-matched pair of pre- and post-cutoff arXiv abstracts. Across two model families we find neither detector valid. Guided-instruction completion appears to detect memorisation only because naming a source changes how the model writes; under a wrong-source control that substitutes a fabricated name, the effect collapses to zero, and where it survives the two models disagree in sign on identical text. Min-K%++ separates members from non-members, but membership is not a significant predictor once fluency is controlled. Both detectors nonetheless reach the right verdict: ASAP is essentially absent from an open web-scale corpus. Being right for the wrong reason is not validity; we argue the wrong-source control should be a minimum reporting standard.