When Trend Looks Like Recall: An Oracle Control for Contamination Audits of Temporal Benchmarks
Abstract
A contamination audit asks whether a benchmark is already in a model's training data, and its default instrument reads the model through generated text alone: prompt, parse, compare, record a miss as clean. On the dates such an audit certifies clean, a candidate-conditioned log-probability readout, which asks whether the reference is more likely than one real alternative from the same series, can still prefer the reference, so a clean verdict is a property of the readout. But on smooth date-indexed series that second readout cannot be trusted either. On Mauna Loa CO2 a local interpolation oracle that holds the queried year out reaches 1.00 against 8 of 10 distractor constructions, reproducing, without reading the queried value, the matched-versus-mismatched contrast (+0.31 at 1.5B, positive at all three scales) that would otherwise read as recall; the usual mismatched-prompt calibration does not remove it, because interpolation is itself date-conditional. On U.S. unemployment, where the oracle falls short of ceiling, the contrast stays within noise at every scale; mountain elevations, which cannot be interpolated, supply the positive control (+0.41). Temporal contamination claims need an oracle control; without one, neither readout establishes absence.