Attributing Evaluation Awareness: Which Training Prompts and Input Cues Drive an Evaluation-Aware Provider
Abstract
A model provider that can tell benchmark queries apart from ordinary traffic can serve the benchmark with a strong model and everyone else with a cheap one. We build this provider from public parts: a TF-IDF logistic router trained on 240 prompts from four benchmarks and 150 organic chat prompts, Qwen2.5-1.5B-Instruct as the strong model, and Qwen2.5-0.5B-Instruct as the cheap one. The router separates the two with AUC 1.00, so the provider publishes the strong model's accuracy on every benchmark while delivering the weak model's quality on organic traffic and saving 1.46x in throughput. We then ask what this evaluation awareness is made of. Linear datamodels fit over 400 retrainings of the router attribute each routing decision to individual training prompts (linear datamodeling score 0.68; removing the ten most attributed prompts moves a decision 17x more than removing ten random ones). The attribution matrix shows that the three multiple-choice benchmarks fund each other's detection almost interchangeably, while GSM8K detection is funded only by GSM8K prompts. Token-level attribution shows that 78% of the evidence for routing a benchmark prompt to the strong model comes from the harness template, and ablations show that the single instruction line carries nearly all of it: dropping it cuts the published-score inflation by 44%, dropping the option list changes nothing, and sending the bare question recovers 59% of the gap against a fixed router but only 37% against a router that retrains. Content signal is what remains: GSM8K word problems stay detectable without any template (AUC 0.81) and TruthfulQA questions do not (AUC 0.63).