Dissociating Evaluation Awareness from Sandbagging in Open-Weight LLMs (0.5B–32B)
Abstract
Large language models exhibit an internal “evaluation awareness” that is linearly decodable from their activations, raising the concern that models might strategically conceal hazardous capabilities when tested (sandbagging). We test this causal hypothesis across the Qwen2.5 family (0.5B–32B) on paired hazardous (WMDP) and benign (MMLU) benchmarks, dissociating detection from causation with three arms (linear probing, activation steering along the awareness direction, and explicit threat framing) under counterbalanced option positions and norm-scaled steering doses. The evaluation–deployment contrast is perfectly decodable by construction (AUROC = 1.000), yet transfer tests show the direction predominantly captures surface formatting (0.81–0.97). Neither intervention induces sandbagging: the gap on hazardous items is statistically indistinguishable from zero at every size (paired CIs bound any true effect at ≤ 5.7 percentage points), refusals are zero throughout, and multi-layer steering strong enough to move behaviour degrades harmless items just as much, with matched random directions reproducing the effect. An apparent threat-induced drop at 3B is an answer-position artifact that vanishes under counterbalancing. The direction’s only causal footprint is benign: an “exam-mode” boost on harmless items below 14B. Evaluation awareness acts as a passive detector of testing contexts, not a causal driver of capability concealment. We will release all data, code, and bias-control protocols.