What Do Evaluation Awareness Metrics Measure
Abstract
Studying evaluation awareness in large language models (LLMs) is critical in order to predict bad behavior in deployment as LLMs can often distinguish between deployment and evaluation contexts. We study how evaluation awareness is measured by introducing a framework that distinguishes two provisional constructs: \textit{context identification}, whether a model can infer that it is being evaluated, and \textit{functional availability}, whether that information is available to its ongoing computation. We evaluate the construct validity of nine metrics spanning black-box and white-box approaches. Agreement is weak even among metrics assigned to the same construct; the strongest convergence is between a linear probe and SAE readout trained on the same evaluation/deployment labels. Moreover, these classifiers achieve high accuracy even though the same labels are highly predictable from the input text alone, showing that successful discrimination does not by itself establish an internal evaluation-awareness representation. Our results suggest that current metrics are not interchangeable and that progress requires validating both the measurements used to study evaluation awareness and the conceptualization those measurements operationalize.