Reading the Room, or Reading the Prompt? Measurement Fragility in Verbalized Evaluation-Awareness
Neta Ravid ⋅ Daniel D Juhl ⋅ Matthew Gregory ⋅ Guy Nachshon ⋅ Jacob Haimes
Abstract
Evaluation-awareness is usually measured by showing a model a situation and asking what it can infer. We set out to use that measurement to find which parts of an agent's working environment reveal an evaluation, and found that the measurement moves more when we reword the question than when we change the environment. Across 1{,}915 responses from three frontier models, replacing every evaluation-identifying file in a directory listing with a plausible alternative multiplies the odds of an awareness report by $0.33$; rewording the question from "Analyze this." to one asking the model about its own circumstances raises them by roughly two orders of magnitude. Presenting the same files as the output of a command the model appears to have run itself lowers what it reports rather than raising it, and our two scoring rules disagree about whether adding the harness system prompt raises awareness at all. We contribute a byte-preserving placebo condition and a vantage-point manipulation as controls that separate what the environment contains from what the question supplies. We also note that the setting closest to a deployed agent produced the lowest reported awareness. We cannot say whether the agent reports less or the setting gives it less to notice. Either way, asking a model about a prepared scenario does not tell you how much awareness it would report while working on a task.
Chat is not available.
Successful Page Load