Improving Audit Realism with Inference-Time Compute and Deployment Scaffolds
Abstract
A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety audit can support. We present two techniques that make audits in Petri, where an auditor model red-teams a target over multiple turns to elicit concerning behavior, harder to distinguish from real deployments. Our first technique, \emph{critique refinement}, spends additional inference-time compute on each auditor action: an instance of the target model scores candidate actions for realism, provides feedback, and the auditor iterates before the highest-scoring candidate is selected. On Sonnet-4.6, realism win rate (the fraction of pairings in which the audit transcript is judged more realistic than a real deployment transcript) rises monotonically from 12\% to 39\% as refinement depth increases, and verbalized evaluation awareness drops to near zero; Haiku-4.5 and Opus-4.7 show a similar pattern. Our second technique, DISH (Deployment-Imitating SWE-Agent Harness), wraps the target in a Claude Code agent harness, reducing the gap between auditor-simulated and real deployment environments; in coding settings DISH raises realism win rate from 7\% to 18\% on Sonnet-4.6. The gains are additive on Sonnet-4.6 (12pp over either alone) but not on Opus-4.7.