What Does a Passed Evaluator Audit Establish? Observation Sufficiency for Post-Repair Claims
Mateo PETEL
Abstract
A repaired evaluator can pass the audit that motivated the repair and leave a distinct assurance property unresolved. We study when post-repair observations are sufficient to support claims beyond the audited failure set. Let $\mathcal R(\mathcal A)$ be the evaluator configurations compatible with stated assumptions $\mathcal A$, $\mathcal O$ the complete declared audit observation, and $H$ a target property. The audit identifies $H$ only when observationally equivalent configurations cannot disagree on $H$. We instantiate failure with an executable finite countermodel: a source-local repair and a joint source-and-held-out repair produce exactly the same public $(\mathrm{accept},\mathrm{score},\mathrm{reason})$ transcript on all declared source witnesses and a benign control and differ on clearance of a separately declared held-out witness set. Any statistic computed solely from that transcript leaves the target distinction unresolved. A separate finite invariance audit closes another local obligation exactly. In a companion study, three named GPT-5.6 evaluators complete 1,152 trials on 64 fresh arithmetic problems. For \texttt{gpt-5.6-terra}, source-authority errors change from 34/64 to 3/64 and clean errors from 2/64 to 1/64 under a targeted prompt intervention; held-out plausible-reasoning errors change from 8/64 to 4/64 and remain nonzero. Sol and Luna fall outside the same diagnostic gate. The formal result is an exact constructed certificate; the model study contributes a finite, model-specific companion observation scoped to the study protocol.
Chat is not available.
Successful Page Load