Calibrated Artifact Auditing of an Agentic Physics Workflow
Abstract
Agentic AI systems produce scientific results through long autonomous workflows. The verification layers they carry are their own, namely self-assessments, plausibility checks and generated reports. We show that these layers can carry no signal. In the public artifact corpus of an autonomous detector-design agent, all 64 preserved physics-verification records return \texttt{plausible: true} with empty issue lists. Those records span runs whose headline results include a retyped constant reported as a 204.62\% improvement, a hardcoded detection efficiency of 1.0 and a placeholder narrated as physics. We propose \emph{calibrated artifact auditing}. Every claim is re-derived from the most primitive surviving evidence instead of the agent's intermediates or self-reports. A failure diagnosis is accepted only if it reproduces the claimed number to all quoted digits and survives an empirically measured false-positive rate on decoy targets. The auditor itself is calibrated by fault injection, so it ships with measured detection profiles and a characterized blind spot in place of an implicit completeness claim. A five-stage implementation audits the full eighteen-run corpus (3{,}142 numeric claims) in under an hour on a laptop, rediscovers every manually established failure mechanism from primitive evidence, corrects one of them, and exposes two analyses that consumed stale data existing nowhere in the recorded provenance. The last finding rests on a moment-inversion technique that reconstructs an analyzed file's summary statistics from archived error formulas. We state the method's provenance preconditions and release the pipeline.