Execution-Grounded Measurement of Verification Blindness in Multi-Agent LLM Systems
Atharv Santosh ⋅ Rohan Sashank Babbellapati ⋅ Phil Mui
Abstract
Verifier agents are widely used to check artifacts produced by LLM coding systems, but which faults they reliably detect, and whether their specific diagnoses can be trusted, is unclear. Evaluating one language model with another is also circular, since the evaluator may share the weaknesses of the system it evaluates. We address this with DEFECTBED, an execution-grounded benchmark that establishes machine-checkable ground truth, categorizes defects by how much evidence is available in the artifact text, and pairs defective artifacts with matched clean controls. We evaluate 14{,}713 episodes across eight verifier configurations and two testbeds. Verifiers distinguish defective from correct artifacts when the fault is visible in the text, with Youden's $J \geq 0.90$ across T0--T2, while their ability to localize the fault falls from $0.764$ at T0 to $0.218$ at T2. Clean controls expose a substantial false-alarm problem: at the cross-artifact tier verifiers report defects on 59.3\% of correct artifacts, so sensitivity alone overstates performance by a wide margin. Among affirmative claims on defective artifacts, confidence is substantially better calibrated for partial diagnostic correctness than for exact fault identification, with expected calibration error of $0.101$ against the former and $0.745$ against the latter at T2. Within tool-enabled runs, executed episodes show higher strict credit at T2 ($+0.071$) and T3 ($+0.117$), little difference at T0--T1, and lower credit at T4 ($-0.176$); this comparison is observational, and T3 carries a visible-cue confound. We recommend treating verifier alarms as triggers for further inspection and pairing them with execution-grounded validation, rather than relying on them for fault localization.
Chat is not available.
Successful Page Load