LLM Judges Verify Presence, Not Absence
Abstract
LLM judges score generated text at scale, and the benchmarks that measure them ask whether what a text says is correct, not whether what it should say is there. We measure that second capability where it is consequential and, unusually, constructible: ambient AI scribes writing clinical notes, where published human audits find omission the dominant error class. The standard check is an LLM judge, and we ask whether it detects omissions. Public corpora could not supply the ground truth - every clinician reference note we audited is materially discrepant with its own transcript - so we release 500 single-error pairs built from transcript-derived, audited fact sheets - 298 in which a named fact is certainly absent, graded by severity and by surviving trace, against 202 added-or-altered controls. Across eight judge designs, added or altered content is detected at 0.79-0.94 paired discrimination (0.5 is a coin flip) against 0.50-0.63 on omissions, and on single notes all eight flag notes with omissions about as often as perfect ones. Wording changes, voting, and prompt optimisation move the operating point without creating detection. What recovers detection is restructuring the task - list the facts the transcript establishes, then verify each against the note as a presence check. Two methods reach it separately, a per-fact pipeline and an optimiser-evolved prompt running the same procedure in one call, and they trade off. The pipeline's flags name the missing fact and its severity at 2.7% false alarms; the single call detects more notes (36.9% against 24.6%, p = 0.002 on the evaluation set) at 6.2% false alarms and roughly a tenth of the cost per note (a third amortised). On real vendor notes from a companion census no benchmark threshold transfers, though the single call re-calibrated still detects more than the best of the eight at half its false-alarm rate, and three quarters of the pipeline's flags on verified-omission notes name the fact the census verified. Omissions whose restatement survives elsewhere in the note defeat both. A physician author validated 70 items in a sitting that covered six stages, three of them structurally blinded, and sided with the per-fact pipeline on 10 of 10 judge disagreements; an independent clinician has since graded the rubric blind, agreeing to within a grade. We release the benchmark, prompts and judgements.