Objectify, Don’t Delete: Complaint-Derived Evaluation of Ambient Clinical Notes
Abstract
Evaluating the quality of AI-generated clinical notes presents a fundamental challenge: standard metrics like ROUGE measure lexical similarity to reference notes, while expert checklists encode what committees believe matters. Neither captures what clinicians actually complain about. Recent work has shown that evaluation criteria can be mined from user feedback and filtered to criteria that LLM judges can reliably enforce. We argue that this is where most of the value is lost. The criteria that fail enforceability are disproportionately the ones clinicians raise most often because high-frequency complaints are stated as several questions glued together. No binary yes/no judgment can grade such composites as posed. We introduce an objectification step that rewrites such criteria into logically equivalent sets of determinate decisions. We demonstrate the approach through an evaluation pipeline that mines criteria from complaints, objectifies them, and deploys LLM judges against both planted defects and production notes. Across eleven planted defects in gold, objectified judges localize violations with median recall of 0.95, while reference metrics show changes <=0.030. An ablation reveals the necessity of objectification: posed as composites, the same judges localize every defect but also false-flag 70% of unperturbed reference notes (Youden's J 0.30 vs 0.72). We release ACI-Defect (a perturbation benchmark), clinician-reviewed instance-level labels showing quality gaps in the ACI-bench gold dataset (32 of 40 notes contain hallucinations; 40 of 40 omit discussed facts), and grounding judge verdicts. On 1,446 production notes, aggregate instance-rate burden falls monotonically as the clinician rating rises, supporting the design principle that defect evaluation must operate at the span level, not the note level.