Clinical Scribe Audits Measure Note Length and the Reader, Not Only the System
Abstract
An audit is how a clinical scribe is cleared for use: clinicians read generated notes and count what the note invented and what it left out. Those counts rank systems and decide purchases, and they are read as measurements of the system. We re-analyse a public corpus of clinician annotations to ask what else is in the counts: 1,140 ratings of notes from ten variants of one 2022 summariser reading manual transcripts of 57 acted consultations, plus 285 ratings of clinician-written notes. Two things are, and both can be measured: how much the system wrote, and who read it. Within a system, one added line carries +0.128 fabrications (95% CI [+0.057, +0.190]) and -0.158 omissions [-0.235, -0.088], and the total an audit records does not detectably move — though a reader catches omissions at about half the rate of fabrications, and the cancellation does not survive correcting for that. The omission half is a conditioning effect: unadjusted the association is absent, and only holding encounter, reader and system fixed makes it negative. For fabrications the slope survives a blind control the corpus already contained: the same clinicians rated human-written notes shuffled among the machine ones, and no such slope appears there, with a pooled arm-by-length interaction of +0.149 [+0.066, +0.232]. The same test on omissions is inconclusive, so that half is uncontrolled. The reader is the second thing. One clinician's count agrees with another's at an intraclass correlation of 0.50 for fabrications and 0.27 for critical omissions. Scoring a single note at a reliability of 0.8 takes four readers, or eleven for critical omissions, and this corpus's engineered spread of systems makes both underestimates. Within one panel, more notes sharpen a ranking; only more readers sharpen a rate. Report both error counts with the note-length distribution, and size the reader panel to the number you intend to quote.