Actionable Feedback: Improving Model Capabilities From User Edit Data on Clinical Documentation
Abstract
Generative AI has transformed clinical documentation practices, giving rise to new paradigms of human-AI collaboration. Applications such as ambient scribes, which draft notes from patient-clinician conversations, and systems that summarize electronic health records, routinely involve a clinician in the loop, with AI tools producing an initial draft and clinicians performing edits to finalize the document. In principle, these edits provide an abundant and scalable source of post-deployment supervision. However, in practice, not all edits provide actionable feedback: for instance, naively fine-tuning on edits that add information unsupported by the input context risks teaching the model to hallucinate. In this paper, to study these practical challenges, we first develop a method for generating semi-synthetic clinician-edit data, given that real clinician edits on deployed documentation tools are often not publicly available. Then, we introduce simple filtering methods for identifying actionable edits. Our methods for generating and filtering edits are based on a taxonomy of clinician-editing behaviors, informed by data from a real-world deployment of an AI clinical documentation product. We demonstrate empirically, on a discharge-summary generation task based on MIMIC-III, that naive fine-tuning on edits indeed increases ungrounded content, while progressively filtering toward grounded edits enables improvement on clinician-alignment metrics while retaining faithfulness to the input context.