DUC-Bench: A Guideline-Grounded Benchmark for Evaluating Evidence-Responsive Decision Updating in Medical Large Language Models
Abstract
Clinical decision support requires more than producing a correct answer from a fixed vignette: a reliable model must revise its recommendation appropriately when the evidential basis of that recommendation changes. Yet current medical LLM benchmarks largely evaluate decisions at a fixed information state, leaving evidence-conditioned revision underexplored. We introduce DUC-Bench, a guideline-grounded benchmark for Decision Update Consistency (DUC), evaluating whether model revisions are warranted in direction, magnitude, and scope. DUC-Bench contains 450 staged instances derived from 150 clinical decision routes and reviewed by two medical experts. Each instance pairs an initial clinical recommendation with a controlled evidence update spanning contradictory, complicating, and uncertainty-inducing regimes. Warranted responses distinguish four clinically meaningful transitions: Maintain, Modify, Replace, and Suspend, while matched variants separate responsiveness to valid evidence from susceptibility to weak evidence and user-assertion framing. Across six models, four-way transition accuracy ranges from 43.6-55.8%, while collapsing the task to change-versus-no-change inflates apparent performance by 17.1-26.5 percentage points. Only one model exhibits statistically supported warrant sensitivity, and weak-evidence resistance ranges from 1.4% to 56.9%. Models with similar aggregate accuracy nevertheless exhibit distinct revision policies, including under-revision, indiscriminate updating, framing sensitivity, and modal-transition collapse. These findings show that evidence-responsive decision updating is a distinct capability that endpoint-oriented medical evaluation can substantially obscure. Code: https://anonymous.4open.science/r/DUC-Bench/