Early Diagnostic Commitment in Large Language Model Clinical Reasoning
Abstract
Sound diagnostic reasoning requires weighing a leading diagnosis against other plausible differentials and is essential to safe clinical practice. As Large Language Models (LLMs) are integrated into clinical workflows, clinicians find themselves increasingly responsible for verifying the veracity of LLM diagnostic recommendations. Long-form chain-of-thought produced by reasoning LLMs has purportedly offered a visible avenue for auditing the intermediate computation models perform before reaching a conclusion. Indeed, chain-of-thought traces often express factual knowledge recall and continued deliberation. However, it remains unclear whether this apparent deliberation corresponds to a continuous revision of the model's final diagnostic conclusion. We introduce a framework for measuring diagnostic commitment, the earliest position in a chain-of-thought response to a clinical vignette, where further reasoning does not meaningfully change the model's distribution of diagnoses. To generate this empirical distribution of diagnoses we resample continuations at various depths of the reasoning trace and quantify changes in this distribution using total variation distance. Importantly, this captures changes in probability mass distributed across all diagnoses as opposed to just the primary diagnosis. We apply this framework to eight open source reasoning models across free-response vignettes from MedQA, MedMCQA and New England Journal of Medicine Clinicopathological cases, covering 1,146 base traces and over 1.7 million resampled reasoning traces. We find that models frequently engage in diagnosis commitment well before the reasoning trace is complete with 36% committing before halfway, 61% before three-quarters of the way through and 96.1% of traces assigning the largest probability to the same diagnosis it ultimately emits. Transition to diagnostic commitment is often abrupt and sentences first introducing the eventual diagnosis disproportionately drive this collapse in diagnosis distribution. The distribution of diagnoses does not change meaningfully after commitment yet reasoning continues to mention differentials at close to the rate before commitment. Our findings reveal a critical disparity between the visible deliberation expressed in the chain-of-thought and how long a model's diagnostic distribution remains open to revision suggesting that evaluating reasoning based on semantic content alone provides a misleading account of how models reach clinical conclusions.