Do Clinical LLMs Update Beliefs Like Clinicians? Trajectory-Based Evaluation of Longitudinal Inconsistency Detection
Abstract
Clinical notes are longitudinal: a statement that appears ordinary in today’s note may be wrong only because it contradicts a dated fact elsewhere in the patient’s record. We study this as longitudinal inconsistency detection, asking whether a target claim is consistent, possibly inconsistent, or definitely inconsistent with prior notes. Rather than evaluate only the final label, we evaluate the belief trajectory by which a model reaches it. We introduce the Trajectory Belief Updater (TBU), a local neuro-symbolic method that uses a quantized open-weight LLM only to judge individual target-claim/prior-fact pairs, then combines the resulting signals within each dated note and accumulates note updates with a deterministic logit-space updater. This separation yields an exact order-invariance guarantee for the accumulated belief: fixed evidence, note grouping, and pairwise tags give the same final belief regardless of the incidental order in which note updates are accumulated. On a 775-case benchmark constructed from oncology notes and MIMIC-IV-Note, we evaluate four quantized open-weight backbones. Per-fact decomposition provides the largest three-class gains when holistic performance is weak and improves the binary screening endpoint for every backbone. Under a fixed alternative ordering of identical evidence, a conventional sequential LLM updater changes its final label on 10–40% of cases, whereas TBU changes only by floating-point noise. Trajectory-aware metrics reveal that most backbones identify the direction of evidence correctly but under-react to strong contradictions, localizing the error primarily to evidence-strength calibration. The result is a compact, auditable evaluation framework for local clinical LLMs: not merely whether a model flags an inconsistency, but whether it updates its belief for the right reasons.