Evidence Is Not a Vote: Categorical Stability and Confidence Sensitivity Under Conflicting Clinical Evidence
Abstract
Clinical language models may encounter evidence sources that differ in methodological strength and direction, yet it remains unclear how repeated lower-tier contradiction affects model behavior when an evidence hierarchy is explicitly supplied. We developed a controlled hierarchy-maintenance benchmark that holds an explicitly designated higher-tier systematic-review anchor fixed while adding one, two, or four synthetic lower-tier items that contradict it, together with a matched aligned condition and an anchor-last condition. Categorical performance was near ceiling under the anchor-only condition in an independently frozen 260-case Gemini-3.6-flash replication cohort (259/260; 99.62%). Against this high-performing baseline, contrary to our prespecified hypothesis, no dose-dependent increase in Evidence Dilution Rate (EDR) was observed: EDR was 2/259 (0.77%), 1/259 (0.39%), and 1/259 (0.39%) after one, two, and four conflicting items. The four events arose from two cases and were all shifts to insufficient; no opposite-direction reversals were observed. Model-reported confidence nevertheless responded systematically to evidence conflict. The prespecified anchor-to-four-conflict shift was −4.64 points on the 0–100 confidence scale (95% paired bootstrap CI, −5.13 to −4.16). In a post-hoc matched comparison, confidence was 6.51 points lower under four conflicting than four aligned lower-tier items (95% CI, −7.07 to −5.95), with lower confidence in 224/260 paired cases. A post-hoc GPT-OSS-120B sensitivity analysis showed the same qualitative pattern, with no EDR events through four conflicting items in 40 baseline-correct cases and a conflict-versus-aligned difference of −5.75 points (95% CI, −7.45 to −3.60). Thus, within an explicitly instructed evidence hierarchy, categorical judgments were highly stable across the tested conflict range while model-reported confidence remained sensitive to evidence polarity. These findings motivate evaluating behavioral responses to evidence conflict alongside categorical correctness, without implying autonomous evidence appraisal or calibrated uncertainty.