A Referral Is Not an Abstention: What Medical LLM Deferral Rates Actually Measure
Abstract
Medical language models increasingly answer clinical questions with a referral to in-person care, sometimes with a caveat and, rarely, a refusal. Evaluations and safety benchmarks report the rate of such deferral as if it measured the model's reliability, and a referral does resemble an abstention: the model declines to have the last word. We argue that it is not one. In current models, deferral responds to the surface framing of the request, not to whether the model knows the answer. We ground this position in a controlled probe: 309 USMLE-style items, each presented with identical clinical content in three frames (third-person exam question, first-person urgent request for advice, first-person urgent request for information), to five models whose per-item competence we measured separately with forced-choice sampling. Hedging rises from 0–6\% in the exam frame to 90–100\% in the advice frame and 17–79\% in the information frame, yet within a frame it is flat across the model's own competence bins wherever the comparison is powered, and unchanged under pooled-difficulty and open-ended-error stratifications. The models' unhedged errors concentrate exactly where competence is low, and the hedge does not move there. A deferral rate therefore measures a policy's sensitivity to framing, not calibrated abstention, and should not be reported, benchmarked, or regulated as though it were: evaluations should condition deferral on measured competence, and deployment guidance should stop reading a referral as evidence about the answer it accompanies.