Do Language Models Know Whom to Trust? A Human-Grounded Evaluation of Source-Sensitive Calibration
Abstract
Large language models are increasingly deployed in settings where they must not only interpret information, but decide whether that information is reliable enough to support belief, recommendation, or action. We introduce a human-grounded evaluation suite for \emph{source-sensitive calibration}: the ability to use source reliability as a decision-relevant variable. The benchmark manipulates source reliability while holding propositional content constant, and evaluates humans and LLMs across production, likelihood-based scoring, prompted generation, and action-guiding uptake. Humans show robust source sensitivity: speakers adjust linguistic choices according to source reliability, and hearers strongly reduce action-readiness for low-trust information. LLMs show a weaker and less stable pattern. In production, apparent source sensitivity depends on scoring protocol, prompting format, tokenization artifacts, base-rate preferences, and output compliance. In uptake, models partially register explicit reliability cues, especially in the expanded 200-item evaluation, but the effect is much smaller than in humans and remains poorly calibrated. These results show that source-grounded reasoning cannot be evaluated by accuracy or surface-form choice alone. Models used for decision support should also be tested on whether they convert source information into calibrated judgments about when information is reliable enough to guide action.