Trust Me or Check Me? Self-Reported Confidence, Verification, and Decision Authority in Language Models
Abstract
When should users trust an AI answer, and when should they verify it independently? We study whether a language model’s own reported confidence can reliably govern that decision as the cost of being wrong rises. On a deterministically selected 500-question sample from MMLU-Pro across four model families, we freeze each answer and reported probability of correctness, then vary confidence visibility, human versus AI decision authority, and error cost, yielding 32,000 matched verification decisions. Making confidence visible produces large but model-specific shifts: Claude verifies more, GPT-5.6 Sol and Grok 4.20 verify less, and Gemini shifts toward more verification mainly at higher costs. For GPT-5.6 Sol, visible confidence reduces disagreement with the policy implied by its reported probability from 19.6–27.2% to 0.2–2.2% and nearly eliminates stake-monotonicity violations. Yet at the highest error cost, wrong answers left unverified rise from roughly 32–33% to 54–56%. Thus, a model can become much more consistent with its stated confidence while becoming worse at catching its own errors. Self-reported confidence is therefore not merely an uncertainty report; it can act as a downstream control signal for when human or external scrutiny enters the loop.