LLMs Can Verify What They Still Accept
Abstract
Language models can correctly judge an answer to be wrong in a separate checking task and still adopt that same answer 93–100% of the time. Understanding when models accept or reject information from tools, retrievers, other agents, and users is therefore a distinct problem from understanding whether they can evaluate that information correctly. We call this decision epistemic arbitration. Across over 10 million trials on 12 language models from the Qwen, Gemma, Mistral, and Llama families, we separate what a model can generate, what it can check, and what it actually uses on the same candidates. Checking capability does not govern use. Models systematically favor answers that align with their latent priors, and these arbitration policies vary substantially across models and domains. Causal interventions show that verification evidence can be represented and even verbalized without governing the final answer. Reliability with external information is not just a matter of capability, tool accuracy, or verification ability, but instead requires understanding how external information is trusted and used.