Large Language Model Failures from Hallucination to Homogenization Are Different Facets of Miscalibration
Abstract
This paper argues that calibration should become a primary evaluation and optimization target for Large Language Models (LLMs), because diverse failures, from confident hallucinations to collapsed diversity to brittle safety refusals, are best understood as different facets of a single underlying problem: miscalibration. We propose a unified framework organized around four types of miscalibration: probabilistic (the model cannot match requested probability distributions), semantic (confidence is misaligned with factual correctness), distributional (output diversity collapses to stereotypical modes), and metacognitive (the model fails to assess its own competence). We argue that foregrounding calibration as a diagnostic lens and evaluation target is essential for building models that are not merely capable, but genuinely reliable. We call for calibration metrics to become a standard part of benchmark reporting, for training paradigms that preserve uncertainty, and for interaction designs that surface model confidence to downstream decision-makers.