Many Benign Errors are Better than A Few Severe Ones: Evaluating Hallucination Severity
Abstract
Factuality evaluation in summarization has largely focused on detecting hallucinations as binary phenomena---labeling content as either supported or unsupported by the source. However, modern LLMs increasingly generate nuanced, partially grounded content whose factual impact can vary widely—from benign inferences to severe distortions. To enable more precise assessment, we propose a framework for annotating hallucination \emph{severity} along two interpretable dimensions: \emph{implausibility} (how unreasonable a span is given the source or general knowledge) and \emph{impact} (its potential to mislead or cause harm if assumed true). Applying this framework to obtain annotations on model generations across domains, we find that many hallucinations are both plausible and benign, indicating that not all unsupported content is equally consequential. We further show that severity-aware evaluation can meaningfully alter model rankings, favoring models that generate more frequent yet benign hallucinations over those producing fewer but more detrimental errors—distinctions that conventional evaluation settings fail to capture. Finally, we evaluate LLMs as severity judges: while their span detection remains weak, their severity ratings align closely with human judgments, making them promising candidates for scaling fine-grained annotation. Using severity as a diagnostic lens, we find that LLMs preferentially detect obvious errors but miss plausible, high-impact ones---precisely the cases that matter most for deployment. Together, these findings argue for a shift toward severity-aware factuality evaluation.