The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
Abstract
In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often evaluated with hallucination detection benchmarks using open-domain question answering (QA) datasets that contain questions and corresponding short reference answers. Currently, these benchmarks use an LLM to generate answers to questions within the QA dataset. Then, various automated strategies are applied to label these answers as hallucinated or not by comparing them with the reference answers. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness --- whether the answer is fully supported by the reference --- and factual correctness --- whether the answer is free from contradictions and factually false specific claims. In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question--answer pairs spanning three commonly used QA datasets and three generator models. Our human annotations target answer-level factual correctness. As automated label generation strategies, we evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants. We compare the effect of using either a faithfulness-oriented prompt, which asks the judge model to treat any answer content unsupported by the reference as a hallucination, or a factual-correctness prompt, which explicitly distinguishes between absence from the reference and factual error. Our experiments reveal substantial disagreement both among different automated label-generation strategies and between these automated labels and human annotations. Many strategies also exhibit strong biases towards certain error types. For most LLM judges, replacing the faithfulness-oriented prompt with the factual-correctness prompt significantly improves agreement with human annotations, showing that automated hallucination labels depend strongly on how the target criterion is specified. For open-domain QA benchmarks targeting factual correctness, relying on strict reference faithfulness can introduce systematic measurement bias. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.