Hallucination Detection in Black-Box LLMs Using Complementary Uncertainty Signals
Abstract
Hallucination detection matters in enterprise workflows because a missed fabrication can cost more than an additional human review. When no trusted context or reference document is available, we study two signals accessible through black-box model APIs: semantic entropy and uncertainty derived from token log-probabilities. Their failure modes can be complementary: semantic entropy becomes uninformative when responses form one semantic cluster, while token uncertainty can miss consistently confident errors. We extend token uncertainty across sampled responses through \textbf{TopK}, evaluate \textbf{CoCoA}, which combines target-response confidence with semantic dissimilarity, and propose two supervised methods: \textbf{Gated}, which routes single-cluster cases to aggregated token features, and \textbf{Stacked}, which learns jointly from semantic uncertainty and a broader set of token features. Across seven benchmarks and four models, Stacked leads nearly half of the model--dataset evaluations, while TopK and CoCoA remain competitive without training labels, although both require threshold calibration. No method is universally strongest. We therefore evaluate performance at different false-positive-rate budgets, sensitivity to generation and calibration choices, and variation across dataset characteristics.