Grounded or Fabricated? Unsupervised Detection of LLM Hallucinations via Contextualized Influence on Response Embeddings
Abstract
Hallucinations in large language models (LLMs) that are plausible-sounding but factually incorrect or unsupported pose a major challenge for deploying these models in high-stakes applications such as medical diagnosis, legal reasoning, and knowledge-based question answering. Existing detection methods primarily rely on either generating multiple responses to check consistency or supervised training on labeled hallucinations. Multiple-response approaches tend to be computationally expensive and may be less practical for real-time scenarios, while supervised methods are limited to known hallucination types and cannot generalize to unseen cases. To address these limitations, we propose an efficient unsupervised hallucination detection method. Our approach measures the embedding discrepancy between the LLM’s response considered independently and the response contextualized with the question. A large discrepancy indicates that the question significantly influences the response, suggesting it is grounded, whereas a small discrepancy signals a higher likelihood of hallucination. This method is efficient, interpretable, and does not require labeled hallucinations. Extensive experiments across multiple tasks demonstrate that our approach consistently achieves superior or comparable detection performance while reducing the detection cost compared to the multiple-response methods.