Beyond Uniform Detection: Adaptive Hallucination Detection for RAG Across Response Regimes
Abstract
Existing hallucination detectors for retrieval-augmented generation (RAG), whether based on model outputs (e.g., likelihood) or internal representations, typically apply a uniform detection strategy across responses. However, we find that different detection signals are informative in different response regimes, making a uniform strategy insufficient for both short- and long-form generation. In short-form tasks such as factoid QA, hallucination typically appears as selecting an incorrect final answer among multiple context-relevant candidates. In this regime, model-output-based signals that discriminate among plausible answers are particularly important. We therefore propose a bidirectional likelihood measure that evaluates logical consistency based on the retrieved evidence and a reasoning step generated after the response. In contrast, long-form descriptive responses are more prone to gradual drift: as generation proceeds, the model increasingly relies on its own prior generations and parametric knowledge, while unsupported continuations may remain locally plausible, making output-based cues less informative. For this regime, we measure context--knowledge conflict by tracking the directional alignment of context-grounding versus internal-knowledge contributions in the hidden states. Building on these regime-specific signals, we introduce ARGUS, an adaptive hallucination detector that emphasizes the more effective signal according to the response regime. Experiments on multiple RAG benchmarks show that our method achieves strong performance across both short- and long-form settings.