Hypergraph Modeling of Transformer Attention for Hallucination Detection
Mark Junjie Li ⋅ zhishun liu ⋅ Wei Wang ⋅ Yuanshan Lu ⋅ yangjingxi ⋅ Xianyu Bao ⋅ Jun Li ⋅ Sunjie Huang
Abstract
Large Language Models (LLMs) often generate factually unsupported content known as hallucination. However, existing detection approaches rely on handcrafted heuristics or isolated attention signals, limiting their ability to capture higher-order dependencies. In this work, we propose $\textbf{AttnHyper}$, a framework that represents a Transformer's attention patterns as a hypergraph for hallucination detection. Tokens are treated as nodes, while hyperedges connect groups of tokens that co-occur across heads and layers, with features derived from attention weights to preserve higher-order interactions beyond pairwise graphs. Based on this representation, the method formulates hallucination detection as a hypergraph learning problem and employs a hypergraph neural network (HGNN) to extract structural cues. It consistently outperforms state-of-the-art methods on the RAGTruth benchmark, achieving an average gain of +5.08 AUROC over the strongest baseline across diverse settings and architectures, while also demonstrating strong zero-shot transfer. Importantly, it incurs minimal overhead at inference time, enabling efficient deployment. Our results demonstrate that hypergraph-based attention modeling provides a more expressive and reliable signal for hallucination detection.
Chat is not available.
Successful Page Load