Sample Transform Cost-Based Training-Free Hallucination Detector for Large Language Models
Abstract
Hallucinations remain a major barrier to the trustworthy deployment of large language models (LLMs). We study hallucination detection from the perspective of prompt-conditioned response distributions. For a fixed prompt, an LLM induces a distribution over possible responses; when the model is uncertain or hallucinates, this distribution often becomes more complex and less internally consistent. However, this distribution is unknown, and its samples are variable-length token sequences rather than fixed-dimensional points, making direct complexity estimation difficult. To address this challenge, we propose a training-free detector based on sample-to-sample transform costs in representation space. Specifically, we represent each sampled response as an empirical distribution over generated-token hidden embeddings and compute pairwise Wasserstein distances between responses. The resulting Wasserstein distance matrix characterizes the cost structure of transforming one sampled response into another. From this matrix, we derive two complementary hallucination signals: \textbf{AvgWD}, which measures the average transform cost, and \textbf{EigenWD}, which captures the spectral complexity of the transform-cost structure. Importantly, we extend the proposed detector beyond white-box access by using an accessible auxiliary model to construct representation-space consistency signals for black-box LLMs, substantially broadening the practical applicability of our method. Experiments across five open-source LLMs and four benchmarks show that AvgWD and EigenWD consistently achieve competitive or superior AUROC compared with strong training-free uncertainty baselines. These results suggest that distributional complexity in token-level representation space provides an effective signal for hallucination detection.