Evidential Semantic Uncertainty Decomposition for Large Language Models
Abstract
Free-form language generation poses a challenge for uncertainty quantification because responses do not belong to a fixed global label space. Semantic-entropy methods address this issue by sampling responses, clustering them by meaning, and estimating uncertainty over prompt-specific semantic classes. However, they mainly measure semantic dispersion and do not provide a principled decomposition into aleatoric and epistemic components, which in our setting correspond to ambiguity among plausible semantic meanings and insufficiency of model support, respectively. Moreover, cluster frequencies depend on the sampling budget and should not be treated as evidential strength. We propose a prompt-level evidential probing framework for semantic uncertainty decomposition in LLM generation. For each prompt, sampled responses define a dynamic semantic frame and provide initial support for the corresponding semantic classes. A lightweight evidence-retention probe learns how much of this support should be retained as reliable semantic evidence. The retained evidence induces a subjective opinion in which dissonance captures conflict among supported semantic alternatives and vacuity captures insufficient retained evidence. To train the probe, we introduce Augmented Semantic Uncertainty Cross-Entropy (AS-UCE), an LLM-specific evidential objective that augments the prompt-specific semantic frame with an unsupported class, allowing unreliable or hallucinated support to be routed away from observed semantic clusters during training. Experiments across four LLM backbones show that our framework consistently improves task-relevant AUROC and AUPR, with average AUROC gains of 7.4\% for OOD detection and 4\% for ambiguity detection over the competitive baselines, while requiring fewer sampled responses than perturbation-based decomposition methods.