Matchability-Aware Conformal Prediction for Open-Ended Language Model Generation
Ilia Azizi
Abstract
Large language models are increasingly deployed for open-ended question answering, yet their answers rarely come with calibrated uncertainty. We propose CAELUM, a framework that builds semantic prediction sets and separates aleatoric from epistemic uncertainty in their construction. CAELUM samples answers under prompt-ensembled clarifications, clusters them by semantic equivalence, estimates aleatoric and epistemic components via kernel Von Neumann entropy, and calibrates cluster-APS nonconformity scores with split conformal prediction. A per-instance weight $w_\lambda(x)=1+U_{\mathrm{ale}}(x)+\lambda\,U_{\mathrm{epi}}(x)$ controls how epistemic uncertainty enters the set, and a per-cell coefficient $\lambda^\star$ is fit on a held-out split to minimise prediction-set size. Our central observation is that, in open-ended generation, coverage on cluster-APS scores is silently capped by \emph{matchability}: the fraction of items whose reference answer appears in any sampled cluster. We make this dependence verifiable with a Mondrian split-conformal bound on a learned stratum, where matchability is predicted from pre-generation question features by a lightweight classifier. Across $18$ (model, dataset) cells and four split seeds on three open-weight 4--9B LLMs, the learned classifier lifts predicted-matchable coverage from $0.61$ overall to $0.71$ on non-fallback rows ($0.82$ vs $0.63$ when oracle-fallback rows are included as upper bounds), and the classifier transfers across LLMs at AUC $0.78$ (in-sample $0.80$). The cell-level $\lambda^\star$ doubles as a regime diagnostic: large values flag epistemic dominance (retrieve, abstain) and $\lambda^\star\to 0$ flags aleatoric dominance (clarify), with a per-instance log-ratio $R(x)$ extending the read-out to individual queries.
Chat is not available.
Successful Page Load