Calibration Free Uncertainty Quantification for LLMs
Kevin Guo ⋅ Bradley Malin
Abstract
Uncertainty quantification (UQ) for large language models (LLMs) involves repeated sampling for a fixed prompt. For statistical guarantees, current approaches, such as conformal prediction, rely on a labeled calibration set and a fixed sampling budget which yield only marginal coverage. To address this, we introduce Anytime-Valid Self-Consistency (AVSC), a calibration-free method to estimate predictive uncertainty in generative models. AVSC constructs anytime-valid e-processes to cover query response distributions from repeated sampling alone. This approach guarantees per-query coverage, while using 36% fewer samples than fixed-sample baselines, and reaching target interval widths on all queries.
Chat is not available.
Successful Page Load