Black-Box Uncertainty Quantification for Large Language Models via Ensemble-of-Ensembles
Abstract
Reliable use of large language models (LLMs) requires uncertainty estimates that separate ambiguity intrinsic to a prompt from genuine knowledge gaps. Bayesian and deep-ensemble methods deliver this aleatoric/epistemic decomposition in principle, but are computationally prohibitive at LLM scale and require white-box access. Existing black-box methods, in contrast, summarize sample variability into a single consistency scalar that cannot make the distinction. We introduce a two-level ensemble that bridges this gap. Stochastic decoding of a fixed prompt forms the inner ensemble; meaning-preserving semantic perturbations of the prompt form the outer ensemble. Variability is measured in a continuous embedding space, and the law of total covariance gives an exact AU/EU split that requires only black-box sampling access. On top of this estimator, we contribute three theoretical results that ground the framework: (i) a finite-sample bias identity for nested Monte Carlo with a closed-form bias-corrected epistemic estimator; (ii) a Bayesian bridge bounding the gap between perturbation-population AU/EU and Bayesian targets through three local conditions that we estimate as diagnostics from the data the estimator already collects; and (iii) a separation result identifying a class of failure modes that any consistency-only black-box method provably misses while perturbation-based EU detects. Across five short-form QA benchmarks and four instruction-tuned LLMs spanning two families and two scales, the estimator matches or surpasses strong white-box baselines; controlled paired interventions on AmbigQA and Natural Questions confirm the AU/EU split tracks distinct, actionable sources of failure; and a long-form pilot on the NQ long-answer split shows the framework extends to claim-level factuality without retraining. Across five short-form QA benchmarks and four instruction-tuned LLMs spanning two families and two scales, the estimator matches or surpasses strong white-box baselines; controlled paired interventions on AmbigQA and Natural Questions confirm the AU/EU split tracks distinct, actionable sources of failure; and a long-form pilot on the NQ long-answer split shows the framework extends to claim-level factuality without retraining.