SimplexUQ: An Evaluation Framework and Benchmark for Conformal Uncertainty on Simplex-Valued Predictions
Liang You ⋅ Hengyu Shi ⋅ Dongwen Ou
Abstract
Conformal prediction can be marginally valid on simplex-valued outputs while still allocating coverage unevenly across prediction space, over-protecting easy regions and leaving hard regions under-covered. This matters because many modern predictors output compositions, including class-probability vectors, topic mixtures, spectral abundances, cell-type fractions, age-label distributions, and emotion mixtures. We argue that allocation quality must be evaluated in its own right, and introduce $\textbf{SimplexUQ}$ to make this possible. SimplexUQ contributes a diagnostic protocol for measuring coverage-allocation failures, $\textbf{SimplexTasks-12}$ (a benchmark with six controlled synthetic regimes and six fixed-predictor real tasks), and a systematic comparison of global, group-wise, normalization-based, exact, and leave-one-out conformal wrappers. Across the benchmark, no wrapper dominates under all stratifications and tasks, though group-wise calibration is competitive across many settings. Global calibration can satisfy nominal marginal coverage while producing severe worst-stratum failures; for example, on CIFAR-10, global calibration attains 0.900 marginal coverage but leaves the worst entropy stratum at 0.542 coverage, whereas group-wise calibration reduces the max disparity from 0.358 to 0.022. Group-wise calibration is strongest when heterogeneity is coarse and aligned with a partition, whereas normalization and leave-one-out-style methods are more competitive under smooth heterogeneity. On several tasks, even the best wrapper fails to recover acceptable worst-stratum coverage. The benchmark standardizes tasks, stratifications, metrics, compute reporting, and artifact packaging, with code and rebuild instructions for restricted assets. Marginal validity alone is therefore insufficient for simplex-valued uncertainty quantification: wrapper choice should be matched to the heterogeneity structure of the task, with allocation, efficiency, and compute considered jointly.
Chat is not available.
Successful Page Load