Ensemble Selective Classification
Abstract
Selective classification (SC) improves model reliability by allowing a classifier to abstain on low-confidence inputs. In the practical post-hoc setting, where the base classifier is fixed and often accessed as a black box, many confidence scores have been proposed, yet no single score is consistently best across models, datasets, and deployment shifts. We address this limitation by proposing, to our knowledge, the first ensemble framework for post-hoc SC, which learns to combine a diverse set of existing confidence scores using a small labeled calibration set drawn from the deployment distribution. We cast this problem as learning to separate correct from incorrect predictions of the base classifier, and train the ensemble selector with SoftRank-AURC, a differentiable soft-rank plug-in objective for the area under the risk-coverage curve (AURC), the standard evaluation metric for SC. We also study hinge and SELE losses as surrogate training variants, and provide an oracle-style guarantee showing that the SoftRank-AURC ensemble is competitive with the best single transformed base score in hindsight. Across 15 image and NLP classification tasks, with additional LLM multiple-choice experiments, the proposed method is consistently strong and often outperforms the best individual score, particularly under distribution shift, establishing ensemble post-hoc SC as a simple and effective way to improve reliability without retraining the base classifier.