BILD: A Bilinear Interaction Learnable Discriminator for Post-Hoc Misclassification Detection
Abstract
In safety-critical applications, deep neural networks must identify when their predictions are unreliable. Existing uncertainty methods either require multiple forward passes or collapse predictive distributions into a single confidence score, discarding informative relationships between classes. We introduce BILD, a model-agnostic, post-hoc framework for misclassification detection based on a Bilinear Interaction Learnable Discriminator. BILD learns pairwise interactions between class probabilities through a bilinear scoring function optimized directly for detecting misclassifications. It is single-pass, black-box compatible, and requires no architectural modifications. Across four benchmark datasets, BILD consistently achieves state-of-the-art FPR95. On Tiny-ImageNet with DenseNet, it reduces FPR95 by 8.13 percentage points (p.p.) over the previous best pairwise interaction method, and by 0.52 p.p. over the strongest competing baseline on Swin-Tiny, ImageNet-1K. These results demonstrate that modeling learned interactions between classes provides an effective and scalable signal for misclassification detection.