The Aggregate Monotonicity Illusion: Monotonic Uncertainty Regularization under Progressive Information
Abstract
Many real-world predictive systems acquire information progressively, with richer observations incurring greater acquisition or computational cost. In clinical imaging, for example, additional examinations can provide richer diagnostic evidence but may also incur greater monetary cost, acquisition time, radiation exposure, and use of limited clinical resources. Predictive uncertainty is therefore a natural signal for deciding whether the available information is sufficient or whether more should be acquired. We ask a basic question underlying such decisions: does uncertainty itself behave consistently as progressively richer information becomes available? We uncover an Aggregate Monotonicity Illusion: standard cross-entropy models become more accurate and less uncertain on average as information increases, yet remain highly non-monotonic for individual samples. Across five progressively informative resolution levels, 91.4% of Imagenette samples and 77.8% of ImageNet-1k samples exhibit at least one entropy-monotonicity violation despite favorable aggregate trends. The same sample-level inconsistency also appears in a substantially different clinical setting on NIH ChestX-ray14: in a three-pathology multilabel task, standard training produces roughly 55.9% sample-level entropy violations. These results reveal a consequential discrepancy between how uncertainty is commonly summarized and how it is ultimately used: models may be monotone in aggregate while remaining non-monotone by instance. To address this failure, we introduce a family of Monotonic Uncertainty Regularization (MUR) objectives that explicitly couple predictions across information levels and can be applied as lightweight second-stage fine-tuning to already-trained predictors. MUR-U directly regularizes scalar uncertainty ordering, whereas MUR-AU couples full predictive distributions, jointly regularizing predictive structure and uncertainty. On Imagenette, MUR-U and MUR-AU reduce entropy-violation rates from 91.4% to 0.7% and 4.7%, respectively; on ImageNet-1k, from 77.8% to 3.5% and 29.1%. On ChestX-ray14, both MUR variants reduce the roughly 56% violation rate to near 0%, while maintaining comparable predictive performance. In contrast, reshaping confidence independently within each information level does not achieve comparable cross-level coherence. Finally, we evaluate whether this property matters when uncertainty is used operationally. Under low-information-first adaptive acquisition, MUR models generally achieve stronger cost--performance trade-offs than standard cross-entropy training, allowing more samples to be handled with less informative inputs while maintaining predictive quality. Together, our results expose a simple but important mismatch in progressive prediction systems: uncertainty is evaluated in aggregate, but acted upon per instance. Explicitly learning cross-level uncertainty structure makes confidence more coherent and more useful for adaptive, cost-aware decision making.