Learning Semantically Coherent Calibration Groups
Abstract
In clinical decision support, a model's confidence directly shapes whether a practitioner trusts or overrides its prediction, so miscalibration that varies across patient subgroups can silently mask diagnostic bias even when a model appears well-calibrated overall. Post-hoc calibration methods adjust model outputs to align with true outcome probabilities, and multicalibration extends this notion by enforcing calibration across multiple groups within the data. But in settings where groups are unknown, existing approaches either offer no interpretable structure to explain why a model is miscalibrated for a given group, or produce groupings that appear coherent without reflecting the model's actual calibration need, leaving practitioners unable to trust what those groups reveal about risk. We introduce Semantically Coherent Calibration (SCC), a method for learning interpretable groups for multicalibration. Our key insight is that coherent groups should contain samples that are similar in the embedding space as well as in their calibration needs, allowing practitioners to identify and interpret the clinical sources of a model's miscalibration rather than just correcting for it. We evaluate SCC on vision-language models across multiple medical imaging datasets and model architectures, demonstrating that SCC achieves superior calibration performance and group structure compared to existing grouping methods, with up to 79\% reduction in Expected Calibration Error (ECE).