Calibrating AlphaFold confidence scores for interface quality
Abstract
AlphaFold confidence scores are widely used to rank protein-complex predictions and to select subsets of models for downstream applications. Although these scores provide strong discriminatory power for distinguishing prediction quality, their numerical values were never intended to represent the probability that a prediction satisfies a given structural quality criterion. In this work, we calibrate commonly used confidence scores to estimate the probability that a predicted complex exceeds a chosen DockQ threshold. On antibody-antigen and general multimer datasets, post-hoc isotonic calibration consistently improves agreement between predicted probabilities and observed DockQ success rates while preserving the underlying score ordering. The calibrated probabilities also provide more accurate estimates of the fraction of selected models expected to pass a given structural quality threshold. Although the mapping is fitted within a given prediction distribution and therefore transfers imperfectly across prediction domains, we show that reliable estimates are recovered from a small number of labeled target-domain samples. Because the calibration is entirely post hoc and requires no underlying AlphaFold model fine-tuning, it provides a straightforward and customizable way to convert existing model quality rankings into calibrated probabilities of passing a chosen structural quality threshold.