Beyond Confidence: A Representation Benchmark for Protein Binder Prioritization
Abstract
Structure-prediction confidence is widely used to filter designed protein binders but does not directly measure experimental binding. We benchmark frozen representations for classification after this filter. Six public sources yield a primary cohort of 2,881 non-immunoglobulin candidates passing an AlphaFold 3 (AF3) iPAEmin ≤ 2.5 Å gate, with 756 positives across 43 target-sequence clusters. We design and train compact adapters for ESMC sequence embeddings, AF2/AF3 single and pair activations, explicit geometry, dMaSIF surfaces, and Rosetta features, evaluated alongside scalar confidence on the same five grouped out-of-fold (OOF) splits. AF3 mature pair reaches area under the precision–recall curve (AUPRC) 0.614, versus 0.534 for a three-score confidence model and 0.498 for direct ipSAEmin; the tested sequence and explicit-interface readouts span 0.374–0.498. AF3 mature single (0.600) and AF3 refined single (0.618) achieve comparable point estimates. The results support the utility of task-specific readouts of internal activations for post-filter prioritization; broader utility requires prospective validation.