Does Benchmark Correlation Predict Selection Quality for Protein Design?
Saanvi S Subramanian
Abstract
Scientific machine-learning benchmarks often report rank correlation between model scores and measured outcomes. Protein-design campaigns, however, act on only the highest-ranked candidates, so global ranking quality may not reflect the value of the subset that is finally selected. We test how closely these quantities agree across 70 units from four independent corpora of measured protein function. The relationship weakens as selection becomes more selective: squared rank association with top-10\% selection value is $0.50$, falling to $0.16$ at top-1\%, and three of 41 assays in one corpus produce worse-than-random top-1\% selections despite positive reported correlations. The mapping also varies across benchmark families. At the same reported correlation, SSMuLA libraries yield $+0.265$ more selection value than ProteinGym assays ($p=0.0001$), and a second scorer reproduces a similar offset of $+0.252$ ($p=0.0003$). This benchmark dependence also limits attempts to predict selection quality from model scores alone. Score skew predicts selection value within one benchmark family but reverses sign when transferred to another and again within a fourth corpus. The mismatch is not universal: in a stronger-signal supervised setting, selection performance improves by $+0.035$ alongside a $+0.031$ improvement in correlation. These results show that global correlation can be informative without reliably determining selection quality at sharp cutoffs or across benchmark families. We conclude that evaluations intended to support candidate selection should report selection performance at the cutoff users will act on and test whether the correlation-to-selection relationship transfers across benchmark regimes.
Chat is not available.
Successful Page Load