Can Folding Models Tell Binders from Bluffers? Evidence from POISK: The Patent-Derived Antibody Dataset
Daria Tupikina ⋅ Andrea Roncoli ⋅ Alexander Bujotzek ⋅ Brennan Abanades Kenyon
Abstract
Structure prediction models are increasingly deployed as zero-shot digital screens in antibody design, operating under the assumption that high folding confidence implies true binding. However, rigorous validation of this premise has been limited by small test sets, narrow target diversity, and the risk of data leakage from training-adjacent structures. Here we introduce POISK: Patent-extracted Organization of Immunoglobulin Sequence Knowledge, a large-scale dataset derived from the US patent literature, comprising 89074 antibody-antigen records extracted from 10110 patents using an LLM-based pipeline. From this resource, we curate a balanced benchmark set of 22494 complexes: 11247 patent-validated binders without a known resolved structure paired with putative non-binders constructed from sequence-similar human repertoire antibodies. Using this dataset, we benchmark six leading structure prediction models (AlphaFold-Multimer, Boltz-2, Intellifold v2, OpenFold3, Protenix-v1, and Protenix-v2) on their ability to discriminate binders from decoys. The best-performing model, Protenix-v2, achieves 87% pairwise accuracy on the most confident half of predictions and correctly identifies the true binder from a pool of 50 candidates in 32% of trials; this represents a substantial enrichment over random selection, yet remains far from reliable single-candidate identification. Structural consensus across models also discriminates binders from decoys, with pairwise accuracy reaching 73.1% when three or more models converge on a pose DockQ $\geq$ 0.23, covering around 54.2% of pairs). Yet epitope localization remains poor: the best model contacts annotated epitope residues in only 45% of cases, indicating that high confidence reflects learned statistical associations rather than accurate physical binding modes. Our results suggest that current folding models provide useful but limited signal for antibody screening and should be complemented by orthogonal methods. We release POISK as a community resource for future benchmarking and model training.
Chat is not available.
Successful Page Load