Benchmarking sequence-to-ensemble predictors on UNICORNEdb, a UniProt-grouped database of PDB-derived conformational ensembles
Abstract
Following the striking success of ML-driven static protein structure prediction, the challenge of protein conformational flexibility and thermodynamic ensemble prediction is rapidly coming to the foreground, and with it the need for rigorous benchmarks. Yet, obtaining ground-truth conformational states at scale remains highly challenging. Past approaches are either small and hand-curated or built by clustering PDB chains, which conflates genuine conformational variation with sequence-induced differences across near-homologs. Here we introduce UNICORNE, a database of 50,000+ protein conformational ensembles centered on UniProt and PDB chain/entity annotations to assemble per-protein ensembles of experimentally resolved conformations without relying on clustering. From our database, we derive UNICORNE-BENCH, a set of 839 multi-conformation proteins on which we extensively benchmark seven recent sequence-conditioned predictors spanning generative ensemble models and MSA-perturbation methods. We find that the ensemble-generation problem is far from solved and identify several shared failure modes: even the strongest tool fully recovers only 8.2\% of ensembles, performance degrades with conformational diversity, and tools fail on overlapping rather than complementary subsets of targets. Across the benchmark, we find that MSA-perturbation methods modestly outperform direct generative approaches. UNICORNE provides a biologically grounded reference needed to guide the next generation protein structure prediction. The full database and leaderboard are available at https://unicornedb.org (anonymized), and source code is available at https://anonymous.4open.science/r/unicornebenchneurips_source-2327/.