From Biased Structural Benchmarks to Generalizable RNA–Protein Interface Prediction
Abstract
The binding of RNA to a protein is a rare event in which both molecules undergo conformational changes, especially RNA, which is known to be highly flexible. One limitation in computational biology is the biased and limited dataset of protein–RNA complexes, which partly reflects these biological constraints as well as the literature’s greater focus on protein binding to more stable molecules, such as chemical compounds. This paper discusses co-folding models, such as AlphaFold3, and current docking methods that are applied to protein–RNA complexes. These methods are typically evaluated using benchmark averages whose composition mirrors the deposited structures on which they were trained or optimized, raising the question of whether they truly generalize. Here, we ask a narrower question focused on the region that carries the biological signal: the interface or binding site. We therefore construct a benchmark that provides a more stringent test of generalization by covering distinct groups and distributions within protein–RNA data. We show that co-folding performance on RNA is not only concentrated in ribosomal complexes but also varies with related factors, such as experimental modality. The same models are uninformative on in vivo CLIPdb datasets. In contrast, docking is accurate only when the supplied conformations are close to the bound state. We then propose RPFace, an unbiased model that predicts the binding preference of every protein residue and RNA nucleotide directly from sequence, requires no experimental structure, template, or multiple sequence alignments (MSA), and remains stable across different structural and experimental conditions. Code and data will be made publicly available upon acceptance.