Pyrite: A Benchmark for Protein-Ligand Pose Ranking Agents
Abstract
Knowing how a ligand binds to its target is central to small molecule drug discovery, but since co-crystal structures cannot be obtained for every ligand of interest, binding modes must be predicted computationally. Modern co-folding models, paired with diversity-promoting sampling, now produce candidate pose sets too large for chemists to triage by hand, making the ability to \textit{rank} poses essential. Yet most prior work evaluates whether models can recall a ground-truth pose, not whether they can rank candidates, and existing ranking methods perform little better than random selection. We make three contributions toward improving protein-ligand pose ranking. First, we introduce \textbf{Pyrite}, a curated pairwise pose ranking benchmark spanning a diverse and challenging region of protein and chemical space. Its design combines scalar physics features with pairwise semantic comparisons. Expert chemists with a structure viewer and precomputed physics outputs perform near chance on \textbf{Pyrite}. We then design \textbf{Pyrite-Rep}, a pose representation for language models that allows them to outperform expert chemist accuracy on our pose ranking task. Finally, we introduce a rubric-based GRPO post-training strategy with the pose representation for open-weight language models (e.g. Qwen3.8-27B) that closes this gap, matching frontier model performance.