Can an LLM Agent Select the Right Pose? Analysis of Agentic Pose Selection for PXR Fragments
Abstract
Can an LLM-based agentic approach close the gap between generating good protein-ligand structure predictions and selecting the correct one? We study this using 184 recent structures from the OpenADMET PXR (Pregnane X Receptor) Blind Challenge Structure Track. Diverse sampling with Boltz and Boltz-Perturb builds pools in which oracle selection reaches 2.13 Å mean ligand RMSD with 104/184 ligands having below 2.0 Å RMSD while ranking by model confidence recovers only 45/184, a persistent gap between generation and selection. The pools in fact contain poses good enough to win the competition: per-target best poses average 0.73 lddt-pli against 0.61 for the challenge's top rank entry, though selecting based on model confidence averages only 0.54. We evaluate two families of agentic selection against this gap: agents that compose literature- and confidence-based scoring functions grounded in the PXR data pool, and a data-driven agent (RnP Ranker) that discovers ranking metrics from empirical feedback on an independent 113-target benchmark with no PXR knowledge and transfers zero-shot. Neither beats confidence-only selection, degrading the number of sub-2.0 Å ligand poses from 45 to a range of 20–31 and 35-45, respectively. Across both families we identify two recurring failure modes: unsupported weighting of correct domain knowledge, and deploying a scoring function without validating it against the available baseline. Nevertheless, agentic methods provide a complementary signal, whose union with confidence-based selection outperforms top-5 and top-10 confidence-based selection. Together, these results show that agents can extract correct domain knowledge while turning that knowledge into discriminative scoring remains the open challenge for building agentic scientific systems — especially in flexible, promiscuous binding pockets where correct and incorrect poses satisfy the same key contacts.