Evaluating Agentic Language Models at Crystallographic Protein Model Building
Abstract
Determining the three-dimensional structures of proteins is crucial for understanding their function. While AlphaFold has revolutionized computational structure prediction, experimental methods are still essential for resolving protein--ligand interactions, conformational dynamics, and multimeric structures. X-ray crystallography remains the gold standard for high-resolution structure determination, but the process of inferring atomic structural coordinates from diffraction data, known as model building, consumes significant manual time by human experts. The key challenge with model building is that 3D coordinates and Fourier phases must be estimated simultaneously over multiple rounds. In this work, we evaluate the performance of modern language models at this task, constructing a new environment and harness tested across a variety of model building examples from the Protein Data Bank (PDB). We find that language models are consistently outperformed by Phenix AutoBuild, a classical automated baseline, and that human-derived models from the PDB achieve better metrics than those produced by automated techniques. These findings indicate that expert human judgment, which can explore locally suboptimal solutions to arrive at higher quality final models, remains a critical component of the model building process. Our verifiable environment and harness offer a foundation for closing this gap and eventually post-training agentic policies using experimental data and human feedback.