Can Self-Improving Agents Recover Cofolding Failures?
Abstract
Co-folding models predict protein--ligand complexes directly from sequence and ligand chemistry, but no single model dominates, and each ranks its own samples with a confidence score that is calibrated only against itself. A near-native pose is often present somewhere in the pool of predictions without being the one any model puts first, which makes recovery a selection problem rather than a generation problem. We ask whether an LLM agent can perform this selection and improve its own procedure for doing so. A Solver agent chooses one complete complex from a fixed pool of precomputed candidates using only reference-free evidence, consulting engines in stages so that the cost of a decision is measured in engines consulted. A separate Improver agent reads Solver traces and reference-based feedback on the training split and rewrites the selection procedure itself: its instructions and analysis code. Model weights, the candidate pool, and the evaluator never change, revisions are kept only after a fresh training rerun, and the procedure is frozen on validation before a single held-out evaluation on a cluster-based split of Runs N' Poses. Across three rounds, mean lDDT-PLI rises from 0.695 to 0.794 while the mean engine budget falls from 1.28 to 1.21. The frozen procedure therefore reaches the level of Boltz-2's native Top-1 ranking (0.795) without surpassing it, and remains behind cheaper unsupervised heuristics such as a consensus medoid. We report this as evidence that bounded, procedure-level self-improvement is learnable in an expensive scientific setting, not that it yet beats native confidence.