Does reflective LLM feedback survive a wet-lab-scale label budget?
Haoxuan Zeng ⋅ Xin Luo
Abstract
Language-model controllers that read rich natural-language feedback beat scalar-reward optimizers when evaluation is cheap. Improving a biological model is not cheap: a mutational-scanning campaign reveals about one plate of labels per round, and the signal that decides whether an edit helped is computed from those same labels. We test whether reflective feedback survives this regime. An LLM reading a cross-validated diagnostic report, two feedback ablations, random editing, local search and fixed recipes improve a protein-fitness predictor through one closed action space spanning curation, pseudo-labeling, training recipe and acquisition, with identical labels, gate and random streams, on $36$ ProteinGym assays, $3$ seeds and two test splits. The pre-registered comparison is a null: the full-report LLM does not beat random edits of matched size on sealed-test rank correlation ($\Delta z=-0.008$, $95\%$ CI $-0.034$ to $+0.018$), nor on a harder position-disjoint split. It does retrieve more of the top $10\%$ of variants ($+0.032$ recall, $p<0.001$), but a fixed adaptive acquisition rule with no LLM does exactly as well. The best arm is the LLM \emph{without} the action menu ($+0.034$ Fisher $z$ over no editing, $q<10^{-4}$), because the schema rejects most of its guesses and confines it to the default model family; on the harder split a fixed literature recipe beats the LLM. At this budget, conservatism rather than reflection is what pays.
Chat is not available.
Successful Page Load