Does Feedback Format Change LLM-Guided Biological Search?
Riya Danait ⋅ Julia Manso
Abstract
Large language model agents can select experiments over sequential biological rounds, but studies disagree about how strongly their choices depend on observed outcomes. We ask whether adding redundant row-level outcome labels changes feedback use in closed-loop biological search. We fix a Sonnet 5 route through NVIDIA Inference, task instructions, assay budget, and biological data, then run a 12-block pilot on a genome-wide CRISPR knockout screen. Each block crosses two displays, grouped headings alone and the same headings plus row labels, with three feedback conditions: correct outcomes, coherent outcome pairs shuffled within each batch, or outcomes withheld. All six campaigns share the same first-round proposal. The primary outcome is distinct true hits first found after feedback begins per 100 assay slots. Post-feedback hit yield was higher with correct than shuffled feedback under both displays, by 2.69 hits per 100 assay slots with grouped headings and 1.48 with row labels. Contrary to our hypothesis, the layout interaction was $-1.20$ (95\% interval $[-2.93, 0.54]$; sign-flip $p = 0.230$). Correct feedback also outperformed withholding, whereas shuffled feedback preserved aggregate batch yield. This exploratory pilot shows that correct gene–outcome mappings outperformed within-batch shuffled mappings, but provides no evidence that row labels increased that advantage.
Chat is not available.
Successful Page Load