PermuLoop: Causal Evidence That LLM Discovery Agents Use Experimental Feedback
Shreyasi Swain ⋅ Muad Mohamud ⋅ Elijah Daniel
Abstract
Lab-in-the-loop agents should learn from earlier outcomes. Yet three recent studies disagree on whether large language model (LLM) agents do so in sequential gene discovery. We introduce PermuLoop, a permutation-controlled harness, and run 141 campaigns of a tool-free LLM agent, at three capability tiers of one model family, on six screen tasks from five published CRISPR screens. After each round, the agent sees the true outcomes, a within-round shuffle of them, or (at the two smaller tiers) nothing. A round-1 placebo, run before any feedback, was null ($-1.2$ percentage points, pp; 95% CI $[-3.1, +0.8]$). In rounds 2–5, true feedback added a modest $+3.0$ pp (95% CI $[+2.1, +3.9]$) to permuted-feedback hit rates of 9–12% per round. This effect varied more across screens ($+1.0$ to $+7.3$ pp) than across tiers ($+2.8$ to $+3.3$ pp; two-way ANOVA: screen $p=0.0002$, tier $p=0.82$). Tier instead tracked calibration. Expected calibration error fell from 39.3 pp (Small) to 10.7 pp (Large), but the share of proposals outside the candidate pool stayed near 15% (6–7% without the one focused library). On five of six screens, the best tier beat a DepMap-informed Gaussian-process baseline in lift over the base hit rate. Claims that an agent learns from outcomes thus need a permuted control on several screens, and weaker tiers' stated confidence should not be read as a hit probability.
Chat is not available.
Successful Page Load