Separating Intervention Search from Response Prediction in Closed-Loop Biological Agents
Abstract
A closed-loop biological agent chooses interventions, observes their responses, and uses those observations to choose subsequent experiments and a final intervention. Does finding a successful intervention also show that feedback improved its predictions of interventions it has not tried? We measure these outcomes separately under an eight-assay budget by scoring held-out intervention forecasts before and after feedback and scoring the final intervention independently. Across 232 tasks from 39 published Boolean networks, replacing an intervention-independent mean predictor with ridge regression lowers held-out squared prediction error by .056 (95% interval [−.064, −.049]) while assays and final choices remain fixed. Conversely, changing only the final-choice rule lowers intervention success by 5.0 percentage points [−9.4, −.7] while predictions remain fixed. The intervention-independent mean nevertheless achieves the lowest prediction error among six procedures while distinguishing no interventions, so we additionally measure whether forecasts remain aligned with the correct interventions after their identities are shuffled. In an LLM study, maintaining a hypothesis note increases this alignment on named simulation tasks, but the difference shrinks when gene names are hidden or runs requiring action fallbacks are excluded. In retrospective replay of six human T-cell cytokine screens and K562 combined CRISPRa perturbations, hypothesis notes do not consistently reduce final prediction error relative to observation notes given identical assay histories. These results show that intervention success and response prediction answer different questions about a closed-loop biological agent and should be reported separately; they do not establish a general benefit from hypothesis notes or prospective biological discovery.