Do Repeated LLM Decisions Improve CRISPR Hit Discovery?
Abstract
Does an LLM that benefits from experimental feedback also choose better experiments than an informed numerical policy? We study retrospective CRISPR hit discovery, where a nearest-neighbour procedure expands five proposed genes into a batch. Our numerical rival samples centres from observed hits while retaining the same initial batch, features, expansion procedure and test budget. Across six models and six readouts from four sources, correct outcomes improve LLM yield over withheld outcomes by 18.18 hits per 640 tests, yet the rival finds 9.98 more hits. Complete-campaign checks with expanded response budgets and a native JSON interface retain this pooled distinction. The native-interface model requires no output recovery, but wins on one source and loses on another. Initialization also changes which controller performs better. Finally, LLM proposals contain some useful candidates that a hypothetical selector with access to future outcomes could exploit, but the tested selectors do not establish a gain over equally sized numerical candidate pools. These conditional findings separate feedback use from the added yield of repeated LLM control. Evaluating that added yield requires an alternative that retains both outcome information and the numerical procedure used to turn proposals into experiments.