Do LLM Proposers Improve Budgeted Search for Ionizable Lipids?
Abstract
Efficient lipid design under limited experimental budgets is a central challenge in computational chemistry and drug discovery. We introduce LEAF (Lipid Evolutionary Agentic Framework), a matched-budget protocol for testing whether language-model proposers improve combinatorial molecular search relative to classical baselines.The study uses a 30,044-candidate ionizable-lipid space, a frozen Gaussian-process oracle fitted to experimental potency labels, and a separate active-learning surrogate. Four language models evaluated over eight seeds and three splits show no significant improvement over genetic-algorithm or random-search baselines. For Llama 8B, 79.1\% of acquisition-selected queries repeat previously evaluated archive members. Removing duplicates reduces the performance gap on some splits but increases proposal cost and can prevent budget exhaustion. Larger models reduce duplication without establishing superiority or equivalence. Exploratory scored prompts and stricter novelty controls yield positive median differences against the genetic algorithm, yet wide intervals and shared-seed dependence preclude stronger conclusions. A validated single-assay reanalysis preserves the chemotype-cold deficit at a smaller numerical scale, and exploratory expected hypervolume improvement does not reverse the result. LEAF isolates proposal quality as the experimental variable and documents a concrete failure mode of zero-shot LLM proposers under controlled budgets.