CHOICE-Bench: Benchmarking Biologically Informed Target Scoring in AI Agents for CHO Cell Line Engineering
Abstract
Biologics manufacturing underpins a growing share of modern pharma, but AI methods have focused far more on drug discovery than on biomanufacturing. Chinese hamster ovary (CHO) cells are the main mammalian host for therapeutic protein production, yet CHO cell line engineering remains heavily empirical because direct perturbation evidence is sparse, scattered and rarely relevant to the biomanufacturing related phenotypes such as titer. This makes CHO a demanding test case for AI agents: the model must retrieve literature, extract perturbation claims, infer the biological meaning of measured readouts, and decide an intervention direction. We introduce CHOICE-Bench, a benchmark for evidence-conditional target scoring, together with Chinese Hamster Ovary Intelligent Cell-line Engineering (CHOICE), a biology-guided AI agent that reasons over a CHO biology knowledge graph for target scoring. In this setting, the agents assign a fixed positive or negative direction only where the relationship has a stable one, and report conditionality (\texttt{context dependent}) where the appropriate direction varies on cellular or process context. The benchmark combines an end-to-end task framed as the question a cell line engineer would ask, with stage tasks for source retrieval, claim extraction, pathway identification, pathway-to-phenotype weight elicitation and recovery of held-out direct evidence from pathway evidence alone. Across the benchmark, general-purpose generative AI such as Claude and GPT failed mostly on judgments requiring biological and CHO-specific knowledge. The largest performance gaps appeared in direction inference, biological pathway assignment and context-dependent scoring rather than on verbatim evidence extraction. In blind curation by four human experts of 361 extracted claims, direction was the most common extraction error, occurring in approximately one in six claims across all general AI agents tested. With the model and corpus fixed, adding twelve CHO expert-curated extraction rules increased direction accuracy from 66.7\% to 80.0\%. CHOICE addressed additional error modes that the rule-augmented prompt did not resolve. In pathway identification from literature evidence, four blinded domain experts judged CHOICE acceptable for 81.1\% of adjudicable readouts randomly sampled, indicating strong performance. In pathway-to-phenotype scoring, Claude Code and Codex showed low run-to-run stability when deciding whether a relationship should receive a fixed sign or remain context dependent. Results from this benchmark show that general-purpose AI agents are less reliable for cell line development because they do not consistently apply the correct biological constraints required for CHO cells. CHOICE addresses this gap through biologically informed reasoning over a CHO-specific biological knowledge graph, thereby enabling more reliable target scoring for biomanufacturing.