Evaluating and Improving AI Agents for Sequencing Protocol Understanding
Abstract
Existing scientific-agent benchmarks do not test whether agents can reconstruct the exact molecular structure and workflow of sequencing protocols, which is criti- cal for analysis of customized sequencing assays. We introduce LibStructBench, a document-grounded benchmark for reconstructing the molecular workflow of sequencing-library generation. Our experiments show that frontier agents are often approximately correct but remain unreliable at exact reconstruction, with discrep- ancies in molecular-state representation and workflow connectivity rather than document parsing. Performance patterns were influenced more by the underlying LLM than agent harness. Neither additional external knowledge nor autonomous accumulation of rules and cross-protocol examples produced reliable gains. These findings highlight the need for more controlled and verifiable methods for scientific reasoning over molecular processes in complex experimental protocols.