Harness Design for Grounded Hypothesis Generation: Lessons from a Deployed Multi-Tool Agent in Industrial Strain Engineering
Aiham Taleb ⋅ Zainab Afolabi ⋅ Luisa Bertoli
Abstract
Gains from LLM agents for biological discovery are usually attributed to the backbone model. From a multi-month industrial deployment of a grounded hypothesis-generation agent over proprietary strain-engineering data, we report the opposite emphasis, and one finding we did not expect: every harness mechanism that bought faithfulness cost novelty in the same direction. Almost all of our engineering effort went into the harness rather than the model, and the properties scientists named as reasons to trust the output were harness properties. The agent orchestrates nine tool families (thirteen callable tools) over a multi-million-abstract public literature corpus, a set of internal project documents, an experimental results warehouse, and a metabolic knowledge graph of order $10^6$ edges. We describe the mechanisms that bore weight, among them a public reference graph with private engineered entities built into it, guarded text-to-query synthesis with three distinct failure channels, tool-scoped entity naming, and a namespace knowledge firewall, naming for each the owning literature and the failure signature we saw in the field. Four deployed failure modes follow, the most dangerous being a private construct identifier that collides exactly with an established public gene symbol, so the model answers fluently about the wrong protein. Expert ratings from working strain engineers ($N{=}16$ sessions, six dimensions) place five dimensions at $\ge$4.5/5 and novelty below 4. We argue that gap is a design property of grounded harnesses rather than a limitation of the model.
Chat is not available.
Successful Page Load