Position: Life-Science Agents Should Put the Symbolic Layer in Charge Where Mistakes Are Expensive
Abstract
AI agents now propose hypotheses, write analysis code and increasingly operate laboratory instruments. Their designs assume that a failed attempt is cheap. Wet-lab biology breaks this assumption: experiments are few and costly, execution is physical and often irreversible, and an unverified claim is paid for in months of follow-up work. Every tool-using agent is already neuro-symbolic, since it acts through symbolic interfaces such as code, tool calls and protocols, so the design question is where the symbolic layer holds authority. We argue that it should hold authority wherever mistakes are expensive: hypotheses should live in explicit models that choose the next experiment, protocols should be executable programs that are compiled and checked before an instrument moves, and claims should pass a verification stack whose depth grows with the cost of acting on them. We make this operational with five levels of symbolic authority and a rule for when an action must be gated, and we audit 49 published agents for biology, chemistry and behavioral science against these levels. Symbolic authority is growing where agents execute protocols but remains rare where they choose experiments and report claims, and no LLM-era agent places it at all three points. We map each gap to mechanisms that already work in isolation, state falsifiable predictions with metrics to test them, and answer the strongest alternative views.