Order Integrity for Agentic DNA Synthesis Under Prompt Injection
Abstract
Laboratories are beginning to delegate DNA synthesis ordering to language-model agents. Screening synthesized sequences against sequences of concern is the established safeguard at that step, and because it validates a particular artifact, it assumes that the sequence under review is the sequence that will be made. An agent that reads untrusted text does not satisfy that assumption, because an indirect prompt injection, in which instructions hidden in data the agent reads are followed as though they came from the user, can cause it to order a construct other than the one that was planned and screened. We call this a failure of order integrity. Using a bioinformatics agent with simulated ordering, we compared four conditions over a complete grid of 676 episodes, consisting of no monitor, a static rule baseline, plan-bound monitoring, and plan-bound monitoring whose monitor also reads untrusted content. Plan-bound monitoring requires the agent to commit to a plan before it reads any untrusted data, after which a separate monitor that never sees that data blocks any action departing from the plan. On the redirected order the undefended agent ordered the attacker's construct in 27 of 30 trials and the rule baseline did so in 22 of 30, whereas plan-bound monitoring blocked the order in all 30 trials, with an exact McNemar p of 4.8e-7 against the baseline. The redirected order evades the baseline because it satisfies every static rule, including the requirement that a screen precede an order. On attacks that a static rule can express the two defenses are indistinguishable, and on benign episodes both block the same nine of 183 calls with no loss of task completion. Fabricated results defeat all four conditions equally.