Biological Identity Failures in Scientific Agents: Measurement and Identity-Bound Execution
Abstract
A scientific agent can call the correct biological database, receive a syntactically valid record, and still compute about the wrong entity. This paper studies that failure across biologically typed, snapshot-bound identity transitions, where identifier bytes may legitimately change across registries and from gene to protein. BioIdentityBench contains 159 source-resolved tasks over 32 human targets and six perturbation conditions, including one genuinely ambiguous prompt assigned a clarify-or-abstain policy. Chronological replay of 2,452 retained executions identifies 335 trajectories with at least one recorded, admitted, stage-invalid scientific-endpoint dispatch; 241 labels differ from the released last-event scorer. After excluding the 14 ambiguous-task executions, the unique-target aggregate is 323/2,438. On a matched 630-trajectory direct-interface slice, wrong-target execution ranges from 5.6% to 48.4% across five model tiers and is non-monotonic; no scaling-law claim is made. In a frozen-registry joinable subset, 30 wrong-return events across 21 trajectories change the canonical protein accession in all 30 comparisons and recorded sequence length in 29, without establishing a downstream scientific outcome. A complementary systems result specifies scoped references bound to task, run, snapshot, role, and operation: under exclusive broker mediation and a fixed commitment, an unlicensed sibling substitution is inexpressible. Sixteen scoped-reference tests exercise the reference implementation; no model trajectory was rerun through it. The result is an event-sourced measurement protocol and a bounded identity-execution contract, not a new entity linker or a claim of autonomous biological discovery.