Silent Operationalization Drift: Prose Scientific Frameworks are Underdetermined Specifications for Agentic Drug Discovery
Hamza Farooq
Abstract
A prose scientific framework is an underdetermined specification: an agent executing it must bind every soft instruction to a concrete operationalization at runtime. In a pre-registered study on a production-scale biomedical knowledge graph, agents ranked drug targets under four arms that vary framework representation and data access. The load-bearing comparison sets prose-plus-tools (A3) against an executable, versioned reasoning model (A4). The two arms agree on every checkable binding, with operationalization fidelity 1.00 for both and a difference of 0.00 (95% CI $[0,0]$), yet they recommend near-disjoint portfolios, with top-10 Jaccard 0.053 in all seeds. One underdetermined aggregation rule, resolved defensibly and disclosed, inverts the conclusion. Because the drift is systematic, with re-walk reproduction 1.0, re-running never reveals it. The divergence replicates on three model families. On DeepSeek-V4-Pro the ambiguous binding is a per-seed coin-flip that alone switches the portfolio between disjoint from the executable answer (Jaccard 0.053) and identical to it (1.000). Under substrate mutation the live-query arm detected 14/14 real perturbations with 0/10 false alarms (Fisher $p \approx 5 \times 10^{-7}$), while the executable arm's precomputed report was stale on 14/14, which shows that versioned reports need an explicit flag for when the underlying data has changed. Controlled adversarial probing elicited zero confabulations on any arm, and our preferred arm fails in its own ways: it faithfully serves a biologically implausible ranking without flagging it. Fidelity is necessary but not sufficient. Decision divergence is the metric that sees the problem.
Chat is not available.
Successful Page Load