What a Tool-Choosing Molecular Agent Should Retain: Support, Feedback and Tool Budget
Abstract
We present an audit framework for evaluating scientific tool routers and apply it to molecular string repair; we test no LLM and make no biological-endpoint or deployment claim. A tool router is a policy that orders a finite tool menu from observable input features, calls tools in that order, stops at the first output passing a runtime check, and is charged for every call, including an optional expensive one; it is retained if its order persists across episodes. Runtime acceptance is parser validity; reference recovery is scored offline with privileged labels. On 1,184 held-out one-error cases, a contextual static order recovers 119 references against an evaluator-only envelope of 120, and neither retained rankings nor an online validity bandit improves on it. The core held-out result is pre-declared. On 2,000 cases whose parents and scaffolds are absent from the neural checkpoint's training archive, identity supervision mostly buys calls to a pretrained neural corrector: at an equal call budget, identity training adds 1.35 percentage points over neural repair alone, whereas the unmatched 45.6-point contrast with a validity-trained order comes with 2,000 versus 89 neural calls. A pre-declared replication on 1,744 fresh cases gives the same pattern (0.6 points at equal calls) but is memorization-exposed: 1,690 of its parents occur in that archive. Validity feedback is indifferent to which tool recovers the reference. Before crediting retention, evaluations should match support, feedback target and tool-call budget.