Persistent Procedure Memory in Biomedical Analysis: Selection and Correction
Abstract
Analysis agents reuse libraries of validated procedures, involving three decisions: whether to reuse a stored procedure, how much to search before accepting an update, and what to rebuild after a source correction. We give a component-level evaluation protocol: a controller retains identifiers from a catalogue of 64 executable classification procedures across nine public biomedical tables, with static-default, reset and random-search controls and no language model. No memory policy establishes a gain: the primary contrast is -0.70 percentage points [-4.01, 1.91], and the static default has the highest observed mean. A post-hoc decomposition locates the loss. The retrieved procedure starts 4.30–7.12 points below the default, search recovers 2.76–3.98, and each policy ends 1.13–3.14 points below it, while pooled logged changes average +3.42 points against their own start. Validation and audit rankings agree moderately (mean Spearman 0.51). A canonical +2.06-point gain from raising search from four to 64 candidates is not reproduced: over 30 pre-declared fresh partitions the two primary contrasts are -0.40 and +0.17 points, both within a post-hoc ±2-point margin on these nine datasets, and a same-source transfer test shows no memory gain, and a planted better procedure is detected against static (+13.52 points) but found equally by reset. In a constructed-fault cache test, a dependency ledger and a root fingerprint each remove all 236 stale disagreements, a missing edge leaves them, and repair lowers balanced accuracy on three datasets. Science-agent procedure memory should be judged against a no-memory default and a reset control on outcomes the agent never saw.