Source Verification and Retained Transfer in Biomedical Workflow Memory
Abstract
We stress-test claims of retained improvement in a reproducible audit of retained tabular workflow selection on 12 public biomedical cohorts grouped into eight dependent task families, the effective unit of evidence; target and retention families are excluded from memory construction. At a target budget of two, memory of independently checked source winners has mean Brier loss 0.11205, compared with 0.10832 for paired score memory; the difference is unresolved and favors checked winners in two of eight families. On 40 fresh split seeds (split replications, not new tasks) under a pre-declared, frozen protocol, a matched ablation that holds the winner-count representation fixed finds the verification gate practically equivalent to size-matched random admission and to no gate: every 90% interval lies within ±0.005 Brier, the gate’s own pre-declared margin (pre-declared MDE 0.007–0.015 Brier), which is coarse for retention. This is consistent with representation, not verification, contributing to the gap to score memory. We also place the gate inside VerifyGuard, a scripted (not an LLM) sequential loop spending a score-access budget on discovery and verification before committing a model. Against the same loop without verification, no improvement is detected and all point estimates are unfavourable (+0.00828 Brier; Holm p = 0.375 over eight families), consistent with the guard reverting to the default. Of 109 admitted memory updates, 19 worsen a separate retention family, although the mean change is −0.00120. Source verification, transfer, and preservation are distinct endpoints. The study characterizes a finite computational workflow library, not language-model agents or laboratory campaigns.