Did Feedback Cause the Gain? A Factored Replay Audit for Molecular Acquisition
Abstract
A rising hit curve across design–make–test–analyze rounds does not identify whether truthful feedback, inherited learner state, repeated selection, or query budget caused the gain. We introduce a factored replay audit crossing faithful versus molecule-misaligned outcomes with warm-start versus cold-reset cumulative replay, while matching the pool, initial observations, features, learner, acquisition rule, and budget; frozen-ranking and random arms provide no-update and budget anchors. After correcting a candidate/audit partition defect and rerunning every frozen seed, 2,580 trajectories over five public tasks gave an Enamine10k faithful-versus-sham alignment effect of 5.35 hits per 100 post-initial queries (95% interval 4.94–5.77). Faithful adaptive replay also beat frozen ranking by 3.19 (2.86–3.55). Alignment effects were positive on four additional tasks (2.91–15.47) and with an MLP. Target-block nulls met equivalence margins; teachers were strongly positive, and a pre-frozen fresh-seed Enamine sensitivity replicated both behaviors. Warm-start state was heterogeneous and slightly harmful on the primary SGD task. Public campaign traces lacked matched arms and state lineage. Thus, under stationary finite-pool replay, feedback mattered and faithful molecule–outcome alignment yielded a substantial matched policy effect; this is not prospective wet-lab or therapeutic evidence.