What Ten Seeds Can and Cannot Resolve About Replay Strategy Selection Under Distribution Shift
Abstract
Replay-strategy comparisons under non-stationarity — a design choice every continually adapting agent embeds — are normally reported as a ranking over policies, plus a claim that the ranking depends on the metric. We evaluate 8 adaptation policies at a fixed update and replay-sample budget across 5 drifting environments at ten seeds (five in the process-control case), and ask which of those claims the data can carry. The ordering fails, though not for the usual reason: interval overlap is not a test of a difference, and an exact within-seed permutation test that accounts for the extremes being selected post hoc separates best from worst in 3 of 5 environments. What fails is everything finer — 2 of 35 adjacent leaderboard steps separate, and the modal winner keeps rank one in as little as 0.47 of seed resamples. A coarse contrast between recency-biased and non-recency replay survives in 2 of 5 environments under an exact paired test and multiplicity correction. It survives where the ranking does not because a seed effect common to every policy carries up to 0.91 of the return variance, which pairing removes, and because a sorted leaderboard places its near-ties adjacent, which pairing cannot rescue. The metric-dependence claim needs a reference this literature lacks: two metrics that are both sampling noise disagree at chance, so we compare each cross-metric flip rate to the floor implied by the metrics' own split-half self-disagreement. Only 2 of 24 comparisons clear it, and those two are a single finding once a redundancy between two of the metrics is priced. A power analysis shows the empty cells rule out gross disagreement while leaving modest disagreement open — an upper bound, not agreement. Two domain findings do survive: purely recent replay sits significantly above the other adaptive policies on one environment and significantly below them on another, and equalising update and sample budgets still leaves wall-clock differing by 1.21–1.45× in our implementation, in the direction that favours the policies already ahead.