When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity
Abstract
Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) maintains full-history and discounted Beta states for each arm. An anytime-valid e-process authorizes the discounted state, after which a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta–Bernoulli prior-predictive stationary model, the probability that e-ATS ever departs from OTS is bounded by a chosen error level, without fitted thresholds. Removing authorization increased mean normalized dynamic pseudo-regret by 38.4% on the registered suite but reduced it by 7.5% on the literature-derived replay suite. Thus, evidence controls when adaptation begins, while its benefit depends on the environment.