Virtual Double Oracle: Faster Convergence by Leveraging Strategies That Were Watching All Along
Ruilan Wang ⋅ Francisco Aristi Reina ⋅ Katerina Papadaki
Abstract
Double Oracle (DO) and PSRO approximate Nash equilibria in two-player zero-sum games by iteratively expanding a strategy population with best responses. Online Double Oracle (ODO) interleaves multiplicative weights updates (MWU) with discovery and obtains an $O(\sqrt{k \ln k / T})$ rate. We identify the \emph{reinitialisation} step as an underexplored design choice: at each expansion, MWU-based methods reset newly discovered strategies to a uniform prior. We introduce Virtual Double Oracle (VDO), which exploits a bilinear identity---a strategy's full hindsight loss history collapses to its utility against the time-averaged opponent---to initialise each new strategy with the MWU state it would have had from the start. Conditional on the discovered population, VDO recovers the fixed-population rate $O(\sqrt{\ln k / T})$ asymptotically, removing ODO's leading-order $\sqrt{k}$ window-decomposition penalty. Discounted VDO adds a single parameter $\beta$ to control finite-horizon prior strength. Experiments across matrix games, poker, and Goofspiel validate both contributions: pure VDO confirms the asymptotic rate, and discounted VDO gives the strongest finite-horizon performance.
Chat is not available.
Successful Page Load