Stabilizing RL+Search for Imperfect-Information Extensive-Form Games
Yinghao Li
Abstract
RL+Search is the computational backbone of modern AI systems for imperfect-information extensive-form games. The Recursive Belief-based Learning (ReBeL) framework already mitigates one source of training fragility by sampling the pivot iteration uniformly from the iteration budget; we identify a second, hitherto unaddressed source: the budget itself is held fixed across all subgames, creating a rigid coupling between solver maturity and outer-loop learning that propagates variance through the recursive self-play and produces noisy, seed-dependent training trajectories. We propose \textbf{Stochastic Horizon Annealing (SHA)}, which randomizes only this remaining static degree of freedom: at each subgame the total CFR iteration count is drawn from a Gaussian centered at the expected budget. Combined with ReBeL's existing uniform pivot, the procedure smooths the learning objective across a continuum of solving depths and behaves as an implicit ensemble at zero additional cost. SHA is a one-line change to any ReBeL-style trainer. We give a complete theoretical treatment with three results, all proved in the main text: (i) SHA preserves the $O(N^{-1/2})$ convergence rate of linear CFR; (ii) the per-subgame target variance is bounded by $C_\beta^2\sigma^2/(4N^3)$, vanishing rapidly with the budget; and (iii) the recursive variance of the value-network targets scales as $\Theta(D)$ for SHA versus $\Theta(D^2)$ for the fixed-horizon ReBeL, where $D$ is the recursion depth. On Liar's Dice 1$\times$4f, 1$\times$5f, and 1$\times$6f (30 seeds each) at the ReBeL reference setting of depth-2 subgames and 1024 search iterations, SHA reduces mean final exploitability by 9--18\% and shrinks the 95\% confidence interval across seeds by 34--63\%, with larger games benefiting more as predicted by the theory, establishing stability as a first-class, theoretically supported contribution for RL+Search on imperfect-information games.
Chat is not available.
Successful Page Load