Implicit-Euler Value Iteration: Long-Horizon Planning via Stable Integration of the Bellman Residual Flow
Nimrod De La Vega ⋅ Amir-massoud Farahmand
Abstract
Value iteration (VI) is the unit-step explicit-Euler discretization of a continuous-time *Bellman residual flow* $\dot V=\operatorname{BR}(V)$, and its standard linear dependence on the effective horizon $1/(1-\gamma)$ can be viewed as reflecting the step-size restriction that explicit integration imposes on stiff dynamics. We propose *Implicit-Euler Value Iteration* (IE-VI): a planning routine equivalent to the implicit-Euler discretization of the same flow, in which each step is unconditionally stable and contracts at rate $1/(1+h(1-\gamma))$ for every step size $h>0$. The implicit step requires solving an equation in the true Bellman operator, which we realize at finite cost by *iterative defect correction*: a short sequence of cheap planning rounds in an approximate model $\hat{\mathcal{P}} \approx \mathcal{P}$ (a simulator or a learned dynamics model), with the true model queried once per round to refresh a residual correction term. For Policy Evaluation and Control, IE-VI converges to the value function of $\mathcal{P}$ for any $\hat{\mathcal{P}}$, however inaccurate---in contrast to model-based planners, which generally converge to a biased fixed point. When the model error shrinks at a sublinear-in-horizon rate, IE-VI attains polylogarithmic iteration complexity in $1/(1-\gamma)$, an exponential improvement over VI; outside this regime its complexity stays within a logarithmic factor of VI. We also give an adaptive variant A-IE-VI, which removes the need to know the model-error scale, and a Dyna-style sample-based variant IE-Dyna for tabular RL.
Chat is not available.
Successful Page Load