A Refined Sample-Complexity Analysis of Robust Policy Optimization under Decaying Actor Stepsizes
Swetha Ganesh ⋅ Vaneet Aggarwal
Abstract
Robust policy-gradient methods provide a principled framework for learning under model uncertainty, but their sample efficiency is often limited by the cost of robust policy evaluation. Unlike standard Bellman updates, robust Bellman operators are nonlinear, so sample-based critic estimates can introduce systematic error that is difficult to control. Existing finite-sample guarantees therefore often require solving the robust critic to high accuracy before each actor update. However, practical robust actor-critic methods typically use small or decaying actor stepsizes rather than aggressive increasing updates, making repeated high-accuracy critic solves particularly costly. Conservative actor updates are especially natural in robust constrained MDPs, where the optimization objective may switch between reward improvement and constraint correction during learning. We show that this high-accuracy critic requirement is overly conservative for KL-robust policy optimization. Rather than controlling the full first-order critic error, such as $\mathbb E \lVert\widehat V_t - V^{\pi_t}\rVert$, the actor analysis only requires control of the conditional critic bias, $\lVert\mathbb E[\widehat V_t \mid \mathcal F_t] - V^{\pi_t}\rVert$. We prove that, in the relevant local regime, this bias scales quadratically with the critic estimation error. Consequently, an $O(\varepsilon)$ actor error only requires the critic mean-square error to be $O(\varepsilon)$, rather than requiring the root-mean-square error to be $O(\varepsilon)$. Combining this bias-based analysis with robust natural policy-gradient updates, we establish an $\widetilde O(\varepsilon^{-3})$ sample-complexity guarantee for model-free discounted robust MDPs with small or decaying actor stepsizes. We further show that the same analysis extends to primal surrogate methods for robust constrained MDPs under KL uncertainty, improving the state of the art for this setting.
Chat is not available.
Successful Page Load