High Probability Risk Control for Online Policy Learning
Abstract
We study risk control for online policy learning, where an agent adaptively collects data and aims to learn a new policy where the risk is below a tolerance threshold at every round with a high probability. Such policy updates and data collection induce feedback covariate shift (FCS): the observed data depend on the history through the evolving policies and data-collection process. FCS brings two substantive problems for risk control: 1) efficient risk evaluation for each candidate policy and 2) valid policy selection with high probability risk control. To address these, we propose an efficient permutation-based weighted risk evaluator based on the Sinkhorn algorithm and select policies whose estimated risk upper confidence bounds (UCB) fall below the target threshold, with the UCB obtained through a proposed Gaussian bootstrap method. Theoretically, we prove that our risk evaluator's worst-case variance is no greater than that of the standard importance-weighted evaluator. Also, our policy control method yields an asymptotically valid simultaneous risk bound and provides asymptotic high-probability risk control for the selected policy. Empirically, our method maintains the target risk level, achieves lower risk than baseline methods, has lower variance in risk estimation, and enables reliable policy selection with high-probability risk control.