Distributionally-Robust Policy Learning from Observational Data
Debmalya Mandal
Abstract
Policy learning from observational data is fundamental to applications ranging from personalized medicine to welfare-program targeting, but standard approaches optimize empirical performance and may fail when the deployment distribution differs from the training distribution. We study \emph{distributionally-robust policy learning}: given observational data $\\{(X_i,W_i,Y_i)\\}_{i=1}^n$ with binary treatment $W$, we seek a policy $\pi \in \Pi$ that maximizes the worst-case policy value over an $f$-divergence ball of radius $\rho$ around the data-generating distribution. We propose an estimator built on cross-fitted doubly-robust scores for the conditional treatment effect, paired with a DRO solver that handles \texttt{KL}, $\chi^2$, and \texttt{CVaR} ambiguity sets. We establish two convergence guarantees. First, the estimated policy attains the standard $\widetilde{O}(n^{-1/2})$ rate on its DRO regret, and inherits the doubly-robust property from its non-DRO counterpart -- the rate continues to hold whenever either the outcome model or the propensity model is consistently estimated, even at slow nonparametric rates. We then propose a sample-splitting bias-correction algorithm derived from the Lagrangian dual, and show that under a Tsybakov-style margin condition on the worst-case treatment-effect contrast, the estimator attains a fast rate of $\widetilde{O}(n^{-(1+\alpha)/(2+\alpha)})$ , interpolating between $\widetilde{O}(1/\sqrt{n})$ and $\widetilde{O}(1/n)$ as the margin parameter $\alpha$ ranges over $[0,\infty)$. To our knowledge, this is the first work to incorporate a Tsybakov margin condition into the analysis of distributionally-robust policy learning, and the first to obtain fast rates in this setting. Empirically, on semi-synthetic experiments calibrated to a welfare-to-work program, distributional robustness yields meaningful improvements over empirical welfare maximization when training and test distributions differ structurally, while remaining competitive when no shift is present.
Chat is not available.
Successful Page Load