Action-Level Behavior Policy Optimization for Variance-Reduced Policy Gradients
Abstract
Policy gradient methods may suffer from high variance, which limits their sample efficiency. Existing variance-reduction techniques largely focus on estimator design under a fixed sampling policy. In contrast, we ask whether the action-sampling distribution itself, often left at the on-policy default, can be optimized to reduce policy gradient variance. We show that this default is in fact not variance-optimal even when the state distribution is held on-policy. To exploit this, we introduce Action-Level Behavior Policy Optimization (BPO), a framework that treats the behavior policy as an action-level proposal distribution at each on-policy state, while keeping the state distribution itself on-policy. For a one-step importance-sampled policy gradient estimator, we sharply decompose its variance into an irreducible state term and a controllable action term, and show that the latter can be minimized in closed form via BPO. Under exact critic and sampling oracles, this yields a strictly tighter nonconvex SGD bound than on-policy sampling whenever the score-weighted critic is action-dependent. We then extend the analysis to approximate critics and parametrized behavior policies updated by a single stochastic step, and verify the resulting sample-efficiency gains empirically. These results identify action-level BPO as a principled mechanism for improving the sample efficiency of policy gradient methods.