Value-Priced Uncertainty: A Family of PPO-Compatible Exploration Bonuses
qili shen ⋅ Xuanhong Chen ⋅ Ang He ⋅ Dake Zhang ⋅ Kairui Feng
Abstract
Proximal Policy Optimization (PPO) is value-driven only after experience has been collected: Generalized Advantage Estimation reinforces sampled actions according to their bootstrapped advantage, but the exploration process that produced those samples is still governed mainly by policy-space noise and entropy regularization. As a result, PPO does not explicitly ask, before sampling, which uncertain transitions would most improve its advantage estimates. We argue that PPO exploration should be formulated as *value-priced uncertainty*: uncertainty should be explored when it is both large and able to change the advantage. We propose **VALU** (**V**alue-**A**ware **L**atent **U**ncertainty), a family of PPO-compatible exploration bonuses of the form b(s,a) = P(s,a) U(s,a). The price $ P(s,a) = \lVert \nabla_{z'} \tilde V_\psi(\hat z') \rVert $ measures how strongly uncertainty in the predicted next latent can perturb the advantage, while \(U(s,a)\) is an uncertainty quantity instantiated by non-parametric effective counts or learned prediction errors such as RND. Our theory derives the value-pricing term from a unified advantage-uncertainty view: an uncertainty coordinate induces an optimistic advantage correction through a dual norm. This yields count-based UCB and MCR members, corresponding respectively to optimism over latent-transition uncertainty and one-sample shrinkage of the optimistic advantage envelope. Across six standard MuJoCo continuous-control tasks, VALU substantially improves PPO's sample efficiency and final return, and achieves the best PPO-compatible performance on most tasks while remaining competitive on the others, compared with strong exploration baselines including RND, ICM, RE3, VCSE, and OPPO. Ablations over uncertainty estimators, schedules, latent encoders, and bonus components show that the gains come from pricing uncertainty by local value sensitivity rather than from any single implementation choice.
Chat is not available.
Successful Page Load