Compose Your Oracles: Off Policy Improvement with Aggregated Guidance
Abstract
Learning from multiple imperfect oracles is a practical way to improve reinforcement learning (RL) under limited interaction. However, most contemporary robust multi-oracle methods are on-policy and equire frequent fresh rollouts, which limits sample efficiency in domains where data collection is expensive. We propose Oracle-Aggregated Policy Improvement (OPI), an off-policy framework for policy improvement with multiple suboptimal oracles. OPI augments the oracle set with the learner, constructs an oracle-guided behavior policy through candidate-action pooling and source-aware critic scoring, and trains the learner using a generalized advantage-weighted behavior cloning objective combined with deterministic policy gradients. The method is motivated by a tractable proxy for the max-advantage objective and a policy-improvement relative to a strong oracle baseline. We evaluate OPI on diverse tasks across MetaWorld, the DeepMind Control Suite, and customized Box2D environments with heterogeneous oracles. Across sparse-reward manipulation, dense-reward locomotion, and customized control settings, OPI is strongest in sparse-reward and complementary-oracle regimes and remains competitive with representative baselines in dense-reward environments. Additional analyses show that oracle composition is most beneficial when oracle skills are complementary and that estimating oracle-specific action values is particularly helpful in sparse-reward regimes