Train by Efficiency, not by Difficulty: Target-Directed Curricular RL via Gain Proxies
Abstract
Developmental psychology and pedagogy have inspired many ways of sequencing practice towards competence, curriculum learning and autocurricula among them. We focus on target-directed curricula, where practice is sequenced to learn a particular target task, similar to how babies learn to walk, rather than their open-ended free play. Pedagogy measures and controls tasks' \textit{difficulty} to sequence a curriculum; while tuning difficulty is core to gradual skill development, we argue \textit{efficiency} matters more when learning a difficult target under tight resources. We price practice tasks by \emph{net gain} --- the future learning cost they save towards target competence, minus their own cost. Modelling a curriculum as two reinforcement learning agents in a teacher--student loop, greedy selection of practice tasks by net gain reduces to guaranteed policy improvement. Paradoxically, computing net gain requires entire budget rollouts, undermining the very learning efficiency the framework seeks to optimize. We propose an amortized net gain critic, a value estimator that sidesteps this cost. On a benchmark where net gain is computable, we stress-test seven existing methods and our critic against control baselines such as direct-target learning and oracle access to net gain as a ground truth. Preliminary results show our approach produces significantly more efficient curricula and reaches target competence in most seeds, where rival teachers inspired by difficulty tuning and performance progression do not. We hope insight from the developmental sciences community will elevate this method into a protocol for efficient learning at scale.