Curriculum Learning with Compositional Data
Jonas Roth ⋅ Konstantin Nikolaou ⋅ Christian Holm
Abstract
Curricula can reduce pretraining costs, but no principle predicts in advance when one will help. We study this question on the Random Hierarchy Model (RHM), a generative model of hierarchically compositional data whose structure is known. Instead of restricting, our curriculum reweights: it starts from a strongly skewed synonym distribution and broadens it to the target distribution, holding the set of reachable sequences fixed. Across RHM configurations and curriculum end times, we find that the curriculum can speed training up (by up to $2.1\\times$, to the same final loss) but also slow it down (by up to $10\\times$). The advantage peaks long after the curriculum has ended, at $\sim 25\\%$ of training for a curriculum active over the first $10\\%$. Both findings have one explanation. On this data the loss descends in a staircase, one drop per level of the hierarchy, and the curriculum can bring such a drop forward: the advantage is largest while the baseline still waits for its own drop. From this we derive a criterion, based only on the data model and the training budget, that predicts whether a curriculum will help or hurt. It does so for every configuration and end time we test.
Chat is not available.
Successful Page Load