xRRMs: Explorative Recurrent Reasoning Models
Abstract
Recurrent Reasoning Models (RRMs) such as Hierarchical Reasoning Models (HRM) and Tiny Recursive Models (TRM) solve combinatorial reasoning problems like Sudoku and Game-of-24 with a few million parameters and roughly a thousand training examples. In principle, RRMs are looped transformers, that employ several computational loops called deep supervision steps, for which the loss is calculated and backpropagated. This strategy forces trajectories of RRMs to approach fast to a fixed-point, which should represent a correct solution of the problem. However, combinatorial reasoning problems can benefit from extensive exploration, which is the core element of many symbolic solvers, such as backtracking. Thus, we hypothesize that RRMs can be improved by encouraging them to explore more candidate solutions while maintaining the deep supervision concept that enabled effective training of RRMs. We propose Explorative Recurrent Reasoning Models (xRRM), which encourage exploration by randomly sampling different initial states and injecting noise with a learned variance. This leads to multiple trajectories in parallel. During training, only the trajectory closest to a valid solution receives a gradient signal. We compare xRRM against other RRMs on Sudoku-Extreme, which has only one correct solution and Game-of-24, which can have multiple correct solutions. In both cases, xRRM outperforms the baselines, even with the same compute budget during training and inference.