Improving the Exploration Behavior of Generative Recurrent Reasoning models with xRRM
Abstract
Recurrent Reasoning Models (RRMs) including HRM and TRM achieve strong performance on combinatorial reasoning tasks such as Sudoku and Maze while remaining parameter efficient. HRM and TRM are in principle looped transformers in which a shared module is applied repeatedly to refine a neural state. Conceptually the computation is grounded in a fixed-point equation. Training is unrolled with deep supervision and the use of stop-gradient. Recently, stochastic variants of RRMs have been proposed (e.g., EqR, PTRM). Beyond enabling generative modeling, stochasticity provides a natural mechanism for improving exploration of the solution space. In particular, these variants induce stochasticity by randomizing the initial latent states and injecting noise during recurrent iterations. However, their gains rely on exploring multiple trajectories at inference. This raises the question of how to encourage exploration more directly within the iterative fixed-point dynamics. We address this with xRRM, a Recurrent Reasoning Model trained with a strategy adapted from generative modeling: at each deep-supervision step, the loss is evaluated for all sampled trajectories, but gradients are computed only from the trajectory whose output is closest to the correct solution, effectively implementing a winner-takes-all objective and restricting backpropagation to the best-matching trajectory. On Sudoku-Extreme, xRRM achieves an average fully solved rate of more than 90 % at inference time with a single evaluation candidate, improving over EqR by more than 20 %. Test-time scaling with 32 evaluation increased this value to more than 99\%. To evaluate xRRM as a conditional generative model, we evaluate it on Game-of-24, a benchmark with non-unique solutions. xRRM likewise outperforms all other RRMs by a large margin.