BiRRM: Bidirectional Token Mixing for Recurrent Reasoning Models
Leon A. Hofmeister-Akkad ⋅ Andreas Mayr ⋅ Günter Klambauer
Abstract
Recurrent Reasoning Models (RRMs) achieve strong performance on structured reasoning tasks by repeatedly applying a shared reasoning network, but existing models rely on self-attention or dense sequence-mixing MLPs for cross-token communication. We investigate whether linear-time recurrent sequence mixers can instead serve as the sole token-mixing mechanism. Across mLSTM, GDN-2, and Mamba-3, straightforward causal replacements obtain $0$% exact accuracy on both Sudoku-Extreme and Maze-Hard. This failure motivates BiRRM, a simple weight-tied bidirectional construction that processes both sequence directions and fuses their representations. Applying this construction consistently recovers strong performance across all three mixer families: on Sudoku-Extreme, BiRRM variants reach $91.83$%—$93.05$% exact accuracy, surpassing a self-attention baseline at $68.57$%. On Maze-Hard, BiRRM-mLSTM and BiRRM-Mamba-3 reach $71.46$% and $75.86$%, respectively, compared with $63.12$% for self-attention. BiRRM retains sequence-length-independent parameterization and linear sequence-compute scaling, showing that a common weight-tied bidirectional construction can make otherwise failing causal recurrent mixers effective as the sole token mixers in RRMs.
Chat is not available.
Successful Page Load