Iterative Nonlinear Computation Underlying Abstract Reasoning
Abstract
Universal Transformers (UTs) show strong performance on abstract reasoning tasks such as ARC-AGI, yet the specific sources of their performance gains remain underexplored. We conduct systematic ablations over UT variants and find that performance improvements on abstract reasoning are driven primarily by (i) the recurrent inductive bias induced by parameter sharing across depth and (ii) strong nonlinearity inside the Transformer block, rather than by elaborate architectural heuristics. Motivated by this finding, we propose the Universal Reasoning Model (URM). First, we propose ConvSwiGLU, which augments the feed-forward block with a channel-wise short convolution to effectively enhance nonlinearity of UT. Second, we adopt Truncated Backpropagation Through Loops (TBPTL) to eliminate noise from early loop when training with many loops, by freezing the gradients of early loops. Our approach substantially improves abstract reasoning performance, achieving 57.5% pass@1 on ARC-AGI 1 and 16.0% pass@1 on ARC-AGI 2. More importantly, we pretrain a 8-billion-parameter large language model based on the URM architecture from scratch, and the results show that URM still substantially outperforms the baseline across all reasoning tasks.