Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
Abstract
Depth-recurrence promises improved latent reasoning by sharing parameters across depths. However, prior work still relies on partially fixed layer stacks and overlooks the bottleneck of constant hidden-size; moreover, it largely lacks rigorously resource-matched ablations. To address this, we introduce depth-recurrent attention mixtures (Dreamer), a modular architecture with a single recurring layer. Concretely, we mix sequence attention, depth attention, and sparse expert attention. This alleviates the hidden-size bottleneck, decouples scaling dimensions, and achieves effective yet efficient depth-recurrence. Across natural language reasoning benchmarks at up to ~2B parameters, Dreamer requires on average ~3x fewer training tokens for the same accuracy as FLOP-, parameter-, and memory-matched modern Transformers, and outperforms ~2x larger Transformers given same training tokens. Analysis further reveals 2-11x greater expert selection diversity than conventional MoEs, highlighting the flexible knowledge sharing across depths.