Reinforced Fast Weights via Next-Sequence Prediction
Hee Seung Hwang ⋅ Xindi Wu ⋅ Sanghyuk Chun ⋅ Zhiwei Deng ⋅ Olga Russakovsky
Abstract
Fast weight architectures offer a promising alternative to standard transformers for long-context modeling by replacing the KV-cache with a recurrently updated fixed-size memory. Despite this architectural shift, they are typically trained with the same next-token prediction (NTP) objective as standard transformers. This creates a mismatch: fast weight models rely on an evolving memory state to support future predictions, while NTP provides only token-level supervision for the immediate next token. We address this mismatch with next-sequence prediction (NSP), a sequence-level extension of NTP that trains models to produce coherent multi-token continuations from their recurrent memory state. To optimize this sequence-level objective, we propose ReFINE ($\textbf{Re}$inforced $\textbf{F}$ast we$\textbf{I}$ghts via $\textbf{N}$ext s$\textbf{E}$quence prediction). ReFINE selects informative token positions based on prediction entropy, generates multi-token rollouts, assigns sequence-level rewards, and optimizes the model with Group Relative Policy Optimization (GRPO). ReFINE is applicable throughout the training lifecycle of pre-trained language models: mid-training, post-training, and test-time training. Experiments on LaCT-760M and DeltaNet-1.3B show that ReFINE consistently outperforms NTP-based supervised fine-tuning in needle-in-a-haystack retrieval, long-context question answering, and diverse tasks in LongBench. ReFINE provides an effective and versatile framework for improving long-context modeling in fast weight architectures.
Chat is not available.
Successful Page Load