State-Resolving Attention for Length-Extrapolating Transformers
Abstract
Long-context failures are often framed as retrieval failures: a model must find a relevant token among many distractors. However, many practical long-context settings are better understood as state-tracking problems: dialogue histories contain corrections to earlier preferences, code traces repeatedly assign values to variables, document histories revise previous claims, and tool-use trajectories update the current task state. In these settings, a model must not only identify relevant tokens, but also resolve which relevant occurrence is currently valid. We propose Sieve Attention, a position-encoding-free attention mechanism for length-extrapolating Transformers. SRA decomposes attention into two operations: content filtering and temporal state resolution. For each query, SRA first applies a sparse content filter over the full prefix to identify candidate state updates, and then applies a sequential hazard allocation process only over the selected candidates. This resolves repeated updates according to their relative order without relying on external positional encodings, allowing irrelevant intervening tokens to be ignored before recency is applied. We formalize the mechanism and show that its final attention support is contained in the sparse candidate set, that filtered distractors receive exactly zero mass, and that attention ratios among selected candidates are independent of global sequence length and absolute position. Experiments on repeated associative recall and copy tasks show strong extrapolation far beyond the training context length. On long-context evaluation and language-model pretraining, SRA improves long-context behavior while remaining competitive with standard Transformer baselines. These results suggest that content-first, time-second state resolution is a useful inductive bias for long-context models that must track evolving information.