Context as Low-Rank Weights: Bounded Parametric Dynamic Memory for Unbounded Context
Abstract
Transformers rely on a growing key–value (KV) cache to store context, causing linear memory growth and limited generalization beyond the training window. We show that contextual influence can instead be viewed as updates to a low-dimensional subspace in the model’s weight space, meaning context can be represented as a bounded, low-rank parameter update rather than token-level storage. Motivated by this, we propose PARADE, which internalizes context as parametric memory. It compresses past inputs into a fixed set of weight-space basis vectors and composes them via query-dependent routing into per-query weight updates, enabling streaming inference with constant memory and effectively unbounded context. Experiments show that PARADE matches full attention on long-context tasks and outperforms efficient baselines, especially when relevant information lies far beyond the attention window, while improving as context exceeds the model’s training length.