Selectivity Makes State-Space Models Universal Sequence Approximators
Abstract
Selective state-space models such as Mamba replace fixed linear state transitions by input-dependent transitions, yielding recurrent sequence models that retain linear-time inference while adapting their memory dynamics to the current token. We show that even a simple form of this selective memory update has universal approximation power. Specifically, for every compact input domain, finite horizon, and continuous causal sequence-to-sequence map, a clocked diagonal selective state-space scan with strictly contractive gates and a linear-in-state readout uniformly approximates the target map. The proof is constructive: input-dependent diagonal propagation generates tensor-product partition features over prefixes, and the linear-in-state readout combines these features with sampled target values. Thus the high-order interactions required for universality are produced inside the selective scan itself, rather than by attention, nonlinear hidden-state dynamics, or a universal multilayer perceptron applied after a sequence encoder. These results identify selectivity as a mathematically sufficient mechanism for expressive linear-time state-space sequence modeling. Selectivity lets the model decide, token by token, which parts of the past should be preserved, faded, or combined with the present. The universal approximation theorem shows that this adaptive use of memory is sufficient to reproduce any well-behaved causal sequence rule over a fixed horizon.