Spectral Measures of Mamba Conductances Predict and Shape Effective Receptive Fields
Zy Li
Abstract
Mamba's selective state-space model (SSM) achieves long-range dependence through input-dependent gating, yet existing analyses offer limited insight into *why* a particular parameter configuration produces a particular effective receptive field (ERF). We define per-step *conductances* $G_t = I - \exp(A\,\Delta_t)$ from the zero-order-hold discretization and construct a signed spectral measure whose atoms are the per-dimension conductances weighted by the corresponding input-output residues. The impulse response at lag $\tau$ is a discrete Laplace transform of this measure, and the ERF is governed by the per-channel squared impulse response weighted by output amplitude. Heavy tails arise from the mixture across state dimensions, not from temporal randomness within any single dimension. We extract conductances from trained Mamba models and report three empirical findings. First, replacing the signed measure with its total-variation envelope overestimates the ERF by more than three orders of magnitude. Signed cancellations among residues are what make the prediction quantitative. Second, the framework generalizes beyond the original single-layer validation setting. Applied to held-out $A$-initializations, a two-layer model trained from scratch, and all $48$ layers of a pretrained Mamba-370M, the spectral measure tracks the frozen-parameter ERF within $0.85$-$1.5{\times}$ on $41$ of $48$ pretrained layers, with log-log slopes matching to within $0.05$ on $35$ layers and within $0.20$ on all $48$. Third, varying the initialization of $A$ causally shifts the ERF tail slope as the spectral prediction anticipates ($-0.92$ for slow init vs. $-2.6$ for default), and this shift translates to task performance. Across two long-context tasks evaluated over three random seeds, the slow initialization is the most reliable condition: it solves MQAR in $3/3$ seeds (vs. $1/3$ for both default and fast), and attains the highest mean selective-copy accuracy on every distance bucket ($0.68\pm0.13$ overall vs. $0.34\pm0.15$ for fast at distances up to $1500$ tokens, with the ranking preserved in a single-seed extension out to $4000$ tokens). The spectral measure thus provides a controllable design lever for long-range memory. The closed-form prediction reproduces the in-sample frozen-parameter ERF at ratio $1.00$ by construction, which we use as a pipeline consistency check.
Chat is not available.
Successful Page Load