Linear approximations to HMM filtering
Abstract
Complex sequence models, such as transformers and state space models (SSMs), learn to represent latent belief states when trained on next-token prediction. Surprisingly, we find that in previously studied tasks, this phenomenon can also be captured by linear recurrent models. This raises a natural question: when are linear models sufficient for optimal belief state approximation? We study this question in the canonical setting of hidden Markov models (HMMs), where Bayes-optimal prediction requires nonlinear filtering. We characterize the full class of HMMs whose filtering dynamics are exactly realizable by linear recurrent neural networks (RNNs). More generally, for arbitrary HMMs, we establish a dissipation relation which shows that approximation error decays at a rate determined by intrinsic properties of the underlying HMM. We validate our theoretical results through experiments on both randomly sampled HMMs and constructed adversarial examples. Together, these findings clarify when linear sequence models suffice for optimal inference and when nonlinearity is fundamentally necessary.