Memory Determines Learning Direction: A Theory of Gradient-Based Optimization in State Space Models
Abstract
State space models have shown strong potential to outperform Transformers, but their learning mechanisms remain poorly understood, particularly the conditions for successful parameter updates that preserve teacher information. Prior work attributes the difficulty of such updates to the vanishing gradient during backpropagation through time, in which error signals decay exponentially through repeated gradient multiplication. In this work, by expressing error signals explicitly as functions of inputs, we show that the conventional vanishing gradient framework is insufficient to explain the loss of information. Using this representation of error, we reveal that the memory of the input encoded in teacher information does not decay exponentially, and we analytically demonstrate that the loss of supervisory information cannot be alleviated through training. This result theoretically confirms the importance of the initial weights of SSMs and suggests that the recurrent layer does not require training. By performing experiments on both linear SSMs and a representative nonlinear SSM, including language modeling tasks, we confirm that appropriate initialization enables us to fix the recurrence, leading to more stable training of the entire network and higher performance than the conventional training scheme. Our results partially elucidate the learning mechanisms of SSMs and are expected to contribute to the development of improved RNN-based models.