Activation Functions Shape Token Synchronization in Stochastic Transformer Dynamics
Nikita Karagodin
Abstract
We study how the MLP activation function shapes the long-time geometry of token representations in a deep transformer. We work in the stochastic transformer dynamics of [Agazzi et.al. 2026], in which tokens evolve as interacting particles on the sphere under deterministic attention and a common MLP noise whose covariance is the activation's NNGP kernel. For each of ReLU, ReLU², ReLU³, GELU, SiLU, and tanh, we identify the dimensions at which common noise stops driving tokens to full synchronization and sustains non-degenerate spread. We further prove that, even in the presence of attention, ReLU noise admits dispersion ceilings close to consensus, while ReLU$^2$ noise prevents complete token collapse under an explicit attention strength condition.
Chat is not available.
Successful Page Load