Rethinking Softmax Attention: Polynomial Activations for Transformers
Abstract
Softmax attention is typically viewed as effective because it produces a normalized probability distribution over input tokens. In this paper, we challenge this view by showing that a key mechanism behind softmax attention is its implicit control of the Frobenius norm of the attention matrix, which stabilizes training. Motivated by this observation, we study alternative attention activations, focusing on polynomial maps that provide a similar regularization effect without satisfying the usual softmax constraints. Our theoretical analysis shows that certain polynomial activations can replace softmax despite violating positivity, row-wise normalization, and sparsity. Extensive experiments across transformer applications show that these alternatives achieve strong performance, suggesting that the success of softmax attention is not inherently tied to its probabilistic interpretation.