CP-MLPs: A Tensor-Rank Theory of Tied and Untied MLP Blocks
Md Rifat Arefin ⋅ Farzaneh Heidari ⋅ Irina Rish ⋅ Guillaume Rabusseau
Abstract
Transformer feed-forward blocks combine several design choices: squared activations, gating, and tied versus untied parameter sharing-whose separate contributions to expressivity remain poorly understood. We introduce \emph{CP-MLPs}, a tensor framework in which each multiplicative unit contributes a rank-1 interaction with three roles: detecting an input feature, reading a value, and writing to the output. Tied gated blocks such as ReLU$^2$ force the detector and value roles to share the same direction; untied gated blocks such as ReGLU and SwiGLU let them differ. We show that this distinction yields exact, degree-wise tensor-rank separations: untied detector-value units are strictly more width-efficient for indefinite quadratic interactions and role-asymmetric monomials, while tied units are structurally matched to symmetric pure-power targets. Tying is therefore not inherently better or worse where its benefit depends on the structure of the target. In experiments, models below the required width hit an irreducible approximation floor where the loss does not improve with extra optimization budget. The detector-value advantage also persists in trained gated MLPs at matched parameter count.
Chat is not available.
Successful Page Load