Mixtures of Embedding Dimension Experts
George Retsinas ⋅ Patrick Judd ⋅ Chong Yu ⋅ Mikail Khona ⋅ Tijmen Blankevoort ⋅ Mohammad Shoeybi ⋅ Jeff Pool
Abstract
We introduce Mixtures of Embedding Dimension Experts (MoDEs), a novel alternative to standard linear layers. MoDEs push the concept of Mixtures of Experts (MoEs) to the embedding dimension; rather than atomically routing each token to $K$ of $E$ total experts, each embedding dimension in each token is routed to $N$ of $M$ total embedding dimension experts. This formulation results in a standard sparse-vector $\times$ dense matrix workload for single-token processing, and its $N$:$M$ structure exposes a sparse operand that is exploitable by future kernels for multi-token processing. We show empirically that applying our design to a single linear layer within standard MLP blocks exposes new Pareto-optimality for dense transformers in quality vs. FLOPs space. We compose MoDEs with standard MoEs and find the combination offers better scaling in quality vs. expert parameters compared to token-level experts alone. Finally, we demonstrate that MoDEs offer a realized 2\% and projected 16-35\% speedups for single-token decoding workloads.
Chat is not available.
Successful Page Load