Basis-Mediated Bilinear Attention: A New Method for Greatly Reducing Query--Key Pathway Parameters
Wei Chen ⋅ Wenhao Jiang ⋅ Yiying Yang
Abstract
Large language models and their vision counterparts have achieved remarkable capabilities, yet their wider deployment is increasingly constrained by the computational and memory demands of Transformer architectures. Much recent work improves Transformer efficiency, but the query-key pathway is still typically parameterized by fully independent per-head projections, which place a substantial parameter burden on attention. Motivated by the view that token selection may admit a shared low-dimensional geometry, we propose $\textbf{Basis-Mediated Bilinear Attention} (\textbf{BMB})$, a training-time reparameterization of the query-key pathway in which all heads within a layer share a latent basis while retaining head-specific interactions. A practical explicit variant, BMB-UV, materializes per-head query and key tensors through a shared basis and lightweight head-specific factors, reducing the query-key parameter count from $2d^2$ to $dr+2Hrs$. We evaluate our method family on ViT-Base image classification on ImageNet-1K and on GPT-2 small pretraining on the C4 RealNewsLike subset, and compare it against recent query-key parameter-sharing and low-rank baselines. Relative to standard multi-head attention, our method family reduces query-key parameters by up to $ \textbf{95.83} $%$ $ while remaining competitive with the baseline on both benchmarks.
Chat is not available.
Successful Page Load