Ordinary Least Squares as an Attention Mechanism
Abstract
I show that ordinary least squares (OLS) predictions can be rewritten as a restricted attention module, akin to those at the heart of transformer architectures. This connection emerges by reframing OLS as a similarity-based prediction rule operating in a learned embedding space. From this perspective, least squares does not estimate coefficients per se, but instead selects an embedding that minimizes squared prediction error by matching training and test vectors through inner products — mapping naturally onto the query-key-value structure of attention. I then extend the framework to dimensionality reduction, nonlinearity, and connections to time series econometrics. Monte Carlo simulations and real-data experiments on UCI/OpenML benchmarks show that a direct implementation of nonlinear Attention Regression performs competitively against standard machine learning baselines. In the other direction, within a transformer architecture for tabular data, an attention block can be replaced by an explicit regression on polynomial-expanded features, matching or exceeding the standard transformer's predictive accuracy at a fraction of its parameter count.