GEM: Interpretable Language Models via Geometric Embedding Alignment
Abstract
Large language models encode rich representations in their hidden states, yet these representations remain largely opaque and disconnected from the model's vocabulary. We introduce Geometric Embedding Mixture (GEM), a pre-training regularization method that aligns transformer hidden states with the geometry of the vocabulary embedding space. GEM utilizes a self-supervised objective to encourage the final hidden state to approximate a probability-weighted mixture of token embeddings. This effectively minimizes the free energy of the models and pushes the representation into the convex hull of the vocabulary. We pre-train models at three scales (135M, 360M, and 760M) and demonstrate that GEM yields superior predictive uncertainty compared to standard training, reducing perplexity by up to 53\% on benchmarks. Furthermore, we show that this geometric structural constraint mitigates representation anisotropy and unlocks \emph{native interpretability}. Unlike baseline models, GEM enables direct layer-wise decoding of intermediate hidden states and better causal faithfulness using the untuned language head, revealing coherent semantic trajectories without the need for auxiliary probes. Our work demonstrates that imposing geometric structure during pre-training improves both performance and transparency, offering a path toward language models that are interpretable by construction.