Hyperbolic Language Models: From Zipf to Compute-Optimal Scaling
Abstract
Compute-optimal scaling laws have become the central organizing tool of language-model training, yet they have been measured exclusively under Euclidean readouts. We give the first compute-optimal scaling-law characterization of hyperbolic language modeling. Trained from scratch on OpenWebText across a paired 5×4 grid spanning 86M to 1.32B parameters and 66M to 8.4B tokens, our hyperbolic LM HypTip follows a markedly different scaling law than its Euclidean counterpart. Fitting the Hoffmann form L = E + A·N^(−α) + B·D^(−β) to each track yields a compute-optimal data-to-parameter ratio of D*/N* = 8.4 for the hyperbolic readout versus 20.8 for the Euclidean baseline. The hyperbolic readout therefore reallocates compute by 2.5× toward parameter-heavy training, winning 17 of 20 paired cells across the grid. This shift is mechanistically grounded: across all 20 trained checkpoints, the per-row Lorentzian radius of the LM head correlates with token log-frequency at Spearman ρ down to −0.83, and the strength of this Zipfian token hierarchy scales with training-token count rather than parameter count, indicating that the geometry-as-prior emerges causally during training rather than being read off a frozen Euclidean checkpoint as in prior post-hoc analyses. Held-out evaluation on WikiText-103 perplexity, LAMBADA, and PIQA preserves the same advantage, and a cross-architecture replication on GPT-2 and DeepSeek-V3 confirms that LM-head geometry, not body details, drives the effect. Output-layer geometry alone therefore reshapes the compute-optimal scaling law of language modeling and brings hyperbolic LMs into the same predictive framework as the Euclidean models that have shaped the field.