Per-Symbol Exposure, Not Vocabulary Size, Governs Symbol Learnability in Transformers
Syed Shaaz
Abstract
What determines whether a transformer learns the mapping for a symbol? In controlled experiments on vocabularies of $V=8{,}192$ to $131{,}072$ symbols, mapping accuracy is set by independent sequences seen per symbol: training trajectories collapse into one band across this $16\times$ range of $V$, and the per-symbol requirement declines with $V$ (from $\sim 21$ to $\sim 10$ sequences per symbol); larger vocabularies are cheaper per symbol, not harder. These measurements adjudicate a live dispute: a recent peer-reviewed paper argues that the LM head is a gradient bottleneck that renders even trivial patterns unlearnable at large $V$, on the evidence of a synthetic benchmark (SpamLang). The declining requirement is the opposite sign of a bottleneck that tightens with $V/d$, and the published protocol supplies $4.9$ sequences per symbol, terminating before the rise completes, so its runs read as failed but are merely early. Separately, the published task itself is non-diagnostic: at the stated configuration our implementation solves it to accuracy $1.0$ within $5\%$ of the stated budget (under default initialization the tied embedding-head geometry solves it at initialization), and removing weight tying alone reduces accuracy to near-chance. We adjudicate only this synthetic leg; the paper's theory, large-scale convergence results, and gradient measurements are untouched.
Chat is not available.
Successful Page Load