Revisiting the Shape Convention of Transformer Language Models
FengTing Liao ⋅ Guan-Ting Yi ⋅ Tzu-Quan Lin ⋅ Meng-Hsi Chen ⋅ Da-shan Shiu
Abstract
Foundation models under tight compute, memory, and latency budgets need efficient allocation of Transformer capacity across depth, width, and attention, yet dense Transformers still fix this via a narrow-wide-narrow feed-forward network (FFN) convention. We study Hourglass Transformers, replacing the FFN with residual stacks of hourglass (wide-narrow-wide) sub-MLPs and decoupling attention width from a widened residual stream, exposing a depth-width trade-off: fewer layers at matched parameter budgets. Across 113M to 8B parameters, Hourglass Transformers match conventional quality while cutting training compute by $8.7\%$ at matched downstream accuracy (906M--8B). At the $\sim$1B (906M) scale, this yields a $50\%$ KV-cache reduction and up to $1.93\times$ faster 64k-context decoding; at 8B, after long-context extension, Hourglass also outperforms its baseline from 4k to 64k tokens, positioning hourglass shape as a practical lever for on-device-conscious Transformer design.
Chat is not available.
Successful Page Load