TrunkFish: Making Model Width Incrementally Refinable
Abstract
Width is a central axis for scaling LLM families, but widening a standard transformer typically changes the hidden representation rather than continuing the smaller-width computation. We introduce Trunkfish, a recipe for converting model architectures into a hierarchy of models with incrementally refinable width (a Trunkfish hierarchy) by enforcing causal hidden dimensions and cumulative prefix readouts: new width may depend on old width, but old width may not depend on new width. This enables schooling, which trains all models in a hierarchy from forward/backward passes of only the largest model, and restart-free cascades, which enable adaptively scaling up inference compute without wasting compute already spent on smaller models. Trunkfish improves the validation-loss/active-parameter frontier over independently trained standard hierarchies under matched hierarchy-training FLOPs, with substantial training-FLOP savings for both full-width and prefix-heavy schooling. In cascade simulations over the same trained hierarchy, exact width continuation also meaningfully reduces matched-loss expected active-parameter traffic relative to restart cascades.