Below One Bit: Training Route and Budget Shape Robustness Under On-Device Storage Constraints
Abstract
Sub-1-bit weight quantization fits language models into tight on-device storage, yet such models are judged almost entirely on clean text. We ask whether the route to those bits changes robustness to corrupted input. We train byte-level models from scratch (0.86M to 10.03M parameters, six seeds each) and reach the same sub-1-bit format two ways: native quantization-aware training (QAT), or post-training quantization of a full-precision model, then fine-tuning in the format (PTQ+FT). With equal tokens trained in the format, the native model degrades less on Python code in every setting and seed (averaged over nine character corruptions), and again on an independent Python corpus. PTQ+FT nonetheless stays the better model. On natural language the native advantage is far smaller, and no setting in the grid passes a seed-level test. The token budget sets the size and sign of the advantage: quadrupling the native model's tokens reverses it at both sizes tested, and equal total training tokens leave only small residuals of either sign. A PTQ+FT checkpoint nearest the native model's clean quality still degrades more; at the smallest size, 44 percent of the native advantage remains. On an Apple M5 Max the packed model is 8.8 to 22.9 times smaller than fp16 but, without a fused sub-1-bit kernel, up to 16.1 times slower at batch 1 than unpacked weights. Bit width alone does not describe a compressed on-device model; the training route and token budget belong beside it.