HiFloat4 Format for Language Model Pre-training on Ascend NPUs
Abstract
Training large foundation models at low numerical precision is one of the most promising directions for reducing the compute and memory cost of modern AI. Recent 4-bit floating-point formats such as MXFP4 and NVFP4 can be applied to linear GEMM operations in LLMs, but their limited dynamic range introduces numerical instability that prior work addresses by stacking stabilization mechanisms, typically executed at higher precision and partially eroding the efficiency gains that motivate FP4. In this work, we argue that numerical format design is itself a first-class lever for stable FP4 training, and present the first systematic study of FP4 LLM pretraining on energy-efficient Huawei Ascend NPUs. We compare the recently proposed HiFloat4 (HiF4) format against MXFP4 across both dense (OpenPangu-1B, Llama3-8B) and Mixture-of-Experts (Qwen3-MoE-30B) architectures, executing all linear and expert GEMMs in FP4. HiF4's hierarchical scaling provides enough representational headroom that a single stabilization step suffices to keep the relative loss within 1\% of a full-precision baseline — roughly half the gap of MXFP4, which requires three such mechanisms — while incurring less than 1\% average degradation on downstream tasks. Our results suggest that stable, accurate FP4 training does not require an ever-growing stack of stabilization techniques; it requires the right numerical format.