HiFloat4 Format for Language Model Pre-training on Ascend NPUs
Training large foundation models at low numerical precision is one of the most promising directions for reducing the compute and memory cost of modern AI. Recent 4-bit floating-point formats such as MXFP4 and NVFP4 can be applied to linear GEMM operations in LLMs, but their limited dynamic range introduces numerical instability that prior work addresses by stacking stabilization mechanisms, typically executed at higher precision and partially eroding the efficiency gains that motivate FP4. In this work, we argue that numerical format design is itself a first-class lever for stable FP4 training, and present the first systematic study of FP4 LLM pretraining on energy-efficient Huawei Ascend NPUs. We compare the recently proposed HiFloat4 (HiF4) format against both MXFP4 and NVFP4, holding one recipe fixed across all three formats across dense (OpenPangu-1B, Llama3-8B) and Mixture-of-Experts (Qwen3-MoE-30B) architectures and executing all linear and expert GEMMs in FP4. At matched storage --- NVFP4 and HiF4 both spend 4.5 bits per value --- the three formats differ far more in what they require before they will train at all than in final accuracy. HiF4 reaches a relative loss of 1.55\% with no stabilization, below fully stabilized MXFP4 (1.79\%) and below NVFP4 carrying the per-tensor scaling it cannot train without (2.00\%); NVFP4 diverges under every combination of stochastic rounding and Hadamard transform we tried. Our results suggest that stable, accurate FP4 training does not require an ever-growing stack of stabilization techniques; it requires the right numerical format.