paper-with-me

홈 › Papers

Pretraining large language models with MXFP4 on Native FP4 Hardware

2026-05-11 · Musa Cim, Poovaiah Palangappa, Miro Hodak, Ravi Dwivedula, Meena Arunachalam, Mahmut Taylan Kandemir arxiv

Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable? We address this question through a controlled study of MXFP4 quantization in transformer training, progressively enabling FP4 across forward propagation (Fprop), activation gradients (Dgrad), and weight gradients (Wgrad) while holding all other factors fixed. In full pretraining of Llama 3.1-8B on the C4 dataset, we observe that quantizing Wgrad is the primary driver of convergence degradation, whereas FP4 in Fprop and Dgrad alone introduces only modest additional token requirements. To interpret this behavior, we evaluate both structured and stochastic interventions under a controlled experimental setting. We find that stochastic rounding and randomized Hadamard rotations fail to stabilize training once Wgrad is quantized, whereas deterministic Hadamard rotations consistently restore stable optimization. These results suggest that FP4 training instability is driven by structured micro-scaling errors along sensitive gradient paths, rather than by insufficient stochasticity. We run experiments with native MXFP4 support on AMD Instinct MI355X GPUs, enabling controlled investigation of these effects without reliance on software emulation.

📄 PDF Abstract BibTeX arXiv:2605.09825

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction

2026-01-30 · Jatin Chhugani, Geonhwa Jeong, Bor-Yiing Su, Yunjie Pan 외 arxiv

Large Language Models (LLMs) have intensified the need for low-precision formats that enable efficient, large-scale inference. The Open Compute Project (OCP) Microscaling (MX) standard is attractive due to its favorable …

MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving

2025-10-16 · Jungi Lee, Junyong Park, Soohyun Cha, Jaehoon Cho 외 arxiv

Reduced-precision data formats are crucial for cost-effective serving of large language models (LLMs). While numerous reduced-precision formats have been introduced thus far, they often require intrusive modifications to…

Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs

2026-03-03 · Wuyue Zhang, Chongdong Huang, Chunbo You, Cheng Gu 외 arxiv

Training large-scale Mixture-of-Experts (MoE) models is bottlenecked by activation memory and expert-parallel communication, yet FP4 training remains impractical on Hopper-class GPUs without native MXFP4 or NVFP4 support…

Block Rotation is All You Need for MXFP4 Quantization

2025-11-06 · Yuantian Shao, Peisong Wang, Yuanteng Chen, Chang Xu 외 arxiv

Large language models (LLMs) have achieved remarkable success, but their rapidly growing scale imposes prohibitive costs in memory, computation, and energy. Post-training quantization (PTQ) is a promising solution for ef…

TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization

2026-05-19 · Zukang Xu, Xing Hu, Dawei Yang arxiv

As Large Language Models (LLMs) advance toward practical deployment, the Microscaling FP4 (MXFP4) format has emerged as a cornerstone for next-generation low-bit inference, owing to its ability to balance high dynamic ra…