paper-with-me

홈 › Papers

Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs

2026-03-03 · Wuyue Zhang, Chongdong Huang, Chunbo You, Cheng Gu, Fengjuan Wang, Mou Sun arxiv

Training large-scale Mixture-of-Experts (MoE) models is bottlenecked by activation memory and expert-parallel communication, yet FP4 training remains impractical on Hopper-class GPUs without native MXFP4 or NVFP4 support. In this work, we present a training recipe that enables MXFP4 efficiency for MoE models on Hopper architectures without native 4-bit computation support. A central challenge is to integrate FP4 into an existing BF16/FP8 hybrid training pipeline without incurring costly precision round-trips (e.g., FP4 $\leftrightarrow$ BF16 $\leftrightarrow$ FP8). We address this challenge by introducing direct FP8-to-FP4 quantization and de-quantization, together with scaling-aware FP4 row-wise to column-wise conversion, enabling FP4 activations and expert-parallel communication with minimal overhead. Core MoE computations are executed in FP8, while activations and expert-parallel communication are compressed using MXFP4, achieving substantial memory and bandwidth savings without degrading convergence. At the 671B parameter scale, our method achieves end-to-end training performance comparable to strong FP8 baselines, while reducing peak activation memory by 14.8\% (11.8 GB) and improving training throughput by 12.5\%, from 1157 to 1302 tokens per GPU per second. These results show that FP4 efficiency can be practically realized for large-scale MoE training through careful software-hardware co-design, even without native FP4 Tensor Core support.

📄 PDF Abstract BibTeX arXiv:2603.02731

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study

2024-09-06 · Jianwei Zhu, Hang Yin, Peng Deng, Aline Almeida 외

This report evaluates the performance impact of enabling Trusted Execution Environments (TEE) on NVIDIA Hopper GPUs for large language model (LLM) inference tasks. We benchmark the overhead introduced by TEE mode across …

CPUGPULanguage ModelingLanguage Modelling+1

Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs

2026-04-25 · Divakar Kumar Yadav, Tian Zhao, Deepak Kumar arxiv

NVIDIA's CUDA Tile (CuTile) introduces a Python-based, tile-centric abstraction for GPU kernel development that aims to simplify programming while retaining Tensor Core and Tensor Memory Accelerator (TMA) efficiency on m…

MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production

2025-05-16 · Chao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong 외

We present MegaScale-MoE, a production system tailored for the efficient training of large-scale mixture-of-experts (MoE) models. MoE emerges as a promising architecture to scale large language models (LLMs) to unprecede…

Mixture-of-Experts

SHOPPER: A Probabilistic Model of Consumer Choice with Substitutes and Complements

2017-11-09 · Francisco J. R. Ruiz, Susan Athey, David M. Blei

We develop SHOPPER, a sequential probabilistic model of shopping data. SHOPPER uses interpretable components to model the forces that drive how a customer chooses products; in particular, we designed SHOPPER to capture h…

counterfactual

SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations

2025-12-16 · Wentao Guo, Mayank Mishra, Xinle Cheng, Ion Stoica 외 arxiv

Mixture of Experts (MoE) models have emerged as the de facto architecture for scaling up language models without significantly increasing the computational cost. Recent MoE models demonstrate a clear trend towards high e…