paper-with-me

홈 › Papers

MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling

2025-11-08 · Yu Zhang, Hui-Ling Zhen, Mingxuan Yuan, Bei Yu arxiv

Training large language models with FP8 formats offers significant efficiency gains. However, the reduced numerical precision of FP8 poses challenges for stable and accurate training. Current frameworks preserve training performance using mixed-granularity quantization, i.e., applying per-group quantization for activations and per-tensor/block quantization for weights. While effective, per-group quantization requires scaling along the inner dimension of matrix multiplication, introducing additional dequantization overhead. Moreover, these frameworks often rely on just-in-time scaling to dynamically adjust scaling factors based on the current data distribution. However, this online quantization is inefficient for FP8 training, as it involves multiple memory reads and writes that negate the performance benefits of FP8. To overcome these limitations, we propose MOSS, a novel FP8 training framework that ensures both efficiency and numerical stability. MOSS introduces two key innovations: (1) a two-level microscaling strategy for quantizing sensitive activations, which balances precision and dequantization cost by combining a high-precision global scale with compact, power-of-two local scales; and (2) automatic scaling for weights in linear layers, which eliminates the need for costly max-reduction operations by predicting and adjusting scaling factors during training. Leveraging these techniques, MOSS enables efficient FP8 training of a 7B parameter model, achieving performance comparable to the BF16 baseline while achieving up to 34% higher training throughput.

📄 PDF Abstract BibTeX arXiv:2511.05811

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Microscaling Floating Point Formats for Large Language Models

2025-10-02 · Marco Cococcioni, Dario Pagani, Federico Rossi arxiv

The increasing computational and memory demands of large language models (LLMs) necessitate innovative approaches to optimize resource usage without compromising performance. This paper leverages microscaling floating-po…

Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models

2024-12-15 · Yun-Chen Lo, Gu-Yeon Wei, David Brooks

As cutting-edge large language models (LLMs) continue to transform various industries, their fast-growing model size and sequence length have led to memory traffic and capacity challenges. Recently, AMD, Arm, Intel, Meta…

MMLUQuantization

Microscaling Data Formats for Deep Learning

2023-10-16 · Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall 외

Narrow bit-width data formats are key to reducing the computational and storage costs of modern deep learning applications. This paper evaluates Microscaling (MX) data formats that combine a per-block scaling factor with…

Deep LearningFriction

Is Finer Better? The Limits of Microscaling Formats in Large Language Models

2026-01-26 · Andrea Fasoli, Monodeep Kar, Chi-Chun Liu, Swagath Venkataramani 외 arxiv

Microscaling data formats leverage per-block tensor quantization to enable aggressive model compression with limited loss in accuracy. Unlocking their potential for efficient training and inference necessitates hardware-…

Model Compression

Post Training Quantization of Large Language Models with Microscaling Formats

2024-05-12 · Sayeh Sharify, Utkarsh Saxena, Zifei Xu, Wanzin Yazar 외

Large Language Models (LLMs) have distinguished themselves with outstanding performance in complex language modeling tasks, yet they come with significant computational and storage challenges. This paper explores the pot…

Language ModelingLanguage ModellingQuantization