paper-with-me

홈 › Papers

MoR: Mixture Of Representations For Mixed-Precision Training

2025-12-28 · Bor-Yiing Su, Peter Dykas, Mike Chrzanowski, Jatin Chhugani arxiv

Mixed-precision training is a crucial technique for scaling deep learning models, but successful mixedprecision training requires identifying and applying the right combination of training methods. This paper presents our preliminary study on Mixture-of-Representations (MoR), a novel, per-tensor and sub-tensor level quantization framework that dynamically analyzes a tensor's numerical properties to select between a variety of different representations. Based on the framework, we have proposed and experimented concrete algorithms that choose dynamically between FP8 and BF16 representations for both per-tensor and sub-tensor level granularities. Our universal approach is designed to preserve model quality across various quantization partition strategies and datasets. Our initial findings show that this approach can achieve state-of-the-art results with 98.38% of tensors quantized to the FP8 format. This work highlights the potential of dynamic, property-aware quantization while preserving model quality. We believe this approach can generally improve the robustness of low precision training, as demonstrated by achieving FP8 accuracies that are on par with existing approaches without the need for fine-grain partitioning, or can be used in combination with other training methods to improve the leverage of even lower precision number formats such as NVFP4.

📄 PDF Abstract BibTeX arXiv:2512.22804

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MoPEQ: Mixture of Mixed Precision Quantized Experts

2025-09-02 · Krishna Teja Chitty-Venkata, Jie Ye, Murali Emani arxiv

Large Language and Vision Models using a Mixture-of-Experts (MoE) architecture pose significant challenges for deployment due to their computational and memory demands. Mixed Precision Quantization assigns different prec…

Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees

2026-04-07 · Mohammed Nowaz Rabbani Chowdhury, Kaoutar El Maghraoui, Hsinyu Tsai, Naigang Wang 외 arxiv

Sparse Mixture-of-Experts (MoE) allows scaling of language and vision models efficiently by activating only a small subset of experts per input. While this reduces computation, the large number of parameters still incurs…

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

2025-05-09 · Haojie Duanmu, Xiuhong Li, Zhihang Yuan, Size Zheng 외

Mixture-of-Experts (MoE) models face deployment challenges due to their large parameter counts and computational demands. We explore quantization for MoE models and highlight two key insights: 1) linear blocks exhibit va…

Mixture-of-ExpertsQuantizationSensitivity

A Sparse Non-Parametric Approach for Single Channel Separation of Known Sounds

2009-12-01 · NeurIPS 2009 12 · Paris Smaragdis, Madhusudana Shashanka, Bhiksha Raj

In this paper we present an algorithm for separating mixed sounds from a monophonic recording. Our approach makes use of training data which allows us to learn representations of the types of sounds that compose the m…

LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes

2025-01-07 · CVPR 2025 1 · Xiang Xu, Lingdong Kong, Hui Shuai, Liang Pan 외

LiDAR data pretraining offers a promising approach to leveraging large-scale, readily available datasets for enhanced data utilization. However, existing methods predominantly focus on sparse voxel representation, overlo…

Mixture-of-ExpertsRepresentation Learning