paper-with-me

홈 › Papers

With Shared Microexponents, A Little Shifting Goes a Long Way

2023-02-16 · Bita Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, Lei Shao, Gaurav Kolhe, Dimitry Melts, Jasmine Klar, Renee L'Heureux, Matt Perry, Doug Burger, Eric Chung, Zhaoxia Deng, Sam Naghshineh, Jongsoo Park, Maxim Naumov

This paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-point and block floating-point. MX utilizes multiple levels of quantization scaling with ultra-fine scaling factors based on shared microexponents in the hardware. The effectiveness of MX is demonstrated on real-world models including large-scale generative pretraining and inferencing, and production-scale recommendation systems.

📄 PDF Abstract BibTeX arXiv:2302.08007

Code (1)

rocm/tensorcast pytorch

Tasks

QuantizationRecommendation Systems

Similar Papers 제목 키워드 기반

Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models

2024-12-15 · Yun-Chen Lo, Gu-Yeon Wei, David Brooks

As cutting-edge large language models (LLMs) continue to transform various industries, their fast-growing model size and sequence length have led to memory traffic and capacity challenges. Recently, AMD, Arm, Intel, Meta…

MMLUQuantization

Predicting the longevity of resources shared in scientific publications

2022-03-24 · Daniel E. Acuna, Jian Jian, Tong Zeng, Lizhen Liang 외

Research has shown that most resources shared in articles (e.g., URLs to code or data) are not kept up to date and mostly disappear from the web after some years (Zeng et al., 2019). Little is known about the factors tha…

Articles

CLULEX at SemEval-2021 Task 1: A Simple System Goes a Long Way

2021-08-01 · SEMEVAL 2021 · Greta Smolenska, Peter Kolb, Sinan Tang, Mironas Bitinis 외

This paper presents the system we submitted to the first Lexical Complexity Prediction (LCP) Shared Task 2021. The Shared Task provides participants with a new English dataset that includes context of the target word. We…

Feature EngineeringLexical Complexity PredictionPredictionWord Embeddings

Play It Cool: Dynamic Shifting Prevents Thermal Throttling

2022-06-22 · Yang Zhou, Feng Liang, Ting-Wu Chin, Diana Marculescu

Machine learning (ML) has entered the mobile era where an enormous number of ML models are deployed on edge devices. However, running common ML models on edge devices continuously may generate excessive heat from the com…

CPU

BART goes multilingual: The UniTN / Essex submission to the CoNLL-2012 Shared Task

2012-07-01 · WS 2012 7 · Olga Uryupina, Aless Moschitti, ro, Massimo Poesio
Coreference Resolution