paper-with-me

홈 › Papers

NUPES : Non-Uniform Post-Training Quantization via Power Exponent Search

2023-08-10 · Edouard Yvinec, Arnaud Dapogny, Kevin Bailly

Deep neural network (DNN) deployment has been confined to larger hardware devices due to their expensive computational requirements. This challenge has recently reached another scale with the emergence of large language models (LLMs). In order to reduce both their memory footprint and latency, a promising technique is quantization. It consists in converting floating point representations to low bit-width fixed point representations, usually by assuming a uniform mapping onto a regular grid. This process, referred to in the literature as uniform quantization, may however be ill-suited as most DNN weights and activations follow a bell-shaped distribution. This is even worse on LLMs whose weight distributions are known to exhibit large, high impact, outlier values. In this work, we propose an improvement over the most commonly adopted way to tackle this limitation in deep learning models quantization, namely, non-uniform quantization. NUPES leverages automorphisms to preserve the scalar multiplications. Such transformations are derived from power functions. However, the optimization of the exponent parameter and weight values remains a challenging and novel problem which could not be solved with previous post training optimization techniques which only learn to round up or down weight values in order to preserve the predictive function. We circumvent this limitation with a new paradigm: learning new quantized weights over the entire quantized space. Similarly, we enable the optimization of the power exponent, i.e. the optimization of the quantization operator itself during training by alleviating all the numerical instabilities. The resulting predictive function is compatible with integer-only low-bit inference. We show the ability of the method to achieve state-of-the-art compression rates in both, data-free and data-driven configurations.

📄 PDF Abstract BibTeX arXiv:2308.05600

Code (0)

등록된 구현이 없습니다.

Tasks

Quantization

Similar Papers 제목 키워드 기반

AdaLog: Post-Training Quantization for Vision Transformers with Adaptive Logarithm Quantizer

2024-07-17 · Zhuguanyu Wu, Jiaxin Chen, Hanwen Zhong, Di Huang 외

Vision Transformer (ViT) has become one of the most prevailing fundamental backbone networks in the computer vision community. Despite the high accuracy, deploying it in real applications raises critical challenges inclu…

Instance Segmentationobject-detectionObject DetectionQuantization

HPTQ: Hardware-Friendly Post Training Quantization

2021-09-19 · Hai Victor Habi, Reuven Peretz, Elad Cohen, Lior Dikstein 외

Neural network quantization enables the deployment of models on edge devices. An essential requirement for their hardware efficiency is that the quantizers are hardware-friendly: uniform, symmetric, and with power-of-two…

object-detectionObject DetectionPose EstimationQuantization+1

Q-Rater: Non-Convex Optimization for Post-Training Uniform Quantization

2021-05-05 · Byeongwook Kim, Dongsoo Lee, Yeonju Ro, Yongkweon Jeon 외

Various post-training uniform quantization methods have usually been studied based on convex optimization. As a result, most previous ones rely on the quantization error minimization and/or quadratic approximations. Such…

Quantization

PoTPTQ: A Two-step Power-of-Two Post-training for LLMs

2025-07-16 · Xinyu Wang, Vahid Partovi Nia, Peng Lu, Jerry Huang 외 arxiv

Large Language Models (LLMs) have demonstrated remarkable performance across various natural language processing (NLP) tasks. However, their deployment is challenging due to the substantial computational resources requir…

Hybrid and Non-Uniform quantization methods using retro synthesis data for efficient inference

2020-12-26 · Tej pratap GVSL, Raja Kumar

Existing quantization aware training methods attempt to compensate for the quantization loss by leveraging on training data, like most of the post-training quantization methods, and are also time consuming. Both these me…

Quantization