paper-with-me

홈 › Papers

What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation

2024-03-11 · Zhuocheng Gong, Jiahao Liu, Jingang Wang, Xunliang Cai, Dongyan Zhao, Rui Yan

Quantization has emerged as a promising technique for improving the memory and computational efficiency of large language models (LLMs). Though the trade-off between performance and efficiency is well-known, there is still much to be learned about the relationship between quantization and LLM performance. To shed light on this relationship, we propose a new perspective on quantization, viewing it as perturbations added to the weights and activations of LLMs. We call this approach "the lens of perturbation". Using this lens, we conduct experiments with various artificial perturbations to explore their impact on LLM performance. Our findings reveal several connections between the properties of perturbations and LLM performance, providing insights into the failure cases of uniform quantization and suggesting potential solutions to improve the robustness of LLM quantization. To demonstrate the significance of our findings, we implement a simple non-uniform quantization approach based on our insights. Our experiments show that this approach achieves minimal performance degradation on both 4-bit weight quantization and 8-bit quantization for weights and activations. These results validate the correctness of our approach and highlight its potential to improve the efficiency of LLMs without sacrificing performance.

📄 PDF Abstract BibTeX arXiv:2403.06408

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyQuantization

Similar Papers 제목 키워드 기반

The Uniqueness of LLaMA3-70B Series with Per-Channel Quantization

2024-08-27 · Minghai Qin

We have observed a distinctive quantization-related behavior in the LLaMA3/3.1-70B models that is absent in both the LLaMA2-70B and LLaMA3/3.1/3.2-1B/3B/8B/405B models. Quantization is a crucial technique for deploying l…

Quantization

SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs

2025-12-05 · Ruixuan Huang, Hao Zeng, Hantao Huang, Jinyuan Shi 외 arxiv

Post-training quantization (PTQ) plays a crucial role in the democratization of large language models (LLMs). However, existing low-bit quantization and sparsification techniques are difficult to balance accuracy and eff…

NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models

2026-02-06 · Hyochan Chong, Dongkyu Kim, Changdong Kim, Minseop Choi arxiv

Weight-only quantization has become a standard approach for efficiently serving large language models (LLMs). However, existing methods fail to efficiently compress models to binary (1-bit) levels, as they either require…

FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference

2025-04-19 · Coleman Hooper, Charbel Sakr, Ben Keller, Rangharajan Venkatesan 외

Quantization is a powerful tool to improve large language model (LLM) inference efficiency by utilizing more energy-efficient low-precision datapaths and reducing memory footprint. However, accurately quantizing LLM weig…

Large Language ModelQuantization

Dynamic Stashing Quantization for Efficient Transformer Training

2023-03-09 · Guo Yang, Daniel Lo, Robert Mullins, Yiren Zhao

Large Language Models (LLMs) have demonstrated impressive performance on a range of Natural Language Processing (NLP) tasks. Unfortunately, the immense amount of computations and memory accesses required for LLM training…

Quantization