paper-with-me

홈 › Papers

GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance

2025-05-11 · Jinuk Kim, Marwa El Halabi, Wonpyo Park, Clemens JS Schaefer, Deokjae Lee, Yeonhong Park, Jae W. Lee, Hyun Oh Song

Post-training quantization is a key technique for reducing the memory and inference latency of large language models by quantizing weights and activations without requiring retraining. However, existing methods either (1) fail to account for the varying importance of hidden features to the end loss or, when incorporating end loss, (2) neglect the critical interactions between model weights. To address these limitations, we propose GuidedQuant, a novel quantization approach that integrates gradient information from the end loss into the quantization objective while preserving cross-weight dependencies within output channels. GuidedQuant consistently boosts the performance of state-of-the-art quantization methods across weight-only scalar, weight-only vector, and weight-and-activation quantization. Additionally, we introduce a novel non-uniform scalar quantization algorithm, which is guaranteed to monotonically decrease the quantization objective value, and outperforms existing methods in this category. We release the code at https://github.com/snu-mllab/GuidedQuant.

📄 PDF Abstract BibTeX arXiv:2505.07004

Code (1)

snu-mllab/guidedquant 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelQuantization

Similar Papers 제목 키워드 기반

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models

2025-01-31 · Alina Shutova, Vladimir Malinovskii, Vage Egiazarian, Denis Kuznedelev 외

Efficient real-world deployments of large language models (LLMs) rely on Key-Value (KV) caching for processing and generating long outputs, reducing the need for repetitive computation. For large contexts, Key-Value cach…

GPUQuantization

Post-Training Weighted Quantization of Neural Networks for Language Models

2021-01-01 · Se Jung Kwon, Dongsoo Lee, Yongkweon Jeon, Byeongwook Kim 외

As a practical model compression technique, parameter quantization is effective especially for language models associated with a large memory footprint. Neural network quantization is usually performed to reduce quantiza…

Model CompressionQuantization

Post-Training Quantization for Vision Transformer

2021-06-27 · NeurIPS 2021 12 · Zhenhua Liu, Yunhe Wang, Kai Han, Siwei Ma 외

Recently, transformer has achieved remarkable performance on a variety of computer vision applications. Compared with mainstream convolutional neural networks, vision transformers are often of sophisticated architectures…

DiversityQuantization

When Quantization Affects Confidence of Large Language Models?

2024-05-01 · Irina Proskurina, Luc Brun, Guillaume Metzler, Julien Velcin

Recent studies introduced effective compression techniques for Large Language Models (LLMs) via post-training quantization or low-bit weight representation. Although quantized weights offer storage efficiency and allow f…

Language ModelingLanguage ModellingQuantization

LeanQuant: Accurate Large Language Model Quantization with Loss-Error-Aware Grid

2024-07-14 · Tianyi Zhang, Anshumali Shrivastava

Large language models (LLMs) have numerous applications across various domains, but their high computational and memory demands pose significant deployment challenges. Weight quantization is an effective technique for re…

GPULanguage ModelingLanguage ModellingLarge Language Model+1