paper-with-me

홈 › Papers

RSQ: Learning from Important Tokens Leads to Better Quantized LLMs

2025-03-03 · Yi-Lin Sung, Prateek Yadav, Jialu Li, Jaehong Yoon, Mohit Bansal

Layer-wise quantization is a key technique for efficiently compressing large models without expensive retraining. Previous methods typically quantize the weights of each layer by "uniformly" optimizing the layer reconstruction loss across all output tokens. However, in this paper, we demonstrate that better-quantized models can be obtained by prioritizing learning from important tokens (e.g. which have large attention scores). Building on this finding, we propose RSQ (Rotate, Scale, then Quantize), which (1) applies rotations (orthogonal transformation) to the model to mitigate outliers (those with exceptionally large magnitude), (2) scales the token feature based on its importance, and (3) quantizes the model using the GPTQ framework with the second-order statistics computed by scaled tokens. To compute token importance, we explore both heuristic and dynamic strategies. Based on a thorough analysis of all approaches, we adopt attention concentration, which uses attention scores of each token as its importance, as the best approach. We demonstrate that RSQ consistently outperforms baseline methods across multiple downstream tasks and three model families: LLaMA3, Mistral, and Qwen2.5. Additionally, models quantized with RSQ achieve superior performance on long-context tasks, further highlighting its effectiveness. Lastly, RSQ demonstrates generalizability across various setups, including different model sizes, calibration datasets, bit precisions, and quantization methods.

📄 PDF Abstract BibTeX arXiv:2503.01820

Code (1)

ylsung/rsq 공식 구현

Tasks

Quantization

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens

2024-11-26 · Xu Ouyang, Tao Ge, Thomas Hartvigsen, Zhisong Zhang 외

We reveal that low-bit quantization favors undertrained large language models (LLMs) by observing that models with larger sizes or fewer training tokens experience less quantization-induced degradation (QiD) when applyin…

Quantization

Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models

2025-01-30 · Qika Lin, Tianzhe Zhao, Kai He, Zhen Peng 외

Due to the presence of the natural gap between Knowledge Graph (KG) structures and the natural language, the effective integration of holistic structural information of KGs with Large Language Models (LLMs) has emerged a…

Instruction FollowingKnowledge GraphsLink PredictionTriple Classification

IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact

2024-03-02 · Ruikang Liu, Haoli Bai, Haokun Lin, Yuening Li 외

Large language models (LLMs) excel in natural language processing but demand intensive computation. To mitigate this, various quantization methods have been explored, yet they compromise LLM performance. This paper unvei…

Language ModelingLanguage ModellingLarge Language ModelQuantization

Fitting Is Not Enough: Smoothness in Extremely Quantized LLMs

2026-05-09 · Yuzhuang Xu, Xu Han, Yuxuan Li, Pengzhan Li 외 arxiv

Large language models (LLMs) achieve strong performance but incur high deployment costs, motivating extremely low-bit but lossy quantization. Existing quantization algorithms mainly focus on improving the numerical accur…

Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models

2025-03-20 · Keda Tao, Haoxuan You, Yang Sui, Can Qin 외

Video large language models (VideoLLMs) have demonstrated the capability to process longer video inputs and enable complex reasoning and analysis. However, due to the thousands of visual tokens from the video frames, key…

Quantization