paper-with-me

홈 › Papers

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization

2026-05-06 · Zhikai Li, Zhen Dong, Xuewen Liu, Jing Zhang, Qingyi Gu arxiv

Large Language Models (LLMs) have demonstrated remarkable capabilities. However, their massive parameter scale leads to significant resource consumption and latency during inference. Post-training weight-only quantization offers a promising solution by reducing model size and accelerating token generation through alleviating the memory-bound issue. Nevertheless, the presence of inherent systematic outliers in weights continues to be a major obstacle. While existing methods, such as scaling and rotation, attempt to address this issue, the performance remains unsatisfactory. In this paper, we propose Outlier Self-Absorption Quantization (OSAQ), which performs additive weight suppression guided by the second-order low-rank property for low-bit weight-only quantization of LLMs. Specifically, we observe that the Hessian exhibits low-rank consistency across different inputs, with certain directions consistently showing vanishing curvature. Leveraging this property, we identify a stable null space of the Hessian and then construct an additive weight transformation by linearly combining the vectors within this null space, thereby suppressing weight outliers without affecting the task loss. This additive transformation can be absorbed into the weights offline, requiring no inter-layer transformations and introducing no inference overhead. Moreover, the construction is efficiently achieved by a closed-form solution, without resource-intensive training or iterative procedures. Extensive experiments demonstrate that OSAQ effectively suppresses outliers and enhances low-bit quantization performance. For instance, in 2-bit quantization, OSAQ, when integrated with GPTQ, achieves over 40% lower perplexity compared to vanilla GPTQ.

📄 PDF Abstract BibTeX arXiv:2605.04738

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models

2025-06-16 · Junho Yoon, Geom Lee, Donghyeon Jeon, Inho Kang 외

Quantization has been widely studied as an effective technique for reducing the memory requirement of large language models (LLMs), potentially improving the latency time as well. Utilizing the characteristic of rotation…

Quantization

TopGQ: Fast GNN Post-Training Quantization Leveraging Topology Information

2026-08-31 · Dain Kwon, Kanghyun Choi, Hyeyoon Lee, Sunjong Park 외 arxiv

Existing GNN quantization methods suffer from considerable quantization overhead, which severely limits their practical usage in real-world scenarios. To this end, we present TopGQ, an accurate post-training GNN quantiza…

Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization

2024-04-04 · Aniruddha Nrusimha, Mayank Mishra, Naigang Wang, Dan Alistarh 외

We consider the problem of accurate quantization for language models, where both the weights and activations are uniformly quantized to 4 bits per parameter, the lowest bitwidth format natively supported by GPU hardware.…

GPULanguage ModelingLanguage ModellingQuantization

Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling

2023-04-18 · Xiuying Wei, Yunchen Zhang, Yuhang Li, Xiangguo Zhang 외

Post-training quantization~(PTQ) of transformer language models faces significant challenges due to the existence of detrimental outliers in activations. We observe that these outliers are concentrated in specific channe…

Quantization

CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts

2026-04-12 · Xiangyang Yin, Xingyu Liu, Tianhua Xia, Bo Bao 외 arxiv

Outliers have emerged as a fundamental bottleneck in preserving accuracy for low-precision large models, particularly within Mixture-of-Experts (MoE) architectures that are increasingly central to large-scale language mo…