paper-with-me

홈 › Papers

On the Importance of a Multi-Scale Calibration for Quantization

2026-02-07 · Seungwoo Son, Ingyu Seong, Junhan Kim, Hyemi Jang, Yongkweon Jeon arxiv

Post-training quantization (PTQ) is a cornerstone for efficiently deploying large language models (LLMs), where a small calibration set critically affects quantization performance. However, conventional practices rely on random sequences of fixed length, overlooking the variable-length nature of LLM inputs. Input length directly influences the activation distribution and, consequently, the weight importance captured by the Hessian, which in turn affects quantization outcomes. As a result, Hessian estimates derived from fixed-length calibration may fail to represent the true importance of weights across diverse input scenarios. We propose MaCa (Matryoshka Calibration), a simple yet effective method for length-aware Hessian construction. MaCa (i) incorporates multi-scale sequence length information into Hessian estimation and (ii) regularizes each sequence as an independent sample, yielding a more stable and fruitful Hessian for accurate quantization. Experiments on state-of-the-art LLMs (e.g., Qwen3, Gemma3, LLaMA3) demonstrate that MaCa consistently improves accuracy under low bit quantization, offering a lightweight enhancement compatible with existing PTQ frameworks. To the best of our knowledge, this is the first work to systematically highlight the role of multi-scale calibration in LLM quantization.

📄 PDF Abstract BibTeX arXiv:2602.07465

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rethinking Practical and Efficient Quantization Calibration for Vision-Language Models

2026-02-08 · Zhenhao Shang, Haizhao Jing, Guoting Wei, Haokui Zhang 외 arxiv

Post-training quantization (PTQ) is a primary approach for deploying large language models without fine-tuning, and the quantized performance is often strongly affected by the calibration in PTQ. By contrast, in vision-l…

Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping

2025-09-03 · Xinzhe Zheng, Zhen-Qun Yang, Zishan Liu, Haoran Xie 외 arxiv

Large Language Models (LLMs) deliver strong performance but are difficult to deploy under tight memory and compute constraints. Low-bit post-training quantization (PTQ) is a promising direction; however, it typically rel…

Intrinsic Structure as a Proxy for Saliency: SVD-Based Weight Preservation for Mixed-Precision Quantization in Large Language Models

2025-12-01 · Shashank Landge, Abhishek Patil, Tejas kamble, Bhushan Buddhivant 외 arxiv

As Large Language Models (LLMs) continue to scale in parameter count, deploying them on commodity hardware has become increasingly challenging. Post-Training Quantization (PTQ) addresses this by reducing the precision of…

AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization

2026-06-03 · Wanqi Yang, Yuexiao Ma, Alexander Conzelmann, Xiawu Zheng 외 arxiv

Mixture-of-Experts (MoE) architectures scale model capacity through sparse expert activation, but their deployment remains memory-bound because all expert weights must reside in memory. Mixed-precision quantization can s…

RSQ: Learning from Important Tokens Leads to Better Quantized LLMs

2025-03-03 · Yi-Lin Sung, Prateek Yadav, Jialu Li, Jaehong Yoon 외

Layer-wise quantization is a key technique for efficiently compressing large models without expensive retraining. Previous methods typically quantize the weights of each layer by "uniformly" optimizing the layer reconstr…

Quantization