paper-with-me

홈 › Papers

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

2025-11-09 · Amit LeVi, Raz Lapid, Rom Himelstein, Chaim Baskin, Ravid Shwartz Ziv, Avi Mendelson arxiv

Many LLM applications require only narrow capabilities, yet standard post-training quantization (PTQ) methods allocate precision without considering the target task. This can waste bits on layers that are less relevant to the task signal while over-compressing layers that are critical for downstream behavior. We propose Task-Aware Quantization (TAQ), a training-free, weight-only mixed-precision PTQ framework that uses a small set of unlabeled task calibration prompts to allocate higher precision to task-relevant transformer layers under a fixed bit budget. TAQ estimates layer importance from hidden representations and output sensitivity, and we instantiate it with three scoring rules: TAQ-IS, based on activation information and stability; TAQ-KL, based on output-distribution sensitivity under a quantization-noise proxy; and TAQ-O, a label-informed oracle diagnostic for analyzing layer sensitivity. Across several benchmarks, TAQ outperforms task-agnostic baselines such in most settings, with especially strong gains in the accuracy--memory ratio. We further validate that these gains translate to real deployment behavior through hardware throughput and latency measurements, and analyze calibration robustness and residual-stream error propagation. Overall, TAQ turns mixed-precision PTQ from a model-centric compression step into a task-conditioned precision-allocation problem. A reference implementation is available at \href{https://anonymous.4open.science/r/TAQ-9217/README.md}{\includegraphics[height=1em]{imgs/github-mark.png}}.

📄 PDF Abstract BibTeX arXiv:2511.06516

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

2024-03-30 · Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 외

We introduce QuaRot, a new Quantization scheme based on Rotations, which is able to quantize LLMs end-to-end, including all weights, activations, and KV cache in 4 bits. QuaRot rotates LLMs in a way that removes outliers…

Quantization

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models

2026-06-03 · Rayyan Abdalla, Amir Hussein, Min Wu, Dinesh Manocha arxiv

Post-training quantization (PTQ) is critical for the efficient deployment of large language models (LLMs). Recent ultra-low-bit PTQ methods rely on rigid weight-saliency assumptions or position heuristics, introducing su…

Compressed code: the hidden effects of quantization and distillation on programming tokens

2026-01-05 · Viacheslav Siniaev, Iaroslav Chelombitko, Aleksey Komissarov arxiv

Large Language Models (LLMs) have demonstrated exceptional code generation capabilities, yet their token-level mechanisms remain underexplored, particularly in compressed models. Through systematic analysis of programmin…

Code Generation

Interpreting the Effects of Quantization on LLMs

2025-08-22 · Manpreet Singh, Hassan Sajjad arxiv

Quantization offers a practical solution to deploy LLMs in resource-constraint environments. However, its impact on internal representations remains understudied, raising questions about the reliability of quantized mode…

Model Compression

Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs

2024-05-31 · Davide Paglieri, Saurabh Dash, Tim Rocktäschel, Jack Parker-Holder

Post-Training Quantization (PTQ) enhances the efficiency of Large Language Models (LLMs) by enabling faster operation and compatibility with more accessible hardware through reduced memory usage, at the cost of small per…

Quantization