paper-with-me

홈 › Papers

NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models

2026-02-06 · Hyochan Chong, Dongkyu Kim, Changdong Kim, Minseop Choi arxiv

Weight-only quantization has become a standard approach for efficiently serving large language models (LLMs). However, existing methods fail to efficiently compress models to binary (1-bit) levels, as they either require large amounts of data and compute or incur additional storage. In this work, we propose NanoQuant, the first post-training quantization (PTQ) method to compress LLMs to both binary and sub-1-bit levels. NanoQuant formulates quantization as a low-rank binary factorization problem, and compresses full-precision weights to low-rank binary matrices and scales. Specifically, it utilizes an efficient alternating direction method of multipliers (ADMM) solver to precisely initialize latent binary matrices and scales, and then tunes the initialized parameters through a block and model reconstruction process. Consequently, NanoQuant establishes a new Pareto frontier in low-memory post-training quantization, and enables sub-1-bit compression. NanoQuant makes large-scale deployment feasible on consumer hardware. For example, it compresses Llama2-70B by 25.8$\times$ in just 13 hours on a single H100, enabling a 70B model to operate on a consumer 8 GB GPU. Code is available at https://github.com/SamsungLabs/NanoQuant.

📄 PDF Abstract BibTeX arXiv:2602.06694

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Post-Training Weighted Quantization of Neural Networks for Language Models

2021-01-01 · Se Jung Kwon, Dongsoo Lee, Yongkweon Jeon, Byeongwook Kim 외

As a practical model compression technique, parameter quantization is effective especially for language models associated with a large memory footprint. Neural network quantization is usually performed to reduce quantiza…

Model CompressionQuantization

The Uneven Impact of Post-Training Quantization in Machine Translation

2025-08-28 · Benjamin Marie, Atsushi Fujita arxiv

Quantization is essential for deploying large language models (LLMs) on resource-constrained hardware, but its implications for multilingual tasks remain underexplored. We conduct the first large-scale evaluation of post…

Machine Translation

R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization

2025-11-21 · Jiayi Chen, Jieqi Shi, Jing Huo, Chen Wu arxiv

The rapid progress of Large Language Models (LLMs) has brought substantial computational and memory demands, spurring the adoption of low-bit quantization. While 8-bit and 4-bit formats have become prevalent, extending q…

Question Answering

LCQ: Low-Rank Codebook based Quantization for Large Language Models

2024-05-31 · Wen-Pu Cai, Wu-Jun Li

Large language models~(LLMs) have recently demonstrated promising performance in many tasks. However, the high storage and computational cost of LLMs has become a challenge for deploying LLMs. Weight quantization has bee…

Model CompressionQuantization

When Quantization Affects Confidence of Large Language Models?

2024-05-01 · Irina Proskurina, Luc Brun, Guillaume Metzler, Julien Velcin

Recent studies introduced effective compression techniques for Large Language Models (LLMs) via post-training quantization or low-bit weight representation. Although quantized weights offer storage efficiency and allow f…

Language ModelingLanguage ModellingQuantization