paper-with-me

홈 › Papers

BPDQ: Bit-Plane Decomposition Quantization on a Variable Grid for Large Language Models

2026-02-04 · Junyu Chen, Jungang Li, Jing Xiong, Wenjie Wang, Qingyao Yang, He Xiao, Zhen Li, Taiqiang Wu, Mengzhao Chen, Zhen Peng, Chaofan Tao, Long Shi, Hongxia Yang, Ngai Wong arxiv

Large language model inference is often bounded by memory footprint and bandwidth in resource-constrained deployments, making quantization fundamental to efficient serving. While post-training quantization (PTQ) maintains high fidelity at 4-bit, it deteriorates at 2-3 bits. In essence, existing methods enforce a shape-invariant quantization grid (e.g., the fixed uniform intervals of UINT2) for each group, severely restricting the feasible set for error minimization. To address this, we propose Bit-Plane Decomposition Quantization (BPDQ), which constructs a variable quantization grid via bit-planes and scalar coefficients, and iteratively refines them using second-order information while progressively compensating for quantization errors to minimize output discrepancy. In the 2-bit regime, BPDQ enables serving Qwen2.5-72B on a single RTX 3090 with 83.85\% GSM8K accuracy (vs. 90.83\% at 16-bit). Moreover, we theoretically show that the variable grid expands the feasible set, and that the quantization process consistently aligns with the optimization objective in Hessian-induced geometry. The code is available at https://github.com/KingdalfGoodman/BPDQ.

📄 PDF Abstract BibTeX arXiv:2602.04163

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reconstruction with prior support information and non-Gaussian constraints

2024-10-09 · Xiaotong Liu, Yiyu Liang

In this study, we introduce a novel model, termed the Weighted Basis Pursuit Dequantization ($\omega$-BPDQ$_p$), which incorporates prior support information by assigning weights on the $\ell_1$ norm in the $\ell_1$ mini…

Towards Lossless Implicit Neural Representation via Bit Plane Decomposition

2025-02-28 · CVPR 2025 1 · Woo Kyoung Han, Byeonghun Lee, Hyunmin Cho, Sunghoon Im 외

We quantify the upper bound on the size of the implicit neural representation (INR) model from a digital perspective. The upper bound of the model size increases exponentially as the required bit-precision increases. To …

Image CompressionQuantization

Progressive Neural Image Compression with Nested Quantization and Latent Ordering

2021-02-04 · Yadong Lu, Yinhao Zhu, Yang Yang, Amir Said 외

We present PLONQ, a progressive neural image compression scheme which pushes the boundary of variable bitrate compression by allowing quality scalable coding with a single bitstream. In contrast to existing learned varia…

Image CompressionQuantization

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor

2026-05-19 · Xiaocan Li, Shiliang Wu, Zheng Shen arxiv

MXFP4 arithmetic can dramatically accelerate reinforcement learning (RL) post-training of large language models (LLMs), yet the quantization error introduces severe accuracy degradation. Existing work treats the quantiza…

Reinforcement Learning

Quantization Aware Factorization for Deep Neural Network Compression

2023-08-08 · Daria Cherniuk, Stanislav Abukhovich, Anh-Huy Phan, Ivan Oseledets 외

Tensor decomposition of convolutional and fully-connected layers is an effective way to reduce parameters and FLOP in neural networks. Due to memory and power consumption limitations of mobile or embedded devices, the qu…

Neural Network CompressionQuantizationTensor Decomposition