HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
We introduce HBLLM, a wavelet-enhanced high-fidelity $1$-bit post-training quantization method for Large Language Models (LLMs). By leveraging Haar wavelet transforms to enhance expressive capacity through frequency decomposition, HBLLM significantly improves quantization fidelity while maintaining minimal overhead. This approach features two innovative structure-aware grouping strategies: (1) frequency-aware multi-parameter intra-row grouping and (2) $\ell_2$-norm-based saliency-driven column selection. For non-salient weights, a shared mean is employed across quantization groups within each frequency band to optimize storage efficiency. Experiments conducted on the OPT and LLaMA models demonstrate that HBLLM achieves state-of-the-art performance in $1$-bit quantization, attaining a perplexity of $6.71$ on LLaMA$2$-$13$B with an average weight storage of only $1.08$ bits. Code available at: https://github.com/Yeyke/HBLLM.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Wavelet-denoising on hardware devices with Perfect Reconstruction, low latency and adaptive thresholding
This paper introduces a wavelet denoising architecture with adaptive thresholding for real-time 1D-systems and without the use of external memories for storing input data or wavelet coefficients. The Discrete Wavelet Tra…
DenoisingQuantizationGEWDiff: Geometric Enhanced Wavelet-based Diffusion Model for Hyperspectral Image Super-resolution
Improving the quality of hyperspectral images (HSIs), such as through super-resolution, is a crucial research area. However, generative modeling for HSIs presents several challenges. Due to their high spectral dimensiona…
Image Super-ResolutionWavelet Feature Maps Compression for Image-to-Image CNNs
Convolutional Neural Networks (CNNs) are known for requiring extensive computational resources, and quantization is among the best and most common methods for compressing them. While aggressive quantization (i.e., less t…
Depth EstimationNeural Network CompressionQuantizationSemantic SegmentationHaar Wavelet Feature Compression for Quantized Graph Convolutional Networks
Graph Convolutional Networks (GCNs) are widely used in a variety of applications, and can be seen as an unstructured version of standard Convolutional Neural Networks (CNNs). As in CNNs, the computational cost of GCNs fo…
Feature CompressionNode ClassificationPoint Cloud ClassificationQuantization+1Noise Homogenization via Multi-Channel Wavelet Filtering for High-Fidelity Sample Generation in GANs
In the generator of typical Generative Adversarial Networks (GANs), a noise is inputted to generate fake samples via a series of convolutional operations. However, current noise generation models merely relies on the inf…