paper-with-me

홈 › Papers

VecQ: Minimal Loss DNN Model Compression With Vectorized Weight Quantization

2020-05-18 · Cheng Gong, Yao Chen, Ye Lu, Tao Li, Cong Hao, Deming Chen

Quantization has been proven to be an effective method for reducing the computing and/or storage cost of DNNs. However, the trade-off between the quantization bitwidth and final accuracy is complex and non-convex, which makes it difficult to be optimized directly. Minimizing direct quantization loss (DQL) of the coefficient data is an effective local optimization method, but previous works often neglect the accurate control of the DQL, resulting in a higher loss of the final DNN model accuracy. In this paper, we propose a novel metric called Vector Loss. Based on this new metric, we develop a new quantization solution called VecQ, which can guarantee minimal direct quantization loss and better model accuracy. In addition, in order to speed up the proposed quantization process during model training, we accelerate the quantization process with a parameterized probability estimation method and template-based derivation calculation. We evaluate our proposed algorithm on MNIST, CIFAR, ImageNet, IMDB movie review and THUCNews text data sets with numerical DNN models. The results demonstrate that our proposed quantization solution is more accurate and effective than the state-of-the-art approaches yet with more flexible bitwidth support. Moreover, the evaluation of our quantized models on Saliency Object Detection (SOD) tasks maintains comparable feature extraction quality with up to 16$\times$ weight size reduction.

📄 PDF Abstract BibTeX arXiv:2005.08501

Code (1)

GongCheng1919/VecQ 공식 구현 tf

Tasks

Model Compressionobject-detectionObject DetectionQuantization

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models

2026-01-08 · Maanas Taneja, Purab Shingvi arxiv

The key-value (KV) cache in large language models presents a significant memory bottleneck during inference, growing linearly with sequence length and often exceeding the memory footprint of model weights themselves. We …

ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs

2026-03-28 · Jinwu Yang, Jiaan Wu, Zedong Liu, Xinyang Ma 외 arxiv

The rapid scaling of Large Language Models presents significant challenges for their deployment and inference, particularly on resource-constrained specialized AI hardware accelerators such as Huawei's Ascend NPUs, where…

Model Compression

Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method

2025-08-20 · Suleyman Olcay Polat, Poli A. Nemkova, Mark V. Albert arxiv

Model distillation enables the transfer of knowledge from large-scale models to compact student models, facilitating deployment in resource-constrained environments. However, conventional distillation approaches often su…

Dimensionality ReductionKnowledge DistillationModel CompressionData Augmentation

Decoding billions of integers per second through vectorization

2012-09-10 · Daniel Lemire, Leonid Boytsov

In many important applications -- such as search engines and relational database systems -- data is stored in the form of arrays of integers. Encoding and, most importantly, decoding of these arrays consumes considerable…

CPU

Model compression as constrained optimization, with application to neural nets. Part II: quantization

2017-07-13 · Miguel Á. Carreira-Perpiñán, Yerlan Idelbayev

We consider the problem of deep neural net compression by quantization: given a large, reference net, we want to quantize its real-valued weights using a codebook with $K$ entries so that the training loss of the quantiz…

BinarizationModel CompressionQuantization