paper-with-me

홈 › Papers

Fixed-point Quantization of Convolutional Neural Networks for Quantized Inference on Embedded Platforms

2021-02-03 · Rishabh Goyal, Joaquin Vanschoren, Victor van Acht, Stephan Nijssen

Convolutional Neural Networks (CNNs) have proven to be a powerful state-of-the-art method for image classification tasks. One drawback however is the high computational complexity and high memory consumption of CNNs which makes them unfeasible for execution on embedded platforms which are constrained on physical resources needed to support CNNs. Quantization has often been used to efficiently optimize CNNs for memory and computational complexity at the cost of a loss of prediction accuracy. We therefore propose a method to optimally quantize the weights, biases and activations of each layer of a pre-trained CNN while controlling the loss in inference accuracy to enable quantized inference. We quantize the 32-bit floating-point precision parameters to low bitwidth fixed-point representations thereby finding optimal bitwidths and fractional offsets for parameters of each layer of a given CNN. We quantize parameters of a CNN post-training without re-training it. Our method is designed to quantize parameters of a CNN taking into account how other parameters are quantized because ignoring quantization errors due to other quantized parameters leads to a low precision CNN with accuracy losses of up to 50% which is far beyond what is acceptable. Our final method therefore gives a low precision CNN with accuracy losses of less than 1%. As compared to a method used by commercial tools that quantize all parameters to 8-bits, our approach provides quantized CNN with averages of 53% lower memory consumption and 77.5% lower cost of executing multiplications for the two CNNs trained on the four datasets that we tested our work on. We find that layer-wise quantization of parameters significantly helps in this process.

📄 PDF Abstract BibTeX arXiv:2102.02147

Code (1)

rishgoyal/FXPQuantCNN 공식 구현 tf

Tasks

image-classificationImage ClassificationQuantization

Similar Papers 제목 키워드 기반

Quantizing deep convolutional networks for efficient inference: A whitepaper

2018-06-21 · Raghuraman Krishnamoorthi

We present an overview of techniques for quantizing convolutional neural networks for inference with integer weights and activations. Per-channel quantization of weights and per-layer quantization of activations to 8-bit…

Quantization

Quantized Memory-Augmented Neural Networks

2017-11-10 · Seongsik Park, Seijoon Kim, Seil Lee, Ho Bae 외

Memory-augmented neural networks (MANNs) refer to a class of neural network models equipped with external memory (such as neural Turing machines and memory networks). These neural networks outperform conventional recurre…

Quantization

Adaptive Precision Training (AdaPT): A dynamic fixed point quantized training approach for DNNs

2021-07-28 · Lorenz Kummer, Kevin Sidak, Tabea Reichmann, Wilfried Gansterer

Quantization is a technique for reducing deep neural networks (DNNs) training and inference times, which is crucial for training in resource constrained environments or applications where inference is time critical. Stat…

Quantization

AQD: Towards Accurate Fully-Quantized Object Detection

2020-07-14 · CVPR 2021 1 · Peng Chen, Jing Liu, Bohan Zhuang, Mingkui Tan 외

Network quantization allows inference to be conducted using low-precision arithmetic for improved inference efficiency of deep neural networks on edge devices. However, designing aggressively low-bit (e.g., 2-bit) quanti…

Image ClassificationObjectobject-detectionObject Detection+1

Fixed-point optimization of deep neural networks with adaptive step size retraining

2017-02-27 · Sungho Shin, Yoonho Boo, Wonyong Sung

Fixed-point optimization of deep neural networks plays an important role in hardware based design and low-power implementations. Many deep neural networks show fairly good performance even with 2- or 3-bit precision when…

Quantization