paper-with-me

홈 › Papers

HIGHLY EFFICIENT 8-BIT LOW PRECISION INFERENCE OF CONVOLUTIONAL NEURAL NETWORKS

2019-05-01 · ICLR 2019 5 · Haihao Shen, Jiong Gong, Xiaoli Liu, Guoming Zhang, Ge Jin, and Eric Lin

High throughput and low latency inference of deep neural networks are critical for the deployment of deep learning applications. This paper presents a general technique toward 8-bit low precision inference of convolutional neural networks, including 1) channel-wise scale factors of weights, especially for depthwise convolution, 2) Winograd convolution, and 3) topology-wise 8-bit support. We experiment the techniques on top of a widely-used deep learning framework. The 8-bit optimized model is automatically generated with a calibration process from FP32 model without the need of fine-tuning or retraining. We perform a systematical and comprehensive study on 18 widely-used convolutional neural networks and demonstrate the effectiveness of 8-bit low precision inference across a wide range of applications and use cases, including image classification, object detection, image segmentation, and super resolution. We show that the inference throughput and latency are improved by 1.6X and 1.5X respectively with minimal within 0.6%1to no loss in accuracy from FP32 baseline. We believe the methodology can provide the guidance and reference design of 8-bit low precision inference for other frameworks. All the code and models will be publicly available soon.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage ClassificationImage Segmentationobject-detectionObject DetectionSemantic SegmentationSuper-Resolution

Similar Papers 제목 키워드 기반

Highly Efficient 8-bit Low Precision Inference of Convolutional Neural Networks with IntelCaffe

2018-05-04 · Jiong Gong, Haihao Shen, Guoming Zhang, Xiaoli Liu 외

High throughput and low latency inference of deep neural networks are critical for the deployment of deep learning applications. This paper presents the efficient inference techniques of IntelCaffe, the first Intel optim…

Deep LearningModel Optimization

PositNN: Training Deep Neural Networks with Mixed Low-Precision Posit

2021-04-30 · Gonçalo Raposo, Pedro Tomás, Nuno Roma

Low-precision formats have proven to be an efficient way to reduce not only the memory footprint but also the hardware resources and power consumption of deep learning computations. Under this premise, the posit numerica…

HOBFLOPS CNNs: Hardware Optimized Bitslice-Parallel Floating-Point Operations for Convolutional Neural Networks

2020-07-11 · James Garland, David Gregg

Convolutional neural networks (CNNs) are typically trained using 16- or 32-bit floating-point (FP) and researchers show that low-precision floating-point (FP) can be highly effective for inference. Low-precision FP can b…

SeerNet: Predicting Convolutional Neural Network Feature-Map Sparsity Through Low-Bit Quantization

2019-06-01 · CVPR 2019 6 · Shijie Cao, Lingxiao Ma, Wencong Xiao, Chen Zhang 외

In this paper we present a novel and general method to accelerate convolutional neural network (CNN) inference by taking advantage of feature map sparsity. We experimentally demonstrate that a highly quantized version of…

Quantization

FQ-Conv: Fully Quantized Convolution for Efficient and Accurate Inference

2019-12-19 · Bram-Ernst Verhoef, Nathan Laubeuf, Stefan Cosemans, Peter Debacker 외

Deep neural networks (DNNs) can be made hardware-efficient by reducing the numerical precision of the weights and activations of the network and by improving the network's resilience to noise. However, this gain in effic…

Quantization