paper-with-me

홈 › Papers

HG-Caffe: Mobile and Embedded Neural Network GPU (OpenCL) Inference Engine with FP16 Supporting

2019-01-03 · Zhuoran Ji

Breakthroughs in the fields of deep learning and mobile system-on-chips are radically changing the way we use our smartphones. However, deep neural networks inference is still a challenging task for edge AI devices due to the computational overhead on mobile CPUs and a severe drain on the batteries. In this paper, we present a deep neural network inference engine named HG-Caffe, which supports GPUs with half precision. HG-Caffe provides up to 20 times speedup with GPUs compared to the original implementations. In addition to the speedup, the peak memory usage is also reduced to about 80%. With HG-Caffe, more innovative and fascinating mobile applications will be turned into reality.

📄 PDF Abstract BibTeX arXiv:1901.00858

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Similar Papers 제목 키워드 기반

Tuning of Mixture-of-Experts Mixed-Precision Neural Networks

2022-09-29 · Fabian Tschopp

Deep learning has become a useful data analysis method, however mainstream adaption in distributed computer software and embedded devices has been low so far. Often, adding deep learning inference in mainstream applicati…

image-classificationImage ClassificationMixture-of-Experts

Comparison and Benchmarking of AI Models and Frameworks on Mobile Devices

2020-05-07 · Chunjie Luo, Xiwen He, Jianfeng Zhan, Lei Wang 외

Due to increasing amounts of data and compute resources, deep learning achieves many successes in various domains. The application of deep learning on the mobile and embedded devices is taken more and more attentions, be…

BenchmarkingDiversityvalid

FeCaffe: FPGA-enabled Caffe with OpenCL for Deep Learning Training and Inference on Intel Stratix 10

2019-11-18 · Ke He, Bo Liu, Yu Zhang, Andrew Ling 외

Deep learning and Convolutional Neural Network (CNN) have becoming increasingly more popular and important in both academic and industrial areas in recent years cause they are able to provide better accuracy and result i…

CPUDeep LearningGPU

Compressed Learning of Deep Neural Networks for OpenCL-Capable Embedded Systems

2019-05-20 · Sangkyun Lee, Jeonghyun Lee

Deep neural networks (DNNs) have been quite successful in solving many complex learning problems. However, DNNs tend to have a large number of learning parameters, leading to a large memory and computation requirement. I…

Model Compression

Highly Efficient 8-bit Low Precision Inference of Convolutional Neural Networks with IntelCaffe

2018-05-04 · Jiong Gong, Haihao Shen, Guoming Zhang, Xiaoli Liu 외

High throughput and low latency inference of deep neural networks are critical for the deployment of deep learning applications. This paper presents the efficient inference techniques of IntelCaffe, the first Intel optim…

Deep LearningModel Optimization