paper-with-me

Papers

A Streaming Accelerator for Deep Convolutional Neural Networks with Image and Feature Decomposition for Resource-limited System Applications

2017-09-15 · Yuan Du, Li Du, Yilei Li, Junjie Su, Mau-Chung Frank Chang

Deep convolutional neural networks (CNN) are widely used in modern artificial intelligence (AI) and smart vision systems but also limited by computation latency, throughput, and energy efficiency on a resource-limited scenario, such as mobile devices, internet of things (IoT), unmanned aerial vehicles (UAV), and so on. A hardware streaming architecture is proposed to accelerate convolution and pooling computations for state-of-the-art deep CNNs. It is optimized for energy efficiency by maximizing local data reuse to reduce off-chip DRAM data access. In addition, image and feature decomposition techniques are introduced to optimize memory access pattern for an arbitrary size of image and number of features within limited on-chip SRAM capacity. A prototype accelerator was implemented in TSMC 65 nm CMOS technology with 2.3 mm x 0.8 mm core area, which achieves 144 GOPS peak throughput and 0.8 TOPS/W peak energy efficiency.

📄 PDF Abstract BibTeX arXiv:1709.05116

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Reconfigurable Streaming Deep Convolutional Neural Network Accelerator for Internet of Things

2017-07-08 · Li Du, Yuan Du, Yilei Li, Mau-Chung Frank Chang

Convolutional neural network (CNN) offers significant accuracy in image detection. To implement image detection using CNN in the internet of things (IoT) devices, a streaming hardware accelerator is proposed. The propose…

A Column Streaming-Based Convolution Engine and Mapping Algorithm for CNN-based Edge AI accelerators

2021-09-15 · Weison Lin, Tughrul Arslan

Edge AI accelerators have been emerging as a solution for near customers' applications in areas such as unmanned aerial vehicles (UAVs), image recognition sensors, wearable devices, robotics, and remote sensing satellite…

Data Streaming and Traffic Gathering in Mesh-based NoC for Deep Neural Network Acceleration

2021-08-01 · Binayak Tiwari, Mei Yang, Xiaohang Wang, Yingtao Jiang

The increasing popularity of deep neural network (DNN) applications demands high computing power and efficient hardware accelerator architecture. DNN accelerators use a large number of processing elements (PEs) and on-ch…

Detection of prostate cancer in whole-slide images through end-to-end training with image-level labels

2020-06-05 · Hans Pinckaers, Wouter Bulten, Jeroen van der Laak, Geert Litjens

Prostate cancer is the most prevalent cancer among men in Western countries, with 1.1 million new diagnoses every year. The gold standard for the diagnosis of prostate cancer is a pathologists' evaluation of prostate tis…

GPUMultiple Instance Learningwhole slide images

Mixed-TD: Efficient Neural Network Accelerator with Layer-Specific Tensor Decomposition

2023-06-08 · Zhewen Yu, Christos-Savvas Bouganis

Neural Network designs are quite diverse, from VGG-style to ResNet-style, and from Convolutional Neural Networks to Transformers. Towards the design of efficient accelerators, many works have adopted a dataflow-based, in…

Efficient Neural NetworkQuantizationTensor Decomposition