paper-with-me

Papers

A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms

2023-06-27 · Cristina Silvano, Daniele Ielmini, Fabrizio Ferrandi, Leandro Fiorin, Serena Curzel, Luca Benini, Francesco Conti, Angelo Garofalo, Cristian Zambelli, Enrico Calore, Sebastiano Fabio Schifano, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Nicola Petra, Davide De Caro, Luciano Lavagno, Teodoro Urso, Valeria Cardellini, Gian Carlo Cardarilli, Robert Birke, Stefania Perri

Recent trends in deep learning (DL) have made hardware accelerators essential for various high-performance computing (HPC) applications, including image classification, computer vision, and speech recognition. This survey summarizes and classifies the most recent developments in DL accelerators, focusing on their role in meeting the performance demands of HPC applications. We explore cutting-edge approaches to DL acceleration, covering not only GPU- and TPU-based platforms but also specialized hardware such as FPGA- and ASIC-based accelerators, Neural Processing Units, open hardware RISC-V-based accelerators, and co-processors. This survey also describes accelerators leveraging emerging memory technologies and computing paradigms, including 3D-stacked Processor-In-Memory, non-volatile memories like Resistive RAM and Phase Change Memories used for in-memory computing, as well as Neuromorphic Processing Units, and Multi-Chip Module-based accelerators. Furthermore, we provide insights into emerging quantum-based accelerators and photonics. Finally, this survey categorizes the most influential architectures and technologies from recent years, offering readers a comprehensive perspective on the rapidly evolving field of deep learning acceleration.

📄 PDF Abstract BibTeX arXiv:2306.15552

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningGPUimage-classificationImage Classificationspeech-recognitionSpeech RecognitionSurvey

Similar Papers 제목 키워드 기반

A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures

2023-11-29 · Serena Curzel, Fabrizio Ferrandi, Leandro Fiorin, Daniele Ielmini 외

Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in the design of heterogeneous High-Performance Computing (HPC) and edge platforms, le…

Deep LearningHigh-Level SynthesisSurvey

Resistive Neural Hardware Accelerators

2021-09-08 · Kamilya Smagulova, Mohammed E. Fouda, Fadi Kurdahi, Khaled Salama 외

Deep Neural Networks (DNNs), as a subset of Machine Learning (ML) techniques, entail that real-world data can be learned and that decisions can be made in real-time. However, their wide adoption is hindered by a number o…

Benchmarking

Mind the Gap: Revealing Inconsistencies Across Heterogeneous AI Accelerators

2025-10-30 · Elliott Wen, Sean Ma, Ewan Tempero, Jens Dietrich 외 arxiv

While NVIDIA remains the dominant provider of AI accelerators within cloud data center, emerging vendors such as AMD, Intel, Mac, and Huawei offer cost-effective alternatives with claims of compatibility and performance.…

Precision-aware Latency and Energy Balancing on Multi-Accelerator Platforms for DNN Inference

2023-06-08 · Matteo Risso, Alessio Burrello, Giuseppe Maria Sarda, Luca Benini 외

The need to execute Deep Neural Networks (DNNs) at low latency and low power at the edge has spurred the development of new heterogeneous Systems-on-Chips (SoCs) encapsulating a diverse set of hardware accelerators. How …

Quantization

MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices

2024-10-11 · Mohamed Amine Hamdi, Francesco Daghero, Giuseppe Maria Sarda, Josse Van Delm 외

Streamlining the deployment of Deep Neural Networks (DNNs) on heterogeneous edge platforms, coupling within the same micro-controller unit (MCU) instruction processors and hardware accelerators for tensor computations, i…