paper-with-me

홈 › Papers

Efficient Deep Learning Using Non-Volatile Memory Technology

2022-06-27 · Ahmet Inci, Mehmet Meric Isgenc, Diana Marculescu

Embedded machine learning (ML) systems have now become the dominant platform for deploying ML serving tasks and are projected to become of equal importance for training ML models. With this comes the challenge of overall efficient deployment, in particular low power and high throughput implementations, under stringent memory constraints. In this context, non-volatile memory (NVM) technologies such as STT-MRAM and SOT-MRAM have significant advantages compared to conventional SRAM due to their non-volatility, higher cell density, and scalability features. While prior work has investigated several architectural implications of NVM for generic applications, in this work we present DeepNVM++, a comprehensive framework to characterize, model, and analyze NVM-based caches in GPU architectures for deep learning (DL) applications by combining technology-specific circuit-level models and the actual memory behavior of various DL workloads. DeepNVM++ relies on iso-capacity and iso-area performance and energy models for last-level caches implemented using conventional SRAM and emerging STT-MRAM and SOT-MRAM technologies. In the iso-capacity case, STT-MRAM and SOT-MRAM provide up to 3.8x and 4.7x energy-delay product (EDP) reduction and 2.4x and 2.8x area reduction compared to conventional SRAM, respectively. Under iso-area assumptions, STT-MRAM and SOT-MRAM provide up to 2.2x and 2.4x EDP reduction and accommodate 2.3x and 3.3x cache capacity when compared to SRAM, respectively. We also perform a scalability analysis and show that STT-MRAM and SOT-MRAM achieve orders of magnitude EDP reduction when compared to SRAM for large cache capacities. DeepNVM++ is demonstrated on STT-/SOT-MRAM technologies and can be used for the characterization, modeling, and analysis of any NVM technology for last-level caches in GPUs for DL applications.

📄 PDF Abstract BibTeX arXiv:2206.13601

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningGPU

Similar Papers 제목 키워드 기반

AUTOMATIC ROOM LIGHT CONTROLLER MANAGEMENT SYSTEM.

2025-06-25 · Zenodo 2025 6 · Kamal Acharya

The AT89S51 is a low-power, high- performance CMOS 8-bit microcontroller with 4K bytes of In-System Programmable Flash memory. The device is manufactured using Atmel’s high-density non-volatile memory technology and is c…

4kCPUManagement

Memory-Oriented Design-Space Exploration of Edge-AI Hardware for XR Applications

2022-06-08 · Vivek Parmar, Syed Shakib Sarwar, Ziyun Li, Hsien-Hsin S. Lee 외

Low-Power Edge-AI capabilities are essential for on-device extended reality (XR) applications to support the vision of Metaverse. In this work, we investigate two representative XR workloads: (i) Hand detection and (ii) …

CPUHand DetectionQuantization

MRAM Co-designed Processing-in-Memory CNN Accelerator for Mobile and IoT Applications

2018-11-26 · Baohua Sun, Daniel Liu, Leo Yu, Jay Li 외

We designed a device for Convolution Neural Network applications with non-volatile MRAM memory and computing-in-memory co-designed architecture. It has been successfully fabricated using 22nm technology node CMOS Si proc…

A Three-terminal Non-Volatile Ferroelectric Switch with an Insulator-Metal Transition Channel

2021-08-27 · Jaykumar Vaidya, R S Surya Kanthi, Shamiul Alam, Nazmul Amin 외

Ferroelectrics offer a promising materials platform to realize energy-efficient non-volatile memory technology with the FeFET-based implementations being one of the most area-efficient ferroelectric memory architectures.…

Low-Rank Training of Deep Neural Networks for Emerging Memory Technology

2020-09-08 · Albert Gural, Phillip Nadeau, Mehul Tikekar, Boris Murmann

The recent success of neural networks for solving difficult decision tasks has incentivized incorporating smart decision making "at the edge." However, this work has traditionally focused on neural network inference, rat…

Computational EfficiencyDecision MakingFederated Learning