paper-with-me

Papers

FPSA: A Full System Stack Solution for Reconfigurable ReRAM-based NN Accelerator Architecture

2019-01-28 · Yu Ji, Youyang Zhang, Xinfeng Xie, Shuangchen Li, Peiqi Wang, Xing Hu, YouHui Zhang, Yuan Xie

Neural Network (NN) accelerators with emerging ReRAM (resistive random access memory) technologies have been investigated as one of the promising solutions to address the \textit{memory wall} challenge, due to the unique capability of \textit{processing-in-memory} within ReRAM-crossbar-based processing elements (PEs). However, the high efficiency and high density advantages of ReRAM have not been fully utilized due to the huge communication demands among PEs and the overhead of peripheral circuits. In this paper, we propose a full system stack solution, composed of a reconfigurable architecture design, Field Programmable Synapse Array (FPSA) and its software system including neural synthesizer, temporal-to-spatial mapper, and placement & routing. We highly leverage the software system to make the hardware design compact and efficient. To satisfy the high-performance communication demand, we optimize it with a reconfigurable routing architecture and the placement & routing tool. To improve the computational density, we greatly simplify the PE circuit with the spiking schema and then adopt neural synthesizer to enable the high density computation-resources to support different kinds of NN operations. In addition, we provide spiking memory blocks (SMBs) and configurable logic blocks (CLBs) in hardware and leverage the temporal-to-spatial mapper to utilize them to balance the storage and computation requirements of NN. Owing to the end-to-end software system, we can efficiently deploy existing deep neural networks to FPSA. Evaluations show that, compared to one of state-of-the-art ReRAM-based NN accelerators, PRIME, the computational density of FPSA improves by 31x; for representative NNs, its inference performance can achieve up to 1000x speedup.

📄 PDF Abstract BibTeX arXiv:1901.09904

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion

2025-06-05 · Akide Liu, Zeyu Zhang, Zhexin Li, Xuehai Bai 외

Diffusion generative models have become the standard for producing high-quality, coherent video content, yet their slow inference speeds and high computational demands hinder practical deployment. Although both quantizat…

DenoisingQuantizationVideo Generation

Learning Stackable and Skippable LEGO Bricks for Efficient, Reconfigurable, and Variable-Resolution Diffusion Modeling

2023-10-10 · Huangjie Zheng, Zhendong Wang, Jianbo Yuan, Guanghan Ning 외

Diffusion models excel at generating photo-realistic images but come with significant computational costs in both training and sampling. While various techniques address these computational challenges, a less-explored is…

Image Generation

RAPID: Reconfigurable, Adaptive Platform for Iterative Design

2026-02-06 · Zi Yin, Fanhong Li, Shurui Zheng, Jia Liu arxiv

Developing robotic manipulation policies is iterative and hypothesis-driven: researchers test tactile sensing, gripper geometries, and sensor placements through real-world data collection and training. Yet even minor end…

Bifrost: End-to-End Evaluation and Optimization of Reconfigurable DNN Accelerators

2022-04-26 · Axel Stjerngren, Perry Gibson, José Cano

Reconfigurable accelerators for deep neural networks (DNNs) promise to improve performance such as inference latency. STONNE is the first cycle-accurate simulator for reconfigurable DNN inference accelerators which allow…

HiAER-Spike: Hardware-Software Co-Design for Large-Scale Reconfigurable Event-Driven Neuromorphic Computing

2025-03-20 · Gwenevere Frank, Gopabandhu Hota, Keli Wang, Abhinav Uppal 외

In this work, we present HiAER-Spike, a modular, reconfigurable, event-driven neuromorphic computing platform designed to execute large spiking neural networks with up to 160 million neurons and 40 billion synapses - rou…

Cloud Computing