paper-with-me

홈 › Papers

Reconfigurable co-processor architecture with limited numerical precision to accelerate deep convolutional neural networks

2021-08-21 · Sasindu Wijeratne, Sandaruwan Jayaweera, Mahesh Dananjaya, Ajith Pasqual

Convolutional Neural Networks (CNNs) are widely used in deep learning applications, e.g. visual systems, robotics etc. However, existing software solutions are not efficient. Therefore, many hardware accelerators have been proposed optimizing performance, power and resource utilization of the implementation. Amongst existing solutions, Field Programmable Gate Array (FPGA) based architecture provides better cost-energy-performance trade-offs as well as scalability and minimizing development time. In this paper, we present a model-independent reconfigurable co-processing architecture to accelerate CNNs. Our architecture consists of parallel Multiply and Accumulate (MAC) units with caching techniques and interconnection networks to exploit maximum data parallelism. In contrast to existing solutions, we introduce limited precision 32 bit Q-format fixed point quantization for arithmetic representations and operations. As a result, our architecture achieved significant reduction in resource utilization with competitive accuracy. Furthermore, we developed an assembly-type microinstructions to access the co-processing fabric to manage layer-wise parallelism, thereby making re-use of limited resources. Finally, we have tested our architecture up to 9x9 kernel size on Xilinx Virtex 7 FPGA, achieving a throughput of up to 226.2 GOp/S for 3x3 kernel size.

📄 PDF Abstract BibTeX arXiv:2109.03040

Code (0)

등록된 구현이 없습니다.

Tasks

Quantization

Methods 이 논문이 사용한 방법론

VirTex VirText, or Visual representations from Textual annotations is a pretraining approach using semantically dense captions to learn visual representations. First a ConvNet…

Similar Papers 제목 키워드 기반

Reconfigurable Radar Signal Processing Accelerator for Integrated Sensing and Communication System

2023-03-03 · Aakanksha Tewari, Shragvi Sidharth Jha, Akanksha Sneh, Sumit J Darak 외

IEEE 802.11ad-based integrated sensing and communications (ISAC) have been identified as a potential solution for enabling next-generation intelligent transportation systems in the millimeter wave (mmW) spectrum. The rad…

Integrated sensing and communicationISACJoint Radar-Communication

Implementation of high-efficiency, lightweight residual spiking neural network processor based on field-programmable gate arrays

2025-12-09 · Hou Yue, Xiang Shuiying, Zou Tao, Huang Zhiquan 외 arxiv

With the development of hardware-optimized deployment of spiking neural networks (SNNs), SNN processors based on field-programmable gate arrays (FPGAs) have become a research hotspot due to their efficiency and flexibili…

Self-learning photonic signal processor with an optical neural network chip

2019-02-18

Photonic signal processing is essential in the optical communication and optical computing. Numerous photonic signal processors have been proposed, but most of them exhibit limited reconfigurability and automaticity. A f…

Self-Learning

ReDON: Recurrent Diffractive Optical Neural Processor with Reconfigurable Self-Modulated Nonlinearity

2026-02-27 · Ziang Yin, Qi Jing, Raktim Sarma, Rena Huang 외 arxiv

Diffractive optical neural networks (DONNs) have demonstrated unparalleled energy efficiency and parallelism by processing information directly in the optical domain. However, their computational expressivity is constrai…

A Reconfigurable Multiplier Architecture for Error-Resilient Applications in RISC-V Core

2026-05-09 · Pragun Jaswal, L. Hemanth Krishna, B. Srinivasu arxiv

Neural Networks (NNs) have been widely adopted due to their outstanding efficacy and adaptability across computer vision and deep learning applications. The optimization of NNs is necessary to enable their deployment on …