FPGA-based Acceleration of Neural Network for Image Classification using Vitis AI
In recent years, Convolutional Neural Networks (CNNs) have been widely adopted in computer vision. Complex CNN architecture running on CPU or GPU has either insufficient throughput or prohibitive power consumption. Hence, there is a need to have dedicated hardware to accelerate the computation workload to solve these limitations. In this paper, we accelerate a CNN for image classification with the CIFAR-10 dataset using Vitis-AI on Xilinx Zynq UltraScale+ MPSoC ZCU104 FPGA evaluation board. The work achieves 3.33-5.82x higher throughput and 3.39-6.30x higher energy efficiency than CPU and GPU baselines. It shows the potential to extract 2D features for downstream tasks, such as depth estimation and 3D reconstruction.
Code (0)
등록된 구현이 없습니다.
Tasks
3D ReconstructionCPUDepth EstimationGPUimage-classificationImage ClassificationSimilar Papers 제목 키워드 기반
Evaluating Four FPGA-accelerated Space Use Cases based on Neural Network Algorithms for On-board Inference
Space missions increasingly deploy high-fidelity sensors that produce data volumes exceeding onboard buffering and downlink capacity. This work evaluates FPGA acceleration of neural networks (NNs) across four space use c…
Real-Time Semantic Segmentation of Aerial Images Using an Embedded U-Net: A Comparison of CPU, GPU, and FPGA Workflows
This study introduces a lightweight U-Net model optimized for real-time semantic segmentation of aerial images, targeting the efficient utilization of Commercial Off-The-Shelf (COTS) embedded computing platforms. We main…
CPUGPUReal-Time Semantic SegmentationSemantic SegmentationFINN-GL: Generalized Mixed-Precision Extensions for FPGA-Accelerated LSTMs
Recurrent neural networks (RNNs), particularly LSTMs, are effective for time-series tasks like sentiment analysis and short-term stock prediction. However, their computational complexity poses challenges for real-time de…
Sentiment AnalysisStock Predictionhls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware
We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field…
Red grape detection with accelerated artificial neural networks in the FPGA's programmable logic
Robots usually slow down for canning to detect objects while moving. Additionally, the robot's camera is configured with a low framerate to track the velocity of the detection algorithms. This would be constrained while …