paper-with-me

홈 › Papers

CuFP: An HLS Library for Customized Floating-Point Operators

2024-07-18 · Electronics 2024 7 · Hajizadeh, Fahimeh; Ould-Bachir, Tarek; David, JP

High-Level Synthesis (HLS) tools have revolutionized FPGA application development by providing a more efficient and streamlined approach, significantly impacting digital design methodologies. Despite the capability of FPGAs to customize numerical representations in data paths, most HLS projects have focused on fixed-point precision, while floating-point representations remain limited to vendor-provided single, double, and half-precision formats. This paper proposes a customized floating-point library compatible with HLS to address these limitations. This library allows programmers to define the number of exponent and mantissa bits at compile time, providing greater flexibility and enabling the use of mixed precision. Moreover, this library includes optimized implementations of common components such as vector summation (VSUM), dot-product (DP), and matrix-vector multiplication (MVM). Results demonstrate that the proposed library reduces latency and resource utilization compared to vendor IP blocks, particularly in VSUM, DP, and MVM operations. For example, the mvm operation involving a 32 × 32 matrix, using vendor IP requires 22 clock cycles, whereas CuFP completes the same task in just 7 clock cycles, using approximately 60% fewer DSPs, 10% fewer LUTs, and 60% fewer FFs.

📄 PDF Abstract BibTeX

Code (1)

FahimeHajizadeh/Custom-Float-HLS

Tasks

High-Level Synthesis

Methods 이 논문이 사용한 방법론

Library 설명 없음

Similar Papers 제목 키워드 기반

Enhancing CuFP Library with Self-Alignment Technique

2025-03-24 · Computers 2025 3 · Hajizadeh, Fahimeh; Ould-Bachir, Tarek; David, JP

High-Level Synthesis (HLS) tools have transformed FPGA development by streamlining digital design and enhancing efficiency. Meanwhile, advancements in semiconductor technology now support the integration of hundreds of f…

High-Level Synthesis

MCTensor: A High-Precision Deep Learning Library with Multi-Component Floating-Point

2022-07-18 · Tao Yu, Wentao Guo, Jianan Canal Li, Tiancheng Yuan 외

In this paper, we introduce MCTensor, a library based on PyTorch for providing general-purpose and high-precision arithmetic for DL training. MCTensor is used in the same way as PyTorch Tensor: we implement multiple basi…

Deploying Customized Data Representation and Approximate Computing in Machine Learning Applications

2018-06-03 · Mahdi Nazemi, Massoud Pedram

Major advancements in building general-purpose and customized hardware have been one of the key enablers of versatility and pervasiveness of machine learning models such as deep neural networks. To sustain this ubiquitou…

BIG-bench Machine Learning

Development of Quantized DNN Library for Exact Hardware Emulation

2021-06-15 · Masato Kiyama, Motoki Amagasaki, Masahiro Iida

Quantization is used to speed up execution time and save power when runnning Deep neural networks (DNNs) on edge devices like AI chips. To investigate the effect of quantization, we need performing inference after quanti…

Quantization

Training DNNs with Hybrid Block Floating Point

2018-04-04 · NeurIPS 2018 12 · Mario Drumond, Tao Lin, Martin Jaggi, Babak Falsafi

The wide adoption of DNNs has given birth to unrelenting computing requirements, forcing datacenter operators to adopt domain-specific accelerators to train them. These accelerators typically employ densely packed full p…