paper-with-me

홈 › Papers

A&B BNN: Add&Bit-Operation-Only Hardware-Friendly Binary Neural Network

2024-03-06 · CVPR 2024 1 · Ruichen Ma, Guanchao Qiao, Yian Liu, Liwei Meng, Ning Ning, Yang Liu, Shaogang Hu

Binary neural networks utilize 1-bit quantized weights and activations to reduce both the model's storage demands and computational burden. However, advanced binary architectures still incorporate millions of inefficient and nonhardware-friendly full-precision multiplication operations. A&B BNN is proposed to directly remove part of the multiplication operations in a traditional BNN and replace the rest with an equal number of bit operations, introducing the mask layer and the quantized RPReLU structure based on the normalizer-free network architecture. The mask layer can be removed during inference by leveraging the intrinsic characteristics of BNN with straightforward mathematical transformations to avoid the associated multiplication operations. The quantized RPReLU structure enables more efficient bit operations by constraining its slope to be integer powers of 2. Experimental results achieved 92.30%, 69.35%, and 66.89% on the CIFAR-10, CIFAR-100, and ImageNet datasets, respectively, which are competitive with the state-of-the-art. Ablation studies have verified the efficacy of the quantized RPReLU structure, leading to a 1.14% enhancement on the ImageNet compared to using a fixed slope RLeakyReLU. The proposed add&bit-operation-only BNN offers an innovative approach for hardware-friendly network architecture.

📄 PDF Abstract BibTeX arXiv:2403.03739

Code (1)

ruichen0424/ab-bnn 공식 구현 pytorch

Tasks

Image Classification

Similar Papers 제목 키워드 기반

AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs

2025-10-12 · Gunho Park, Jeongin Bae, Beomseok Kwon, Byeongwook Kim 외 arxiv

The deployment of large language models (LLMs) is increasingly constrained by memory and latency bottlenecks, motivating the need for quantization techniques that flexibly balance accuracy and efficiency. Recent work has…

COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference

2025-04-22 · Ye Qiao, Zhiheng Chen, Yian Wang, Yifan Zhang 외

Transformer-based models have demonstrated superior performance in various fields, including natural language processing and computer vision. However, their enormous model size and high demands in computation, memory, an…

Edge-computing

BitSplit-Net: Multi-bit Deep Neural Network with Bitwise Activation Function

2019-03-23 · Hyungjun Kim, Yulhwa Kim, Sungju Ryu, Jae-Joon Kim

Significant computational cost and memory requirements for deep neural networks (DNNs) make it difficult to utilize DNNs in resource-constrained environments. Binary neural network (BNN), which uses binary weights and bi…

GPU

Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at Scale

2026-08-28 · Andrea Ceni, Gianluca Milano, Carlo Ricciardi, Claudio Gallicchio arxiv

Reservoir Computing (RC) designs Recurrent Neural Networks around a fixed, i.e., untrained, recurrent layer, and is a natural candidate for neuromorphic hardware. Memristive-friendly reservoirs derive the neuron dynamics…

Dedicated Inference Engine and Binary-Weight Neural Networks for Lightweight Instance Segmentation

2025-01-03 · Tse-Wei Chen, Wei Tao, Dongyue Zhao, Kazuhiro Mima 외

Reducing computational costs is an important issue for development of embedded systems. Binary-weight Neural Networks (BNNs), in which weights are binarized and activations are quantized, are employed to reduce computati…

Instance SegmentationSemantic Segmentation