paper-with-me

홈 › Papers

ReCU: Reviving the Dead Weights in Binary Neural Networks

2021-03-23 · ICCV 2021 10 · Zihan Xu, Mingbao Lin, Jianzhuang Liu, Jie Chen, Ling Shao, Yue Gao, Yonghong Tian, Rongrong Ji

Binary neural networks (BNNs) have received increasing attention due to their superior reductions of computation and memory. Most existing works focus on either lessening the quantization error by minimizing the gap between the full-precision weights and their binarization or designing a gradient approximation to mitigate the gradient mismatch, while leaving the "dead weights" untouched. This leads to slow convergence when training BNNs. In this paper, for the first time, we explore the influence of "dead weights" which refer to a group of weights that are barely updated during the training of BNNs, and then introduce rectified clamp unit (ReCU) to revive the "dead weights" for updating. We prove that reviving the "dead weights" by ReCU can result in a smaller quantization error. Besides, we also take into account the information entropy of the weights, and then mathematically analyze why the weight standardization can benefit BNNs. We demonstrate the inherent contradiction between minimizing the quantization error and maximizing the information entropy, and then propose an adaptive exponential scheduler to identify the range of the "dead weights". By considering the "dead weights", our method offers not only faster BNN training, but also state-of-the-art performance on CIFAR-10 and ImageNet, compared with recent methods. Code can be available at https://github.com/z-hXu/ReCU.

📄 PDF Abstract BibTeX arXiv:2103.12369

Code (3)

z-hXu/ReCU 공식 구현 pytorch
stevetsui/rbonn pytorch
stevetsui/rebnn pytorch

Tasks

BinarizationQuantization

Methods 이 논문이 사용한 방법론

Weight Standardization Weight Standardization is a normalization technique that smooths the loss landscape by standardizing the weights in convolutional layers. Different from the previous…

Similar Papers 제목 키워드 기반

HadamRNN: Binary and Sparse Ternary Orthogonal RNNs

2025-01-28 · Armand Foucault, Franck Mamalet, François Malgouyres

Binary and sparse ternary weights in neural networks enable faster computations and lighter representations, facilitating their use on edge devices with limited computational power. Meanwhile, vanilla RNNs are highly sen…

Binarization

Learning Recurrent Binary/Ternary Weights

2018-09-28 · ICLR 2019 5 · Arash Ardakani, Zhengyun Ji, Sean C. Smithson, Brett H. Meyer 외

Recurrent neural networks (RNNs) have shown excellent performance in processing sequence data. However, they are both complex and memory intensive due to their recursive nature. These limitations make RNNs difficult to e…

Language ModelingLanguage Modelling

Recursive Binary Neural Network Learning Model with 2.28b/Weight Storage Requirement

2017-09-15 · Tianchan Guan, Xiaoyang Zeng, Mingoo Seok

This paper presents a storage-efficient learning model titled Recursive Binary Neural Networks for sensing devices having a limited amount of on-chip data storage such as < 100's kilo-Bytes. The main idea of the proposed…

ClassificationGeneral Classification

Recursive Binary Neural Network Learning Model with 2-bit/weight Storage Requirement

2018-01-01 · ICLR 2018 1 · Tianchan Guan, Xiaoyang Zeng, Mingoo Seok

This paper presents a storage-efficient learning model titled Recursive Binary Neural Networks for embedded and mobile devices having a limited amount of on-chip data storage such as hundreds of kilo-Bytes. The main idea…

Action DetectionActivity DetectionGeneral Classification

Bayesian Sparsification of Recurrent Neural Networks

2017-07-31 · Ekaterina Lobacheva, Nadezhda Chirkova, Dmitry Vetrov

Recurrent neural networks show state-of-the-art results in many text analysis tasks but often require a lot of memory to store their weights. Recently proposed Sparse Variational Dropout eliminates the majority of the we…

Language ModelingLanguage ModellingSentiment Analysis