paper-with-me

홈 › Papers

SURGE: Surrogate Gradient Adaptation in Binary Neural Networks

2026-05-09 · Haoyu Huang, Boyu Liu, Linlin Yang, Yanjing Li, Yuguang Yang, Xuhui Liu, Canyu Chen, Zhongqian Fu, Baochang Zhang arxiv

The training of Binary Neural Networks (BNNs) is fundamentally based on gradient approximation for non-differentiable binarization operations (e.g., sign function). However, prevailing methods including the Straight-Through Estimator (STE) and its improved variants, rely on hand-crafted designs that suffer from gradient mismatch problem and information loss induced by fixed-range gradient clipping. To address this, we propose SURrogate GradiEnt Adaptation (SURGE), a novel learnable gradient compensation framework with theoretical grounding. SURGE mitigates gradient mismatch through auxiliary backpropagation. Specifically, we design a Dual-Path Gradient Compensator (DPGC) that constructs a parallel full-precision auxiliary branch for each binarized layer, decoupling gradient flow via output decomposition during backpropagation. DPGC enables bias-reduced gradient estimation by leveraging the full-precision branch to estimate components beyond STE's first-order approximation. To further enhance training stability, we introduce an Adaptive Gradient Scaler (AGS) based on an optimal scale factor to dynamically balance inter-branch gradient contributions via norm-based scaling. Experiments on image classification, object detection, and language understanding tasks demonstrate that SURGE performs best over state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2605.10989

Code (0)

등록된 구현이 없습니다.

Tasks

Image ClassificationObject Detection

Similar Papers 제목 키워드 기반

SurGE: Surrogate Gradient-guided Evolution for Co-design of Legged Robots with Parallel Elasticity

2026-06-20 · Yulun Zhuang, Yue Qin, Justin Lu, Zelin Shen 외 arxiv

Co-design of legged robots with elastic elements is challenging due to the non-differentiability of contact dynamics and mechanism engagement. This paper presents SurGE, a framework that computes surrogate gradients of t…

A surrogate loss function for optimization of $F_β$ score in binary classification with imbalanced data

2021-04-03 · Namgil Lee, Heejung Yang, Hojin Yoo

The $F_\beta$ score is a commonly used measure of classification performance, which plays crucial roles in classification tasks with imbalanced data sets. However, the $F_\beta$ score cannot be used as a loss function by…

Binary ClassificationClassificationGeneral Classification

Elucidating the theoretical underpinnings of surrogate gradient learning in spiking neural networks

2024-04-23 · Julia Gygax, Friedemann Zenke

Training spiking neural networks to approximate universal functions is essential for studying information processing in the brain and for neuromorphic computing. Yet the binary nature of spikes poses a challenge for dire…

Relation

A generalized neural tangent kernel for surrogate gradient learning

2024-05-24 · Luke Eilers, Raoul-Martin Memmesheimer, Sven Goedeke

State-of-the-art neural network training methods depend on the gradient of the network function. Therefore, they cannot be applied to networks whose activation functions do not have useful derivatives, such as binary and…

A Framework for Flexible Peak Storm Surge Prediction

2022-04-27 · Benjamin Pachev, Prateek Arora, Carlos del-Castillo-Negrete, Eirik Valseth 외

Storm surge is a major natural hazard in coastal regions, responsible both for significant property damage and loss of life. Accurate, efficient models of storm surge are needed both to assess long-term risk and to guide…

ManagementPrediction