paper-with-me

홈 › Papers

Beyond Discreteness: Finite-Sample Analysis of Straight-Through Estimator for Quantization

2025-05-23 · Halyun Jeong, Jack Xin, Penghang Yin

Training quantized neural networks requires addressing the non-differentiable and discrete nature of the underlying optimization problem. To tackle this challenge, the straight-through estimator (STE) has become the most widely adopted heuristic, allowing backpropagation through discrete operations by introducing surrogate gradients. However, its theoretical properties remain largely unexplored, with few existing works simplifying the analysis by assuming an infinite amount of training data. In contrast, this work presents the first finite-sample analysis of STE in the context of neural network quantization. Our theoretical results highlight the critical role of sample size in the success of STE, a key insight absent from existing studies. Specifically, by analyzing the quantization-aware training of a two-layer neural network with binary weights and activations, we derive the sample complexity bound in terms of the data dimensionality that guarantees the convergence of STE-based optimization to the global minimum. Moreover, in the presence of label noises, we uncover an intriguing recurrence property of STE-gradient method, where the iterate repeatedly escape from and return to the optimal binary weights. Our analysis leverages tools from compressed sensing and dynamical systems theory.

📄 PDF Abstract BibTeX arXiv:2505.18113

Code (0)

등록된 구현이 없습니다.

Tasks

compressed sensingQuantization

Similar Papers 제목 키워드 기반

Learning with risks based on M-location

2020-12-04 · Matthew J. Holland

In this work, we study a new class of risks defined in terms of the location and deviation of the loss distribution, generalizing far beyond classical mean-variance risk functions. The class is easily implemented as a wr…

Stochastic Optimization

Beyond Uncertainty Sets: Leveraging Optimal Transport to Extend Conformal Predictive Distribution to Multivariate Settings

2025-11-19 · Eugene Ndiaye arxiv

Conformal prediction (CP) constructs uncertainty sets for model outputs with finite-sample coverage guarantees. A candidate output is included in the prediction set if its non-conformity score is not considered extreme r…

Exact Discrete Stochastic Simulation with Deep-Learning-Scale Gradient Optimization

2026-02-23 · Jose M. G. Vilar, Leonor Saiz arxiv

Exact stochastic simulation of continuous-time Markov chains (CTMCs) is essential when discreteness and noise drive system behavior, but the hard categorical event selection in Gillespie-type algorithms blocks gradient-b…

On Finite-sample Concentration of Median of Incomplete U-Statistics

2026-05-30 · Nong Minh Hieu arxiv

Median-of-means (MoM) is a powerful technique that theoretically enables near sub-Gaussian finite-sample rate for parameter estimation when the underlying data distribution is heavy-tailed (e.g., assumed to have only two…

Discretized Integrated Gradients for Explaining Language Models

2021-08-31 · EMNLP 2021 11 · Soumya Sanyal, Xiang Ren

As a prominent attribution-based explanation algorithm, Integrated Gradients (IG) is widely adopted due to its desirable explanation axioms and the ease of gradient computation. It measures feature importance by averagin…

Feature ImportanceSentiment AnalysisSentiment Classification