paper-with-me

홈 › Papers

Towards Scalable Backpropagation-Free Gradient Estimation

2025-11-05 · Daniel Wang, Evan Markou, Dylan Campbell arxiv

While backpropagation--reverse-mode automatic differentiation--has been extraordinarily successful in deep learning, it requires two passes (forward and backward) through the neural network and the storage of intermediate activations. Existing gradient estimation methods that instead use forward-mode automatic differentiation struggle to scale beyond small networks due to the high variance of the estimates. Efforts to mitigate this have so far introduced significant bias to the estimates, reducing their utility. We introduce a gradient estimation approach that reduces both bias and variance by manipulating upstream Jacobian matrices when computing guess directions. It shows promising results and has the potential to scale to larger networks, indeed performing better as the network width is increased. Our understanding of this method is facilitated by analyses of bias and variance, and their connection to the low-dimensional structure of neural network gradients.

📄 PDF Abstract BibTeX arXiv:2511.03110

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sparse Backpropagation for MoE Training

2023-10-01 · Liyuan Liu, Jianfeng Gao, Weizhu Chen

One defining characteristic of Mixture-of-Expert (MoE) models is their capacity for conducting sparse computation via expert routing, leading to remarkable scalability. However, backpropagation, the cornerstone of deep l…

Machine Translation

Fast Second-Order Stochastic Backpropagation for Variational Inference

2015-09-09 · Kai Fan, Ziteng Wang, Jeff Beck, James Kwok 외

We propose a second-order (Hessian or Hessian-free) based optimization method for variational inference inspired by Gaussian backpropagation, and argue that quasi-Newton optimization can be developed as well. This is acc…

regressionVariational Inference

Fast Second Order Stochastic Backpropagation for Variational Inference

2015-12-01 · NeurIPS 2015 12 · Kai Fan, Ziteng Wang, Jeff Beck, James Kwok 외

We propose a second-order (Hessian or Hessian-free) based optimization method for variational inference inspired by Gaussian backpropagation, and argue that quasi-Newton optimization can be developed as well. This is ac…

regressionVariational Inference

Probabilistic Unrolling: Scalable, Inverse-Free Maximum Likelihood Estimation for Latent Gaussian Models

2023-06-05 · Alexander Lin, Bahareh Tolooshams, Yves Atchadé, Demba Ba

Latent Gaussian models have a rich history in statistics and machine learning, with applications ranging from factor analysis to compressed sensing to time series analysis. The classical method for maximizing the likelih…

compressed sensingTime SeriesTime Series Analysis

Gradient-free variational learning with conditional mixture networks

2024-08-29 · Conor Heins, Hao Wu, Dimitrije Markovic, Alexander Tschantz 외

Balancing computational efficiency with robust predictive performance is crucial in supervised learning, especially for critical applications. Standard deep learning models, while accurate and scalable, often lack probab…

Computational EfficiencyMixture-of-ExpertsUncertainty QuantificationVariational Inference