paper-with-me

홈 › Papers

Stochastic Particle Gradient Descent for Infinite Ensembles

2017-12-14 · Atsushi Nitanda, Taiji Suzuki

The superior performance of ensemble methods with infinite models are well known. Most of these methods are based on optimization problems in infinite-dimensional spaces with some regularization, for instance, boosting methods and convex neural networks use $L^1$-regularization with the non-negative constraint. However, due to the difficulty of handling $L^1$-regularization, these problems require early stopping or a rough approximation to solve it inexactly. In this paper, we propose a new ensemble learning method that performs in a space of probability measures, that is, our method can handle the $L^1$-constraint and the non-negative constraint in a rigorous way. Such an optimization is realized by proposing a general purpose stochastic optimization method for learning probability measures via parameterization using transport maps on base models. As a result of running the method, a transport map to output an infinite ensemble is obtained, which forms a residual-type network. From the perspective of functional gradient methods, we give a convergence rate as fast as that of a stochastic optimization method for finite dimensional nonconvex problems. Moreover, we show an interior optimality property of a local optimality condition used in our analysis.

📄 PDF Abstract BibTeX arXiv:1712.05438

Code (0)

등록된 구현이 없습니다.

Tasks

Ensemble LearningStochastic Optimization

Methods 이 논문이 사용한 방법론

Early Stopping Early Stopping is a regularization technique for deep neural networks that stops training when parameter updates no longer begin to yield improves on a validation set. In…

Similar Papers 제목 키워드 기반

Beyond Propagation of Chaos: A Stochastic Algorithm for Mean Field Optimization

2025-03-17 · Chandan Tankala, Dheeraj M. Nagaraj, Anant Raj

Gradient flow in the 2-Wasserstein space is widely used to optimize functionals over probability distributions and is typically implemented using an interacting particle system with $n$ particles. Analyzing these algorit…

Limit Theorems for Stochastic Gradient Descent with Infinite Variance

2024-10-21 · Jose Blanchet, Aleksandar Mijatović, Wenhao Yang

Stochastic gradient descent is a classic algorithm that has gained great popularity especially in the last decades as the most common approach for training models in machine learning. While the algorithm has been well-st…

regression

Stochastic Approximation Algorithms for Systems of Interacting Particles

2023-09-21 · NeurIPS 2023 11

Interacting particle systems have proven highly successful in various machine learning tasks, including approximate Bayesian inference and neural network optimization. However, the analysis of these systems often relies …

Mean-field Langevin dynamics: Time-space discretization, stochastic gradient, and variance reduction

2023-09-21 · NeurIPS 2023 11

The mean-field Langevin dynamics (MFLD) is a nonlinear generalization of the Langevin dynamics that incorporates a distribution-dependent drift, and it naturally arises from the optimization of two-layer neural networks …

Convergence of mean-field Langevin dynamics: Time and space discretization, stochastic gradient, and variance reduction

2023-06-12 · Taiji Suzuki, Denny Wu, Atsushi Nitanda

The mean-field Langevin dynamics (MFLD) is a nonlinear generalization of the Langevin dynamics that incorporates a distribution-dependent drift, and it naturally arises from the optimization of two-layer neural networks …