paper-with-me

Papers

Error bounds for particle gradient descent, and extensions of the log-Sobolev and Talagrand inequalities

2024-03-04 · Rocco Caprio, Juan Kuntz, Samuel Power, Adam M. Johansen

We prove non-asymptotic error bounds for particle gradient descent (PGD)~(Kuntz et al., 2023), a recently introduced algorithm for maximum likelihood estimation of large latent variable models obtained by discretizing a gradient flow of the free energy. We begin by showing that, for models satisfying a condition generalizing both the log-Sobolev and the Polyak--{\L}ojasiewicz inequalities (LSI and P{\L}I, respectively), the flow converges exponentially fast to the set of minimizers of the free energy. We achieve this by extending a result well-known in the optimal transport literature (that the LSI implies the Talagrand inequality) and its counterpart in the optimization literature (that the P{\L}I implies the so-called quadratic growth condition), and applying it to our new setting. We also generalize the Bakry--\'Emery Theorem and show that the LSI/P{\L}I generalization holds for models with strongly concave log-likelihoods. For such models, we further control PGD's discretization error, obtaining non-asymptotic error bounds. While we are motivated by the study of PGD, we believe that the inequalities and results we extend may be of independent interest.

📄 PDF Abstract BibTeX arXiv:2403.02004

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Finite-Particle Rates for Regularized Stein Variational Gradient Descent

2026-02-05 · Ye He, Krishnakumar Balasubramanian, Sayan Banerjee, Promit Ghosal arxiv

We derive finite-particle rates for the regularized Stein variational gradient descent (R-SVGD) algorithm introduced by He et al. (2024) that corrects the constant-order bias of the SVGD by applying a resolvent-type prec…

Efficient displacement convex optimization with particle gradient descent

2023-02-09 · Hadi Daneshmand, Jason D. Lee, Chi Jin

Particle gradient descent, which uses particles to represent a probability measure and performs gradient descent on particles in parallel, is widely used to optimize functions of probability measures. This paper consider…

Beyond Propagation of Chaos: A Stochastic Algorithm for Mean Field Optimization

2025-03-17 · Chandan Tankala, Dheeraj M. Nagaraj, Anant Raj

Gradient flow in the 2-Wasserstein space is widely used to optimize functionals over probability distributions and is typically implemented using an interacting particle system with $n$ particles. Analyzing these algorit…

Learning Operators by Regularized Stochastic Gradient Descent with Operator-valued Kernels

2025-04-25 · Jia-Qi Yang, Lei Shi

This paper investigates regularized stochastic gradient descent (SGD) algorithms for estimating nonlinear operators from a Polish space to a separable Hilbert space. We assume that the regression operator lies in a vecto…

Decoder

Superpolynomial Lower Bounds for Learning One-Layer Neural Networks using Gradient Descent

2020-06-22 · ICML 2020 1 · Surbhi Goel, Aravind Gollakota, Zhihan Jin, Sushrut Karmalkar 외

We prove the first superpolynomial lower bounds for learning one-layer neural networks with respect to the Gaussian distribution using gradient descent. We show that any classifier trained using gradient descent with res…