paper-with-me

Papers

Finite-Sum Optimization: A New Perspective for Convergence to a Global Solution

2022-02-07 · Lam M. Nguyen, Trang H. Tran, Marten van Dijk

Deep neural networks (DNNs) have shown great success in many machine learning tasks. Their training is challenging since the loss surface of the network architecture is generally non-convex, or even non-smooth. How and under what assumptions is guaranteed convergence to a \textit{global} minimum possible? We propose a reformulation of the minimization problem allowing for a new recursive algorithmic framework. By using bounded style assumptions, we prove convergence to an $\varepsilon$-(global) minimum using $\mathcal{\tilde{O}}(1/\varepsilon^3)$ gradient computations. Our theoretical foundation motivates further study, implementation, and optimization of the new algorithmic framework and further investigation of its non-standard bounded style assumptions. This new direction broadens our understanding of why and under what circumstances training of a DNN converges to a global minimum.

📄 PDF Abstract BibTeX arXiv:2202.03524

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

New Perspective on the Global Convergence of Finite-Sum Optimization

2021-09-29 · Lam M. Nguyen, Trang H. Tran, Marten van Dijk

Deep neural networks (DNNs) have shown great success in many machine learning tasks. Their training is challenging since the loss surface of the network architecture is generally non-convex, or even non-smooth. How and u…

Global Convergence of Policy Gradient Methods to (Almost) Locally Optimal Policies

2019-06-19 · Kaiqing Zhang, Alec Koppel, Hao Zhu, Tamer Başar

Policy gradient (PG) methods are a widely used reinforcement learning methodology in many applications such as video games, autonomous driving, and robotics. In spite of its empirical success, a rigorous understanding of…

Autonomous DrivingPolicy Gradient MethodsReinforcement Learning

Generalization bound of globally optimal non-convex neural network training: Transportation map estimation by infinite dimensional Langevin dynamics

2020-07-11 · NeurIPS 2020 12 · Taiji Suzuki

We introduce a new theoretical framework to analyze deep learning optimization with connection to its generalization error. Existing frameworks such as mean field theory and neural tangent kernel theory for neural networ…

Sharp global convergence guarantees for iterative nonconvex optimization: A Gaussian process perspective

2021-09-20 · Kabir Aladin Chandrasekher, Ashwin Pananjady, Christos Thrampoulidis

We consider a general class of regression models with normally distributed covariates, and the associated nonconvex problem of fitting these models from data. We develop a general recipe for analyzing the convergence of …

parameter estimationRetrieval

Global Convergence of Adjoint-Optimized Neural PDEs

2025-06-16 · Konstantin Riedl, Justin Sirignano, Konstantinos Spiliopoulos

Many engineering and scientific fields have recently become interested in modeling terms in partial differential equations (PDEs) with neural networks. The resulting neural-network PDE model, being a function of the neur…