paper-with-me

홈 › Papers

Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees

2023-05-06 · Nachuan Xiao, Xiaoyin Hu, Xin Liu, Kim-Chuan Toh

In this paper, we present a comprehensive study on the convergence properties of Adam-family methods for nonsmooth optimization, especially in the training of nonsmooth neural networks. We introduce a novel two-timescale framework that adopts a two-timescale updating scheme, and prove its convergence properties under mild assumptions. Our proposed framework encompasses various popular Adam-family methods, providing convergence guarantees for these methods in training nonsmooth neural networks. Furthermore, we develop stochastic subgradient methods that incorporate gradient clipping techniques for training nonsmooth neural networks with heavy-tailed noise. Through our framework, we show that our proposed methods converge even when the evaluation noises are only assumed to be integrable. Extensive numerical experiments demonstrate the high efficiency and robustness of our proposed methods.

📄 PDF Abstract BibTeX arXiv:2305.03938

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Gradient Clipping One difficulty that arises with optimization of deep neural networks is that large parameter gradients can lead an SGD optimizer to update…

Similar Papers 제목 키워드 기반

Adam-family Methods with Decoupled Weight Decay in Deep Learning

2023-10-13 · Kuangyu Ding, Nachuan Xiao, Kim-Chuan Toh

In this paper, we investigate the convergence properties of a wide class of Adam-family methods for minimizing quadratically regularized nonsmooth nonconvex optimization problems, especially in the context of training no…

Deep Learning

Adam Converges in Nonsmooth Nonconvex Optimization

2026-06-21 · Zijian Liu arxiv

Adam is one of the most widely implemented and influential modern optimizers. Why is it effective across different optimization problems in practice? This question arguably lies at the center of the optimization communit…

Developing Lagrangian-based Methods for Nonsmooth Nonconvex Optimization

2024-04-15 · Nachuan Xiao, Kuangyu Ding, Xiaoyin Hu, Kim-Chuan Toh

In this paper, we consider the minimization of a nonsmooth nonconvex objective function $f(x)$ over a closed convex subset $\mathcal{X}$ of $\mathbb{R}^n$, with additional nonsmooth nonconvex constraints $c(x) = 0$. We d…

Adam with model exponential moving average is effective for nonconvex optimization

2024-05-28 · Kwangjun Ahn, Ashok Cutkosky

In this work, we offer a theoretical analysis of two modern optimization techniques for training large and complex models: (i) adaptive optimization algorithms, such as Adam, and (ii) the model exponential moving average…

A Novel Convergence Analysis for Algorithms of the Adam Family

2021-12-07 · Zhishuai Guo, Yi Xu, Wotao Yin, Rong Jin 외

Since its invention in 2014, the Adam optimizer has received tremendous attention. On one hand, it has been widely used in deep learning and many variants have been proposed, while on the other hand their theoretical con…

Bilevel Optimization