paper-with-me

홈 › Papers

Charged Point Normalization: An Efficient Solution to the Saddle Point Problem

2016-09-29 · Armen Aghajanyan

Recently, the problem of local minima in very high dimensional non-convex optimization has been challenged and the problem of saddle points has been introduced. This paper introduces a dynamic type of normalization that forces the system to escape saddle points. Unlike other saddle point escaping algorithms, second order information is not utilized, and the system can be trained with an arbitrary gradient descent learner. The system drastically improves learning in a range of deep neural networks on various data-sets in comparison to non-CPN neural networks.

📄 PDF Abstract BibTeX arXiv:1609.09522

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Power of Normalization: Faster Evasion of Saddle Points

2016-11-15 · Kfir. Y. Levy

A commonly used heuristic in non-convex optimization is Normalized Gradient Descent (NGD) - a variant of gradient descent in which only the direction of the gradient is taken into account and its magnitude ignored. We an…

Tensor Decomposition

Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points

2025-09-21 · Naoya Yamamoto, Juno Kim, Taiji Suzuki arxiv

Wasserstein gradient flow (WGF) is a common method to perform optimization over the space of probability measures. While WGF is guaranteed to converge to a first-order stationary point, for nonconvex functionals the conv…

The Landscape of Matrix Factorization Revisited

2020-02-27 · Hossein Valavi, Sulin Liu, Peter J. Ramadge

We revisit the landscape of the simple matrix factorization problem. For low-rank matrix factorization, prior work has shown that there exist infinitely many critical points all of which are either global minima or stric…

Last-Iterate Convergence of Saddle-Point Optimizers via High-Resolution Differential Equations

2021-12-27 · Tatjana Chavdarova, Michael I. Jordan, Manolis Zampetakis

Several widely-used first-order saddle-point optimization methods yield an identical continuous-time ordinary differential equation (ODE) that is identical to that of the Gradient Descent Ascent (GDA) method when derived…

Generalized Gradient Flows with Provable Fixed-Time Convergence and Fast Evasion of Non-Degenerate Saddle Points

2022-12-07 · Mayank Baranwal, Param Budhraja, Vishal Raj, Ashish R. Hota

Gradient-based first-order convex optimization algorithms find widespread applicability in a variety of domains, including machine learning tasks. Motivated by the recent advances in fixed-time stability theory of contin…