paper-with-me

홈 › Papers

Theory of Deep Learning IIb: Optimization Properties of SGD

2018-01-07 · Chiyuan Zhang, Qianli Liao, Alexander Rakhlin, Brando Miranda, Noah Golowich, Tomaso Poggio

In Theory IIb we characterize with a mix of theory and experiments the optimization of deep convolutional networks by Stochastic Gradient Descent. The main new result in this paper is theoretical and experimental evidence for the following conjecture about SGD: SGD concentrates in probability -- like the classical Langevin equation -- on large volume, "flat" minima, selecting flat minimizers which are with very high probability also global minimizers

📄 PDF Abstract BibTeX arXiv:1801.02254

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Learning

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Query Optimization Properties of Modified VBS

2019-09-26 · Mieczysław A. Kłopotek, Sławomir T. Wierzchoń

Valuation-Based~System can represent knowledge in different domains including probability theory, Dempster-Shafer theory and possibility theory. More recent studies show that the framework of VBS is also appropriate for …

Unifying Framework for Optimizations in non-boolean Formalisms

2022-06-16 · Yuliya Lierler

Search-optimization problems are plentiful in scientific and engineering domains. Artificial intelligence has long contributed to the development of search algorithms and declarative programming languages geared towards …

Perspectives on Contractivity in Control, Optimization, and Learning

2024-04-17 · Alexander Davydov, Francesco Bullo

Contraction theory is a mathematical framework for studying the convergence, robustness, and modularity properties of dynamical systems and algorithms. In this opinion paper, we provide five main opinions on the virtues …

Re-examination of Bregman functions and new properties of their divergences

2018-03-01 · Daniel Reem, Simeon Reich, Alvaro De Pierro

The Bregman divergence (Bregman distance, Bregman measure of distance) is a certain useful substitute for a distance, obtained from a well-chosen function (the "Bregman function"). Bregman functions and divergences have …

ASGO: Adaptive Structured Gradient Optimization

2025-03-26 · Kang An, Yuxing Liu, Rui Pan, Yi Ren 외

Training deep neural networks is a structured optimization problem, because the parameters are naturally represented by matrices and tensors rather than by vectors. Under this structural representation, it has been widel…

Language ModelingLanguage Modelling