paper-with-me

Papers

Piecewise Strong Convexity of Neural Networks

2018-10-30 · NeurIPS 2019 12 · Tristan Milne

We study the loss surface of a feed-forward neural network with ReLU non-linearities, regularized with weight decay. We show that the regularized loss function is piecewise strongly convex on an important open set which contains, under some conditions, all of its global minimizers. This is used to prove that local minima of the regularized loss function in this set are isolated, and that every differentiable critical point in this set is a local minimum, partially addressing an open problem given at the Conference on Learning Theory (COLT) 2015; our result is also applied to linear neural networks to show that with weight decay regularization, there are no non-zero critical points in a norm ball obtaining training error below a given threshold. We also include an experimental section where we validate our theoretical work and show that the regularized loss function is almost always piecewise strongly convex when restricted to stochastic gradient descent trajectories for three standard image classification problems.

📄 PDF Abstract BibTeX arXiv:1810.12805

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage ClassificationLearning Theory

Methods 이 논문이 사용한 방법론

Weight Decay 설명 없음
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Piecewise convexity of artificial neural networks

2016-07-17 · Blaine Rister, Daniel L. Rubin

Although artificial neural networks have shown great promise in applications including computer vision and speech recognition, there remains considerable practical and theoretical difficulty in optimizing their parameter…

global-optimizationspeech-recognitionSpeech Recognition

Perfect reconstruction of sparse signals with piecewise continuous nonconvex penalties and nonconvexity control

2019-02-20 · Ayaka Sakata, Tomoyuki Obuchi

We consider compressed sensing formulated as a minimization problem of nonconvex sparse penalties, Smoothly Clipped Absolute deviation (SCAD) and Minimax Concave Penalty (MCP). The nonconvexity of these penalties is cont…

compressed sensing

Piecewise Convex Function Estimation and Model Selection

2018-03-11 · Kurt S. Riedel

Given noisy data, function estimation is considered when the unknown function is known apriori to consist of a small number of regions where the function is either convex or concave. When the regions are known apriori, t…

modelModel Selection

Strong convexity-guided hyper-parameter optimization for flatter losses

2024-02-07 · Rahul Yedida, Snehanshu Saha

We propose a novel white-box approach to hyper-parameter optimization. Motivated by recent work establishing a relationship between flat minima and generalization, we first establish a relationship between the strong con…

Fast networked data selection via distributed smoothed quantile estimation

2024-06-04 · Xu Zhang, Marcos M. Vasconcelos

Collecting the most informative data from a large dataset distributed over a network is a fundamental problem in many fields, including control, signal processing and machine learning. In this paper, we establish a conne…