paper-with-me

Papers

Piecewise convexity of artificial neural networks

2016-07-17 · Blaine Rister, Daniel L. Rubin

Although artificial neural networks have shown great promise in applications including computer vision and speech recognition, there remains considerable practical and theoretical difficulty in optimizing their parameters. The seemingly unreasonable success of gradient descent methods in minimizing these non-convex functions remains poorly understood. In this work we offer some theoretical guarantees for networks with piecewise affine activation functions, which have in recent years become the norm. We prove three main results. Firstly, that the network is piecewise convex as a function of the input data. Secondly, that the network, considered as a function of the parameters in a single layer, all others held constant, is again piecewise convex. Finally, that the network as a function of all its parameters is piecewise multi-convex, a generalization of biconvexity. From here we characterize the local minima and stationary points of the training objective, showing that they minimize certain subsets of the parameter space. We then analyze the performance of two optimization algorithms on multi-convex problems: gradient descent, and a method which repeatedly solves a number of convex sub-problems. We prove necessary convergence conditions for the first algorithm and both necessary and sufficient conditions for the second, after introducing regularization to the objective. Finally, we remark on the remaining difficulty of the global optimization problem. Under the squared error objective, we show that by varying the training data, a single rectifier neuron admits local minima arbitrarily far apart, both in objective value and parameter space.

📄 PDF Abstract BibTeX arXiv:1607.04917

Code (0)

등록된 구현이 없습니다.

Tasks

global-optimizationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Perfect reconstruction of sparse signals with piecewise continuous nonconvex penalties and nonconvexity control

2019-02-20 · Ayaka Sakata, Tomoyuki Obuchi

We consider compressed sensing formulated as a minimization problem of nonconvex sparse penalties, Smoothly Clipped Absolute deviation (SCAD) and Minimax Concave Penalty (MCP). The nonconvexity of these penalties is cont…

compressed sensing

Piecewise Strong Convexity of Neural Networks

2018-10-30 · NeurIPS 2019 12 · Tristan Milne

We study the loss surface of a feed-forward neural network with ReLU non-linearities, regularized with weight decay. We show that the regularized loss function is piecewise strongly convex on an important open set which …

image-classificationImage ClassificationLearning Theory

Piecewise Convex Function Estimation and Model Selection

2018-03-11 · Kurt S. Riedel

Given noisy data, function estimation is considered when the unknown function is known apriori to consist of a small number of regions where the function is either convex or concave. When the regions are known apriori, t…

modelModel Selection

Data-driven forced response analysis with min-max representations of nonlinear restoring forces

2026-03-17 · Akira Saito, Hiromu Fujita arxiv

This paper discusses a novel data-driven nonlinearity identification method for mechanical systems with nonlinear restoring forces such as polynomial, piecewise-linear, and general displacement-dependent nonlinearities. …

Couplings for Andersen Dynamics

2020-09-29 · Nawaf Bou-Rabee, Andreas Eberle

Andersen dynamics is a standard method for molecular simulations, and a precursor of the Hamiltonian Monte Carlo algorithm used in MCMC inference. The stochastic process corresponding to Andersen dynamics is a PDMP (piec…