paper-with-me

홈 › Papers

Small nonlinearities in activation functions create bad local minima in neural networks

2018-02-10 · ICLR 2019 5 · Chulhee Yun, Suvrit Sra, Ali Jadbabaie

We investigate the loss surface of neural networks. We prove that even for one-hidden-layer networks with "slightest" nonlinearity, the empirical risks have spurious local minima in most cases. Our results thus indicate that in general "no spurious local minima" is a property limited to deep linear networks, and insights obtained from linear networks may not be robust. Specifically, for ReLU(-like) networks we constructively prove that for almost all practical datasets there exist infinitely many local minima. We also present a counterexample for more general activations (sigmoid, tanh, arctan, ReLU, etc.), for which there exists a bad local minimum. Our results make the least restrictive assumptions relative to existing results on spurious local optima in neural networks. We complete our discussion by presenting a comprehensive characterization of global optimality for deep linear networks, which unifies other results on this topic.

📄 PDF Abstract BibTeX arXiv:1802.03487

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Automated Design of Linear Bounding Functions for Sigmoidal Nonlinearities in Neural Networks

2024-06-14 · Matthias König, Xiyue Zhang, Holger H. Hoos, Marta Kwiatkowska 외

The ubiquity of deep learning algorithms in various applications has amplified the need for assuring their robustness against small input perturbations such as those occurring in adversarial attacks. Existing complete ve…

From Hard to Soft: Understanding Deep Network Nonlinearities via Vector Quantization and Statistical Inference

2018-10-22 · ICLR 2019 5 · Randall Balestriero, Richard G. Baraniuk

Nonlinearity is crucial to the performance of a deep (neural) network (DN). To date there has been little progress understanding the menagerie of available nonlinearities, but recently progress has been made on understan…

Quantization

Function-Space Optimality of Neural Architectures with Multivariate Nonlinearities

2023-10-05 · Rahul Parhi, Michael Unser

We investigate the function-space optimality (specifically, the Banach-space optimality) of a large class of shallow neural architectures with multivariate nonlinearities/activation functions. To that end, we construct a…

Two-argument activation functions learn soft XOR operations like cortical neurons

2021-10-13 · KiJung Yoon, Emin Orhan, Juhyun Kim, Xaq Pitkow

Neurons in the brain are complex machines with distinct functional compartments that interact nonlinearly. In contrast, neurons in artificial neural networks abstract away this complexity, typically down to a scalar acti…

Vocal Bursts Valence Prediction

Efficient Activation Function Optimization through Surrogate Modeling

2023-01-13 · NeurIPS 2023 11 · Garrett Bingham, Risto Miikkulainen

Carefully designed activation functions can improve the performance of neural networks in many machine learning tasks. However, it is difficult for humans to construct optimal activation functions, and current activation…