paper-with-me

홈 › Papers

The effect of Leaky ReLUs on the training and generalization of overparameterized networks

2024-02-19 · Yinglong Guo, Shaohan Li, Gilad Lerman

We investigate the training and generalization errors of overparameterized neural networks (NNs) with a wide class of leaky rectified linear unit (ReLU) functions. More specifically, we carefully upper bound both the convergence rate of the training error and the generalization error of such NNs and investigate the dependence of these bounds on the Leaky ReLU parameter, $\alpha$. We show that $\alpha =-1$, which corresponds to the absolute value activation function, is optimal for the training error bound. Furthermore, in special settings, it is also optimal for the generalization error bound. Numerical experiments empirically support the practical choices guided by the theory.

📄 PDF Abstract BibTeX arXiv:2402.11942

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…

Similar Papers 제목 키워드 기반

Leaky ReLUs That Differ in Forward and Backward Pass Facilitate Activation Maximization in Deep Neural Networks

2024-10-22 · Christoph Linse, Erhardt Barth, Thomas Martinetz

Activation maximization (AM) strives to generate optimal input stimuli, revealing features that trigger high responses in trained deep neural networks. AM is an important method of explainable AI. We demonstrate that AM …

Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)

2015-11-23 · Djork-Arné Clevert, Thomas Unterthiner, Sepp Hochreiter

We introduce the "exponential linear unit" (ELU) which speeds up learning in deep neural networks and leads to higher classification accuracies. Like rectified linear units (ReLUs), leaky ReLUs (LReLUs) and parametrized …

General ClassificationImage Classification

Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers

2022-03-15 · ICLR 2022 4 · Guodong Zhang, Aleksandar Botev, James Martens

Training very deep neural networks is still an extremely challenging task. The common solution is to use shortcut connections and normalization layers, which are both crucial ingredients in the popular ResNet architectur…

Deep Learning

Exploring the Long-Term Generalization of Counting Behavior in RNNs

2022-11-29 · Nadine El-Naggar, Pranava Madhyastha, Tillman Weyde

In this study, we investigate the generalization of LSTM, ReLU and GRU models on counting tasks over long sequences. Previous theoretical work has established that RNNs with ReLU activation and LSTMs have the capacity fo…

Towards moderate overparameterization: global convergence guarantees for training shallow neural networks

2019-02-12 · Samet Oymak, Mahdi Soltanolkotabi

Many modern neural network architectures are trained in an overparameterized regime where the parameters of the model exceed the size of the training dataset. Sufficiently overparameterized neural network architectures i…