paper-with-me

홈 › Papers

Spurious Local Minima of Deep ReLU Neural Networks in the Neural Tangent Kernel Regime

2018-06-13 · Tohru Nitta

In this paper, we theoretically prove that the deep ReLU neural networks do not lie in spurious local minima in the loss landscape under the Neural Tangent Kernel (NTK) regime, that is, in the gradient descent training dynamics of the deep ReLU neural networks whose parameters are initialized by a normal distribution in the limit as the widths of the hidden layers tend to infinity.

📄 PDF Abstract BibTeX arXiv:1806.04884

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Kaiming Initialization 설명 없음

Similar Papers 제목 키워드 기반

LoRA Training in the NTK Regime has No Spurious Local Minima

2024-02-19 · Uijeong Jang, Jason D. Lee, Ernest K. Ryu

Low-rank adaptation (LoRA) has become the standard approach for parameter-efficient fine-tuning of large language models (LLM), but our theoretical understanding of LoRA has been limited. In this work, we theoretically a…

parameter-efficient fine-tuning

Small nonlinearities in activation functions create bad local minima in neural networks

2018-02-10 · ICLR 2019 5 · Chulhee Yun, Suvrit Sra, Ali Jadbabaie

We investigate the loss surface of neural networks. We prove that even for one-hidden-layer networks with "slightest" nonlinearity, the empirical risks have spurious local minima in most cases. Our results thus indicate …

Spurious Local Minima are Common in Two-Layer ReLU Neural Networks

2017-12-24 · ICML 2018 7 · Itay Safran, Ohad Shamir

We consider the optimization problem associated with training simple ReLU neural networks of the form $\mathbf{x}\mapsto \sum_{i=1}^{k}\max\{0,\mathbf{w}_i^\top \mathbf{x}\}$ with respect to the squared loss. We provide …

Vocal Bursts Valence Prediction

Optimal Rates for Averaged Stochastic Gradient Descent under Neural Tangent Kernel Regime

2020-06-22 · ICLR 2021 1 · Atsushi Nitanda, Taiji Suzuki

We analyze the convergence of the averaged stochastic gradient descent for overparameterized two-layer neural networks for regression problems. It was recently found that a neural tangent kernel (NTK) plays an important …

Neural Tangent Kernels and Fisher Information Matrices for Simple ReLU Networks with Random Hidden Weights

2025-07-24 · Jun'ichi Takeuchi, Yoshinari Takeishi, Noboru Murata, Kazushi Mimura 외 arxiv

Fisher information matrices and neural tangent kernels (NTK) for 2-layer ReLU networks with random hidden weight are argued. We discuss the relation between both notions as a linear transformation and show that spectral …