paper-with-me

홈 › Papers

Learning One-hidden-layer ReLU Networks via Gradient Descent

2018-06-20 · Xiao Zhang, Yaodong Yu, Lingxiao Wang, Quanquan Gu

We study the problem of learning one-hidden-layer neural networks with Rectified Linear Unit (ReLU) activation function, where the inputs are sampled from standard Gaussian distribution and the outputs are generated from a noisy teacher network. We analyze the performance of gradient descent for training such kind of neural networks based on empirical risk minimization, and provide algorithm-dependent guarantees. In particular, we prove that tensor initialization followed by gradient descent can converge to the ground-truth parameters at a linear rate up to some statistical error. To the best of our knowledge, this is the first work characterizing the recovery guarantee for practical learning of one-hidden-layer ReLU networks with multiple neurons. Numerical experiments verify our theoretical findings.

📄 PDF Abstract BibTeX arXiv:1806.07808

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear Widths

2021-01-24 · Quynh Nguyen

We give a simple proof for the global convergence of gradient descent in training deep ReLU networks with the standard square loss, and show some of its improvements over the state-of-the-art. In particular, while prior …

No Spurious Local Minima in a Two Hidden Unit ReLU Network

2018-01-01 · ICLR 2018 1 · Chenwei Wu, Jiajun Luo, Jason D. Lee

Deep learning models can be efficiently optimized via stochastic gradient descent, but there is little theoretical evidence to support this. A key question in optimization is to understand when the optimization landscape…

Vocal Bursts Valence Prediction

Convex Approximation of Two-Layer ReLU Networks for Hidden State Differential Privacy

2024-07-05 · Rob Romijnders, Antti Koskela

The hidden state threat model of differential privacy (DP) assumes that the adversary has access only to the final trained machine learning (ML) model, without seeing intermediate states during training. However, the cur…

A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions

2021-04-01 · Arnulf Jentzen, Adrian Riekert

In this article we study the stochastic gradient descent (SGD) optimization method in the training of fully-connected feedforward artificial neural networks with ReLU activation. The main result of this work proves that …

The Implicit Bias of Minima Stability in Multivariate Shallow ReLU Networks

2023-06-30 · Mor Shpigel Nacson, Rotem Mulayoff, Greg Ongie, Tomer Michaeli 외

We study the type of solutions to which stochastic gradient descent converges when used to train a single hidden-layer multivariate ReLU network with the quadratic loss. Our results are based on a dynamical stability ana…