paper-with-me

홈 › Papers

On the asymptotics of wide networks with polynomial activations

2020-06-11 · Kyle Aitken, Guy Gur-Ari

We consider an existing conjecture addressing the asymptotic behavior of neural networks in the large width limit. The results that follow from this conjecture include tight bounds on the behavior of wide networks during stochastic gradient descent, and a derivation of their finite-width dynamics. We prove the conjecture for deep networks with polynomial activation functions, greatly extending the validity of these results. Finally, we point out a difference in the asymptotic behavior of networks with analytic (and non-linear) activation functions and those with piecewise-linear activations such as ReLU.

📄 PDF Abstract BibTeX arXiv:2006.06687

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Exact asymptotics for phase retrieval and compressed sensing with random generative priors

2019-12-04 · Benjamin Aubin, Bruno Loureiro, Antoine Baker, Florent Krzakala 외

We consider the problem of compressed sensing and of (real-valued) phase retrieval with random measurement matrix. We derive sharp asymptotics for the information-theoretically optimal performance and for the best known …

compressed sensingRetrieval

Exponential Approximation Rates and Parameter Efficiency of Learnable Bernstein Activations

2026-02-04 · Ibrahim Albool, Malak Gamal El-Din, Salma Elmalaki, Yasser Shoukry arxiv

The choice of activation function fundamentally shapes the representational capacity and parameter efficiency of deep neural networks, yet most widely used activations lack rigorous theoretical guarantees on these proper…

The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks

2025-05-23 · Vittorio Erba, Emanuele Troiani, Lenka Zdeborová, Florent Krzakala

We study the high-dimensional asymptotics of empirical risk minimization (ERM) in over-parametrized two-layer neural networks with quadratic activations trained on synthetic data. We derive sharp asymptotics for both tra…

Polynomial, trigonometric, and tropical activations

2025-02-03 · Ismail Khalfaoui-Hassani, Stefan Kesselheim

Which functions can be used as activations in deep neural networks? This article explores families of functions based on orthonormal bases, including the Hermite polynomial basis and the Fourier trigonometric basis, as w…

image-classificationImage ClassificationLanguage ModellingText Generation

Topological complexity of spiked random polynomials and finite-rank spherical integrals

2023-12-19 · Vanessa Piccolo

We study the annealed complexity of a random Gaussian homogeneous polynomial on the $N$-dimensional unit sphere in the presence of deterministic polynomials that depend on fixed unit vectors and external parameters. In p…