paper-with-me

Papers

Sparsity-depth Tradeoff in Infinitely Wide Deep Neural Networks

2023-05-17 · Chanwoo Chun, Daniel D. Lee

We investigate how sparse neural activity affects the generalization performance of a deep Bayesian neural network at the large width limit. To this end, we derive a neural network Gaussian Process (NNGP) kernel with rectified linear unit (ReLU) activation and a predetermined fraction of active neurons. Using the NNGP kernel, we observe that the sparser networks outperform the non-sparse networks at shallow depths on a variety of datasets. We validate this observation by extending the existing theory on the generalization error of kernel-ridge regression.

📄 PDF Abstract BibTeX arXiv:2305.10550

Code (0)

등록된 구현이 없습니다.

Tasks

regression

Methods 이 논문이 사용한 방법론

Gaussian Process Gaussian Processes are non-parametric models for approximating functions. They rely upon a measure of similarity between points (the kernel function) to predict the value for…

Similar Papers 제목 키워드 기반

Convex Sparse Blind Deconvolution

2021-06-13 · Qingyun Sun, David Donoho

In the blind deconvolution problem, we observe the convolution of an unknown filter and unknown signal and attempt to reconstruct the filter and signal. The problem seems impossible in general, since there are seemingly …

Neural Tangent Kernel Analysis of Deep Narrow Neural Networks

2022-02-07 · Jongmin Lee, Joo Young Choi, Ernest K. Ryu, Albert No

The tremendous recent progress in analyzing the training dynamics of overparameterized neural networks has primarily focused on wide networks and therefore does not sufficiently address the role of depth in deep learning…

Revisiting the Neural Tangent Kernel: the role of large width and depth

2025-11-10 · William St-Arnaud, Margarida Carvalho, Golnoosh Farnadi arxiv

Overparameterized fully-connected neural networks have been shown to behave like kernel models when trained with gradient descent, assuming standard scaling conditions on the width, the learning rate, and the parameter i…

Revisiting Deep Information Propagation: Fractal Frontier and Finite-size Effects

2025-08-05 · Giuseppe Alessio D'Inverno, Zhiyuan Hu, Leo Davy, Michael Unser 외 arxiv

Information propagation characterizes how input correlations evolve across layers in deep neural networks. This framework has been well studied using mean-field theory, which assumes infinitely wide networks. However, th…

Analyzing Finite Neural Networks: Can We Trust Neural Tangent Kernel Theory?

2020-12-08 · Mariia Seleznova, Gitta Kutyniok

Neural Tangent Kernel (NTK) theory is widely used to study the dynamics of infinitely-wide deep neural networks (DNNs) under gradient descent. But do the results for infinitely-wide networks give us hints about the behav…

valid