paper-with-me

Papers

Information-theoretic reduction of deep neural networks to linear models in the overparametrized proportional regime

2025-05-06 · Francesco Camilli, Daria Tieplova, Eleonora Bergamin, Jean Barbier

We rigorously analyse fully-trained neural networks of arbitrary depth in the Bayesian optimal setting in the so-called proportional scaling regime where the number of training samples and width of the input and all inner layers diverge proportionally. We prove an information-theoretic equivalence between the Bayesian deep neural network model trained from data generated by a teacher with matching architecture, and a simpler model of optimal inference in a generalized linear model. This equivalence enables us to compute the optimal generalization error for deep neural networks in this regime. We thus prove the "deep Gaussian equivalence principle" conjectured in Cui et al. (2023) (arXiv:2302.00375). Our result highlights that in order to escape this "trivialisation" of deep neural networks (in the sense of reduction to a linear model) happening in the strongly overparametrized proportional regime, models trained from much more data have to be considered.

📄 PDF Abstract BibTeX arXiv:2505.03577

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploiting Non-Linear Redundancy for Neural Model Compression

2020-05-28 · Muhammad A. Shah, Raphael Olivier, Bhiksha Raj

Deploying deep learning models, comprising of non-linear combination of millions, even billions, of parameters is challenging given the memory, power and compute constraints of the real world. This situation has led to r…

modelModel Compression

VC dimension of partially quantized neural networks in the overparametrized regime

2021-10-06 · ICLR 2022 4 · Yutong Wang, Clayton D. Scott

Vapnik-Chervonenkis (VC) theory has so far been unable to explain the small generalization error of overparametrized neural networks. Indeed, existing applications of VC theory to large networks obtain upper bounds on VC…

Fundamental limits of overparametrized shallow neural networks for supervised learning

2023-07-11 · Francesco Camilli, Daria Tieplova, Jean Barbier

We carry out an information-theoretical analysis of a two-layer neural network trained from input-output pairs generated by a teacher network with matching architecture, in overparametrized regimes. Our results come in t…

Training (Overparametrized) Neural Networks in Near-Linear Time

2020-06-20 · Jan van den Brand, Binghui Peng, Zhao Song, Omri Weinstein

The slow convergence rate and pathological curvature issues of first-order gradient methods for training deep neural networks, initiated an ongoing effort for developing faster $\mathit{second}$-$\mathit{order}$ optimiza…

Dimensionality Reductionregression

Linear Recursive Feature Machines provably recover low-rank matrices

2024-01-09 · Adityanarayanan Radhakrishnan, Mikhail Belkin, Dmitriy Drusvyatskiy

A fundamental problem in machine learning is to understand how neural networks make accurate predictions, while seemingly bypassing the curse of dimensionality. A possible explanation is that common training algorithms f…

Dimensionality ReductionLow-Rank Matrix CompletionMatrix Completionregression