paper-with-me

홈 › Papers

Bayesian Interpolation with Deep Linear Networks

2022-12-29 · Boris Hanin, Alexander Zlokapa

Characterizing how neural network depth, width, and dataset size jointly impact model quality is a central problem in deep learning theory. We give here a complete solution in the special case of linear networks with output dimension one trained using zero noise Bayesian inference with Gaussian weight priors and mean squared error as a negative log-likelihood. For any training dataset, network depth, and hidden layer widths, we find non-asymptotic expressions for the predictive posterior and Bayesian model evidence in terms of Meijer-G functions, a class of meromorphic special functions of a single complex variable. Through novel asymptotic expansions of these Meijer-G functions, a rich new picture of the joint role of depth, width, and dataset size emerges. We show that linear networks make provably optimal predictions at infinite depth: the posterior of infinitely deep linear networks with data-agnostic priors is the same as that of shallow networks with evidence-maximizing data-dependent priors. This yields a principled reason to prefer deeper networks when priors are forced to be data-agnostic. Moreover, we show that with data-agnostic priors, Bayesian model evidence in wide linear networks is maximized at infinite depth, elucidating the salutary role of increased depth for model selection. Underpinning our results is a novel emergent notion of effective depth, given by the number of hidden layers times the number of data points divided by the network width; this determines the structure of the posterior in the large-data limit.

📄 PDF Abstract BibTeX arXiv:2212.14457

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian InferenceLearning TheoryModel Selection

Methods 이 논문이 사용한 방법론

Gaussian Process Gaussian Processes are non-parametric models for approximating functions. They rely upon a measure of similarity between points (the kernel function) to predict the value for…

Similar Papers 제목 키워드 기반

Connecting and Comparing Language Model Interpolation Techniques

2019-08-26 · Ernest Pusateri, Christophe Van Gysel, Rami Botros, Sameer Badaskar 외

In this work, we uncover a theoretical connection between two language model interpolation techniques, count merging and Bayesian interpolation. We compare these techniques as well as linear interpolation in three scenar…

Language ModelingLanguage Modellingmodel

Fundamental limits and algorithms for sparse linear regression with sublinear sparsity

2021-01-27 · Lan V. Truong

We establish exact asymptotic expressions for the normalized mutual information and minimum mean-square-error (MMSE) of sparse linear regression in the sub-linear sparsity regime. Our result is achieved by a generalizati…

Bayesian Inferenceregression

Bayesian Reconstruction of Missing Observations

2014-04-23 · Shun Kataoka, Muneki Yasuda, Kazuyuki Tanaka

We focus on an interpolation method referred to Bayesian reconstruction in this paper. Whereas in standard interpolation methods missing data are interpolated deterministically, in Bayesian reconstruction, missing data a…

Faster Kernel Interpolation for Gaussian Processes

2021-01-28 · Mohit Yadav, Daniel Sheldon, Cameron Musco

A key challenge in scaling Gaussian Process (GP) regression to massive datasets is that exact inference requires computation with a dense n x n kernel matrix, where n is the number of data points. Significant work focuse…

Gaussian Processesregression

On permutation symmetries in Bayesian neural network posteriors: a variational perspective

2023-10-16 · NeurIPS 2023 11

The elusive nature of gradient-based optimization in neural networks is tied to their loss landscape geometry, which is poorly understood. However recent work has brought solid evidence that there is essentially no loss …

Combinatorial Optimization