paper-with-me

홈 › Papers

Statistical Physics of Deep Neural Networks: Initialization toward Optimal Channels

2022-12-04 · Kangyu Weng, Aohua Cheng, Ziyang Zhang, Pei Sun, Yang Tian

In deep learning, neural networks serve as noisy channels between input data and its representation. This perspective naturally relates deep learning with the pursuit of constructing channels with optimal performance in information transmission and representation. While considerable efforts are concentrated on realizing optimal channel properties during network optimization, we study a frequently overlooked possibility that neural networks can be initialized toward optimal channels. Our theory, consistent with experimental validation, identifies primary mechanics underlying this unknown possibility and suggests intrinsic connections between statistical physics and deep learning. Unlike the conventional theories that characterize neural networks applying the classic mean-filed approximation, we offer analytic proof that this extensively applied simplification scheme is not valid in studying neural networks as information channels. To fill this gap, we develop a corrected mean-field framework applicable for characterizing the limiting behaviors of information propagation in neural networks without strong assumptions on inputs. Based on it, we propose an analytic theory to prove that mutual information maximization is realized between inputs and propagated signals when neural networks are initialized at dynamic isometry, a case where information transmits via norm-preserving mappings. These theoretical predictions are validated by experiments on real neural networks, suggesting the robustness of our theory against finite-size effects. Finally, we analyze our findings with information bottleneck theory to confirm the precise relations among dynamic isometry, mutual information maximization, and optimal channel properties in deep learning.

📄 PDF Abstract BibTeX arXiv:2212.01744

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Learning

Similar Papers 제목 키워드 기반

Statistical Inference in Tensor Completion: Optimal Uncertainty Quantification and Statistical-to-Computational Gaps

2024-10-15 · Wanteng Ma, Dong Xia

This paper presents a simple yet efficient method for statistical inference of tensor linear forms using incomplete and noisy observations. Under the Tucker low-rank tensor model and the missing-at-random assumption, we …

Uncertainty Quantification

Entropic alternatives to initialization

2021-07-16 · Daniele Musso

Local entropic loss functions provide a versatile framework to define architecture-aware regularization procedures. Besides the possibility of being anisotropic in the synaptic space, the local entropic smoothening of th…

Estimation of Low-Rank Matrices via Approximate Message Passing

2017-11-06 · Andrea Montanari, Ramji Venkataramanan

Consider the problem of estimating a low-rank matrix when its entries are perturbed by Gaussian noise. If the empirical distribution of the entries of the spikes is known, optimal estimators that exploit this knowledge c…

Community Detection

Revisit CP Tensor Decomposition: Statistical Optimality and Fast Convergence

2025-05-29 · Runshi Tang, Julien Chhor, Olga Klopp, Anru R. Zhang

Canonical Polyadic (CP) tensor decomposition is a fundamental technique for analyzing high-dimensional tensor data. While the Alternating Least Squares (ALS) algorithm is widely used for computing CP decomposition due to…

Tensor Decomposition

Blind Over-the-Air Computation and Data Fusion via Provable Wirtinger Flow

2018-11-12 · Jialin Dong, Yuanming Shi, Zhi Ding

Over-the-air computation (AirComp) shows great promise to support fast data fusion in Internet-of-Things (IoT) networks. AirComp typically computes desired functions of distributed sensing data by exploiting superposed d…