paper-with-me

홈 › Papers

Hessian Eigenspectra of More Realistic Nonlinear Models

2021-03-02 · NeurIPS 2021 12 · Zhenyu Liao, Michael W. Mahoney

Given an optimization problem, the Hessian matrix and its eigenspectrum can be used in many ways, ranging from designing more efficient second-order algorithms to performing model analysis and regression diagnostics. When nonlinear models and non-convex problems are considered, strong simplifying assumptions are often made to make Hessian spectral analysis more tractable. This leads to the question of how relevant the conclusions of such analyses are for more realistic nonlinear models. In this paper, we exploit deterministic equivalent techniques from random matrix theory to make a \emph{precise} characterization of the Hessian eigenspectra for a broad family of nonlinear models, including models that generalize the classical generalized linear models, without relying on strong simplifying assumptions used previously. We show that, depending on the data properties, the nonlinear response model, and the loss function, the Hessian can have \emph{qualitatively} different spectral behaviors: of bounded or unbounded support, with single- or multi-bulk, and with isolated eigenvalues on the left- or right-hand side of the bulk. By focusing on such a simple but nontrivial nonlinear model, our analysis takes a step forward to unveil the theoretical origin of many visually striking features observed in more complex machine learning models.

📄 PDF Abstract BibTeX arXiv:2103.01519

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Deeper Look at the Hessian Eigenspectrum of Deep Neural Networks and its Applications to Regularization

2020-12-07 · Adepu Ravi Sankar, Yash Khasbage, Rahul Vigneswaran, Vineeth N Balasubramanian

Loss landscape analysis is extremely useful for a deeper understanding of the generalization ability of deep neural network models. In this work, we propose a layerwise loss landscape analysis where the loss surface at e…

Eigenvalue spectral properties of sparse random matrices obeying Dale's law

2022-12-03 · Isabelle D Harris, Hamish Meffin, Anthony N Burkitt, Andre D. H Peterson

This paper examines the relationship between sparse random network architectures and neural network stability by examining the eigenvalue spectral distribution. Specifically, we generalise classical eigenspectral results…

Does the Data Induce Capacity Control in Deep Learning?

2021-10-27 · Rubing Yang, Jialin Mao, Pratik Chaudhari

We show that the input correlation matrix of typical classification datasets has an eigenspectrum where, after a sharp initial drop, a large number of small eigenvalues are distributed uniformly over an exponentially lar…

Deep LearningGeneralization Bounds

Vanishing Curvature and the Power of Adaptive Methods in Randomly Initialized Deep Networks

2021-06-07 · Antonio Orvieto, Jonas Kohler, Dario Pavllo, Thomas Hofmann 외

This paper revisits the so-called vanishing gradient phenomenon, which commonly occurs in deep randomly initialized neural networks. Leveraging an in-depth analysis of neural chains, we first show that vanishing gradient…

Preserving Deep Representations In One-Shot Pruning: A Hessian-Free Second-Order Optimization Framework

2024-11-27 · Ryan Lucas, Rahul Mazumder

We present SNOWS, a one-shot post-training pruning framework aimed at reducing the cost of vision network inference without retraining. Current leading one-shot pruning methods minimize layer-wise least squares reconstru…