paper-with-me

홈 › Papers

A Deeper Look at the Hessian Eigenspectrum of Deep Neural Networks and its Applications to Regularization

2020-12-07 · Adepu Ravi Sankar, Yash Khasbage, Rahul Vigneswaran, Vineeth N Balasubramanian

Loss landscape analysis is extremely useful for a deeper understanding of the generalization ability of deep neural network models. In this work, we propose a layerwise loss landscape analysis where the loss surface at every layer is studied independently and also on how each correlates to the overall loss surface. We study the layerwise loss landscape by studying the eigenspectra of the Hessian at each layer. In particular, our results show that the layerwise Hessian geometry is largely similar to the entire Hessian. We also report an interesting phenomenon where the Hessian eigenspectrum of middle layers of the deep neural network are observed to most similar to the overall Hessian eigenspectrum. We also show that the maximum eigenvalue and the trace of the Hessian (both full network and layerwise) reduce as training of the network progresses. We leverage on these observations to propose a new regularizer based on the trace of the layerwise Hessian. Penalizing the trace of the Hessian at every layer indirectly forces Stochastic Gradient Descent to converge to flatter minima, which are shown to have better generalization performance. In particular, we show that such a layerwise regularizer can be leveraged to penalize the middlemost layers alone, which yields promising results. Our empirical studies on well-known deep nets across datasets support the claims of this work

📄 PDF Abstract BibTeX arXiv:2012.03801

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Wolkowicz-Styan Upper Bound on the Hessian Eigenspectrum for Cross-Entropy Loss in Nonlinear Smooth Neural Networks

2026-04-11 · Yuto Omae, Kazuki Sakai, Yohei Kakimoto, Makoto Sasaki 외 arxiv

Neural networks (NNs) are central to modern machine learning and achieve state-of-the-art results in many applications. However, the relationship between loss geometry and generalization is still not well understood. The…

Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks

2026-06-27 · Yuto Omae, Kazuki Sakai, Yohei Kakimoto, Makoto Sasaki 외 arxiv

The flatness hypothesis suggests that flatness of the loss landscape, as measured by the eigenvalues of the loss Hessian, correlates with better neural network generalization. While various algorithms reduce these eigenv…

A Teacher-Student Perspective on the Dynamics of Learning Near the Optimal Point

2025-12-17 · Carlos Couto, José Mourão, Mário A. T. Figueiredo, Pedro Ribeiro arxiv

Near an optimal learning point of a neural network, the learning performance of gradient descent dynamics is dictated by the Hessian matrix of the loss function with respect to the network parameters. We characterize the…

Hessian Eigenspectra of More Realistic Nonlinear Models

2021-03-02 · NeurIPS 2021 12 · Zhenyu Liao, Michael W. Mahoney

Given an optimization problem, the Hessian matrix and its eigenspectrum can be used in many ways, ranging from designing more efficient second-order algorithms to performing model analysis and regression diagnostics. Whe…

On Training Implicit Meta-Learning With Applications to Inductive Weighing in Consistency Regularization

2023-10-28 · Fady Rezk

Meta-learning that uses implicit gradient have provided an exciting alternative to standard techniques which depend on the trajectory of the inner loop training. Implicit meta-learning (IML), however, require computing $…

Meta-Learning