paper-with-me

Papers

Optimizing Information-theoretical Generalization Bounds via Anisotropic Noise in SGLD

2021-10-26 · NeurIPS 2021 12 · Bohan Wang, Huishuai Zhang, Jieyu Zhang, Qi Meng, Wei Chen, Tie-Yan Liu

Recently, the information-theoretical framework has been proven to be able to obtain non-vacuous generalization bounds for large models trained by Stochastic Gradient Langevin Dynamics (SGLD) with isotropic noise. In this paper, we optimize the information-theoretical generalization bound by manipulating the noise structure in SGLD. We prove that with constraint to guarantee low empirical risk, the optimal noise covariance is the square root of the expected gradient covariance if both the prior and the posterior are jointly optimized. This validates that the optimal noise is quite close to the empirical gradient covariance. Technically, we develop a new information-theoretical bound that enables such an optimization analysis. We then apply matrix analysis to derive the form of optimal noise covariance. Presented constraint and results are validated by the empirical observations.

📄 PDF Abstract BibTeX arXiv:2110.13750

Code (0)

등록된 구현이 없습니다.

Tasks

Generalization Bounds

Similar Papers 제목 키워드 기반

Optimizing Information-theoretical Generalization Bound via Anisotropic Noise of SGLD

2021-05-21 · NeurIPS 2021 12 · Bohan Wang, Huishuai Zhang, Jieyu Zhang, Qi Meng 외

Recently, the information-theoretical framework has been proven to be able to obtain non-vacuous generalization bounds for large models trained by Stochastic Gradient Langevin Dynamics (SGLD) with isotropic noise. In th…

Generalization Bounds

Shrinkage to Infinity: Reducing Test Error by Inflating the Minimum Norm Interpolator in Linear Models

2025-10-22 · Jake Freeman arxiv

Hastie et al. (2022) found that ridge regularization is essential in high dimensional linear regression $y=β^Tx + ε$ with isotropic co-variates $x\in \mathbb{R}^d$ and $n$ samples at fixed $d/n$. However, Hastie et al. (…

Improving Generalization of Deep Neural Networks by Leveraging Margin Distribution

2018-12-27 · ICLR 2019 5 · Shen-Huan Lyu, Lu Wang, Zhi-Hua Zhou

Recent research has used margin theory to analyze the generalization performance for deep neural networks (DNNs). The existed results are almost based on the spectrally-normalized minimum margin. However, optimizing the …

Representation Learning

Shedding a PAC-Bayesian Light on Adaptive Sliced-Wasserstein Distances

2022-06-07 · Ruben Ohana, Kimia Nadjahi, Alain Rakotomamonjy, Liva Ralaivola

The Sliced-Wasserstein distance (SW) is a computationally efficient and theoretically grounded alternative to the Wasserstein distance. Yet, the literature on its statistical properties -- or, more accurately, its genera…

Generalization Bounds

Information-Theoretic Generalization Bounds for Stochastic Gradient Descent

2021-02-01 · Gergely Neu, Gintare Karolina Dziugaite, Mahdi Haghifam, Daniel M. Roy

We study the generalization properties of the popular stochastic optimization method known as stochastic gradient descent (SGD) for optimizing general non-convex loss functions. Our main contribution is providing upper b…

Generalization BoundsStochastic Optimization