paper-with-me

홈 › Papers

Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures

2020-10-27 · Carlo Baldassi, Enrico M. Malatesta, Matteo Negri, Riccardo Zecchina

We analyze the connection between minimizers with good generalizing properties and high local entropy regions of a threshold-linear classifier in Gaussian mixtures with the mean squared error loss function. We show that there exist configurations that achieve the Bayes-optimal generalization error, even in the case of unbalanced clusters. We explore analytically the error-counting loss landscape in the vicinity of a Bayes-optimal solution, and show that the closer we get to such configurations, the higher the local entropy, implying that the Bayes-optimal solution lays inside a wide flat region. We also consider the algorithmically relevant case of targeting wide flat minima of the (differentiable) mean squared error loss. Our analytical and numerical results show not only that in the balanced case the dependence on the norm of the weights is mild, but also, in the unbalanced case, that the performances can be improved.

📄 PDF Abstract BibTeX arXiv:2010.14761

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Entropic gradient descent algorithms and wide flat minima

2020-06-14 · ICLR 2021 1 · Fabrizio Pittorino, Carlo Lucibello, Christoph Feinauer, Gabriele Perugini 외

The properties of flat minima in the empirical risk landscape of neural networks have been debated for some time. Increasing evidence suggests they possess better generalization capabilities with respect to sharp ones. F…

Unveiling the structure of wide flat minima in neural networks

2021-07-02 · Carlo Baldassi, Clarissa Lauditi, Enrico M. Malatesta, Gabriele Perugini 외

The success of deep learning has revealed the application potential of neural networks across the sciences and opened up fundamental theoretical problems. In particular, the fact that learning algorithms based on simple …

Neighborhood Region Smoothing Regularization for Finding Flat Minima In Deep Neural Networks

2022-01-16 · Yang Zhao, Hao Zhang

Due to diverse architectures in deep neural networks (DNNs) with severe overparameterization, regularization techniques are critical for finding optimal solutions in the huge hypothesis space. In this paper, we propose a…

image-classificationImage Classification

Flat Minima and Generalization: Insights from Stochastic Convex Optimization

2025-11-05 · Matan Schliserman, Shira Vansover-Hager, Tomer Koren arxiv

Understanding the generalization behavior of learning algorithms is a central goal of learning theory. A recently emerging explanation is that learning algorithms are successful in practice because they converge to flat …

Do Flat Minima Improve Sparse Novel View Synthesis?

2025-11-22 · Youngsik Yun, Dongjun Gu, Youngjung Uh arxiv

Despite the success of recent novel view synthesis methods, they tend to struggle in sparse-view settings. This poor generalization to unseen viewpoints is an inherent challenge when training with limited data. To addres…

Novel View Synthesis