paper-with-me

홈 › Papers

Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks

2026-04-10 · Jie Huang, Bruno Loureiro, Stefano Sarao Mannelli arxiv

We study the population loss landscape of two-layer ReLU networks of the form $\sum_{k=1}^K \mathrm{ReLU}(w_k^\top x)$ in a realisable teacher-student setting with Gaussian covariates. We show that local minima admit an exact low-dimensional representation in terms of summary statistics, yielding a sharp and interpretable characterisation of the landscape. We further establish a direct link with one-pass SGD: local minima correspond to attractive fixed points of the dynamics in summary statistics space. This perspective reveals a hierarchical organisation of minima into discrete families and shows how overparameterisation changes their stability and reachability under gradient-based dynamics. In this overparameterised regime, global minima become increasingly accessible, attracting the dynamics and reducing convergence to spurious solutions. Overall, our results reveal intrinsic limitations of common simplifying assumptions, which may miss essential features of the loss landscape even in minimal neural network models.

📄 PDF Abstract BibTeX arXiv:2604.09412

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Local Sharpness: Communication-Efficient Global Sharpness-aware Minimization for Federated Learning

2024-12-04 · CVPR 2025 1 · Debora Caldarola, Pietro Cagnasso, Barbara Caputo, Marco Ciccone

Federated learning (FL) enables collaborative model training with privacy preservation. Data heterogeneity across edge devices (clients) can cause models to converge to sharp minima, negatively impacting generalization a…

Federated Learning

Do Flat Minima Improve Sparse Novel View Synthesis?

2025-11-22 · Youngsik Yun, Dongjun Gu, Youngjung Uh arxiv

Despite the success of recent novel view synthesis methods, they tend to struggle in sparse-view settings. This poor generalization to unseen viewpoints is an inherent challenge when training with limited data. To addres…

Novel View Synthesis

Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware Minimization

2024-05-29 · Ziqing Fan, Shengchao Hu, Jiangchao Yao, Gang Niu 외

In federated learning (FL), the multi-step update and data heterogeneity among clients often lead to a loss landscape with sharper minima, degenerating the performance of the resulted global model. Prevalent federated ap…

Federated Learning

Sharp Minima Can Generalize: A Loss Landscape Perspective On Data

2025-11-06 · Raymond Fan, Bryce Sandlund, Lin Myat Ko arxiv

The volume hypothesis suggests deep learning is effective because it is likely to find flat minima due to their large volumes, and flat minima generalize well. This picture does not explain the role of large datasets in …

Global Dynamics of Heavy-Tailed SGDs in Nonconvex Loss Landscape: Characterization and Control

2025-10-23 · Xingyu Wang, Chang-Han Rhee arxiv

Stochastic gradient descent (SGD) and its variants enable modern artificial intelligence. However, theoretical understanding lags far behind their empirical success. It is widely believed that SGD has a curious ability t…