paper-with-me

홈 › Papers

From Approximation Rates to Loss-Landscape Barrier Decay in Shallow ReLU Networks

2026-02-19 · Saveliy Baturin arxiv

We study pathwise connectivity of sublevel sets for one-hidden-layer ReLU networks with constrained first-layer weights and an $\ell_1$ penalty on the output layer. The data term is assumed convex and globally Lipschitz in the scalar logit. We first give a finite-width construction that connects any two points of a common sublevel through a path controlled by a loss-consistent compression functional and a first-order perturbation term. The proof replaces the quadratic perturbation estimate in the Freeman--Bruna mechanism by a direct Lipschitz bound. Positive homogeneity is then used in a direction that is compatible with the penalty: every active atom is moved monotonically from the unit ball to the unit sphere while its output coefficient is reduced. Sphere covering and cluster merging consequently give $O(m^{-1/(n-1)})$ fixed-level thickening for $n\ge2$, while the one-dimensional two-ray dictionary gives exact connectivity for every $m\ge4$. We also prove internally that the regularized approximation values satisfy $e(l)-e_\infty=O(l^{-1/2})$. More generally, a rate $O(l^{-s})$ transfers to a near-optimal barrier rate $O(m^{-s/((n-1)s+1)})$; under the standing assumptions, this yields the explicit rate $O(m^{-1/(n+1)})$. A theorem-aligned finite-distribution experiment complements the analysis. The primary Huber run yields a maximal best certified upper gap $1.66\times10^{-5}$ over 720 recorded pairs at widths $m\ge16$; a matched binary-cross-entropy rerun and a 720-endpoint dense-representation stress test probe loss robustness and the active cluster-merging mechanism.

📄 PDF Abstract BibTeX arXiv:2602.17596

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

2024-10-07 · Kaiyue Wen, Zhiyuan Li, Jason Wang, David Hall 외

Training language models currently requires pre-determining a fixed compute budget because the typical cosine learning rate schedule depends on the total number of steps. In contrast, the Warmup-Stable-Decay (WSD) schedu…

On permutation symmetries in Bayesian neural network posteriors: a variational perspective

2023-10-16 · NeurIPS 2023 11

The elusive nature of gradient-based optimization in neural networks is tied to their loss landscape geometry, which is poorly understood. However recent work has brought solid evidence that there is essentially no loss …

Combinatorial Optimization

Beyond the Quadratic Approximation: the Multiscale Structure of Neural Network Loss Landscapes

2022-04-24 · Chao Ma, Daniel Kunin, Lei Wu, Lexing Ying

A quadratic approximation of neural network loss landscapes has been extensively used to study the optimization process of these networks. Though, it usually holds in a very small neighborhood of the minimum, it cannot e…

Understanding the Generalization Benefits of Late Learning Rate Decay

2024-01-21 · Yinuo Ren, Chao Ma, Lexing Ying

Why do neural networks trained with large learning rates for a longer time often lead to better generalization? In this paper, we delve into this question by examining the relation between training and testing loss in ne…

Mpemba Effect in Large-Language Model Training Dynamics: A Minimal Analysis of the Valley-River model

2025-07-06 · Sibei Liu, Zhijian Hu arxiv

Learning rate (LR) schedules in large language model (LLM) training often follow empirical templates: warm-up, constant plateau/stable phase, and decay (WSD). However, the mechanistic explanation for this strategy remain…