paper-with-me

홈 › Papers

Zero loss guarantees and explicit minimizers for generic overparametrized Deep Learning networks

2025-02-19 · Thomas Chen, Andrew G. Moore

We determine sufficient conditions for overparametrized deep learning (DL) networks to guarantee the attainability of zero loss in the context of supervised learning, for the $\mathcal{L}^2$ cost and {\em generic} training data. We present an explicit construction of the zero loss minimizers without invoking gradient descent. On the other hand, we point out that increase of depth can deteriorate the efficiency of cost minimization using a gradient descent algorithm by analyzing the conditions for rank loss of the training Jacobian. Our results clarify key aspects on the dichotomy between zero loss reachability in underparametrized versus overparametrized DL.

📄 PDF Abstract BibTeX arXiv:2502.14114

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On non-approximability of zero loss global ${\mathcal L}^2$ minimizers by gradient descent in Deep Learning

2023-11-13 · Thomas Chen, Patricia Muñoz Ewald

We analyze geometric aspects of the gradient descent algorithm in Deep Learning (DL), and give a detailed discussion of the circumstance that in underparametrized DL networks, zero loss minimization can generically not b…

Clustering

$\mathscr{H}$-Consistency Estimation Error of Surrogate Loss Minimizers

2022-05-16 · Pranjal Awasthi, Anqi Mao, Mehryar Mohri, Yutao Zhong

We present a detailed study of estimation errors in terms of surrogate loss estimation errors. We refer to such guarantees as $\mathscr{H}$-consistency estimation error bounds, since they account for the hypothesis set $…

Proximal methods avoid active strict saddles of weakly convex functions

2019-12-16 · Damek Davis, Dmitriy Drusvyatskiy

We introduce a geometrically transparent strict saddle property for nonsmooth functions. This property guarantees that simple proximal algorithms on weakly convex problems converge only to local minimizers, when randomly…

On the Loss Landscape Geometry of Regularized Deep Matrix Factorization: Uniqueness and Sharpness

2026-03-28 · Anil Kamber, Rahul Parhi arxiv

Weight decay is ubiquitous in training deep neural network architectures. Its empirical success is often attributed to capacity control; nonetheless, our theoretical understanding of its effect on the loss landscape and …

Interpretable global minima of deep ReLU neural networks on sequentially separable data

2024-05-11 · Thomas Chen, Patrícia Muñoz Ewald

We explicitly construct zero loss neural network classifiers. We write the weight matrices and bias vectors in terms of cumulative parameters, which determine truncation maps acting recursively on input space. The config…