paper-with-me

홈 › Papers

On non-approximability of zero loss global ${\mathcal L}^2$ minimizers by gradient descent in Deep Learning

2023-11-13 · Thomas Chen, Patricia Muñoz Ewald

We analyze geometric aspects of the gradient descent algorithm in Deep Learning (DL), and give a detailed discussion of the circumstance that in underparametrized DL networks, zero loss minimization can generically not be attained. As a consequence, we conclude that the distribution of training inputs must necessarily be non-generic in order to produce zero loss minimizers, both for the method constructed in [Chen-Munoz Ewald 2023, 2024], or for gradient descent [Chen 2025] (which assume clustering of training data).

📄 PDF Abstract BibTeX arXiv:2311.07065

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

Zero loss guarantees and explicit minimizers for generic overparametrized Deep Learning networks

2025-02-19 · Thomas Chen, Andrew G. Moore

We determine sufficient conditions for overparametrized deep learning (DL) networks to guarantee the attainability of zero loss in the context of supervised learning, for the $\mathcal{L}^2$ cost and {\em generic} traini…

Geometric structure of Deep Learning networks and construction of global ${\mathcal L}^2$ minimizers

2023-09-19 · Thomas Chen, Patricia Muñoz Ewald

In this paper, we explicitly determine local and global minimizers of the $\mathcal{L}^2$ cost function in underparametrized Deep Learning (DL) networks; our main goal is to shed light on their geometric structure and pr…

On the Loss Landscape Geometry of Regularized Deep Matrix Factorization: Uniqueness and Sharpness

2026-03-28 · Anil Kamber, Rahul Parhi arxiv

Weight decay is ubiquitous in training deep neural network architectures. Its empirical success is often attributed to capacity control; nonetheless, our theoretical understanding of its effect on the loss landscape and …

Characterization of Gradient Dominance and Regularity Conditions for Neural Networks

2017-10-18 · Yi Zhou, Yingbin Liang

The past decade has witnessed a successful application of deep learning to solving many challenging problems in machine learning and artificial intelligence. However, the loss functions of deep neural networks (especiall…

Interpretable global minima of deep ReLU neural networks on sequentially separable data

2024-05-11 · Thomas Chen, Patrícia Muñoz Ewald

We explicitly construct zero loss neural network classifiers. We write the weight matrices and bias vectors in terms of cumulative parameters, which determine truncation maps acting recursively on input space. The config…