paper-with-me

홈 › Papers

Geometric structure of shallow neural networks and constructive ${\mathcal L}^2$ cost minimization

2023-09-19 · Thomas Chen, Patricia Muñoz Ewald

In this paper, we approach the problem of cost (loss) minimization in underparametrized shallow neural networks through the explicit construction of upper bounds, without any use of gradient descent. A key focus is on elucidating the geometric structure of approximate and precise minimizers. We consider shallow neural networks with one hidden layer, a ReLU activation function, an ${\mathcal L}^2$ Schatten class (or Hilbert-Schmidt) cost function, input space ${\mathbb R}^M$, output space ${\mathbb R}^Q$ with $Q\leq M$, and training input sample size $N>QM$ that can be arbitrarily large. We prove an upper bound on the minimum of the cost function of order $O(\delta_P)$ where $\delta_P$ measures the signal to noise ratio of training inputs. In the special case $M=Q$, we explicitly determine an exact degenerate local minimum of the cost function, and show that the sharp value differs from the upper bound obtained for $Q\leq M$ by a relative error $O(\delta_P^2)$. The proof of the upper bound yields a constructively trained network; we show that it metrizes a particular $Q$-dimensional subspace in the input space ${\mathbb R}^M$. We comment on the characterization of the global minimum of the cost function in the given context.

📄 PDF Abstract BibTeX arXiv:2309.10370

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

A General Constructive Upper Bound on Shallow Neural Nets Complexity

2025-10-07 · Frantisek Hakl, Vit Fojtik arxiv

We provide an upper bound on the number of neurons required in a shallow neural network to approximate a continuous function on a compact set with a given accuracy. This method, inspired by a specific proof of the Stone-…

Geometric Measurements of the Axiom of Choice in Neural Proof Embeddings

2026-06-26 · Rodrigo Mendoza-Smith arxiv

The axiom of choice has divided the foundations of mathematics for over a century, but the distinction between classical and constructive proofs has remained a philosophical and methodological one. We use Lean 4's kernel…

On the geometric and Riemannian structure of the spaces of group equivariant non-expansive operators

2021-03-03 · Pasquale Cascarano, Patrizio Frosini, Nicola Quercioli, Amir Saki

Group equivariant non-expansive operators have been recently proposed as basic components in topological data analysis and deep learning. In this paper we study some geometric properties of the spaces of group equivarian…

Topological Data Analysis

Geometric structure of Deep Learning networks and construction of global ${\mathcal L}^2$ minimizers

2023-09-19 · Thomas Chen, Patricia Muñoz Ewald

In this paper, we explicitly determine local and global minimizers of the $\mathcal{L}^2$ cost function in underparametrized Deep Learning (DL) networks; our main goal is to shed light on their geometric structure and pr…

Uncovering the Latent Potential of Deep Intermediate Representations

2026-05-21 · Arnesh Batra, Arush Gumber, Aniket Khandelwal, Jashn Khemani 외 arxiv

Foundational Models pretrained on huge amount of data learn representations that evolve across depth, forming a hierarchy of embeddings with distinct semantic content and geometric structure. Contrary to the widespread p…