paper-with-me

홈 › Papers

Characterization of Gradient Dominance and Regularity Conditions for Neural Networks

2017-10-18 · Yi Zhou, Yingbin Liang

The past decade has witnessed a successful application of deep learning to solving many challenging problems in machine learning and artificial intelligence. However, the loss functions of deep neural networks (especially nonlinear networks) are still far from being well understood from a theoretical aspect. In this paper, we enrich the current understanding of the landscape of the square loss functions for three types of neural networks. Specifically, when the parameter matrices are square, we provide an explicit characterization of the global minimizers for linear networks, linear residual networks, and nonlinear networks with one hidden layer. Then, we establish two quadratic types of landscape properties for the square loss of these neural networks, i.e., the gradient dominance condition within the neighborhood of their full rank global minimizers, and the regularity condition along certain directions and within the neighborhood of their global minimizers. These two landscape properties are desirable for the optimization around the global minimizers of the loss function for these neural networks.

📄 PDF Abstract BibTeX arXiv:1710.06910

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Luce Model, Regularity, and Choice Overload

2025-02-28 · Daniele Caliari, Henrik Petri

We characterize regularity (Block & Marschak, 1960) within a novel stochastic model: the General Threshold Luce model [GTLM]. We apply our results to study choice overload, identified by regularity violations that impose…

model

Capacity dependent analysis for functional online learning algorithms

2022-09-25 · Xin Guo, Zheng-Chu Guo, Lei Shi

This article provides convergence analysis of online stochastic gradient descent algorithms for functional linear models. Adopting the characterizations of the slope function regularity, the kernel space capacity, and th…

Prediction

A new characterization of second-order stochastic dominance

2024-02-20 · Yuanying Guan, Muqiao Huang, Ruodu Wang

We provide a new characterization of second-order stochastic dominance, also known as increasing concave order. The result has an intuitive interpretation that adding a risk with negative expected value in adverse scenar…

ManagementPosition

Characterization of Gaussian Universality Breakdown in High-Dimensional Empirical Risk Minimization

2026-04-03 · Chiheb Yaakoubi, Cosme Louart, Malik Tiomoko, Zhenyu Liao arxiv

We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs. By heuristically extending the Convex Gaussian Min-Max Theorem (CGMT) to non-Gaussian settings, we derive an asy…

Regularity of Solutions to Beckmann's Parametric Optimal Transport

2026-03-20 · Hanno Gottschalk, Tobias J. Riedlinger arxiv

Beckmann's problem in optimal transport minimizes the total squared flux in a continuous transport problem from a source to a target distribution. In this article, the regularity theory for solutions to Beckmann's proble…