paper-with-me

홈 › Papers

Generalization error bounds for two-layer neural networks with Lipschitz loss function

2026-04-07 · Jiang Yu Nguwi, Nicolas Privault arxiv

We derive generalization error bounds for the training of two-layer neural networks without assuming boundedness of the loss function, using Wasserstein distance estimates on the discrepancy between a probability distribution and its associated empirical measure, together with moment bounds for the associated stochastic gradient method. In the case of independent test data, we obtain a dimension-free rate of order $O(n^{-1/2} )$ on the $n$-sample generalization error, whereas without independence assumption, we derive a bound of order $O(n^{-1 / ( d_{\rm in}+d_{\rm out} )} )$, where $d_{\rm in}$, $d_{\rm out}$ denote input and output dimensions. Our bounds and their coefficients can be explicitly computed prior to the training of the model, and are confirmed by numerical simulations.

📄 PDF Abstract BibTeX arXiv:2604.06281

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Near Complete Nonasymptotic Generalization Theory For Multilayer Neural Networks: Beyond the Bias-Variance Tradeoff

2025-03-03 · Hao Yu, Xiangyang Ji

We propose a first near complete (that will make explicit sense in the main text) nonasymptotic generalization theory for multilayer neural networks with arbitrary Lipschitz activations and general Lipschitz loss functio…

On Lipschitz Continuity and Smoothness of Loss Functions in Learning to Rank

2014-05-03 · Ambuj Tewari, Sougata Chaudhuri

In binary classification and regression problems, it is well understood that Lipschitz continuity and smoothness of the loss function play key roles in governing generalization error bounds for empirical risk minimizatio…

Binary ClassificationLearning-To-Rank

Beyond Lipschitz: Sharp Generalization and Excess Risk Bounds for Full-Batch GD

2022-04-26 · Konstantinos E. Nikolakakis, Farzin Haddadpour, Amin Karbasi, Dionysios S. Kalogerias

We provide sharp path-dependent generalization and excess risk guarantees for the full-batch Gradient Descent (GD) algorithm on smooth losses (possibly non-Lipschitz, possibly nonconvex). At the heart of our analysis is …

A Hierarchical Sampling Framework for bounding the Generalization Error of Federated Learning

2026-05-05 · Dario Filatrella, Ragnar Thobaben, Mikael Skoglund arxiv

We study expected generalization bounds for the Hierarchical Federated Learning (HFL) setup using Wasserstein distance. We introduce a generalized framework in which data is sampled hierarchically, and we model it with a…

Federated Learning

Generalization bounds for deep convolutional neural networks

2019-05-29 · ICLR 2020 1 · Philip M. Long, Hanie Sedghi

We prove bounds on the generalization error of convolutional networks. The bounds are in terms of the training loss, the number of parameters, the Lipschitz constant of the loss and the distance from the weights to the i…

Generalization Bounds