paper-with-me

홈 › Papers

A Comparative Analysis of the Optimization and Generalization Property of Two-layer Neural Network and Random Feature Models Under Gradient Descent Dynamics

2019-04-08 · Weinan E, Chao Ma, Lei Wu

A fairly comprehensive analysis is presented for the gradient descent dynamics for training two-layer neural network models in the situation when the parameters in both layers are updated. General initialization schemes as well as general regimes for the network width and training data size are considered. In the over-parametrized regime, it is shown that gradient descent dynamics can achieve zero training loss exponentially fast regardless of the quality of the labels. In addition, it is proved that throughout the training process the functions represented by the neural network model are uniformly close to that of a kernel method. For general values of the network width and training data size, sharp estimates of the generalization error is established for target functions in the appropriate reproducing kernel Hilbert space.

📄 PDF Abstract BibTeX arXiv:1904.04326

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Limitations of Neural Collapse for Understanding Generalization in Deep Learning

2022-02-17 · Like Hui, Mikhail Belkin, Preetum Nakkiran

The recent work of Papyan, Han, & Donoho (2020) presented an intriguing "Neural Collapse" phenomenon, showing a structural property of interpolating classifiers in the late stage of training. This opened a rich area of e…

Deep LearningRepresentation Learning

LayerDAG: A Layerwise Autoregressive Diffusion Model for Directed Acyclic Graph Generation

2024-11-04 · Mufei Li, Viraj Shitole, Eli Chien, Changhai Man 외

Directed acyclic graphs (DAGs) serve as crucial data representations in domains such as hardware synthesis and compiler/program optimization for computing systems. DAG generative models facilitate the creation of synthet…

BenchmarkingGraph Generation

From Optimization Dynamics to Generalization Bounds via Łojasiewicz Gradient Inequality

2022-02-22 · Fusheng Liu, Haizhao Yang, Soufiane Hayou, Qianxiao Li

Optimization and generalization are two essential aspects of statistical machine learning. In this paper, we propose a framework to connect optimization with generalization by analyzing the generalization error based on …

BIG-bench Machine LearningGeneralization Boundsregression

Progressive Learning for Systematic Design of Large Neural Networks

2017-10-23 · Saikat Chatterjee, Alireza M. Javid, Mostafa Sadeghi, Partha P. Mitra 외

We develop an algorithm for systematic design of a large artificial neural network using a progression property. We find that some non-linear functions, such as the rectifier linear unit and its derivatives, hold the pro…

Comparative layer-wise analysis of self-supervised speech models

2022-11-08 · Ankita Pasad, Bowen Shi, Karen Livescu

Many self-supervised speech models, varying in their pre-training objective, input modality, and pre-training data, have been proposed in the last few years. Despite impressive successes on downstream tasks, we still hav…

speech-recognitionSpeech RecognitionSpoken Language Understanding