paper-with-me

Papers

Generalization in Deep Learning

2017-10-16 · Kenji Kawaguchi, Leslie Pack Kaelbling, Yoshua Bengio

This paper provides theoretical insights into why and how deep learning can generalize well, despite its large capacity, complexity, possible algorithmic instability, nonrobustness, and sharp minima, responding to an open question in the literature. We also discuss approaches to provide non-vacuous generalization guarantees for deep learning. Based on theoretical observations, we propose new open problems and discuss the limitations of our results.

📄 PDF Abstract BibTeX arXiv:1710.05468

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningOpen-Ended Question Answering

Similar Papers 제목 키워드 기반

Uniform Generalization, Concentration, and Adaptive Learning

2016-08-22 · Ibrahim Alabdulmohsin

One fundamental goal in any learning algorithm is to mitigate its risk for overfitting. Mathematically, this requires that the learning algorithm enjoys a small generalization risk, which is defined either in expectation…

Learning Theory

Low-Dimension-to-High-Dimension Generalization And Its Implications for Length Generalization

2024-10-11 · Yang Chen, Yitao Liang, Zhouchen Lin

Low-Dimension-to-High-Dimension (LDHD) generalization is a special case of Out-of-Distribution (OOD) generalization, where the training data are restricted to a low-dimensional subspace of the high-dimensional testing sp…

Inductive BiasPosition

Understanding Generalization via Set Theory

2023-11-11 · Shiqi Liu

Generalization is at the core of machine learning models. However, the definition of generalization is not entirely clear. We employ set theory to introduce the concepts of algorithms, hypotheses, and dataset generalizat…

Learning Trajectories are Generalization Indicators

2023-04-25 · NeurIPS 2023 11

This paper explores the connection between learning trajectories of Deep Neural Networks (DNNs) and their generalization capabilities when optimized using (stochastic) gradient descent algorithms. Instead of concentratin…

Diversity

Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

2022-01-06 · Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin 외

In this paper we propose to study generalization of neural networks on small algorithmically generated datasets. In this setting, questions about data efficiency, memorization, generalization, and speed of learning can b…

Memorization