paper-with-me

홈 › Papers

Implicit Bias of Gradient Descent on Linear Convolutional Networks

2018-06-01 · NeurIPS 2018 12 · Suriya Gunasekar, Jason Lee, Daniel Soudry, Nathan Srebro

We show that gradient descent on full-width linear convolutional networks of depth $L$ converges to a linear predictor related to the $\ell_{2/L}$ bridge penalty in the frequency domain. This is in contrast to linearly fully connected networks, where gradient descent converges to the hard margin linear support vector machine solution, regardless of depth.

📄 PDF Abstract BibTeX arXiv:1806.00468

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Implicit Bias of Batch Normalization in Linear Models and Two-layer Linear Convolutional Neural Networks

2023-06-20 · Yuan Cao, Difan Zou, Yuanzhi Li, Quanquan Gu

We study the implicit bias of batch normalization trained by gradient descent. We show that when learning a linear model with batch normalization for binary classification, gradient descent converges to a uniform margin …

Binary Classification

Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data

2022-10-13 · Spencer Frei, Gal Vardi, Peter L. Bartlett, Nathan Srebro 외

The implicit biases of gradient-based optimization algorithms are conjectured to be a major factor in the success of modern deep learning. In this work, we investigate the implicit bias of gradient flow and gradient desc…

Vocal Bursts Intensity Prediction

Implicit Bias of (Stochastic) Gradient Descent for Rank-1 Linear Neural Network

2023-09-21 · NeurIPS 2023 11

Studying the implicit bias of gradient descent (GD) and stochastic gradient descent (SGD) is critical to unveil the underlying mechanism of deep learning. Unfortunately, even for standard linear networks in regression se…

Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures

2025-12-23 · Yedi Zhang, Andrew Saxe, Peter E. Latham arxiv

Neural networks trained with gradient descent often learn solutions of increasing complexity over time, a phenomenon known as simplicity bias. Despite being widely observed across architectures, existing theoretical trea…

The Implicit Bias of Gradient Descent on Generalized Gated Linear Networks

2022-02-05 · Samuel Lippl, L. F. Abbott, SueYeon Chung

Understanding the asymptotic behavior of gradient-descent training of deep neural networks is essential for revealing inductive biases and improving network performance. We derive the infinite-time training limit of a ma…

Inductive Bias