paper-with-me

홈 › Papers

Implicit Bias of Gradient Descent on Reparametrized Models: On Equivalence to Mirror Descent

2022-07-08 · Zhiyuan Li, Tianhao Wang, JasonD. Lee, Sanjeev Arora

As part of the effort to understand implicit bias of gradient descent in overparametrized models, several results have shown how the training trajectory on the overparametrized model can be understood as mirror descent on a different objective. The main result here is a characterization of this phenomenon under a notion termed commuting parametrization, which encompasses all the previous results in this setting. It is shown that gradient flow with any commuting parametrization is equivalent to continuous mirror descent with a related Legendre function. Conversely, continuous mirror descent with any Legendre function can be viewed as gradient flow with a related commuting parametrization. The latter result relies upon Nash's embedding theorem.

📄 PDF Abstract BibTeX arXiv:2207.04036

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Implicit Bias of Gradient Descent for Wide Two-layer Neural Networks Trained with the Logistic Loss

2020-02-11 · Lenaic Chizat, Francis Bach

Neural networks trained to minimize the logistic (a.k.a. cross-entropy) loss with gradient-based methods are observed to perform well in many supervised classification tasks. Towards understanding this phenomenon, we ana…

Generalization Bounds

Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal Data

2023-10-29 · NeurIPS 2023 11

The implicit bias towards solutions with favorable properties is believed to be a key reason why neural networks trained by gradient-based optimization can generalize well. While the implicit bias of gradient flow has be…

Implicit Regularization and Convergence for Weight Normalization

2019-11-18 · NeurIPS 2020 12 · Xiaoxia Wu, Edgar Dobriban, Tongzheng Ren, Shanshan Wu 외

Normalization methods such as batch [Ioffe and Szegedy, 2015], weight [Salimansand Kingma, 2016], instance [Ulyanov et al., 2016], and layer normalization [Baet al., 2016] have been widely used in modern machine learning…

General Loss Functions Lead to (Approximate) Interpolation in High Dimensions

2023-03-13 · Kuo-Wei Lai, Vidya Muthukumar

We provide a unified framework that applies to a general family of convex losses across binary and multiclass settings in the overparameterized regime to approximately characterize the implicit bias of gradient descent i…

Vocal Bursts Intensity Prediction

Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data

2022-10-13 · Spencer Frei, Gal Vardi, Peter L. Bartlett, Nathan Srebro 외

The implicit biases of gradient-based optimization algorithms are conjectured to be a major factor in the success of modern deep learning. In this work, we investigate the implicit bias of gradient flow and gradient desc…

Vocal Bursts Intensity Prediction