paper-with-me

홈 › Papers

Gradient Descent for Deep Matrix Factorization: Dynamics and Implicit Bias towards Low Rank

2020-11-27 · Hung-Hsu Chou, Carsten Gieshoff, Johannes Maly, Holger Rauhut

In deep learning, it is common to use more network parameters than training points. In such scenarioof over-parameterization, there are usually multiple networks that achieve zero training error so that thetraining algorithm induces an implicit bias on the computed solution. In practice, (stochastic) gradientdescent tends to prefer solutions which generalize well, which provides a possible explanation of thesuccess of deep learning. In this paper we analyze the dynamics of gradient descent in the simplifiedsetting of linear networks and of an estimation problem. Although we are not in an overparameterizedscenario, our analysis nevertheless provides insights into the phenomenon of implicit bias. In fact, wederive a rigorous analysis of the dynamics of vanilla gradient descent, and characterize the dynamicalconvergence of the spectrum. We are able to accurately locate time intervals where the effective rankof the iterates is close to the effective rank of a low-rank projection of the ground-truth matrix. Inpractice, those intervals can be used as criteria for early stopping if a certain regularity is desired. Wealso provide empirical evidence for implicit bias in more general scenarios, such as matrix sensing andrandom initialization. This suggests that deep learning prefers trajectories whose complexity (measuredin terms of effective rank) is monotonically increasing, which we believe is a fundamental concept for thetheoretical understanding of deep learning.

📄 PDF Abstract BibTeX arXiv:2011.13772

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningDenoising

Methods 이 논문이 사용한 방법론

Early Stopping Early Stopping is a regularization technique for deep neural networks that stops training when parameter updates no longer begin to yield improves on a validation set. In…

Similar Papers 제목 키워드 기반

Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent

2024-10-21 · Santhosh Karnik, Anna Veselovska, Mark Iwen, Felix Krahmer

We provide a rigorous analysis of implicit regularization in an overparametrized tensor factorization problem beyond the lazy training regime. For matrix factorization problems, this phenomenon has been studied in a numb…

Implicit Regularization in Matrix Factorization

2017-05-25 · NeurIPS 2017 12 · Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur 외

We study implicit regularization when optimizing an underdetermined quadratic objective over a matrix $X$ with gradient descent on a factorization of $X$. We conjecture and provide empirical and theoretical evidence that…

Global Convergence of Four-Layer Matrix Factorization under Random Initialization

2025-11-13 · Minrui Luo, Weihang Xu, Xiang Gao, Maryam Fazel 외 arxiv

Gradient descent dynamics on the deep matrix factorization problem is extensively studied as a simplified theoretical model for deep neural networks. Although the convergence theory for two-layer matrix factorization is …

Implicit Bias in Matrix Factorization and its Explicit Realization in a New Architecture

2025-01-27 · Yikun Hou, Suvrit Sra, Alp Yurtsever

Gradient descent for matrix factorization is known to exhibit an implicit bias toward approximately low-rank solutions. While existing theories often assume the boundedness of iterates, empirically the bias persists even…

Implicit Regularization in Deep Tensor Factorization

2021-05-04 · Paolo Milanesi, Hachem Kadri, Stéphane Ayache, Thierry Artières

Attempts of studying implicit regularization associated to gradient descent (GD) have identified matrix completion as a suitable test-bed. Late findings suggest that this phenomenon cannot be phrased as a minimization-no…

Matrix Completion