paper-with-me

Papers

Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers

2019-10-12 · Bubacarr Bah, Holger Rauhut, Ulrich Terstiege, Michael Westdickenberg

We study the convergence of gradient flows related to learning deep linear neural networks (where the activation function is the identity map) from data. In this case, the composition of the network layers amounts to simply multiplying the weight matrices of all layers together, resulting in an overparameterized problem. The gradient flow with respect to these factors can be re-interpreted as a Riemannian gradient flow on the manifold of rank-$r$ matrices endowed with a suitable Riemannian metric. We show that the flow always converges to a critical point of the underlying functional. Moreover, we establish that, for almost all initializations, the flow converges to a global minimum on the manifold of rank $k$ matrices for some $k\leq r$.

📄 PDF Abstract BibTeX arXiv:1910.05505

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fast Global Convergence for Low-rank Matrix Recovery via Riemannian Gradient Descent with Random Initialization

2020-12-31 · Thomas Y. Hou, Zhenzhen Li, Ziyun Zhang

In this paper, we propose a new global analysis framework for a class of low-rank matrix recovery problems on the Riemannian manifold. We analyze the global behavior for the Riemannian optimization with random initializa…

RetrievalRiemannian optimization

Continuous-time Riemannian SGD and SVRG Flows on Wasserstein Probabilistic Space

2024-01-24 · Mingyang Yi, Bohan Wang

Recently, optimization on the Riemannian manifold has provided new insights to the optimization community. In this regard, the manifold taken as the probability measure metric space equipped with the second-order Wassers…

Stochastic Optimization

Riemannian Federated Learning via Averaging Gradient Stream

2024-09-11 · Zhenwei Huang, Wen Huang, Pratik Jawanpuria, Bamdev Mishra

In recent years, federated learning has garnered significant attention as an efficient and privacy-preserving distributed learning paradigm. In the Euclidean setting, Federated Averaging (FedAvg) and its variants are a c…

Federated LearningPrivacy Preserving

Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares

2025-10-20 · Lachlan Ewen MacDonald, Hancheng Min, Leandro Palma, Salma Tarmoun 외 arxiv

Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or ``stable", regime. In contrast, gradient descent on neural networks is frequently per…

The Riemannian Geometry Associated to Gradient Flows of Linear Convolutional Networks

2025-07-08 · El Mehdi Achour, Kathlén Kohn, Holger Rauhut arxiv

We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresponding gradient flow on parameter space …