paper-with-me

홈 › Papers

Global Convergence and Geometric Characterization of Slow to Fast Weight Evolution in Neural Network Training for Classifying Linearly Non-Separable Data

2020-02-28 · Ziang Long, Penghang Yin, Jack Xin

In this paper, we study the dynamics of gradient descent in learning neural networks for classification problems. Unlike in existing works, we consider the linearly non-separable case where the training data of different classes lie in orthogonal subspaces. We show that when the network has sufficient (but not exceedingly large) number of neurons, (1) the corresponding minimization problem has a desirable landscape where all critical points are global minima with perfect classification; (2) gradient descent is guaranteed to converge to the global minima. Moreover, we discovered a geometric condition on the network weights so that when it is satisfied, the weight evolution transitions from a slow phase of weight direction spreading to a fast phase of weight convergence. The geometric condition says that the convex hull of the weights projected on the unit sphere contains the origin.

📄 PDF Abstract BibTeX arXiv:2002.12563

Code (0)

등록된 구현이 없습니다.

Tasks

General Classification

Similar Papers 제목 키워드 기반

Can We Find Nash Equilibria at a Linear Rate in Markov Games?

2023-03-03 · Zhuoqing Song, Jason D. Lee, Zhuoran Yang

We study decentralized learning in two-player zero-sum discounted Markov games where the goal is to design a policy optimization algorithm for either agent satisfying two properties. First, the player does not need to kn…

Over-Parameterization Exponentially Slows Down Gradient Descent for Learning a Single Neuron

2023-02-20 · Weihang Xu, Simon S. Du

We revisit the problem of learning a single neuron with ReLU activation under Gaussian input with square loss. We particularly focus on the over-parameterization setting where the student network has $n\ge 2$ neurons. We…

From Order to Distribution: A Spectral Characterization of Forgetting in Continual Learning

2026-04-15 · Zonghuan Xu, Xingjun Ma arxiv

A central challenge in continual learning is forgetting, the loss of performance on previously learned tasks induced by sequential adaptation to new ones. While forgetting has been extensively studied empirically, rigoro…

Continual Learning

Fast Escape, Slow Convergence: Learning Dynamics of Phase Retrieval under Power-Law Data

2025-11-24 · Guillaume Braun, Bruno Loureiro, Ha Quang Minh, Masaaki Imaizumi arxiv

Scaling laws describe how learning performance improves with data, compute, or training time, and have become a central theme in modern deep learning. We study this phenomenon in a canonical nonlinear model: phase retrie…

Statistical and Computational Guarantees for the Baum-Welch Algorithm

2015-12-27 · Fanny Yang, Sivaraman Balakrishnan, Martin J. Wainwright

The Hidden Markov Model (HMM) is one of the mainstays of statistical modeling of discrete time series, with applications including speech recognition, computational biology, computer vision and econometrics. Estimating a…

Econometricsspeech-recognitionSpeech RecognitionTime Series+1