paper-with-me

Papers

Faster Directional Convergence of Linear Neural Networks under Spherically Symmetric Data

2021-12-01 · NeurIPS 2021 12 · Dachao Lin, Ruoyu Sun, Zhihua Zhang

In this paper, we study gradient methods for training deep linear neural networks with binary cross-entropy loss. In particular, we show global directional convergence guarantees from a polynomial rate to a linear rate for (deep) linear networks with spherically symmetric data distribution, which can be viewed as a specific zero-margin dataset. Our results do not require the assumptions in other works such as small initial loss, presumed convergence of weight direction, or overparameterization. We also characterize our findings in experiments.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Directional Convergence Analysis under Spherically Symmetric Distribution

2021-05-09 · Dachao Lin, Zhihua Zhang

We consider the fundamental problem of learning linear predictors (i.e., separable datasets with zero margin) using neural networks with gradient flow or gradient descent. Under the assumption of spherically symmetric da…

Greedy and Random Quasi-Newton Methods with Faster Explicit Superlinear Convergence

2021-12-01 · NeurIPS 2021 12 · Dachao Lin, Haishan Ye, Zhihua Zhang

In this paper, we follow Rodomanov and Nesterov’s work to study quasi-Newton methods. We focus on the common SR1 and BFGS quasi-Newton methods to establish better explicit (local) superlinear convergence rates. First, ba…

Linear Convergence of the Subspace Constrained Mean Shift Algorithm: From Euclidean to Directional Data

2021-04-29 · Yikun Zhang, Yen-Chi Chen

This paper studies the linear convergence of the subspace constrained mean shift (SCMS) algorithm, a well-known algorithm for identifying a density ridge defined by a kernel density estimator. By arguing that the SCMS al…

lp-Recovery of the Most Significant Subspace among Multiple Subspaces with Outliers

2010-12-18 · Gilad Lerman, Teng Zhang

We assume data sampled from a mixture of d-dimensional linear subspaces with spherically symmetric distributions within each subspace and an additional outlier component with spherically symmetric distribution within the…

FastMMD: Ensemble of Circular Discrepancy for Efficient Two-Sample Test

2014-05-12 · Ji Zhao, Deyu Meng

The maximum mean discrepancy (MMD) is a recently proposed test statistic for two-sample test. Its quadratic time complexity, however, greatly hampers its availability to large-scale applications. To accelerate the MMD ca…

Vocal Bursts Valence Prediction