paper-with-me

홈 › Papers

On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained Features

2022-03-02 · Jinxin Zhou, Xiao Li, Tianyu Ding, Chong You, Qing Qu, Zhihui Zhu

When training deep neural networks for classification tasks, an intriguing empirical phenomenon has been widely observed in the last-layer classifiers and features, where (i) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and (ii) cross-example within-class variability of last-layer activations collapses to zero. This phenomenon is called Neural Collapse (NC), which seems to take place regardless of the choice of loss functions. In this work, we justify NC under the mean squared error (MSE) loss, where recent empirical evidence shows that it performs comparably or even better than the de-facto cross-entropy loss. Under a simplified unconstrained feature model, we provide the first global landscape analysis for vanilla nonconvex MSE loss and show that the (only!) global minimizers are neural collapse solutions, while all other critical points are strict saddles whose Hessian exhibit negative curvature directions. Furthermore, we justify the usage of rescaled MSE loss by probing the optimization landscape around the NC solutions, showing that the landscape can be improved by tuning the rescaling hyperparameters. Finally, our theoretical findings are experimentally verified on practical network architectures.

📄 PDF Abstract BibTeX arXiv:2203.01238

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CP-GAN: Towards a Better Global Landscape of GANs

2019-09-25 · Ruoyu Sun, Tiantian Fang, Alex Schwing

GANs have been very popular in data generation and unsupervised learning, but our understanding of GAN training is still very limited. One major reason is that GANs are often formulated as non-convex-concave min-max opt…

Are All Losses Created Equal: A Neural Collapse Perspective

2022-10-04 · Jinxin Zhou, Chong You, Xiao Li, Kangning Liu 외

While cross entropy (CE) is the most commonly used loss to train deep neural networks for classification tasks, many alternative losses have been developed to obtain better empirical performance. Among them, which one is…

All

How Gradient Descent Separates Data with Neural Collapse: A Layer-Peeled Perspective

2021-05-21 · NeurIPS 2021 12 · Wenlong Ji, Yiping Lu, Yiliang Zhang, Zhun Deng 외

In this paper, we derive a landscape analysis to the surrogate model to study the inductive bias of the neural features and parameters from neural networks with cross-entropy. We show that once the training cross-entropy…

Inductive Bias

Neural Collapse with Normalized Features: A Geometric Analysis over the Riemannian Manifold

2022-09-19 · Can Yaras, Peng Wang, Zhihui Zhu, Laura Balzano 외

When training overparameterized deep networks for classification tasks, it has been widely observed that the learned features exhibit a so-called "neural collapse" phenomenon. More specifically, for the output features o…

Multi-class ClassificationRepresentation LearningRiemannian optimization

A Geometric Analysis of Neural Collapse with Unconstrained Features

2021-05-06 · NeurIPS 2021 12 · Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li 외

We provide the first global optimization landscape analysis of $Neural\;Collapse$ -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of neural networks during the terminal phase …

global-optimization