paper-with-me

홈 › Papers

Geometry of Critical Sets and Existence of Saddle Branches for Two-layer Neural Networks

2024-05-26 · Leyang Zhang, Yaoyu Zhang, Tao Luo

This paper presents a comprehensive analysis of critical point sets in two-layer neural networks. To study such complex entities, we introduce the critical embedding operator and critical reduction operator as our tools. Given a critical point, we use these operators to uncover the whole underlying critical set representing the same output function, which exhibits a hierarchical structure. Furthermore, we prove existence of saddle branches for any critical set whose output function can be represented by a narrower network. Our results provide a solid foundation to the further study of optimization and training behavior of neural networks.

📄 PDF Abstract BibTeX arXiv:2405.17501

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Symmetry & Critical Points for Symmetric Tensor Decomposition Problems

2023-06-13 · Yossi Arjevani, Gal Vinograd

We consider the nonconvex optimization problem associated with the decomposition of a real symmetric tensor into a sum of rank one terms. Use is made of the rich symmetry structure to construct infinite families of criti…

LEMMATensor Decomposition

The loss landscape of deep linear neural networks: a second-order analysis

2021-07-28 · El Mehdi Achour, François Malgouyres, Sébastien Gerchinovitz

We study the optimization landscape of deep linear neural networks with the square loss. It is known that, under weak assumptions, there are no spurious local minima and no local maxima. However, the existence and divers…

Diversity

Uncovering Critical Sets of Deep Neural Networks via Sample-Independent Critical Lifting

2025-05-19 · Leyang Zhang, Yaoyu Zhang, Tao Luo

This paper investigates the sample dependence of critical points for neural networks. We introduce a sample-independent critical lifting operator that associates a parameter of one network with a set of parameters of ano…

Deep Learning without Poor Local Minima

2016-05-23 · NeurIPS 2016 12 · Kenji Kawaguchi

In this paper, we prove a conjecture published in 1989 and also partially address an open problem announced at the Conference on Learning Theory (COLT) 2015. With no unrealistic assumption, we first prove the following s…

Deep LearningLearning Theory

A Theory of Speciation in Generative Diffusion Models on Compact Riemannian Manifolds

2026-08-24 · Alessio Marta, Paola Causin arxiv

Speciation in generative diffusion models denotes the emergence of distinct stable branches during denoising, through which initially undifferentiated trajectories progressively commit to different data classes. In this …