paper-with-me

홈 › Papers

Training invariances and the low-rank phenomenon: beyond linear networks

2022-01-28 · ICLR 2022 4 · Thien Le, Stefanie Jegelka

The implicit bias induced by the training of neural networks has become a topic of rigorous study. In the limit of gradient flow and gradient descent with appropriate step size, it has been shown that when one trains a deep linear network with logistic or exponential loss on linearly separable data, the weights converge to rank-1 matrices. In this paper, we extend this theoretical result to the last few linear layers of the much wider class of nonlinear ReLU-activated feedforward networks containing fully-connected layers and skip connections. Similar to the linear case, the proof relies on specific local training invariances, sometimes referred to as alignment, which we show to hold for submatrices where neurons are stably-activated in all training examples, and it reflects empirical results in the literature. We also show this is not true in general for the full matrix of ReLU fully-connected layers. Our proof relies on a specific decomposition of the network into a multilinear function and another ReLU network whose weights are constant under a certain parameter directional convergence.

📄 PDF Abstract BibTeX arXiv:2201.11968

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Linear Mode Connectivity in Differentiable Tree Ensembles

2024-05-23 · Ryuichi Kanoh, Mahito Sugiyama

Linear Mode Connectivity (LMC) refers to the phenomenon that performance remains consistent for linearly interpolated models in the parameter space. For independently optimized model pairs from different random initializ…

Linear Mode Connectivity

On the detrimental effect of invariances in the likelihood for variational inference

2022-09-15 · Richard Kurle, Ralf Herbrich, Tim Januschowski, Yuyang Wang 외

Variational Bayesian posterior inference often requires simplifying approximations such as mean-field parametrisation to ensure tractability. However, prior work has associated the variational mean-field approximation fo…

Variational Inference

Rank-Based Causal Discovery for Post-Nonlinear Models

2023-02-23 · Grigor Keropyan, David Strieder, Mathias Drton

Learning causal relationships from empirical observations is a central task in scientific research. A common method is to employ structural causal models that postulate noisy functional relations among a set of interacti…

Causal Discovery

Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations

2026-02-05 · Alec S. Xu, Can Yaras, Matthew Asato, Qing Qu 외 arxiv

Recent empirical evidence has demonstrated that the training dynamics of large-scale deep neural networks occur within low-dimensional subspaces. While this has inspired new research into low-rank training, compression, …

Characterizing the invariances of learning algorithms using category theory

2019-05-06 · Kenneth D. Harris

Many learning algorithms have invariances: when their training data is transformed in certain ways, the function they learn transforms in a predictable manner. Here we formalize this notion using concepts from the mathem…

regression