paper-with-me

홈 › Papers

Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations

2026-02-05 · Alec S. Xu, Can Yaras, Matthew Asato, Qing Qu, Laura Balzano arxiv

Recent empirical evidence has demonstrated that the training dynamics of large-scale deep neural networks occur within low-dimensional subspaces. While this has inspired new research into low-rank training, compression, and adaptation, theoretical justification for these dynamics in nonlinear networks remains limited. %compared to deep linear settings. To address this gap, this paper analyzes the learning dynamics of multi-layer perceptrons (MLPs) under gradient descent (GD). We demonstrate that the weight dynamics concentrate within invariant low-dimensional subspaces throughout training. Theoretically, we precisely characterize these invariant subspaces for two-layer networks with smooth nonlinear activations, providing insight into their emergence. Experimentally, we validate that this phenomenon extends beyond our theoretical assumptions. Leveraging these insights, we empirically show there exists a low-rank MLP parameterization that, when initialized within the appropriate subspaces, matches the classification performance of fully-parameterized counterparts on a variety of classification tasks.

📄 PDF Abstract BibTeX arXiv:2602.06208

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Trap of Feature Diversity in the Learning of MLPs

2021-12-02 · Dongrui Liu, Shaobo Wang, Jie Ren, Kangrui Wang 외

In this paper, we focus on a typical two-phase phenomenon in the learning of multi-layer perceptrons (MLPs), and we aim to explain the reason for the decrease of feature diversity in the first phase. Specifically, people…

Diversity

Optimizing Rank for High-Fidelity Implicit Neural Representations

2025-12-16 · Julian McGinnis, Florian A. Hölzl, Suprosanna Shit, Florentin Bieder 외 arxiv

Implicit Neural Representations (INRs) based on vanilla Multi-Layer Perceptrons (MLPs) are widely believed to be incapable of representing high-frequency content. This has directed research efforts towards architectural …

Novel View Synthesis

Kernel-Based Smoothness Analysis of Residual Networks

2020-09-21 · Tom Tirer, Joan Bruna, Raja Giryes

A major factor in the success of deep neural networks is the use of sophisticated architectures rather than the classical multilayer perceptron (MLP). Residual networks (ResNets) stand out among these powerful modern arc…

The late-stage training dynamics of (stochastic) subgradient descent on homogeneous neural networks

2025-02-08 · Sholom Schechtman, Nicolas Schreuder

We analyze the implicit bias of constant step stochastic subgradient descent (SGD). We consider the setting of binary classification with homogeneous neural networks - a large class of deep neural networks with ReLU-type…

Binary Classification

Machine Learning Nonadiabatic Dynamics: Eliminating Phase Freedom of Nonadiabatic Couplings with the State-Intraction State-Averaged Spin-Restricted Ensemble-Referenced Kohn-Sham Approach

2024-10-30 · Sung Wook Moon, Soohaeng Yoo Willow, Tae Hyeon Park, Seung Kyu Min 외

Excited-state molecular dynamics (ESMD) simulations near conical intersections (CIs) pose significant challenges when using machine learning potentials (MLPs). Although MLPs have gained recognition for their integration …