paper-with-me

홈 › Papers

Low-Rank Learning by Design: the Role of Network Architecture and Activation Linearity in Gradient Rank Collapse

2024-02-09 · Bradley T. Baker, Barak A. Pearlmutter, Robyn Miller, Vince D. Calhoun, Sergey M. Plis

Our understanding of learning dynamics of deep neural networks (DNNs) remains incomplete. Recent research has begun to uncover the mathematical principles underlying these networks, including the phenomenon of "Neural Collapse", where linear classifiers within DNNs converge to specific geometrical structures during late-stage training. However, the role of geometric constraints in learning extends beyond this terminal phase. For instance, gradients in fully-connected layers naturally develop a low-rank structure due to the accumulation of rank-one outer products over a training batch. Despite the attention given to methods that exploit this structure for memory saving or regularization, the emergence of low-rank learning as an inherent aspect of certain DNN architectures has been under-explored. In this paper, we conduct a comprehensive study of gradient rank in DNNs, examining how architectural choices and structure of the data effect gradient rank bounds. Our theoretical analysis provides these bounds for training fully-connected, recurrent, and convolutional neural networks. We also demonstrate, both theoretically and empirically, how design choices like activation function linearity, bottleneck layer introduction, convolutional stride, and sequence truncation influence these bounds. Our findings not only contribute to the understanding of learning dynamics in DNNs, but also provide practical guidance for deep learning engineers to make informed design decisions.

📄 PDF Abstract BibTeX arXiv:2402.06751

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Transformer as a hippocampal memory consolidation model based on NMDAR-inspired nonlinearity

2023-09-21 · NeurIPS 2023 11

The hippocampus plays a critical role in learning, memory, and spatial representation, processes that depend on the NMDA receptor (NMDAR). Inspired by recent findings that compare deep learning models to the hippocampus,…

Gated MLPs as Symmetry-Broken Rank-1 Bilinear Attention

2026-06-20 · Nathan Breslow arxiv

We show that the conventional gated MLP can be viewed as a rank-1 approximation to a bilinear attention mechanism with two distinct factors corresponding to the query and the key. We further show that moving the nonlinea…

Empirical Loss Landscape Analysis of Neural Network Activation Functions

2023-06-28 · Anna Sergeevna Bosman, Andries Engelbrecht, Marde Helbig

Activation functions play a significant role in neural network design by enabling non-linearity. The choice of activation function was previously shown to influence the properties of the resulting loss landscape. Underst…

Understanding the Role of Nonlinearity in Training Dynamics of Contrastive Learning

2022-06-02 · Yuandong Tian

While the empirical success of self-supervised learning (SSL) heavily relies on the usage of deep nonlinear models, existing theoretical works on SSL understanding still focus on linear ones. In this paper, we study the …

Contrastive LearningSelf-Supervised Learning

Activation Function Design Sustains Plasticity in Continual Learning

2025-09-26 · Lute Lillo, Nick Cheney arxiv

In independent, identically distributed (i.i.d.) training regimes, activation functions have been benchmarked extensively, and their differences often shrink once model size and optimization are tuned. In continual learn…

Reinforcement LearningContinual Learning