paper-with-me

Papers

Average gradient outer product as a mechanism for deep neural collapse

2024-02-21 · Daniel Beaglehole, Peter Súkeník, Marco Mondelli, Mikhail Belkin

Deep Neural Collapse (DNC) refers to the surprisingly rigid structure of the data representations in the final layers of Deep Neural Networks (DNNs). Though the phenomenon has been measured in a variety of settings, its emergence is typically explained via data-agnostic approaches, such as the unconstrained features model. In this work, we introduce a data-dependent setting where DNC forms due to feature learning through the average gradient outer product (AGOP). The AGOP is defined with respect to a learned predictor and is equal to the uncentered covariance matrix of its input-output gradients averaged over the training dataset. The Deep Recursive Feature Machine (Deep RFM) is a method that constructs a neural network by iteratively mapping the data with the AGOP and applying an untrained random feature map. We demonstrate empirically that DNC occurs in Deep RFM across standard settings as a consequence of the projection with the AGOP matrix computed at each layer. Further, we theoretically explain DNC in Deep RFM in an asymptotic setting and as a result of kernel learning. We then provide evidence that this mechanism holds for neural networks more generally. In particular, we show that the right singular vectors and values of the weights can be responsible for the majority of within-class variability collapse for DNNs trained in the feature learning regime. As observed in recent work, this singular structure is highly correlated with that of the AGOP.

📄 PDF Abstract BibTeX arXiv:2402.13728

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Speeding-Up Back-Propagation in DNN: Approximate Outer Product with Memory

2021-10-18 · Eduin E. Hernandez, Stefano Rini, Tolga M. Duman

In this paper, an algorithm for approximate evaluation of back-propagation in DNN training is considered, which we term Approximate Outer Product Gradient Descent with Memory (Mem-AOP-GD). The Mem-AOP-GD algorithm implem…

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts

2026-05-12 · Sagi Ahrac, Noya Hochwald, Mor Geva arxiv

Sparse Mixture-of-Experts (SMoE) models enable scaling language models efficiently, but training them remains challenging, as routing can collapse onto few experts and auxiliary load-balancing losses can reduce specializ…

Emergence in non-neural models: grokking modular arithmetic via average gradient outer product

2024-07-29 · Neil Mallinar, Daniel Beaglehole, Libin Zhu, Adityanarayanan Radhakrishnan 외

Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accuracy in the training process. It is often …

Revisiting GANs by Best-Response Constraint: Perspective, Methodology, and Application

2022-05-20 · Risheng Liu, Jiaxin Gao, Xuan Liu, Xin Fan

In past years, the minimax type single-level optimization formulation and its variations have been widely utilized to address Generative Adversarial Networks (GANs). Unfortunately, it has been proved that these alternati…

Rank-1 Convolutional Neural Network

2018-08-13 · Hyein Kim, Jungho Yoon, Byeongseon Jeong, Sukho Lee

In this paper, we propose a convolutional neural network(CNN) with 3-D rank-1 filters which are composed by the outer product of 1-D filters. After being trained, the 3-D rank-1 filters can be decomposed into 1-D filters…