paper-with-me

홈 › Papers

Low Rank Gradients and Where to Find Them

2025-10-01 · Rishi Sonthalia, Michael Murray, Guido Montúfar arxiv

This paper investigates low-rank structure in the gradients of the training loss for two-layer neural networks while relaxing the usual isotropy assumptions on the training data and parameters. We consider a spiked data model in which the bulk can be anisotropic and ill-conditioned, we do not require independent data and weight matrices and we also analyze both the mean-field and neural-tangent-kernel scalings. We show that the gradient with respect to the input weights is approximately low rank and is dominated by two rank-one terms: one aligned with the bulk data-residue , and another aligned with the rank one spike in the input data. We characterize how properties of the training data, the scaling regime and the activation function govern the balance between these two components. Additionally, we also demonstrate that standard regularizers, such as weight decay, input noise and Jacobian penalties, also selectively modulate these components. Experiments on synthetic and real data corroborate our theoretical predictions.

📄 PDF Abstract BibTeX arXiv:2510.01303

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LoRA-Pro: Are Low-Rank Adapters Properly Optimized?

2024-07-25 · Zhengbo Wang, Jian Liang, Ran He, Zilei Wang 외

Low-rank adaptation, also known as LoRA, has emerged as a prominent method for parameter-efficient fine-tuning of foundation models. Despite its computational efficiency, LoRA still yields inferior performance compared t…

Code GenerationComputational EfficiencyDialogue Generationimage-classification+4

AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment

2025-11-12 · Ruibo Deng, Duanyu Feng, Wenqiang Lei arxiv

Offline preference optimization offers a simpler and more stable alternative to RLHF for aligning language models. However, their effectiveness is critically dependent on ranking accuracy, a metric where further gains ar…

Low-Rank Learning by Design: the Role of Network Architecture and Activation Linearity in Gradient Rank Collapse

2024-02-09 · Bradley T. Baker, Barak A. Pearlmutter, Robyn Miller, Vince D. Calhoun 외

Our understanding of learning dynamics of deep neural networks (DNNs) remains incomplete. Recent research has begun to uncover the mathematical principles underlying these networks, including the phenomenon of "Neural Co…

Neural Conditional Gradients

2018-03-12 · Patrick Schramowski, Christian Bauckhage, Kristian Kersting

The move from hand-designed to learned optimizers in machine learning has been quite successful for gradient-based and -free optimizers. When facing a constrained problem, however, maintaining feasibility typically requi…

Automatic differentiation for Riemannian optimization on low-rank matrix and tensor-train manifolds

2021-03-27 · Alexander Novikov, Maxim Rakhuba, Ivan Oseledets

In scientific computing and machine learning applications, matrices and more general multidimensional arrays (tensors) can often be approximated with the help of low-rank decompositions. Since matrices and tensors of fix…

Riemannian optimization