paper-with-me

홈 › Papers

Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis

2024-06-24 · Hongkang Li, Meng Wang, Shuai Zhang, Sijia Liu, Pin-Yu Chen

Efficient training and inference algorithms, such as low-rank adaption and model pruning, have shown impressive performance for learning Transformer-based large foundation models. However, due to the technical challenges of the non-convex optimization caused by the complicated architecture of Transformers, the theoretical study of why these methods can be applied to learn Transformers is mostly elusive. To the best of our knowledge, this paper shows the first theoretical analysis of the property of low-rank and sparsity of one-layer Transformers by characterizing the trained model after convergence using stochastic gradient descent. By focusing on a data model based on label-relevant and label-irrelevant patterns, we quantify that the gradient updates of trainable parameters are low-rank, which depends on the number of label-relevant patterns. We also analyze how model pruning affects the generalization while improving computation efficiency and conclude that proper magnitude-based pruning has a slight effect on the testing performance. We implement numerical experiments to support our findings.

📄 PDF Abstract BibTeX arXiv:2406.17167

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Provable Low Rank Plus Sparse Matrix Separation Via Nonconvex Regularizers

2021-09-26 · April Sagan, John E. Mitchell

This paper considers a large class of problems where we seek to recover a low rank matrix and/or sparse vector from some set of measurements. While methods based on convex relaxations suffer from a (possibly large) estim…

Matrix Completion

On the Role of Attention Masks and LayerNorm in Transformers

2024-05-29 · Xinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka 외

Self-attention is the key mechanism of transformers, which are the essential building blocks of modern foundation models. Recent studies have shown that pure self-attention suffers from an increasing degree of rank colla…

On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery

2024-10-17 · Renpu Liu, Ruida Zhou, Cong Shen, Jing Yang

An intriguing property of the Transformer is its ability to perform in-context learning (ICL), where the Transformer can solve different inference tasks without parameter updating based on the contextual information prov…

In-Context Learning

Rethinking Attention with Performers

2020-09-30 · ICLR 2021 1 · Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 외

We introduce Performers, Transformer architectures which can estimate regular (softmax) full-rank-attention Transformers with provable accuracy, but using only linear (as opposed to quadratic) space and time complexity, …

D4RLImage GenerationLanguage ModellingOffline RL

Transformers are Deep Optimizers: Provable In-Context Learning for Deep Model Training

2024-11-25 · Weimin Wu, Maojiang Su, Jerry Yao-Chieh Hu, Zhao Song 외

We investigate the transformer's capability for in-context learning (ICL) to simulate the training process of deep models. Our key contribution is providing a positive example of using a transformer to train a deep neura…

In-Context Learning