paper-with-me

홈 › Papers

Mimetic Initialization of Self-Attention Layers

2023-05-16 · Asher Trockman, J. Zico Kolter

It is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point. We explore the weights of such pre-trained Transformers (particularly for vision) to attempt to find reasons for this discrepancy. Surprisingly, we find that simply initializing the weights of self-attention layers so that they "look" more like their pre-trained counterparts allows us to train vanilla Transformers faster and to higher final accuracies, particularly on vision tasks such as CIFAR-10 and ImageNet classification, where we see gains in accuracy of over 5% and 4%, respectively. Our initialization scheme is closed form, learning-free, and very simple: we set the product of the query and key weights to be approximately the identity, and the product of the value and projection weights to approximately the negative identity. As this mimics the patterns we saw in pre-trained Transformers, we call the technique "mimetic initialization".

📄 PDF Abstract BibTeX arXiv:2305.09828

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mimetic Initialization of MLPs

2026-02-06 · Asher Trockman, J. Zico Kolter arxiv

Mimetic initialization uses pretrained models as case studies of good initialization, using observations of structures in trained weights to inspire new, simple initialization techniques. So far, it has been applied only…

Mimetic Initialization Helps State Space Models Learn to Recall

2024-10-14 · Asher Trockman, Hrayr Harutyunyan, J. Zico Kolter, Sanjiv Kumar 외

Recent work has shown that state space models such as Mamba are significantly worse than Transformers on recall-based tasks due to the fact that their state size is constant with respect to their input sequence length. B…

MambaState Space Models

Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention

2019-08-29 · IJCNLP 2019 11 · Biao Zhang, Ivan Titov, Rico Sennrich

The general trend in NLP is towards increasing model capacity and performance via deeper neural networks. However, simply stacking more layers of the popular Transformer architecture for machine translation results in po…

DecoderMachine TranslationTranslation

Can graphene bilayers be the membrane mimetic materials? "Ion channels" in graphene-based nanostructures

2018-07-23 · Oleg V. Gradov, Margaret A. Gradova

The prospects of application of graphene and related structures as the membrane mimetic materials, capable of reproducing several biomembrane functions up to the certain limit, are analyzed in the series of our papers. T…

Initialization and Regularization of Factorized Neural Layers

2021-05-03 · ICLR 2021 1 · Mikhail Khodak, Neil Tenenholtz, Lester Mackey, Nicolò Fusi

Factorized layers--operations parameterized by products of two or more matrices--occur in a variety of deep learning contexts, including compressed model training, certain types of knowledge distillation, and multi-head …

Knowledge DistillationModel CompressionTensor DecompositionUnsupervised Pre-training