paper-with-me

홈 › Papers

Incorporating Residual and Normalization Layers into Analysis of Masked Language Models

2021-09-15 · EMNLP 2021 11 · Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui

Transformer architecture has become ubiquitous in the natural language processing field. To interpret the Transformer-based models, their attention patterns have been extensively analyzed. However, the Transformer architecture is not only composed of the multi-head attention; other components can also contribute to Transformers' progressive performance. In this study, we extended the scope of the analysis of Transformers from solely the attention patterns to the whole attention block, i.e., multi-head attention, residual connection, and layer normalization. Our analysis of Transformer-based masked language models shows that the token-to-token interaction performed via attention has less impact on the intermediate representations than previously assumed. These results provide new intuitive explanations of existing reports; for example, discarding the learned attention patterns tends not to adversely affect the performance. The codes of our experiments are publicly available.

📄 PDF Abstract BibTeX arXiv:2109.07152

Code (2)

gorokoba560/norm-analysis-of-transformer 공식 구현 pytorch
mt-upc/transformer-contributions pytorch

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Adam 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

On residual network depth

2025-10-03 · Benoit Dherin, Michael Munn arxiv

Deep residual architectures, such as ResNet and the Transformer, have enabled models of unprecedented depth, yet a formal understanding of why depth is so effective remains an open question. A popular intuition, followin…

Progressive Residual Warmup for Language Model Pretraining

2026-03-05 · Tianhao Chen, Xin Xu, Lu Yin, Hao Chen 외 arxiv

Transformer architectures serve as the backbone for most modern Large Language Models, therefore their pretraining stability and convergence speed are of central concern. Motivated by the logical dependency of sequential…

Fixup Initialization: Residual Learning Without Normalization

2019-01-27 · ICLR 2019 5 · Hongyi Zhang, Yann N. Dauphin, Tengyu Ma

Normalization layers are a staple in state-of-the-art deep neural network architectures. They are widely believed to stabilize training, enable higher learning rate, accelerate convergence and improve generalization, tho…

General Classificationimage-classificationImage ClassificationMachine Translation+1

Rethinking the role of normalization and residual blocks for spiking neural networks

2022-03-03 · Shin-ichi Ikegawa, Ryuji Saiin, Yoshihide Sawada, Naotake Natori

Biologically inspired spiking neural networks (SNNs) are widely used to realize ultralow-power energy consumption. However, deep SNNs are not easy to train due to the excessive firing of spiking neurons in the hidden lay…

PyHessian: Neural Networks Through the Lens of the Hessian

2019-12-16 · Zhewei Yao, Amir Gholami, Kurt Keutzer, Michael Mahoney

We present PYHESSIAN, a new scalable framework that enables fast computation of Hessian (i.e., second-order derivative) information for deep neural networks. PYHESSIAN enables fast computations of the top Hessian eigenva…