paper-with-me

홈 › Papers

Quantifying Context Mixing in Transformers

2023-01-30 · Hosein Mohebbi, Willem Zuidema, Grzegorz Chrupała, Afra Alishahi

Self-attention weights and their transformed variants have been the main source of information for analyzing token-to-token interactions in Transformer-based models. But despite their ease of interpretation, these weights are not faithful to the models' decisions as they are only one part of an encoder, and other components in the encoder layer can have considerable impact on information mixing in the output representations. In this work, by expanding the scope of analysis to the whole encoder block, we propose Value Zeroing, a novel context mixing score customized for Transformers that provides us with a deeper understanding of how information is mixed at each encoder layer. We demonstrate the superiority of our context mixing score over other analysis methods through a series of complementary evaluations with different viewpoints based on linguistically informed rationales, probing, and faithfulness analysis.

📄 PDF Abstract BibTeX arXiv:2301.12971

Code (1)

hmohebbi/valuezeroing 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Homophone Disambiguation Reveals Patterns of Context Mixing in Speech Transformers

2023-10-15 · Hosein Mohebbi, Grzegorz Chrupała, Willem Zuidema, Afra Alishahi

Transformers have become a key architecture in speech processing, but our understanding of how they build up representations of acoustic and linguistic structure is limited. In this study, we address this gap by investig…

Decoderspeech-recognitionSpeech Recognition

Deep Hyperspectral Unmixing using Transformer Network

2022-03-31 · Preetam Ghosh, Swalpa Kumar Roy, Bikram Koirala, Behnood Rasti 외

Currently, this paper is under review in IEEE. Transformers have intrigued the vision research community with their state-of-the-art performance in natural language processing. With their superior performance, transforme…

DecoderHyperspectral Image ClassificationHyperspectral Unmixingimage-classification+1

Hardwiring ViT Patch Selectivity into CNNs using Patch Mixing

2023-06-30 · Ariel N. Lee, Sarah Adel Bargal, Janavi Kasera, Stan Sclaroff 외

Vision transformers (ViTs) have significantly changed the computer vision landscape and have periodically exhibited superior performance in vision tasks compared to convolutional neural networks (CNNs). Although the jury…

Data AugmentationInductive Bias

Transformer based Endmember Fusion with Spatial Context for Hyperspectral Unmixing

2024-02-06 · R. M. K. L. Ratnayake, D. M. U. P. Sumanasekara, H. M. K. D. Wickramathilaka, G. M. R. I. Godaliyadda 외

In recent years, transformer-based deep learning networks have gained popularity in Hyperspectral (HS) unmixing applications due to their superior performance. The attention mechanism within transformers facilitates inpu…

Hyperspectral Unmixing

Active Token Mixer

2022-03-11 · Guoqiang Wei, Zhizheng Zhang, Cuiling Lan, Yan Lu 외

The three existing dominant network families, i.e., CNNs, Transformers, and MLPs, differ from each other mainly in the ways of fusing spatial contextual information, leaving designing more effective token-mixing mechanis…

Image ClassificationInstance SegmentationObject DetectionSemantic Segmentation