paper-with-me

홈 › Papers

Synthesizer Based Efficient Self-Attention for Vision Tasks

2022-01-05 · Guangyang Zhu, Jianfeng Zhang, Yuanzhi Feng, Hai Lan

Self-attention module shows outstanding competence in capturing long-range relationships while enhancing performance on vision tasks, such as image classification and image captioning. However, the self-attention module highly relies on the dot product multiplication and dimension alignment among query-key-value features, which cause two problems: (1) The dot product multiplication results in exhaustive and redundant computation. (2) Due to the visual feature map often appearing as a multi-dimensional tensor, reshaping the scale of the tensor feature to adapt to the dimension alignment might destroy the internal structure of the tensor feature map. To address these problems, this paper proposes a self-attention plug-in module with its variants, namely, Synthesizing Tensor Transformations (STT), for directly processing image tensor features. Without computing the dot-product multiplication among query-key-value, the basic STT is composed of the tensor transformation to learn the synthetic attention weight from visual information. The effectiveness of STT series is validated on the image classification and image caption. Experiments show that the proposed STT achieves competitive performance while keeping robustness compared to self-attention in the aforementioned vision tasks.

📄 PDF Abstract BibTeX arXiv:2201.01410

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioningimage-classificationImage Classification

Similar Papers 제목 키워드 기반

Synthesizer: Rethinking Self-Attention in Transformer Models

2020-05-02 · Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan 외

The dot product self-attention is known to be central and indispensable to state-of-the-art Transformer models. But is it really required? This paper investigates the true importance and contribution of the dot product-b…

Abstractive Text SummarizationDialogue GenerationDocument SummarizationLanguage Modeling+6

Synthesizer: Rethinking Self-Attention for Transformer Models

2021-01-01 · Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan 외

The dot product self-attention is known to be central and indispensable to state-of-the-art Transformer models. But is it really required? This paper investigates the true importance and contribution of the dot product-b…

Language ModelingLanguage ModellingMachine TranslationText Generation+1

Portrait4D-v2: Pseudo Multi-View Data Creates Better 4D Head Synthesizer

2024-03-20 · Yu Deng, Duomin Wang, Baoyuan Wang

In this paper, we propose a novel learning approach for feed-forward one-shot 4D head avatar synthesis. Different from existing methods that often learn from reconstructing monocular videos guided by 3DMM, we employ pseu…

Transformer-based End-to-End Speech Recognition with Local Dense Synthesizer Attention

2020-10-23 · Menglong Xu, Shengqiang Li, Xiao-Lei Zhang

Recently, several studies reported that dot-product selfattention (SA) may not be indispensable to the state-of-theart Transformer models. Motivated by the fact that dense synthesizer attention (DSA), which dispenses wit…

speech-recognitionSpeech Recognition

Synthesizer Preset Interpolation using Transformer Auto-Encoders

2022-10-27 · Gwendal Le Vaillant, Thierry Dutoit

Sound synthesizers are widespread in modern music production but they increasingly require expert skills to be mastered. This work focuses on interpolation between presets, i.e., sets of values of all sound synthesis par…