paper-with-me

홈 › Papers

Attention Is Not All You Need Anymore

2023-08-15 · Zhe Chen

In recent years, the popular Transformer architecture has achieved great success in many application areas, including natural language processing and computer vision. Many existing works aim to reduce the computational and memory complexity of the self-attention mechanism in the Transformer by trading off performance. However, performance is key for the continuing success of the Transformer. In this paper, a family of drop-in replacements for the self-attention mechanism in the Transformer, called the Extractors, is proposed. Four types of the Extractors, namely the super high-performance Extractor (SHE), the higher-performance Extractor (HE), the worthwhile Extractor (WE), and the minimalist Extractor (ME), are proposed as examples. Experimental results show that replacing the self-attention mechanism with the SHE evidently improves the performance of the Transformer, whereas the simplified versions of the SHE, i.e., the HE, the WE, and the ME, perform close to or better than the self-attention mechanism with less computational and memory complexity. Furthermore, the proposed Extractors have the potential or are able to run faster than the self-attention mechanism since their critical paths of computation are much shorter. Additionally, the sequence prediction problem in the context of text generation is formulated using variable-length discrete-time Markov chains, and the Transformer is reviewed based on our understanding.

📄 PDF Abstract BibTeX arXiv:2308.07661

Code (2)

fabianwinter93/JAX jax
fabianwinter93/JAX/tree/main/SuperHighPerformanceExtractor jax

Tasks

AllText Generation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Adam 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

HDDL -- A Language to Describe Hierarchical Planning Problems

2019-11-13 · D. Höller, G. Behnke, P. Bercher, S. Biundo 외

The research in hierarchical planning has made considerable progress in the last few years. Many recent systems do not rely on hand-tailored advice anymore to find solutions, but are supposed to be domain-independent sys…

SemEval-2020 Task 8: Memotion Analysis- the Visuo-Lingual Metaphor!

2020-12-01 · SEMEVAL 2020 · Chhavi Sharma, Deepesh Bhageria, William Scott, Srinivas PYKL 외

Information on social media comprises of various modalities such as textual, visual and audio. NLP and Computer Vision communities often leverage only one prominent modality in isolation to study social media. However, c…

Emotion Recognition

SemEval-2020 Task 8: Memotion Analysis -- The Visuo-Lingual Metaphor!

2020-08-09 · Chhavi Sharma, Deepesh Bhageria, William Scott, Srinivas PYKL 외

Information on social media comprises of various modalities such as textual, visual and audio. NLP and Computer Vision communities often leverage only one prominent modality in isolation to study social media. However, t…

Emotion Recognition

Statistically Significant Stopping of Neural Network Training

2021-03-01 · J. K. Terry, Mario Jayakumar, Kusal De Alwis

The general approach taken when training deep learning classifiers is to save the parameters after every few iterations, train until either a human observer or a simple metric-based heuristic decides the network isn't le…

DF26: We Cannot Tell Fake From Real Anymore

2026-09-07 · Severyn Shykula, Andrii Yermakov, Ivan Samarskyi, Dmytro Mishkin 외 hf

We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenari…