paper-with-me

홈 › Papers

FastFormers: Highly Efficient Transformer Models for Natural Language Understanding

2020-10-26 · EMNLP (sustainlp) 2020 11 · Young Jin Kim, Hany Hassan Awadalla

Transformer-based models are the state-of-the-art for Natural Language Understanding (NLU) applications. Models are getting bigger and better on various tasks. However, Transformer models remain computationally challenging since they are not efficient at inference-time compared to traditional approaches. In this paper, we present FastFormers, a set of recipes to achieve efficient inference-time performance for Transformer-based models on various NLU tasks. We show how carefully utilizing knowledge distillation, structured pruning and numerical optimization can lead to drastic improvements on inference efficiency. We provide effective recipes that can guide practitioners to choose the best settings for various NLU tasks and pretrained models. Applying the proposed recipes to the SuperGLUE benchmark, we achieve from 9.8x up to 233.9x speed-up compared to out-of-the-box models on CPU. On GPU, we also achieve up to 12.4x speed-up with the presented methods. We show that FastFormers can drastically reduce cost of serving 100 million requests from 4,223 USD to just 18 USD on an Azure F16s_v2 instance. This translates to a sustainable runtime by reducing energy consumption 6.9x - 125.8x according to the metrics used in the SustaiNLP 2020 shared task.

📄 PDF Abstract BibTeX arXiv:2010.13382

Code (2)

microsoft/fastformers 공식 구현 pytorch
philschmid/knowledge-distillation-transformers-pytorch-sagemaker pytorch

Tasks

CPUGPUKnowledge DistillationNatural Language Understanding

Methods 이 논문이 사용한 방법론

Pruning 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

Understanding Prior Bias and Choice Paralysis in Transformer-based Language Representation Models through Four Experimental Probes

2022-10-03 · Ke Shen, Mayank Kejriwal

Recent work on transformer-based neural networks has led to impressive advances on multiple-choice natural language understanding (NLU) problems, such as Question Answering (QA) and abductive reasoning. Despite these adv…

Decision MakingMultiple-choiceNatural Language UnderstandingQuestion Answering

End-to-End Neural Transformer Based Spoken Language Understanding

2020-08-12 · Martin Radfar, Athanasios Mouchtaris, Siegfried Kunzmann

Spoken language understanding (SLU) refers to the process of inferring the semantic information from audio signals. While the neural transformers consistently deliver the best performance among the state-of-the-art neura…

Spoken Language Understanding

BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques

2024-11-22 · Muhammad Rafsan Kabir, Md. Mohibur Rahman Nabil, Mohammad Ashrafuzzaman Khan

Sentence-level embedding is essential for various tasks that require understanding natural language. Many studies have explored such embeddings for high-resource languages like English. However, low-resource languages li…

Hate Speech DetectionKnowledge DistillationSemantic Textual SimilaritySentence+3

Dual Transformer for Point Cloud Analysis

2021-04-27 · Xian-Feng Han, Yi-Fei Jin, Hui-Xian Cheng, Guo-Qiang Xiao

Following the tremendous success of transformer in natural language processing and image understanding tasks, in this paper, we present a novel point cloud representation learning architecture, named Dual Transformer Net…

3D Point Cloud ClassificationPoint Cloud ClassificationPositionRepresentation Learning

Towards Highly Expressive Machine Learning Models of Non-Melanoma Skin Cancer

2022-07-09 · Simon M. Thomas, James G. Lefevre, Glenn Baxter, Nicholas A. Hamilton

Pathologists have a rich vocabulary with which they can describe all the nuances of cellular morphology. In their world, there is a natural pairing of images and words. Recent advances demonstrate that machine learning m…

BIG-bench Machine Learning