paper-with-me

홈 › Papers

Block Skim Transformer for Efficient Question Answering

2021-01-01 · Yue Guan, Jingwen Leng, Yuhao Zhu, Minyi Guo

Transformer based encoder models have achieved promising results on natural language processing (NLP) task including question answering (QA). Different from sequence classification or language modeling tasks, hidden states at all position are used for the final classification in QA. However, we do not always need all the context to answer the raised question. Following this idea, we proposed Block Skim Transformer (BST) to improve and accelerate the processing of transformer QA models. The key idea of BST is to identify the context that must be further processed and the blocks that could be safely rejected early on during inference. Critically, we learn such information from self attention weights. As a result, the model hidden states are pruned at sequence dimension, achieving significant inference speedup. We also show that such extra training optimization objection also improves model performance. As an plugin to the transformer based QA models, BST is compatible to other model compression methods without changing existing network architectures. BST improves QA models performance on different datasets and achieves $1.6\times$ speedup on $\text{BERT}_{\text{large}}$ model.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingModel CompressionQuestion Answering

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Multi-Head Attention 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Block-Skim: Efficient Question Answering for Transformer

2021-12-16 · Yue Guan, Zhengyi Li, Jingwen Leng, Zhouhan Lin 외

Transformer models have achieved promising results on natural language processing (NLP) tasks including extractive question answering (QA). Common Transformer encoders used in NLP tasks process the hidden states of all i…

Extractive Question-AnsweringQuestion Answering

Transkimmer: Transformer Learns to Layer-wise Skim

2022-05-15 · ACL 2022 5 · Yue Guan, Zhengyi Li, Jingwen Leng, Zhouhan Lin 외

Transformer architecture has become the de-facto model for many machine learning tasks from natural language processing and computer vision. As such, improving its computational efficiency becomes paramount. One of the m…

Computational Efficiency

Read As Human: Compressing Context via Parallelizable Close Reading and Skimming

2026-02-02 · Jiwei Tang, Shilei Liu, Zhicheng Zhang, Qingsong Lv 외 arxiv

Large Language Models (LLMs) demonstrate exceptional capability across diverse tasks. However, their deployment in long-context scenarios is hindered by two challenges: computational inefficiency and redundant informatio…

Contrastive LearningQuestion Answering

Skim-Attention: Learning to Focus via Document Layout

2021-09-02 · Findings (EMNLP) 2021 11 · Laura Nguyen, Thomas Scialom, Jacopo Staiano, Benjamin Piwowarski

Transformer-based pre-training techniques of text and layout have proven effective in a number of document understanding tasks. Despite this success, multimodal pre-training models suffer from very high computational and…

document understandingLanguage ModelingLanguage Modelling

QuALITY: Question Answering with Long Input Texts, Yes!

2021-12-16 · NAACL 2022 7 · Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia 외

To enable building and testing models on long-document comprehension, we introduce QuALITY, a multiple-choice QA dataset with context passages in English that have an average length of about 5,000 tokens, much longer tha…

Multiple-choiceMultiple Choice Question Answering (MCQA)Question Answering