Block Skim Transformer for Efficient Question Answering
Transformer based encoder models have achieved promising results on natural language processing (NLP) task including question answering (QA). Different from sequence classification or language modeling tasks, hidden states at all position are used for the final classification in QA. However, we do not always need all the context to answer the raised question. Following this idea, we proposed Block Skim Transformer (BST) to improve and accelerate the processing of transformer QA models. The key idea of BST is to identify the context that must be further processed and the blocks that could be safely rejected early on during inference. Critically, we learn such information from self attention weights. As a result, the model hidden states are pruned at sequence dimension, achieving significant inference speedup. We also show that such extra training optimization objection also improves model performance. As an plugin to the transformer based QA models, BST is compatible to other model compression methods without changing existing network architectures. BST improves QA models performance on different datasets and achieves $1.6\times$ speedup on $\text{BERT}_{\text{large}}$ model.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingModel CompressionQuestion AnsweringMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Block-Skim: Efficient Question Answering for Transformer
Transformer models have achieved promising results on natural language processing (NLP) tasks including extractive question answering (QA). Common Transformer encoders used in NLP tasks process the hidden states of all i…
Extractive Question-AnsweringQuestion AnsweringTranskimmer: Transformer Learns to Layer-wise Skim
Transformer architecture has become the de-facto model for many machine learning tasks from natural language processing and computer vision. As such, improving its computational efficiency becomes paramount. One of the m…
Computational EfficiencyRead As Human: Compressing Context via Parallelizable Close Reading and Skimming
Large Language Models (LLMs) demonstrate exceptional capability across diverse tasks. However, their deployment in long-context scenarios is hindered by two challenges: computational inefficiency and redundant informatio…
Contrastive LearningQuestion AnsweringSkim-Attention: Learning to Focus via Document Layout
Transformer-based pre-training techniques of text and layout have proven effective in a number of document understanding tasks. Despite this success, multimodal pre-training models suffer from very high computational and…
document understandingLanguage ModelingLanguage ModellingQuALITY: Question Answering with Long Input Texts, Yes!
To enable building and testing models on long-document comprehension, we introduce QuALITY, a multiple-choice QA dataset with context passages in English that have an average length of about 5,000 tokens, much longer tha…
Multiple-choiceMultiple Choice Question Answering (MCQA)Question Answering