paper-with-me

홈 › Papers

Knowledgeable Salient Span Mask for Enhancing Language Models as Knowledge Base

2022-04-17 · Cunxiang Wang, Fuli Luo, Yanyang Li, Runxin Xu, Fei Huang, Yue Zhang

Pre-trained language models (PLMs) like BERT have made significant progress in various downstream NLP tasks. However, by asking models to do cloze-style tests, recent work finds that PLMs are short in acquiring knowledge from unstructured text. To understand the internal behaviour of PLMs in retrieving knowledge, we first define knowledge-baring (K-B) tokens and knowledge-free (K-F) tokens for unstructured text and ask professional annotators to label some samples manually. Then, we find that PLMs are more likely to give wrong predictions on K-B tokens and attend less attention to those tokens inside the self-attention module. Based on these observations, we develop two solutions to help the model learn more knowledge from unstructured text in a fully self-supervised manner. Experiments on knowledge-intensive tasks show the effectiveness of the proposed methods. To our best knowledge, we are the first to explore fully self-supervised learning of knowledge in continual pre-training.

📄 PDF Abstract BibTeX arXiv:2204.07994

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Multi-Head Attention 설명 없음
Residual Connection 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…

Similar Papers 제목 키워드 기반

Salient Span Masking for Temporal Understanding

2023-03-22 · Jeremy R. Cole, Aditi Chaudhary, Bhuwan Dhingra, Partha Talukdar

Salient Span Masking (SSM) has shown itself to be an effective strategy to improve closed-book question answering performance. SSM extends general masked language model pretraining by creating additional unsupervised tra…

AvgLanguage ModelingLanguage ModellingQuestion Answering

Neural Knowledge Bank for Pretrained Transformers

2022-07-31 · Damai Dai, Wenbin Jiang, Qingxiu Dong, Yajuan Lyu 외

The ability of pretrained Transformers to remember factual knowledge is essential but still limited for existing models. Inspired by existing work that regards Feed-Forward Networks (FFNs) in Transformers as key-value me…

Language ModelingLanguage ModellingMachine TranslationQuestion Answering+1

The Effect of Masking Strategies on Knowledge Retention by Language Models

2023-06-12 · Jonas Wallat, Tianyi Zhang, Avishek Anand

Language models retain a significant amount of world knowledge from their pre-training stage. This allows knowledgeable models to be applied to knowledge-intensive tasks prevalent in information retrieval, such as rankin…

Information RetrievalQuestion AnsweringRetrievalWorld Knowledge

Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification

2021-11-16 · ACL ARR September 2021 9 · Anonymous

Tuning pre-trained language models (PLMs) with task-specific prompts has been a promising approach for text classification. Particularly, previous studies suggest that prompt-tuning has remarkable superiority in the low-…

Few-Shot Text ClassificationLanguage ModelingLanguage ModellingMasked Language Modeling+2

Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification

2021-08-04 · ACL 2022 5 · Shengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu 외

Tuning pre-trained language models (PLMs) with task-specific prompts has been a promising approach for text classification. Particularly, previous studies suggest that prompt-tuning has remarkable superiority in the low-…

ClassificationFew-Shot Text ClassificationLanguage ModelingLanguage Modelling+3