paper-with-me

홈 › Papers

A Closer Look at Linguistic Knowledge in Masked Language Models: The Case of Relative Clauses in American English

2020-11-02 · COLING 2020 8 · Marius Mosbach, Stefania Degaetano-Ortlieb, Marie-Pauline Krielke, Badr M. Abdullah, Dietrich Klakow

Transformer-based language models achieve high performance on various tasks, but we still lack understanding of the kind of linguistic knowledge they learn and rely on. We evaluate three models (BERT, RoBERTa, and ALBERT), testing their grammatical and semantic knowledge by sentence-level probing, diagnostic cases, and masked prediction tasks. We focus on relative clauses (in American English) as a complex phenomenon needing contextual information and antecedent identification to be resolved. Based on a naturalistic dataset, probing shows that all three models indeed capture linguistic knowledge about grammaticality, achieving high performance. Evaluation on diagnostic cases and masked prediction tasks considering fine-grained linguistic knowledge, however, shows pronounced model-specific weaknesses especially on semantic knowledge, strongly impacting models' performance. Our results highlight the importance of (a)model comparison in evaluation task and (b) building up claims of model performance and the linguistic knowledge they capture beyond purely probing-based evaluations.

📄 PDF Abstract BibTeX arXiv:2011.00960

Code (1)

uds-lsv/rc-probing 공식 구현 pytorch

Tasks

DiagnosticSentence

Methods 이 논문이 사용한 방법론

American 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
WordPiece 설명 없음
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

Frustratingly Easy Edit-based Linguistic Steganography with a Masked Language Model

2021-04-20 · NAACL 2021 4 · Honai Ueoka, Yugo Murawaki, Sadao Kurohashi

With advances in neural language models, the focus of linguistic steganography has shifted from edit-based approaches to generation-based ones. While the latter's payload capacity is impressive, generating genuine-lookin…

Language ModelingLanguage ModellingLinguistic steganography

MaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining

2022-06-01 · Pengyuan Lyu, Chengquan Zhang, Shanshan Liu, Meina Qiao 외

Text images contain both visual and linguistic information. However, existing pre-training techniques for text recognition mainly focus on either visual representation learning or linguistic knowledge learning. In this p…

DecoderLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)+1

Learning Non-linguistic Skills without Sacrificing Linguistic Proficiency

2023-05-14 · Mandar Sharma, Nikhil Muralidhar, Naren Ramakrishnan

The field of Math-NLP has witnessed significant growth in recent years, motivated by the desire to expand LLM performance to the learning of non-linguistic notions (numerals, and subsequently, arithmetic reasoning). Howe…

Arithmetic ReasoningMath

SentiLARE: Sentiment-Aware Language Representation Learning with Linguistic Knowledge

2019-11-06 · EMNLP 2020 11 · Pei Ke, Haozhe Ji, Siyang Liu, Xiaoyan Zhu 외

Most of the existing pre-trained language representation models neglect to consider the linguistic knowledge of texts, which can promote language understanding in NLP tasks. To benefit the downstream tasks in sentiment a…

Data AugmentationLanguage ModelingLanguage ModellingRepresentation Learning+2

How Vision Affects Language: Comparing Masked Self-Attention in Uni-Modal and Multi-Modal Transformer

2021-06-01 · ACL (mmsr, IWCS) 2021 6 · Nikolai Ilinykh, Simon Dobnik

The problem of interpretation of knowledge learned by multi-head self-attention in transformers has been one of the central questions in NLP. However, a lot of work mainly focused on models trained for uni-modal tasks, e…

Image CaptioningMachine TranslationText GenerationTranslation