A Closer Look at Linguistic Knowledge in Masked Language Models: The Case of Relative Clauses in American English
Transformer-based language models achieve high performance on various tasks, but we still lack understanding of the kind of linguistic knowledge they learn and rely on. We evaluate three models (BERT, RoBERTa, and ALBERT), testing their grammatical and semantic knowledge by sentence-level probing, diagnostic cases, and masked prediction tasks. We focus on relative clauses (in American English) as a complex phenomenon needing contextual information and antecedent identification to be resolved. Based on a naturalistic dataset, probing shows that all three models indeed capture linguistic knowledge about grammaticality, achieving high performance. Evaluation on diagnostic cases and masked prediction tasks considering fine-grained linguistic knowledge, however, shows pronounced model-specific weaknesses especially on semantic knowledge, strongly impacting models' performance. Our results highlight the importance of (a)model comparison in evaluation task and (b) building up claims of model performance and the linguistic knowledge they capture beyond purely probing-based evaluations.
Code (1)
Tasks
DiagnosticSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Frustratingly Easy Edit-based Linguistic Steganography with a Masked Language Model
With advances in neural language models, the focus of linguistic steganography has shifted from edit-based approaches to generation-based ones. While the latter's payload capacity is impressive, generating genuine-lookin…
Language ModelingLanguage ModellingLinguistic steganographyMaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining
Text images contain both visual and linguistic information. However, existing pre-training techniques for text recognition mainly focus on either visual representation learning or linguistic knowledge learning. In this p…
DecoderLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)+1Learning Non-linguistic Skills without Sacrificing Linguistic Proficiency
The field of Math-NLP has witnessed significant growth in recent years, motivated by the desire to expand LLM performance to the learning of non-linguistic notions (numerals, and subsequently, arithmetic reasoning). Howe…
Arithmetic ReasoningMathSentiLARE: Sentiment-Aware Language Representation Learning with Linguistic Knowledge
Most of the existing pre-trained language representation models neglect to consider the linguistic knowledge of texts, which can promote language understanding in NLP tasks. To benefit the downstream tasks in sentiment a…
Data AugmentationLanguage ModelingLanguage ModellingRepresentation Learning+2How Vision Affects Language: Comparing Masked Self-Attention in Uni-Modal and Multi-Modal Transformer
The problem of interpretation of knowledge learned by multi-head self-attention in transformers has been one of the central questions in NLP. However, a lot of work mainly focused on models trained for uni-modal tasks, e…
Image CaptioningMachine TranslationText GenerationTranslation