paper-with-me

홈 › Papers

Curriculum: A Broad-Coverage Benchmark for Linguistic Phenomena in Natural Language Understanding

2022-01-16 · ACL ARR January 2022 1 · Anonymous

In the age of large transformer language models, linguistic benchmarks play an important role in diagnosing models' abilities and limitations on natural language understanding. However, current benchmarks show some significant shortcomings. In particular, they do not provide insight into how well a language model captures distinct linguistic phenomena essential for language understanding and reasoning. In this paper, we introduce Curriculum, a new large-scale NLI benchmark for evaluation on broad-coverage linguistic phenomena. We show that our benchmark for linguistic phenomena serves as a more difficult challenge for current state-of-the-art models. Our experiments also provide insight into the limitation of existing benchmark datasets. In addition, we find that sequential training on selected linguistic phenomena effectively improves generalizing performance on adversarial NLI under limited training examples.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingNatural Language Understanding

Similar Papers 제목 키워드 기반

Curriculum: A Broad-Coverage Benchmark for Linguistic Phenomena in Natural Language Understanding

2022-04-13 · NAACL 2022 7 · Zeming Chen, Qiyue Gao

In the age of large transformer language models, linguistic evaluation play an important role in diagnosing models' abilities and limitations on natural language understanding. However, current evaluation methods show so…

Language ModelingLanguage ModellingNatural Language Understanding

BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources

2026-04-20 · Raghvendra Kumar, Devankar Raj, Sriparna Saha arxiv

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates r…

Domain Generalization

Interpretability of Language Models via Task Spaces

2024-06-10 · Lucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke Hupkes

The usual way to interpret language models (LMs) is to test their performance on different benchmarks and subsequently infer their internal processes. In this paper, we present an alternative approach, concentrating on t…

Linguistic Analysis of Pretrained Sentence Encoders with Acceptability Judgments

2019-01-11 · Alex Warstadt, Samuel R. Bowman

Recent work on evaluating grammatical knowledge in pretrained sentence encoders gives a fine-grained view of a small number of phenomena. We introduce a new analysis dataset that also has broad coverage of linguistic phe…

CoLAGeneral ClassificationLinguistic AcceptabilitySentence

ZhoBLiMP: a Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese

2024-11-09 · Yikang Liu, Yeting Shen, Hongao Zhu, Lilong Xu 외

Whether and how language models (LMs) acquire the syntax of natural languages has been widely evaluated under the minimal pair paradigm. However, a lack of wide-coverage benchmarks in languages other than English has con…

Language Acquisition