paper-with-me

홈 › Papers

OCNLI: Original Chinese Natural Language Inference

2020-10-12 · Findings of the Association for Computational Linguistics 2020 · Hai Hu, Kyle Richardson, Liang Xu, Lu Li, Sandra Kuebler, Lawrence S. Moss

Despite the tremendous recent progress on natural language inference (NLI), driven largely by large-scale investment in new datasets (e.g., SNLI, MNLI) and advances in modeling, most progress has been limited to English due to a lack of reliable datasets for most of the world's languages. In this paper, we present the first large-scale NLI dataset (consisting of ~56,000 annotated sentence pairs) for Chinese called the Original Chinese Natural Language Inference dataset (OCNLI). Unlike recent attempts at extending NLI to other languages, our dataset does not rely on any automatic translation or non-expert annotation. Instead, we elicit annotations from native speakers specializing in linguistics. We follow closely the annotation protocol used for MNLI, but create new strategies for eliciting diverse hypotheses. We establish several baseline results on our dataset using state-of-the-art pre-trained models for Chinese, and find even the best performing models to be far outpaced by human performance (~12% absolute performance gap), making it a challenging new resource that we hope will help to accelerate progress in Chinese NLU. To the best of our knowledge, this is the first human-elicited MNLI-style corpus for a non-English language.

📄 PDF Abstract BibTeX arXiv:2010.05444

Code (1)

CLUEbenchmark/OCNLI 공식 구현 tf

Tasks

Natural Language InferenceSentenceTranslation

Similar Papers 제목 키워드 기반

Investigating Transfer Learning in Multilingual Pre-trained Language Models through Chinese Natural Language Inference

2021-06-07 · Findings (ACL) 2021 8 · Hai Hu, He Zhou, Zuoyu Tian, Yiwen Zhang 외

Multilingual transformers (XLM, mT5) have been shown to have remarkable transfer skills in zero-shot settings. Most transfer studies, however, rely on automatically translated resources (XNLI, XQuAD), making it hard to d…

Cross-Lingual TransferNatural Language InferenceTransfer LearningXLM-R

DocNLI: A Large-scale Dataset for Document-level Natural Language Inference

2021-06-17 · Findings (ACL) 2021 8 · Wenpeng Yin, Dragomir Radev, Caiming Xiong

Natural language inference (NLI) is formulated as a unified framework for solving various NLP problems such as relation extraction, question answering, summarization, etc. It has been studied intensively in the past few …

Natural Language InferenceQuestion AnsweringRelation ExtractionSentence

R$^2$F: A General Retrieval, Reading and Fusion Framework for Document-level Natural Language Inference

2022-10-22 · Hao Wang, Yixin Cao, Yangguang Li, Zhen Huang 외

Document-level natural language inference (DOCNLI) is a new challenging task in natural language processing, aiming at judging the entailment relationship between a pair of hypothesis and premise documents. Current datas…

Natural Language InferenceRetrievalSentence

UnNatural Language Inference

2020-12-30 · ACL 2021 5 · Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, Adina Williams

Recent investigations into the inner-workings of state-of-the-art large-scale pre-trained Transformer-based Natural Language Understanding (NLU) models indicate that they appear to know humanlike syntax, at least to some…

Natural Language InferenceNatural Language Understanding

CLUE: A Chinese Language Understanding Evaluation Benchmark

2020-04-13 · COLING 2020 8 · Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li 외

The advent of natural language understanding (NLU) benchmarks for English, such as GLUE and SuperGLUE allows new NLU models to be evaluated across a diverse set of tasks. These comprehensive benchmarks have facilitated a…

General ClassificationMachine Reading ComprehensionNatural Language UnderstandingReading Comprehension+3