paper-with-me

Papers

Asking Crowdworkers to Write Entailment Examples: The Best of Bad Options

2020-10-13 · Asian Chapter of the Association for Computational Linguistics 2020 · Clara Vania, Ruijie Chen, Samuel R. Bowman

Large-scale natural language inference (NLI) datasets such as SNLI or MNLI have been created by asking crowdworkers to read a premise and write three new hypotheses, one for each possible semantic relationships (entailment, contradiction, and neutral). While this protocol has been used to create useful benchmark data, it remains unclear whether the writing-based annotation protocol is optimal for any purpose, since it has not been evaluated directly. Furthermore, there is ample evidence that crowdworker writing can introduce artifacts in the data. We investigate two alternative protocols which automatically create candidate (premise, hypothesis) pairs for annotators to label. Using these protocols and a writing-based baseline, we collect several new English NLI datasets of over 3k examples each, each using a fixed amount of annotator time, but a varying number of examples to fit that time budget. Our experiments on NLI and transfer learning show negative results: None of the alternative protocols outperforms the baseline in evaluations of generalization within NLI or on transfer to outside target tasks. We conclude that crowdworker writing still the best known option for entailment data, highlighting the need for further data collection work to focus on improving writing-based annotation processes.

📄 PDF Abstract BibTeX arXiv:2010.06122

Code (1)

nyu-mll/semi-automatic-nli 공식 구현

Tasks

Natural Language InferenceTransfer Learning

Similar Papers 제목 키워드 기반

What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?

2021-06-01 · ACL 2021 5 · Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt 외

Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding of language, there has been little focu…

Multiple-choiceNatural Language UnderstandingQuestion Answering

Fool Me Twice: Entailment from Wikipedia Gamification

2021-04-10 · NAACL 2021 4 · Julian Martin Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger 외

We release FoolMeTwice (FM2 for short), a large dataset of challenging entailment pairs collected through a fun multi-player game. Gamification encourages adversarial examples, drastically lowering the number of examples…

Retrieval

Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions

2022-05-01 · Mihir Parmar, Swaroop Mishra, Mor Geva, Chitta Baral

In recent years, progress in NLU has been driven by benchmarks. These benchmarks are typically collected by crowdsourcing, where annotators write examples based on annotation instructions crafted by dataset creators. In …

Label Verbalization and Entailment for Effective Zero and Few-Shot Relation Extraction

2021-11-01 · EMNLP 2021 11 · Oscar Sainz, Oier Lopez de Lacalle, Gorka Labaka, Ander Barrena 외

Relation extraction systems require large amounts of labeled examples which are costly to annotate. In this work we reformulate relation extraction as an entailment task, with simple, hand-made, verbalizations of relatio…

Natural Language InferenceRelationRelation Extraction

Label Verbalization and Entailment for Effective Zero- and Few-Shot Relation Extraction

2021-09-08 · Oscar Sainz, Oier Lopez de Lacalle, Gorka Labaka, Ander Barrena 외

Relation extraction systems require large amounts of labeled examples which are costly to annotate. In this work we reformulate relation extraction as an entailment task, with simple, hand-made, verbalizations of relatio…

Natural Language InferenceRelationRelation Extraction