paper-with-me

Papers

Crowdsourcing Beyond Annotation: Case Studies in Benchmark Data Collection

2021-11-01 · EMNLP (ACL) 2021 11 · Alane Suhr, Clara Vania, Nikita Nangia, Maarten Sap, Mark Yatskar, Samuel R. Bowman, Yoav Artzi

Crowdsourcing from non-experts is one of the most common approaches to collecting data and annotations in NLP. Even though it is such a fundamental tool in NLP, crowdsourcing use is largely guided by common practices and the personal experience of researchers. Developing a theory of crowdsourcing use for practical language problems remains an open challenge. However, there are various principles and practices that have proven effective in generating high quality and diverse data. This tutorial exposes NLP researchers to such data collection crowdsourcing methods and principles through a detailed discussion of a diverse set of case studies. The selection of case studies focuses on challenging settings where crowdworkers are asked to write original text or otherwise perform relatively unconstrained work. Through these case studies, we discuss in detail processes that were carefully designed to achieve data with specific properties, for example to require logical inference, grounded reasoning or conversational understanding. Each case study focuses on data collection crowdsourcing protocol details that often receive limited attention in research presentations, for example in conferences, but are critical for research success.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity Recognition

2021-05-31 · ACL 2021 5 · Xin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang 외

Crowdsourcing is regarded as one prospective solution for effective supervised learning, aiming to build large-scale annotated training data by crowd workers. Previous studies focus on reducing the influences from the no…

Domain Adaptationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2

Annotating omission in statement pairs

2017-04-01 · WS 2017 4 · H{\'e}ctor Mart{\'\i}nez Alonso, Amaury Delamaire, Beno{\^\i}t Sagot

We focus on the identification of omission in statement pairs. We compare three annotation schemes, namely two different crowdsourcing schemes and manual expert annotation. We show that the simplest of the two crowdsourc…

Natural Language Inference

CoRefi: A Crowd Sourcing Suite for Coreference Annotation

2020-10-15 · NeurIPS Workshop HAMLETS 2020 12 · Anonymous

Coreference annotation is an important, yet expensive and time consuming, task, which often involved expert annotators trained on complex decision guidelines. To enable cheaper and more efficient annotation, we present C…

A Comparative Study on Annotation Quality of Crowdsourcing and LLM via Label Aggregation

2024-01-18 · Jiyi Li

Whether Large Language Models (LLMs) can outperform crowdsourcing on the data annotation task is attracting interest recently. Some works verified this issue with the average performance of individual crowd workers and L…

Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions

2022-05-01 · Mihir Parmar, Swaroop Mishra, Mor Geva, Chitta Baral

In recent years, progress in NLU has been driven by benchmarks. These benchmarks are typically collected by crowdsourcing, where annotators write examples based on annotation instructions crafted by dataset creators. In …