paper-with-me

홈 › Papers

The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation

2021-09-14 · EMNLP 2021 11 · Marzena Karpinska, Nader Akoury, Mohit Iyyer

Recent text generation research has increasingly focused on open-ended domains such as story and poetry generation. Because models built for such tasks are difficult to evaluate automatically, most researchers in the space justify their modeling choices by collecting crowdsourced human judgments of text quality (e.g., Likert scores of coherence or grammaticality) from Amazon Mechanical Turk (AMT). In this paper, we first conduct a survey of 45 open-ended text generation papers and find that the vast majority of them fail to report crucial details about their AMT tasks, hindering reproducibility. We then run a series of story evaluation experiments with both AMT workers and English teachers and discover that even with strict qualification filters, AMT workers (unlike teachers) fail to distinguish between model-generated text and human-generated references. We show that AMT worker judgments improve when they are shown model-generated output alongside human-generated references, which enables the workers to better calibrate their ratings. Finally, interviews with the English teachers provide deeper insights into the challenges of the evaluation process, particularly when rating model-generated text.

📄 PDF Abstract BibTeX arXiv:2109.06835

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

EasyTurk: A User-Friendly Interface for High-Quality Linguistic Annotation with Amazon Mechanical Turk

2021-04-01 · EACL 2021 2 · Lorenzo Bocchi, Valentino Frasnelli, Alessio Palmero Aprosio

Amazon Mechanical Turk (AMT) has recently become one of the most popular crowd-sourcing platforms, allowing researchers from all over the world to create linguistic datasets quickly and at a relatively low cost. Amazon p…

Speak: A Toolkit Using Amazon Mechanical Turk to Collect and Validate Speech Audio Recordings

2022-06-01 · LREC 2022 6 · Christopher Song, David Harwath, Tuka Alhanai, James Glass

We present Speak, a toolkit that allows researchers to crowdsource speech audio recordings using Amazon Mechanical Turk (MTurk). Speak allows MTurk workers to submit speech recordings in response to a task prompt and sti…

Using Mechanical Turk to Build Machine Translation Evaluation Sets

2014-10-20 · Michael Bloodgood, Chris Callison-Burch

Building machine translation (MT) test sets is a relatively expensive task. As MT becomes increasingly desired for more and more language pairs and more and more domains, it becomes necessary to build test sets for each …

Machine TranslationTranslation

Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent

2017-11-21 · ICLR 2018 1 · Zhilin Yang, Saizheng Zhang, Jack Urbanek, Will Feng 외

Contrary to most natural language processing research, which makes use of static datasets, humans learn language interactively, grounded in an environment. In this work we propose an interactive learning procedure called…

Grounded language learning

A Gold Standard for Scalar Adjectives

2016-05-01 · LREC 2016 5 · Bryan Wilkinson, Oates Tim

We present a gold standard for evaluating scale membership and the order of scalar adjectives. In addition to evaluating existing methods of ordering adjectives, this knowledge will aid in studying the organization of ad…