Understanding Task Design Trade-offs in Crowdsourced Paraphrase Collection
Linguistically diverse datasets are critical for training and evaluating robust machine learning systems, but data collection is a costly process that often requires experts. Crowdsourcing the process of paraphrase generation is an effective means of expanding natural language datasets, but there has been limited analysis of the trade-offs that arise when designing tasks. In this paper, we present the first systematic study of the key factors in crowdsourcing paraphrase collection. We consider variations in instructions, incentives, data domains, and workflows. We manually analyzed paraphrases for correctness, grammaticality, and linguistic diversity. Our observations provide new insight into the trade-offs between accuracy and diversity in crowd responses that arise as a result of task design, providing guidance for future paraphrase generation procedures.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityParaphrase GenerationSimilar Papers 제목 키워드 기반
Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support
AI programming tools enable powerful code generation, and recent prototypes attempt to reduce user effort with proactive AI agents, but their impact on programming workflows remains unexplored. We introduce and evaluate …
Code GenerationFACT: A Diagnostic for Group Fairness Trade-offs
Group fairness, a class of fairness notions that measure how different groups of individuals are treated differently according to their protected attributes, has been shown to conflict with one another, often with a nece…
AttributeDiagnosticFairnessTrade-offs and Guarantees of Adversarial Representation Learning for Information Obfuscation
Crowdsourced data used in machine learning services might carry sensitive information about attributes that users do not want to share. Various methods have been proposed to minimize the potential information leakage of …
AttributeInference AttackRepresentation LearningTrade-offs in Image Generation: How Do Different Dimensions Interact?
Model performance in text-to-image (T2I) and image-to-image (I2I) generation often depends on multiple aspects, including quality, alignment, diversity, and robustness. However, models' complex trade-offs among these dim…
Image GenerationUnderstanding the Impact of Experiment Design for Evaluating Dialogue System Output
Evaluation of output from natural language generation (NLG) systems is typically conducted via crowdsourced human judgments. To understand the impact of how experiment design might affect the quality and consistency of s…
Text Generation