paper-with-me

홈 › Papers

Understanding Task Design Trade-offs in Crowdsourced Paraphrase Collection

2017-04-19 · ACL 2017 7 · Youxuan Jiang, Jonathan K. Kummerfeld, Walter S. Lasecki

Linguistically diverse datasets are critical for training and evaluating robust machine learning systems, but data collection is a costly process that often requires experts. Crowdsourcing the process of paraphrase generation is an effective means of expanding natural language datasets, but there has been limited analysis of the trade-offs that arise when designing tasks. In this paper, we present the first systematic study of the key factors in crowdsourcing paraphrase collection. We consider variations in instructions, incentives, data domains, and workflows. We manually analyzed paraphrases for correctness, grammaticality, and linguistic diversity. Our observations provide new insight into the trade-offs between accuracy and diversity in crowd responses that arise as a result of task design, providing guidance for future paraphrase generation procedures.

📄 PDF Abstract BibTeX arXiv:1704.05753

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityParaphrase Generation

Similar Papers 제목 키워드 기반

Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support

2025-02-25 · Kevin Pu, Daniel Lazaro, Ian Arawjo, Haijun Xia 외

AI programming tools enable powerful code generation, and recent prototypes attempt to reduce user effort with proactive AI agents, but their impact on programming workflows remains unexplored. We introduce and evaluate …

Code Generation

FACT: A Diagnostic for Group Fairness Trade-offs

2020-04-07 · Joon Sik Kim, Jiahao Chen, Ameet Talwalkar

Group fairness, a class of fairness notions that measure how different groups of individuals are treated differently according to their protected attributes, has been shown to conflict with one another, often with a nece…

AttributeDiagnosticFairness

Trade-offs and Guarantees of Adversarial Representation Learning for Information Obfuscation

2019-06-19 · NeurIPS 2020 12 · Han Zhao, Jianfeng Chi, Yuan Tian, Geoffrey J. Gordon

Crowdsourced data used in machine learning services might carry sensitive information about attributes that users do not want to share. Various methods have been proposed to minimize the potential information leakage of …

AttributeInference AttackRepresentation Learning

Trade-offs in Image Generation: How Do Different Dimensions Interact?

2025-07-29 · Sicheng Zhang, Binzhu Xie, Zhonghao Yan, Yuli Zhang 외 arxiv

Model performance in text-to-image (T2I) and image-to-image (I2I) generation often depends on multiple aspects, including quality, alignment, diversity, and robustness. However, models' complex trade-offs among these dim…

Image Generation

Understanding the Impact of Experiment Design for Evaluating Dialogue System Output

2020-07-01 · WS 2020 7 · Sashank Santhanam, Samira Shaikh

Evaluation of output from natural language generation (NLG) systems is typically conducted via crowdsourced human judgments. To understand the impact of how experiment design might affect the quality and consistency of s…

Text Generation