Embracing Error to Enable Rapid Crowdsourcing
Microtask crowdsourcing has enabled dataset advances in social science and machine learning, but existing crowdsourcing schemes are too expensive to scale up with the expanding volume of data. To scale and widen the applicability of crowdsourcing, we present a technique that produces extremely rapid judgments for binary and categorical labels. Rather than punishing all errors, which causes workers to proceed slowly and deliberately, our technique speeds up workers' judgments to the point where errors are acceptable and even expected. We demonstrate that it is possible to rectify these errors by randomizing task order and modeling response latency. We evaluate our technique on a breadth of common labeling tasks such as image verification, word similarity, sentiment analysis and topic classification. Where prior work typically achieves a 0.25x to 1x speedup over fixed majority vote, our approach often achieves an order of magnitude (10x) speedup.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationSentiment AnalysisTopic ClassificationWord SimilaritySimilar Papers 제목 키워드 기반
Embracing Ambiguity: A Comparison of Annotation Methodologies for Crowdsourcing Word Sense Labels
Designing LLM Chains by Adapting Techniques from Crowdsourcing Workflows
LLM chains enable complex tasks by decomposing work into a sequence of subtasks. Similarly, the more established techniques of crowdsourcing workflows decompose complex tasks into smaller tasks for human crowdworkers. Ch…
Regularized Minimax Conditional Entropy for Crowdsourcing
There is a rapidly increasing interest in crowdsourcing for data labeling. By crowdsourcing, a large number of labels can be often quickly gathered at low cost. However, the labels provided by the crowdsourcing workers a…
Rapid Development of a Corpus with Discourse Annotations using Two-stage Crowdsourcing
Embracing Federated Learning: Enabling Weak Client Participation via Partial Model Training
In Federated Learning (FL), clients may have weak devices that cannot train the full model or even hold it in their memory space. To implement large-scale FL applications, thus, it is crucial to develop a distributed lea…
Federated Learning