paper-with-me

홈 › Papers

Toward Annotator Group Bias in Crowdsourcing

2021-10-08 · ACL 2022 5 · Haochen Liu, Joseph Thekinen, Sinem Mollaoglu, Da Tang, Ji Yang, Youlong Cheng, Hui Liu, Jiliang Tang

Crowdsourcing has emerged as a popular approach for collecting annotated data to train supervised machine learning models. However, annotator bias can lead to defective annotations. Though there are a few works investigating individual annotator bias, the group effects in annotators are largely overlooked. In this work, we reveal that annotators within the same demographic group tend to show consistent group bias in annotation tasks and thus we conduct an initial study on annotator group bias. We first empirically verify the existence of annotator group bias in various real-world crowdsourcing datasets. Then, we develop a novel probabilistic graphical framework GroupAnno to capture annotator group bias with a new extended Expectation Maximization (EM) training algorithm. We conduct experiments on both synthetic and real-world datasets. Experimental results demonstrate the effectiveness of our model in modeling annotator group bias in label aggregation and model learning over competitive baselines.

📄 PDF Abstract BibTeX arXiv:2110.08038

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluation of Summarization Systems across Gender, Age, and Race

2021-10-08 · EMNLP (newsum) 2021 11 · Anna Jørgensen, Anders Søgaard

Summarization systems are ultimately evaluated by human annotators and raters. Usually, annotators and raters do not reflect the demographics of end users, but are recruited through student populations or crowdsourcing p…

False Discovery Rate Control and Statistical Quality Assessment of Annotators in Crowdsourced Ranking

2016-05-19 · Qianqian Xu, Jiechao Xiong, Xiaochun Cao, Yuan YAO

With the rapid growth of crowdsourcing platforms it has become easy and relatively inexpensive to collect a dataset labeled by multiple annotators in a short time. However due to the lack of control over the quality of t…

PositionSociology

Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions

2022-05-01 · Mihir Parmar, Swaroop Mishra, Mor Geva, Chitta Baral

In recent years, progress in NLU has been driven by benchmarks. These benchmarks are typically collected by crowdsourcing, where annotators write examples based on annotation instructions crafted by dataset creators. In …

Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets

2019-08-21 · IJCNLP 2019 11 · Mor Geva, Yoav Goldberg, Jonathan Berant

Crowdsourcing has been the prevalent paradigm for creating natural language understanding datasets in recent years. A common crowdsourcing practice is to recruit a small number of high-quality workers, and have them mass…

DiversityNatural Language Understanding

Towards A Reliable Ground-Truth For Biased Language Detection

2021-12-14 · Timo Spinde, David Krieger, Manuel Plank, Bela Gipp

Reference texts such as encyclopedias and news articles can manifest biased language when objective reporting is substituted by subjective writing. Existing methods to detect bias mostly rely on annotated data to train m…

ArticlesBias Detection