paper-with-me

홈 › Papers

What is the ground truth? Reliability of multi-annotator data for audio tagging

2021-04-09 · Irene Martin-Morato, Annamaria Mesaros

Crowdsourcing has become a common approach for annotating large amounts of data. It has the advantage of harnessing a large workforce to produce large amounts of data in a short time, but comes with the disadvantage of employing non-expert annotators with different backgrounds. This raises the problem of data reliability, in addition to the general question of how to combine the opinions of multiple annotators in order to estimate the ground truth. This paper presents a study of the annotations and annotators' reliability for audio tagging. We adapt the use of Krippendorf's alpha and multi-annotator competence estimation (MACE) for a multi-labeled data scenario, and present how MACE can be used to estimate a candidate ground truth based on annotations from non-expert users with different levels of expertise and competence.

📄 PDF Abstract BibTeX arXiv:2104.04214

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Tagging

Similar Papers 제목 키워드 기반

Momresp: A Bayesian Model for Multi-Annotator Document Labeling

2014-05-01 · LREC 2014 5 · Paul Felt, Robbie Haertel, Eric Ringger, Kevin Seppi

Data annotation in modern practice often involves multiple, imperfect human annotators. Multiple annotations can be used to infer estimates of the ground-truth labels and to estimate individual annotator error characteri…

Document Classification

Learning Ambiguity from Crowd Sequential Annotations

2023-01-04 · Xiaolei Lu

Most crowdsourcing learning methods treat disagreement between annotators as noisy labelings while inter-disagreement among experts is often a good indicator for the ambiguity and uncertainty that is inherent in natural …

NERPOSPOS Tagging

Label Curation Using Agentic AI

2026-01-30 · Subhodeep Ghosh, Bayan Divaaniaazar, Md Ishat-E-Rabban, Spencer Clarke 외 arxiv

Data annotation is essential for supervised learning, yet producing accurate, unbiased, and scalable labels remains challenging as datasets grow in size and modality. Traditional human-centric pipelines are costly, slow,…

Beyond Black \& White: Leveraging Annotator Disagreement via Soft-Label Multi-Task Learning

2021-06-01 · NAACL 2021 4 · Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank 외

Supervised learning assumes that a ground truth label exists. However, the reliability of this ground truth depends on human annotators, who often disagree. Prior work has shown that this disagreement can be helpful in t…

Multi-Task Learning

Multi-Label Annotation Aggregation in Crowdsourcing

2017-06-19 · Xuan Wei, Daniel Dajun Zeng, Junming Yin

As a means of human-based computation, crowdsourcing has been widely used to annotate large-scale unlabeled datasets. One of the obvious challenges is how to aggregate these possibly noisy labels provided by a set of het…