paper-with-me

홈 › Papers

OpinionRank: Extracting Ground Truth Labels from Unreliable Expert Opinions with Graph-Based Spectral Ranking

2021-02-11 · Glenn Dawson, Robi Polikar

As larger and more comprehensive datasets become standard in contemporary machine learning, it becomes increasingly more difficult to obtain reliable, trustworthy label information with which to train sophisticated models. To address this problem, crowdsourcing has emerged as a popular, inexpensive, and efficient data mining solution for performing distributed label collection. However, crowdsourced annotations are inherently untrustworthy, as the labels are provided by anonymous volunteers who may have varying, unreliable expertise. Worse yet, some participants on commonly used platforms such as Amazon Mechanical Turk may be adversarial, and provide intentionally incorrect label information without the end user's knowledge. We discuss three conventional models of the label generation process, describing their parameterizations and the model-based approaches used to solve them. We then propose OpinionRank, a model-free, interpretable, graph-based spectral algorithm for integrating crowdsourced annotations into reliable labels for performing supervised or semi-supervised learning. Our experiments show that OpinionRank performs favorably when compared against more highly parameterized algorithms. We also show that OpinionRank is scalable to very large datasets and numbers of label sources, and requires considerably fewer computational resources than previous approaches.

📄 PDF Abstract BibTeX arXiv:2102.05884

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gray Learning from Non-IID Data with Out-of-distribution Samples

2022-06-19 · Zhilin Zhao, Longbing Cao, Chang-Dong Wang

The integrity of training data, even when annotated by experts, is far from guaranteed, especially for non-IID datasets comprising both in- and out-of-distribution samples. In an ideal scenario, the majority of samples w…

Learning Theory

Robust Representation Learning for Unreliable Partial Label Learning

2023-08-31 · Yu Shi, Dong-Dong Wu, Xin Geng, Min-Ling Zhang

Partial Label Learning (PLL) is a type of weakly supervised learning where each training instance is assigned a set of candidate labels, but only one label is the ground-truth. However, this idealistic assumption may not…

Contrastive LearningPartial Label LearningRepresentation LearningWeakly-supervised Learning

Source-free domain adaptation based on label reliability for cross-domain bearing fault diagnosis

2025-03-11 · Wenyi Wu, Hao Zhang, Zhisen Wei, Xiao-Yuan Jing 외

Source-free domain adaptation (SFDA) has been exploited for cross-domain bearing fault diagnosis without access to source data. Current methods select partial target samples with reliable pseudo-labels for model adaptati…

Data AugmentationDomain AdaptationFault DiagnosisSource-Free Domain Adaptation

Unreliable Partial Label Learning with Recursive Separation

2023-02-20 · Yu Shi, Ning Xu, Hua Yuan, Xin Geng

Partial label learning (PLL) is a typical weakly supervised learning problem in which each instance is associated with a candidate label set, and among which only one is true. However, the assumption that the ground-trut…

Partial Label LearningWeakly-supervised Learning

Twitter Language Identification Of Similar Languages And Dialects Without Ground Truth

2017-04-01 · WS 2017 4 · Jennifer Williams, Charlie Dagli

We present a new method to bootstrap filter Twitter language ID labels in our dataset for automatic language identification (LID). Our method combines geo-location, original Twitter LID labels, and Amazon Mechanical Turk…

General ClassificationLanguage IdentificationSentiment Analysis