paper-with-me

Papers

Semi-Crowdsourced Clustering: Generalizing Crowd Labeling by Robust Distance Metric Learning

2012-12-01 · NeurIPS 2012 12 · Jinfeng Yi, Rong Jin, Shaili Jain, Tianbao Yang, Anil K. Jain

One of the main challenges in data clustering is to define an appropriate similarity measure between two objects. Crowdclustering addresses this challenge by defining the pairwise similarity based on the manual annotations obtained through crowdsourcing. Despite its encouraging results, a key limitation of crowdclustering is that it can only cluster objects when their manual annotations are available. To address this limitation, we propose a new approach for clustering, called \textit{semi-crowdsourced clustering} that effectively combines the low-level features of objects with the manual annotations of a subset of the objects obtained via crowdsourcing. The key idea is to learn an appropriate similarity measure, based on the low-level features of objects, from the manual annotations of only a small portion of the data to be clustered. One difficulty in learning the pairwise similarity measure is that there is a significant amount of noise and inter-worker variations in the manual annotations obtained via crowdsourcing. We address this difficulty by developing a metric learning algorithm based on the matrix completion method. Our empirical study with two real-world image data sets shows that the proposed algorithm outperforms state-of-the-art distance metric learning algorithms in both clustering accuracy and computational efficiency.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringComputational EfficiencyMatrix CompletionMetric Learning

Similar Papers 제목 키워드 기반

Inducing Script Structure from Crowdsourced Event Descriptions via Semi-Supervised Clustering

2017-04-01 · WS 2017 4 · Lilian Wanzare, Aless Zarcone, ra, Stefan Thater 외

We present a semi-supervised clustering approach to induce script structure from crowdsourced descriptions of event sequences by grouping event descriptions into paraphrase sets (representing event types) and inducing th…

ClusteringQuestion AnsweringSemantic Role Labeling

Crowdsourced Labeling for Worker-Task Specialization Model

2020-03-21 · Do-Yeon Kim, Hye Won Chung

We consider crowdsourced labeling under a $d$-type worker-task specialization model, where each worker and task is associated with one particular type among a finite set of types and a worker provides a more reliable ans…

ClusteringmodelVocal Bursts Type Prediction

Semi-crowdsourced Clustering with Deep Generative Models

2018-10-29 · NeurIPS 2018 12 · Yucen Luo, Tian Tian, Jiaxin Shi, Jun Zhu 외

We consider the semi-supervised clustering problem where crowdsourcing provides noisy information about the pairwise comparisons on a small subset of data, i.e., whether a sample pair is in the same cluster. We propose a…

ClusteringVariational Inference

Bayesian Crowdsourcing with Constraints

2020-12-20 · Panagiotis A. Traganitis, Georgios B. Giannakis

Crowdsourcing has emerged as a powerful paradigm for efficiently labeling large datasets and performing various learning tasks, by leveraging crowds of human annotators. When additional information is available about the…

Variational Inference

Crowd-Powered Data Mining

2018-06-13 · Chengliang Chai, Ju Fan, Guoliang Li, Jiannan Wang 외

Many data mining tasks cannot be completely addressed by auto- mated processes, such as sentiment analysis and image classification. Crowdsourcing is an effective way to harness the human cognitive ability to process the…

ClusteringGeneral Classificationimage-classificationImage Classification+2