paper-with-me

홈 › Papers

Making Binary Classification from Multiple Unlabeled Datasets Almost Free of Supervision

2023-06-12 · Yuhao Wu, Xiaobo Xia, Jun Yu, Bo Han, Gang Niu, Masashi Sugiyama, Tongliang Liu

Training a classifier exploiting a huge amount of supervised data is expensive or even prohibited in a situation, where the labeling cost is high. The remarkable progress in working with weaker forms of supervision is binary classification from multiple unlabeled datasets which requires the knowledge of exact class priors for all unlabeled datasets. However, the availability of class priors is restrictive in many real-world scenarios. To address this issue, we propose to solve a new problem setting, i.e., binary classification from multiple unlabeled datasets with only one pairwise numerical relationship of class priors (MU-OPPO), which knows the relative order (which unlabeled dataset has a higher proportion of positive examples) of two class-prior probabilities for two datasets among multiple unlabeled datasets. In MU-OPPO, we do not need the class priors for all unlabeled datasets, but we only require that there exists a pair of unlabeled datasets for which we know which unlabeled dataset has a larger class prior. Clearly, this form of supervision is easier to be obtained, which can make labeling costs almost free. We propose a novel framework to handle the MU-OPPO problem, which consists of four sequential modules: (i) pseudo label assignment; (ii) confident example collection; (iii) class prior estimation; (iv) classifier training with estimated class priors. Theoretically, we analyze the gap between estimated class priors and true class priors under the proposed framework. Empirically, we confirm the superiority of our framework with comprehensive experiments. Experimental results demonstrate that our framework brings smaller estimation errors of class priors and better performance of binary classification.

📄 PDF Abstract BibTeX arXiv:2306.07036

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationPseudo Label

Similar Papers 제목 키워드 기반

Binary Classification from Multiple Unlabeled Datasets via Surrogate Set Classification

2021-02-01 · Nan Lu, Shida Lei, Gang Niu, Issei Sato 외

To cope with high annotation costs, training a classifier only from weakly supervised data has attracted a great deal of attention these days. Among various approaches, strengthening supervision from completely unsupervi…

Binary ClassificationClassificationGeneral ClassificationMulti-class Classification

Integrating Distribution Matching into Semi-Supervised Contrastive Learning for Labeled and Unlabeled Data

2026-01-08 · Shogo Nakayama, Masahiro Okuda arxiv

The advancement of deep learning has greatly improved supervised image classification. However, labeling data is costly, prompting research into unsupervised learning methods such as contrastive learning. In real-world s…

Contrastive LearningImage Classification

Community-Based Hierarchical Positive-Unlabeled (PU) Model Fusion for Chronic Disease Prediction

2023-09-06 · Yang Wu, Xurui Li, Xuhong Zhang, Yangyang Kang 외

Positive-Unlabeled (PU) Learning is a challenge presented by binary classification problems where there is an abundance of unlabeled data along with a small number of positive data instances, which can be used to address…

Binary ClassificationData AugmentationDecision MakingDiabetes Prediction+1

Leveraging Labeled and Unlabeled Data for Consistent Fair Binary Classification

2019-06-12 · NeurIPS 2019 12 · Evgenii Chzhen, Christophe Denis, Mohamed Hebiri, Luca Oneto 외

We study the problem of fair binary classification using the notion of Equal Opportunity. It requires the true positive rate to distribute equally across the sensitive groups. Within this setting we show that the fair op…

Binary ClassificationClassificationFairnessGeneral Classification

Mitigating Overfitting in Supervised Classification from Two Unlabeled Datasets: A Consistent Risk Correction Approach

2019-10-20 · Nan Lu, Tianyi Zhang, Gang Niu, Masashi Sugiyama

The recently proposed unlabeled-unlabeled (UU) classification method allows us to train a binary classifier only from two unlabeled datasets with different class priors. Since this method is based on the empirical risk m…

ClassificationGeneral Classification