paper-with-me

홈 › Papers

Mitigating Overfitting in Supervised Classification from Two Unlabeled Datasets: A Consistent Risk Correction Approach

2019-10-20 · Nan Lu, Tianyi Zhang, Gang Niu, Masashi Sugiyama

The recently proposed unlabeled-unlabeled (UU) classification method allows us to train a binary classifier only from two unlabeled datasets with different class priors. Since this method is based on the empirical risk minimization, it works as if it is a supervised classification method, compatible with any model and optimizer. However, this method sometimes suffers from severe overfitting, which we would like to prevent in this paper. Our empirical finding in applying the original UU method is that overfitting often co-occurs with the empirical risk going negative, which is not legitimate. Therefore, we propose to wrap the terms that cause a negative empirical risk by certain correction functions. Then, we prove the consistency of the corrected risk estimator and derive an estimation error bound for the corrected risk minimizer. Experiments show that our proposal can successfully mitigate overfitting of the UU method and significantly improve the classification accuracy.

📄 PDF Abstract BibTeX arXiv:1910.08974

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral Classification

Similar Papers 제목 키워드 기반

Multi-class Classification from Multiple Unlabeled Datasets with Partial Risk Regularization

2022-07-04 · Yuting Tang, Nan Lu, Tianyi Zhang, Masashi Sugiyama

Recent years have witnessed a great success of supervised deep learning, where predictive models were trained from a large amount of fully labeled data. However, in practice, labeling such big data can be very costly and…

Multi-class Classification

ABC: Auxiliary Balanced Classifier for Class-imbalanced Semi-supervised Learning

2021-10-20 · NeurIPS 2021 12 · Hyuck Lee, Seungjae Shin, Heeyoung Kim

Existing semi-supervised learning (SSL) algorithms typically assume class-balanced datasets, although the class distributions of many real-world datasets are imbalanced. In general, classifiers trained on a class-imbalan…

Contrastive Self-supervised Learning for Graph Classification

2020-09-13 · Jiaqi Zeng, Pengtao Xie

Graph classification is a widely studied problem and has broad applications. In many real-world problems, the number of labeled graphs available for training classification models is limited, which renders these models p…

ClassificationData AugmentationGeneral ClassificationGraph Classification+1

Interpolation Consistency Training for Semi-Supervised Learning

2019-03-09 · Vikas Verma, Kenji Kawaguchi, Alex Lamb, Juho Kannala 외

We introduce Interpolation Consistency Training (ICT), a simple and computation efficient algorithm for training Deep Neural Networks in the semi-supervised learning paradigm. ICT encourages the prediction at an interpol…

General ClassificationSemi-Supervised Image Classification

Comparing effectiveness of regularization methods on text classification: Simple and complex model in data shortage situation

2024-02-27 · Jongga Lee, Jaeseung Yim, Seohee Park, Changwon Lim

Text classification is the task of assigning a document to a predefined class. However, it is expensive to acquire enough labeled documents or to label them. In this paper, we study the regularization methods' effects on…

text-classificationText Classification