paper-with-me

홈 › Papers

Evaluating multiple models using labeled and unlabeled data

2025-01-21 · Divya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John Guttag, Bonnie Berger, Emma Pierson

It remains difficult to evaluate machine learning classifiers in the absence of a large, labeled dataset. While labeled data can be prohibitively expensive or impossible to obtain, unlabeled data is plentiful. Here, we introduce Semi-Supervised Model Evaluation (SSME), a method that uses both labeled and unlabeled data to evaluate machine learning classifiers. SSME is the first evaluation method to take advantage of the fact that: (i) there are frequently multiple classifiers for the same task, (ii) continuous classifier scores are often available for all classes, and (iii) unlabeled data is often far more plentiful than labeled data. The key idea is to use a semi-supervised mixture model to estimate the joint distribution of ground truth labels and classifier predictions. We can then use this model to estimate any metric that is a function of classifier scores and ground truth labels (e.g., accuracy or expected calibration error). We present experiments in four domains where obtaining large labeled datasets is often impractical: (1) healthcare, (2) content moderation, (3) molecular property prediction, and (4) image annotation. Our results demonstrate that SSME estimates performance more accurately than do competing methods, reducing error by 5.1x relative to using labeled data alone and 2.4x relative to the next best competing method. SSME also improves accuracy when evaluating performance across subsets of the test distribution (e.g., specific demographic subgroups) and when evaluating the performance of language models.

📄 PDF Abstract BibTeX arXiv:2501.11866

Code (0)

등록된 구현이 없습니다.

Tasks

Molecular Property PredictionProperty Prediction

Similar Papers 제목 키워드 기반

Evaluating the fairness of task-adaptive pretraining on unlabeled test data before few-shot text classification

2024-09-30 · Kush Dubey

Few-shot learning benchmarks are critical for evaluating modern NLP techniques. It is possible, however, that benchmarks favor methods which easily make use of unlabeled text, because researchers can use unlabeled text f…

FairnessFew-Shot LearningFew-Shot Text Classificationtext-classification+1

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

2026-08-21 · Tsao-Lun Chen, Chi-Cheng Fu, Han-Yi E. Chou, Shun-Feng Su arxiv

Pseudo-label-based semi-supervised learning has achieved strong performance due to its simplicity and scalability. However, it is typically developed under a closed-world assumption that unlabeled data are drawn from the…

Re-Evaluating the Impact of Unseen-Class Unlabeled Data on Semi-Supervised Learning Model

2025-03-02 · Rundong He, Yicong Dong, LanZhe Guo, Yilong Yin 외

Semi-supervised learning (SSL) effectively leverages unlabeled data and has been proven successful across various fields. Current safe SSL methods believe that unseen classes in unlabeled data harm the performance of SSL…

Unsupervised Full Constituency Parsing with Neighboring Distribution Divergence

2021-10-29 · Letian Peng, Zuchao Li, Hai Zhao

Unsupervised constituency parsing has been explored much but is still far from being solved. Conventional unsupervised constituency parser is only able to capture the unlabeled structure of sentences. Towards unsupervise…

Constituency ParsingPOSSemantic SimilaritySemantic Textual Similarity

Adaptively Unified Semi-Supervised Dictionary Learning With Active Points

2015-12-01 · ICCV 2015 12 · Xiaobo Wang, Xiaojie Guo, Stan Z. Li

Semi-supervised dictionary learning aims to construct a dictionary by utilizing both labeled and unlabeled data. To enhance the discriminative capability of the learned dictionary, numerous discriminative terms have been…

Dictionary Learning