paper-with-me

Papers

Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the Truth

2020-05-01 · ICLR 2020 1 · Igor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei, Ardavan Saeedi, Yang Liu, Swami Sankaranarayanan, Tomer Gafner, Ben Sternlieb, Patrick Maher, Nathan Silberman

In most machine learning tasks unambiguous ground truth labels can easily be acquired. However, this luxury is often not afforded to many high-stakes, real-world scenarios such as medical image interpretation, where even expert human annotators typically exhibit very high levels of disagreement with one another. While prior works have focused on overcoming noisy labels during training, the question of how to evaluate models when annotators disagree about ground truth has remained largely unexplored. To address this, we propose the discrepancy ratio: a novel, task-independent and principled framework for validating machine learning models in the presence of high label noise. Conceptually, our approach evaluates a model by comparing its predictions to those of human annotators, taking into account the degree to which annotators disagree with one another. While our approach is entirely general, we show that in the special case of binary classification, our proposed metric can be evaluated in terms of simple, closed-form expressions that depend only on aggregate statistics of the labels and not on any individual label. Finally, we demonstrate how this framework can be used effectively to validate machine learning models using two real-world tasks from medical imaging. The discrepancy ratio metric reveals what conventional metrics do not: that our models not only vastly exceed the average human performance, but even exceed the performance of the best human experts in our datasets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningBinary Classification

Similar Papers 제목 키워드 기반

Domain Discrepancy Measure for Complex Models in Unsupervised Domain Adaptation

2019-01-30 · Jongyeong Lee, Nontawat Charoenphakdee, Seiichi Kuroki, Masashi Sugiyama

Appropriately evaluating the discrepancy between domains is essential for the success of unsupervised domain adaptation. In this paper, we first point out that existing discrepancy measures are less informative when comp…

Binary ClassificationClassificationDomain AdaptationGeneral Classification+2

Robustness May Be at Odds with Fairness: An Empirical Study on Class-wise Accuracy

2020-10-26 · Philipp Benz, Chaoning Zhang, Adil Karjauv, In So Kweon

Convolutional neural networks (CNNs) have made significant advancement, however, they are widely known to be vulnerable to adversarial attacks. Adversarial training is the most widely used technique for improving adversa…

Adversarial RobustnessAutonomous DrivingFairness

Re-Benchmarking Pool-Based Active Learning for Binary Classification

2023-06-15 · Po-Yi Lu, Chun-Liang Li, Hsuan-Tien Lin

Active learning is a paradigm that significantly enhances the performance of machine learning models when acquiring labeled data is expensive. While several benchmarks exist for evaluating active learning strategies, the…

Active LearningBenchmarkingBinary ClassificationClassification

Reformulate LLM Reinforcement Learning for Efficient Training under Black-box Discrepancy

2026-06-07 · Jiashun Liu, Runze Liu, Xu Wan, Jing Liang 외 arxiv

Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm, yet it frequently suffers from unpredictable sub-optimum performance or even training collapses. Recent findings attribute these failures to a …

Reinforcement Learning

Discrepancy Minimization Improves Cross-Hospital Robustness in Digital Pathology

2026-05-24 · Ben Vardi, Dana Schonberger, Yuval Friedmann, Zohar Yakhini 외 arxiv

Pathology foundation models (PFMs) have advanced rapidly in recent years and support training classifiers for a range of histopathology tasks. However, their robustness across hospitals remains limited: performance often…

Domain GeneralizationDomain Adaptation