paper-with-me

홈 › Papers

Benchmarking noisy label detection methods

2025-10-17 · Henrique Pickler, Jorge K. S. Kamassury, Danilo Silva arxiv

Label noise is a common problem in real-world datasets, affecting both model training and validation. Clean data are essential for achieving strong performance and ensuring reliable evaluation. While various techniques have been proposed to detect noisy labels (or label errors), there is no clear consensus on optimal approaches. We perform a comprehensive benchmark of detection methods by decomposing them into three fundamental components: gathering strategy (in-sample vs out-of-sample), label disagreement measure, and aggregation method. This decomposition can be applied to many existing detection methods, and enables systematic comparison across diverse approaches. To fairly compare methods, we propose a unified benchmark task: detecting a fraction of training samples equal to the dataset's noise rate. We also introduce a novel metric: the false negative rate at this fixed operating point. Our evaluation spans vision and tabular datasets under both synthetic and real-world noise conditions. We identify that in-sample gathering using average probability aggregation combined with the logit margin as the label disagreement measure achieves the best results across most scenarios. Our findings provide practical guidance for designing new detection methods and selecting techniques for specific applications.

📄 PDF Abstract BibTeX arXiv:2510.16211

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective

2018-03-29 · CVPR 2018 6 · Jing Zhang, Tong Zhang, Yuchao Dai, Mehrtash Harandi 외

The success of current deep saliency detection methods heavily depends on the availability of large-scale supervision in the form of per-pixel labeling. Such supervision, while labor-intensive and not always possible, te…

BenchmarkingSaliency DetectionSaliency PredictionUnsupervised Saliency Detection

Benchmarking Instance-Dependent Label Noise with Controlled Corruptions

2026-06-12 · Shadman Islam, Agustinus Kristiadi, Mostafa Milani arxiv

Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically generate noise through imperfect annotators or classifier raters, leav…

AEON: Adaptive Estimation of Instance-Dependent In-Distribution and Out-of-Distribution Label Noise for Robust Learning

2025-01-23 · Arpit Garg, Cuong Nguyen, Rafael Felix, Yuyuan Liu 외

Robust training with noisy labels is a critical challenge in image classification, offering the potential to reduce reliance on costly clean-label datasets. Real-world datasets often contain a mix of in-distribution (ID)…

Benchmarkingimage-classificationImage Classification

Noisy Ostracods: A Fine-Grained, Imbalanced Real-World Dataset for Benchmarking Robust Machine Learning and Label Correction Methods

2024-12-03 · Jiamian Hu, Yuanyuan Hong, Yihua Chen, He Wang 외

We present the Noisy Ostracods, a noisy dataset for genus and species classification of crustacean ostracods with specialists' annotations. Over the 71466 specimens collected, 5.58% of them are estimated to be noisy (pos…

Benchmarking

Global Wheat Head Dataset 2021: more diversity to improve the benchmarking of wheat head localization methods

2021-05-17 · Etienne David, Mario Serouart, Daniel Smith, Simon Madec 외

The Global Wheat Head Detection (GWHD) dataset was created in 2020 and has assembled 193,634 labelled wheat heads from 4,700 RGB images acquired from various acquisition platforms and 7 countries/institutions. With an as…

BenchmarkingDiversityHead Detection