paper-with-me

Papers Label Error Detection

“Label Error Detection” 태그가 달린 논문 21편 · 필터 해제

Data filtering methods for training language models

2026-05-28 · Egor Shevchenko, Elena Bruches arxiv

Data quality is a critical factor in the effectiveness of machine learning models. Label errors, present even in widely used benchmarks, introduce noise into training data and reduce model generalization. In this work, w…

Linguistic AcceptabilityEmotion ClassificationLabel Error DetectionText Classification

Detecting and refurbishing ground truth errors during training of deep learning-based echocardiography segmentation models

2026-04-14 · Iman Islam, Bram Ruijsink, Andrew J. Reader, Andrew P. King arxiv

Deep learning-based medical image segmentation typically relies on ground truth (GT) labels obtained through manual annotation, but these can be prone to random errors or systematic biases. This study examines the robust…

Medical Image SegmentationLabel Error Detection

A Human-in-the-Loop Label Error Detection Framework Applied to Arabic-Script HTR Datasets

2026-01-23 · Sana Al-azzawi, Elisa Barney, Marcus Liwicki arxiv

Despite recent advances, Handwritten Text Recognition (HTR) for Arabic-script languages still lags behind Latin-script HTR. Part of the problem is dataset quality. To help closing this gap, we propose a two-stage framewo…

Handwritten Text RecognitionLabel Error Detection

Adaptive Label Error Detection: A Bayesian Approach to Mislabeled Data Detection

2026-01-15 · Zan Chaudhry, Noam H. Rotenberg, Brian Caffo, Craig K. Jones 외 arxiv

Machine learning classification systems are susceptible to poor performance when trained with incorrect ground truth labels, even when data is well-curated by expert annotators. As machine learning becomes more widesprea…

Label Error Detection

Hard Samples, Bad Labels: Robust Loss Functions That Know When to Back Off

2025-11-20 · Nicholas Pellegrino, David Szczecina, Paul Fieguth arxiv

Incorrectly labelled training data are frustratingly ubiquitous in both benchmark and specially curated datasets. Such mislabelling clearly adversely affects the performance and generalizability of models trained through…

Label Error Detection

Towards Cross-Modal Error Detection with Tables and Images

2025-10-14 · Olga Ovcharenko, Sebastian Schelter arxiv

Ensuring data quality at scale remains a persistent challenge for large organizations. Despite recent advances, maintaining accurate and consistent data is still complex, especially when dealing with multiple data modali…

Label Error Detection

Learning to Detect Label Errors by Making Them: A Method for Segmentation and Object Detection Datasets

2025-08-25 · Sarina Penquitt, Tobias Riedlinger, Timo Heller, Markus Reischl 외 arxiv

Recently, detection of label errors and improvement of label quality in datasets for supervised learning tasks has become an increasingly important goal in both research and industry. The consequences of incorrectly anno…

Semantic SegmentationInstance SegmentationLabel Error DetectionObject Detection

From Label Error Detection to Correction: A Modular Framework and Benchmark for Object Detection Datasets

2025-08-06 · Sarina Penquitt, Jonathan Klees, Rinor Cakaj, Daniel Kondermann 외 arxiv

Object detection has advanced rapidly in recent years, driven by increasingly large and diverse datasets. However, label errors often compromise the quality of these datasets and affect the outcomes of training and bench…

Label Error DetectionObject Detection

Bias-Aware Mislabeling Detection via Decoupled Confident Learning

2025-07-09 · Yunyi Li, Maria De-Arteaga, Maytal Saar-Tsechansky arxiv

Reliable data is a cornerstone of modern organizational systems. A notable data integrity challenge stems from label bias, which refers to systematic errors in a label, a covariate that is central to a quantitative analy…

Hate Speech DetectionLabel Error Detection

CleanPatrick: A Benchmark for Image Data Cleaning

2025-05-16 · Fabian Gröger, Simone Lionetti, Philippe Gottfrois, Alvaro Gonzalez-Jimenez 외

Robust machine learning depends on clean data, yet current image data cleaning benchmarks rely on synthetic noise or narrow human studies, limiting comparison and real-world relevance. We introduce CleanPatrick, the firs…

BenchmarkingLabel Error DetectionSSIM

Automatic Dataset Construction (ADC): Sample Collection, Data Curation, and Beyond

2024-08-21 · Minghao Liu, Zonglin Di, Jiaheng Wei, Zhongruo Wang 외

Large-scale data collection is essential for developing personalized training data, mitigating the shortage of training data, and fine-tuning specialized models. However, creating high-quality datasets quickly and accura…

Code Generationimage-classificationImage ClassificationLabel Error Detection

LEMoN: Label Error Detection using Multimodal Neighbors

2024-07-10 · Haoran Zhang, Aparna Balagopalan, Nassim Oufattole, Hyewon Jeong 외

Large repositories of image-caption pairs are essential for the development of vision-language models. However, these datasets are often extracted from noisy data scraped from the web, and contain many mislabeled instanc…

Label Error Detection

Improving Label Error Detection and Elimination with Uncertainty Quantification

2024-05-15 · Johannes Jakubik, Michael Vössing, Manil Maskey, Christopher Wölfle 외

Identifying and handling label errors can significantly enhance the accuracy of supervised machine learning models. Recent approaches for identifying label errors demonstrate that a low self-confidence of models with res…

Ensemble Learningimage-classificationImage ClassificationLabel Error Detection+1

AQuA: A Benchmarking Tool for Label Quality Assessment

2023-06-15 · NeurIPS 2023 11 · Mononito Goswami, Vedant Sanil, Arjun Choudhry, Arvind Srinivasan 외

Machine learning (ML) models are only as good as the data they are trained on. But recent studies have found datasets widely used to train and evaluate ML models, e.g. ImageNet, to have pervasive labeling errors. Erroneo…

BenchmarkingLabel Error DetectionModel Selection

Improving Opinion-based Question Answering Systems Through Label Error Detection and Overwrite

2023-06-13 · Xiao Yang, Ahmed K. Mohamed, Shashank Jain, Stanislav Peshterliev 외

Label error is a ubiquitous problem in annotated data. Large amounts of label error substantially degrades the quality of deep learning models. Existing methods to tackle the label error problem largely focus on the clas…

Label Error DetectionMachine Reading ComprehensionQuestion AnsweringReading Comprehension+1

Identifying Label Errors in Object Detection Datasets by Loss Inspection

2023-03-13 · Marius Schubert, Tobias Riedlinger, Karsten Kahl, Daniel Kröll 외

Labeling datasets for supervised object detection is a dull and time-consuming task. Errors can be easily introduced during annotation and overlooked during review, yielding inaccurate benchmarks and performance degradat…

Label Error DetectionObjectobject-detectionObject Detection

The Re-Label Method For Data-Centric Machine Learning

2023-02-09 · Tong Guo

In industry deep learning application, our manually labeled data has a certain number of noisy data. To solve this problem and achieve more than 90 score in dev dataset, we present a simple method to find the noisy data …

Click-Through Rate PredictionDeep LearningLabel Error Detectionobject-detection+2

Identifying Incorrect Annotations in Multi-Label Classification Data

2022-11-25 · Aditya Thyagarajan, Elías Snorrason, Curtis Northcutt, Jonas Mueller

In multi-label classification, each example in a dataset may be annotated as belonging to one or more classes (or none of the classes). Example applications include image (or document) tagging where each possible tag eit…

ClassificationLabel Error DetectionMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION+1

CTRL: Clustering Training Losses for Label Error Detection

2022-08-17 · Chang Yue, Niraj K. Jha

In supervised machine learning, use of correct labels is extremely important to ensure high accuracy. Unfortunately, most datasets contain corrupted labels. Machine learning models trained on such datasets do not general…

ClusteringLabel Error Detection

Automated Detection of Label Errors in Semantic Segmentation Datasets via Deep Learning and Uncertainty Quantification

2022-07-13 · Matthias Rottmann, Marco Reese

In this work, we for the first time present a method for detecting label errors in image datasets with semantic segmentation, i.e., pixel-wise class labels. Annotation acquisition for semantic segmentation datasets is ti…

BenchmarkingLabel Error DetectionSegmentationSemantic Segmentation+1
1–20 / 21 다음 →