paper-with-me

Papers

CLEANANERCorp: Identifying and Correcting Incorrect Labels in the ANERcorp Dataset

2024-08-22 · Mashael Al-Duwais, Hend Al-Khalifa, Abdulmalik Al-Salman

Label errors are a common issue in machine learning datasets, particularly for tasks such as Named Entity Recognition. Such label errors might hurt model training, affect evaluation results, and lead to an inaccurate assessment of model performance. In this study, we dived deep into one of the widely adopted Arabic NER benchmark datasets (ANERcorp) and found a significant number of annotation errors, missing labels, and inconsistencies. Therefore, in this study, we conducted empirical research to understand these errors, correct them and propose a cleaner version of the dataset named CLEANANERCorp. CLEANANERCorp will serve the research community as a more accurate and consistent benchmark.

📄 PDF Abstract BibTeX arXiv:2408.12362

Code (0)

등록된 구현이 없습니다.

Tasks

Missing Labelsnamed-entity-recognitionNamed Entity RecognitionNER

Similar Papers 제목 키워드 기반

An automated method of identifying incorrectly labelled images based on the sequences of loss functions of deep learning networks

2026-07-01 · Zhipeng Zhang, Wenhui Shou, Wengting Ma, Dongjia Xing 외 arxiv

Deep learning is widely applied in medical image analysis, but up to 10% of manually labelled images may be incorrect, degrading model performance. This paper proposes an automated method to identify incorrectly labelled…

CSOT: Curriculum and Structure-Aware Optimal Transport for Learning with Noisy Labels

2023-12-11 · NeurIPS 2023 11 · Wanxing Chang, Ye Shi, Jingya Wang

Learning with noisy labels (LNL) poses a significant challenge in training a well-generalized model while avoiding overfitting to corrupted labels. Recent advances have achieved impressive performance by identifying clea…

DenoisingLearning with noisy labels

Identifying Incorrect Labels in the CoNLL-2003 Corpus

2020-11-01 · CONLL 2020 · Frederick Reiss, Hong Xu, Bryan Cutler, Karthik Muthuraman 외

The CoNLL-2003 corpus for English-language named entity recognition (NER) is one of the most influential corpora for NER model research. A large number of publications, including many landmark works, have used this corpu…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

On-the-fly Historical Handwritten Text Annotation

2017-09-06 · Ekta Vats, Anders Hast

The performance of information retrieval algorithms depends upon the availability of ground truth labels annotated by experts. This is an important prerequisite, and difficulties arise when the annotated ground truth lab…

Information RetrievalRetrievaltext annotation

When Color Constancy Goes Wrong: Correcting Improperly White-Balanced Images

2019-06-01 · CVPR 2019 6 · Mahmoud Afifi, Brian Price, Scott Cohen, Michael S. Brown

This paper focuses on correcting a camera image that has been improperly white-balanced. This situation occurs when a camera's auto white balance fails or when the wrong manual white-balance setting is used. Even after d…

Color Constancy