paper-with-me

홈 › Papers

Budget Sensitive Reannotation of Noisy Relation Classification Data Using Label Hierarchy

2021-12-26 · Akshay Parekh, Ashish Anand, Amit Awekar

Large crowd-sourced datasets are often noisy and relation classification (RC) datasets are no exception. Reannotating the entire dataset is one probable solution however it is not always viable due to time and budget constraints. This paper addresses the problem of efficient reannotation of a large noisy dataset for the RC. Our goal is to catch more annotation errors in the dataset while reannotating fewer instances. Existing work on RC dataset reannotation lacks the flexibility about how much data to reannotate. We introduce the concept of a reannotation budget to overcome this limitation. The immediate follow-up problem is: Given a specific reannotation budget, which subset of the data should we reannotate? To address this problem, we present two strategies to selectively reannotate RC datasets. Our strategies utilize the taxonomic hierarchy of relation labels. The intuition of our work is to rely on the graph distance between actual and predicted relation labels in the label hierarchy graph. We evaluate our reannotation strategies on the well-known TACRED dataset. We design our experiments to answer three specific research questions. First, does our strategy select novel candidates for reannotation? Second, for a given reannotation budget is our reannotation strategy more efficient at catching annotation errors? Third, what is the impact of data reannotation on RC model performance measurement? Experimental results show that our both reannotation strategies are novel and efficient. Our analysis indicates that the current reported performance of RC models on noisy TACRED data is inflated.

📄 PDF Abstract BibTeX arXiv:2112.13320

Code (0)

등록된 구현이 없습니다.

Tasks

RelationRelation Classification

Similar Papers 제목 키워드 기반

Noise in Relation Classification Dataset TACRED: Characterization and Reduction

2023-11-21 · Akshay Parekh, Ashish Anand, Amit Awekar

The overarching objective of this paper is two-fold. First, to explore model-based approaches to characterize the primary cause of the noise. in the RE dataset TACRED Second, to identify the potentially noisy instances. …

ClassificationRelationRelation Classification

Multimodal Large Language Models as Image Classifiers

2026-03-06 · Nikita Kisel, Illia Volkov, Klara Janouskova, Jiri Matas arxiv

Multimodal Large Language Models (MLLM) classification performance depends critically on evaluation protocol and ground truth quality. Studies comparing MLLMs with supervised and vision-language models report conflicting…

Turning silver into gold: error-focused corpus reannotation with active learning

2019-09-01 · RANLP 2019 9 · Pierre Andr{\'e} M{\'e}nard, Antoine Mougeot

While high quality gold standard annotated corpora are crucial for most tasks in natural language processing, many annotated corpora published in recent years, created by annotators or tools, contains noisy annotations. …

Active LearningDocument ClassificationPart-Of-Speech Tagging

Cost-Sensitive Uncertainty-Based Failure Recognition for Object Detection

2024-04-26 · Moussa Kassem Sbeyti, Michelle Karg, Christian Wirth, Nadja Klein 외

Object detectors in real-world applications often fail to detect objects due to varying factors such as weather conditions and noisy input. Therefore, a process that mitigates false detections is crucial for both safety …

Autonomous DrivingObjectobject-detectionObject Detection

Classification with Rejection Based on Cost-sensitive Classification

2020-10-22 · Nontawat Charoenphakdee, Zhenghang Cui, Yivan Zhang, Masashi Sugiyama

The goal of classification with rejection is to avoid risky misclassification in error-critical applications such as medical diagnosis and product inspection. In this paper, based on the relationship between classificati…

ClassificationGeneral ClassificationMedical Diagnosis