Detecting annotation noise in automatically labelled data
We introduce a method for error detection in automatically annotated text, aimed at supporting the creation of high-quality language resources at affordable cost. Our method combines an unsupervised generative model with human supervision from active learning. We test our approach on in-domain and out-of-domain data in two languages, in AL simulations and in a real world setting. For all settings, the results show that our method is able to detect annotation errors with high precision and high recall.
Code (0)
등록된 구현이 없습니다.
Tasks
Active LearningDomain AdaptationLanguage ModelingLanguage ModellingNamed Entity Recognition (NER)Similar Papers 제목 키워드 기반
Self-Learning for Player Localization in Sports Video
This paper introduces a novel self-learning framework that automates the label acquisition process for improving models for detecting players in broadcast footage of sports games. Unlike most previous self-learning appro…
Self-LearningBounding Box Priors for Cell Detection with Point Annotations
The size of an individual cell type, such as a red blood cell, does not vary much among humans. We use this knowledge as a prior for classifying and detecting cells in images with only a few ground truth bounding box ann…
Cell DetectionDo LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
We present judgeWEL, a dataset for named entity recognition (NER) in Luxembourgish, automatically labelled and subsequently verified using large language models (LLM) in a novel pipeline. Building datasets for under-repr…
Pre-trained Language Models as Re-Annotators
Annotation noise is widespread in datasets, but manually revising a flawed corpus is time-consuming and error-prone. Hence, given the prior knowledge in Pre-trained Language Models and the expected uniformity across all …
Contrastive LearningDensity EstimationRelation ExtractionFact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
Political misinformation poses significant challenges to democratic processes, shaping public opinion and trust in media. Manual fact-checking methods face issues of scalability and annotator bias, while machine learning…
ArticlesFact CheckingMisinformation