paper-with-me

Papers

Efficient Annotator Reliability Assessment and Sample Weighting for Knowledge-Based Misinformation Detection on Social Media

2024-10-18 · Owen Cook, Charlie Grimshaw, Ben Wu, Sophie Dillon, Jack Hicks, Luke Jones, Thomas Smith, Matyas Szert, Xingyi Song

Misinformation spreads rapidly on social media, confusing the truth and targeting potentially vulnerable people. To effectively mitigate the negative impact of misinformation, it must first be accurately detected before applying a mitigation strategy, such as X's community notes, which is currently a manual process. This study takes a knowledge-based approach to misinformation detection, modelling the problem similarly to one of natural language inference. The EffiARA annotation framework is introduced, aiming to utilise inter- and intra-annotator agreement to understand the reliability of each annotator and influence the training of large language models for classification based on annotator reliability. In assessing the EffiARA annotation framework, the Russo-Ukrainian Conflict Knowledge-Based Misinformation Classification Dataset (RUC-MCD) was developed and made publicly available. This study finds that sample weighting using annotator reliability performs the best, utilising both inter- and intra-annotator agreement and soft-label training. The highest classification performance achieved using Llama-3.2-1B was a macro-F1 of 0.757 and 0.740 using TwHIN-BERT-large.

📄 PDF Abstract BibTeX arXiv:2410.14515

Code (2)

minieggz/effiara 공식 구현
minieggz/ruc-misinfo 공식 구현 pytorch

Tasks

ClassificationMisinformationNatural Language Inference

Similar Papers 제목 키워드 기반

Efficient Annotator Reliability Assessment with EffiARA

2025-04-01 · Owen Cook, Jake Vasilakes, Ian Roberts, Xingyi Song

Data annotation is an essential component of the machine learning pipeline; it is also a costly and time-consuming process. With the introduction of transformer-based models, annotation at the document level is increasin…

Stop Replacing Noise with Noise: Two-Source Reliability Assessment for Label Correction and Sample Reweighting in Label-Noise Learning

2026-08-04 · Wenxiao Fan, Kan Li arxiv

Refurbishment-based noisy-label learning mixes an observed label with a model-derived pseudo target, typically using one sample-wise cleanliness score to control both branches. This creates a hidden coupling: reducing tr…

On User Interfaces for Large-Scale Document-Level Human Evaluation of Machine Translation Outputs

2021-04-21 · EACL (HumEval) 2021 4 · Roman Grundkiewicz, Marcin Junczys-Dowmunt, Christian Federmann, Tom Kocmi

Recent studies emphasize the need of document context in human evaluation of machine translations, but little research has been done on the impact of user interfaces on annotator productivity and the reliability of asses…

Machine TranslationTranslation

Label Curation Using Agentic AI

2026-01-30 · Subhodeep Ghosh, Bayan Divaaniaazar, Md Ishat-E-Rabban, Spencer Clarke 외 arxiv

Data annotation is essential for supervised learning, yet producing accurate, unbiased, and scalable labels remains challenging as datasets grow in size and modality. Traditional human-centric pipelines are costly, slow,…

Reliability-Aware Prediction via Uncertainty Learning for Person Image Retrieval

2022-10-24 · Zhaopeng Dou, Zhongdao Wang, Weihua Chen, YaLi Li 외

Current person image retrieval methods have achieved great improvements in accuracy metrics. However, they rarely describe the reliability of the prediction. In this paper, we propose an Uncertainty-Aware Learning (UAL) …

Image RetrievalRetrieval