paper-with-me

홈 › Papers

CODA-19: Using a Non-Expert Crowd to Annotate Research Aspects on 10,000+ Abstracts in the COVID-19 Open Research Dataset

2020-05-05 · ACL 2020 7 · Ting-Hao 'Kenneth' Huang, Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Yen-Chia Hsu, C. Lee Giles

This paper introduces CODA-19, a human-annotated dataset that codes the Background, Purpose, Method, Finding/Contribution, and Other sections of 10,966 English abstracts in the COVID-19 Open Research Dataset. CODA-19 was created by 248 crowd workers from Amazon Mechanical Turk within 10 days, and achieved labeling quality comparable to that of experts. Each abstract was annotated by nine different workers, and the final labels were acquired by majority vote. The inter-annotator agreement (Cohen's kappa) between the crowd and the biomedical expert (0.741) is comparable to inter-expert agreement (0.788). CODA-19's labels have an accuracy of 82.2% when compared to the biomedical expert's labels, while the accuracy between experts was 85.0%. Reliable human annotations help scientists access and integrate the rapidly accelerating coronavirus literature, and also serve as the battery of AI/NLP research, but obtaining expert annotations can be slow. We demonstrated that a non-expert crowd can be rapidly employed at scale to join the fight against COVID-19.

📄 PDF Abstract BibTeX arXiv:2005.02367

Code (1)

windx0303/CODA-19 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Good Data, Large Data, or No Data? Comparing Three Approaches in Developing Research Aspect Classifiers for Biomedical Papers

2023-06-07 · Shreya Chandrasekhar, Chieh-Yang Huang, Ting-Hao 'Kenneth' Huang

The rapid growth of scientific publications, particularly during the COVID-19 pandemic, emphasizes the need for tools to help researchers efficiently comprehend the latest advancements. One essential part of understandin…

Phrase Detectives Corpus 1.0 Crowdsourced Anaphoric Coreference.

2016-05-01 · LREC 2016 5 · Jon Chamberlain, Massimo Poesio, Udo Kruschwitz

Natural Language Engineering tasks require large and complex annotated datasets to build more advanced models of language. Corpora are typically annotated by several experts to create a gold standard; however, there are …

text annotation

Annotation Quality in Aspect-Based Sentiment Analysis: A Case Study Comparing Experts, Students, Crowdworkers, and Large Language Model

2026-05-05 · Niklas Donhauser, Jakob Fehle, Nils Constantin Hellwig, Markus Weinberger 외 arxiv

Aspect-Based Sentiment Analysis (ABSA) enables fine-grained opinion analysis by identifying sentiments toward specific aspects or targets within a text. While ABSA has been widely studied for English, research on other l…

Aspect Category Sentiment AnalysisFine-Grained Opinion Analysis

IndoNLI: A Natural Language Inference Dataset for Indonesian

2021-10-27 · EMNLP 2021 11 · Rahmad Mahendra, Alham Fikri Aji, Samuel Louvan, Fahrurrozi Rahman 외

We present IndoNLI, the first human-elicited NLI dataset for Indonesian. We adapt the data collection protocol for MNLI and collect nearly 18K sentence pairs annotated by crowd workers and experts. The expert-annotated d…

Natural Language InferenceSentenceSpatial ReasoningXLM-R

QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines

2020-11-01 · EMNLP 2020 11 · Valentina Pyatkin, Ayal Klein, Reut Tsarfaty, Ido Dagan

Discourse relations describe how two propositions relate to one another, and identifying them automatically is an integral part of natural language understanding. However, annotating discourse relations typically require…

Natural Language UnderstandingSentence