paper-with-me

홈 › Papers

Estimating Agreement by Chance for Sequence Annotation

2024-07-16 · Diya Li, Carolyn Rosé, Ao Yuan, Chunxiao Zhou

In the field of natural language processing, correction of performance assessment for chance agreement plays a crucial role in evaluating the reliability of annotations. However, there is a notable dearth of research focusing on chance correction for assessing the reliability of sequence annotation tasks, despite their widespread prevalence in the field. To address this gap, this paper introduces a novel model for generating random annotations, which serves as the foundation for estimating chance agreement in sequence annotation tasks. Utilizing the proposed randomization model and a related comparison approach, we successfully derive the analytical form of the distribution, enabling the computation of the probable location of each annotated text segment and subsequent chance agreement estimation. Through a combination simulation and corpus-based evaluation, we successfully assess its applicability and validate its accuracy and efficacy.

📄 PDF Abstract BibTeX arXiv:2407.11371

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Consensus Measures for Unstructured Biomedical Text Annotations

2026-08-04 · Pascal Wullschleger, Christian Kreis, Martin A. Walter, Marc Pouly 외 arxiv

Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in advance, annotators prefer open-ended labels, whose agreement is hard …

Natural Language Inference

Establishing Annotation Quality in Multi-label Annotations

2022-10-01 · COLING 2022 10 · Marian Marchal, Merel Scholman, Frances Yung, Vera Demberg

In many linguistic fields requiring annotated data, multiple interpretations of a single item are possible. Multi-label annotations more accurately reflect this possibility. However, allowing for multi-label annotations …

DiPietro-Hazari Kappa: A Novel Metric for Assessing Labeling Quality via Annotation

2022-09-17 · Daniel M. DiPietro, Vivek Hazari

Data is a key component of modern machine learning, but statistics for assessing data label quality remain sparse in literature. Here, we introduce DiPietro-Hazari Kappa, a novel statistical metric for assessing the qual…

Same Words, Different Judgments: How Preferences Vary Across Modalities

2026-02-26 · Aaron Broukhim, Nadir Weibel, Eshin Jolly arxiv

Preference-based reinforcement learning (PbRL) is the dominant framework for aligning AI systems to human preferences. However, evaluation protocols for such data were designed for text and have not been validated for sp…

Reinforcement Learning

Designing and Evaluating a Reliable Corpus of Web Genres via Crowd-Sourcing

2014-05-01 · LREC 2014 5 · Noushin Rezapour Asheghi, Serge Sharoff, Katja Markert

Research in Natural Language Processing often relies on a large collection of manually annotated documents. However, currently there is no reliable genre-annotated corpus of web pages to be employed in Automatic Genre Id…

Information RetrievalMachine Translation