paper-with-me

Papers

k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations

2022-03-24 · ACL 2022 5 · Ka Wong, Praveen Paritosh

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many natural language processing (NLP) applications that rely on aggregate ratings only report the reliability of individual ratings, which is the incorrect unit of analysis. In these instances, the data reliability is under-reported, and a proposed k-rater reliability (kRR) should be used as the correct data reliability for aggregated datasets. It is a multi-rater generalization of inter-rater reliability (IRR). We conducted two replications of the WordSim-353 benchmark, and present empirical, analytical, and bootstrap-based methods for computing kRR on WordSim-353. These methods produce very similar results. We hope this discussion will nudge researchers to report kRR in addition to IRR.

📄 PDF Abstract BibTeX arXiv:2203.12913

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many NLP datasets that rely on aggregate r…

Word Similarity

Reliability Gaps Between Groups in COMPAS Dataset

2023-08-29 · Tim Räz

This paper investigates the inter-rater reliability of risk assessment instruments (RAIs). The main question is whether different, socially salient groups are affected differently by a lack of inter-rater reliability of …

Reliability of deep learning models for anatomical landmark detection: The role of inter-rater variability

2024-11-26 · Soorena Salari, Hassan Rivaz, Yiming Xiao

Automated detection of anatomical landmarks plays a crucial role in many diagnostic and surgical applications. Progresses in deep learning (DL) methods have resulted in significant performance enhancement in tasks relate…

Anatomical Landmark DetectionDiagnostic

Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory

2025-07-26 · Dan Song, Won-Chan Lee, Hong Jiao arxiv

This study investigates the estimation of reliability for large language models (LLMs) in scoring writing tasks from the AP Chinese Language and Culture Exam. Using generalizability theory, the research evaluates and com…

Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach

2026-02-26 · Jodi M. Casabianca, Maggie Beiting-Parrish arxiv

Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psychometric rater models into the AI pipeli…