paper-with-me

Papers

k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many NLP datasets that rely on aggregate ratings only report the reliability of individual ones, which is the incorrect unit of analysis. In these instances, the data reliability is being under-reported. We present empirical, analytical, and bootstrap-based methods for measuring the reliability of aggregate ratings. We call this k-rater reliability (kRR), a multi-rater extension of inter-rater reliability (IRR). We apply these methods to the widely used word similarity benchmark dataset, WordSim. We conducted two replications of the WordSim dataset to obtain an empirical reference point. We hope this discussion will nudge researchers to report kRR, the correct unit of reliability for aggregate ratings, in addition to IRR.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Word Similarity

Similar Papers 제목 키워드 기반

k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations

2022-03-24 · ACL 2022 5 · Ka Wong, Praveen Paritosh

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many natural language processing (NLP) app…

Reliability Gaps Between Groups in COMPAS Dataset

2023-08-29 · Tim Räz

This paper investigates the inter-rater reliability of risk assessment instruments (RAIs). The main question is whether different, socially salient groups are affected differently by a lack of inter-rater reliability of …

Reliability of deep learning models for anatomical landmark detection: The role of inter-rater variability

2024-11-26 · Soorena Salari, Hassan Rivaz, Yiming Xiao

Automated detection of anatomical landmarks plays a crucial role in many diagnostic and surgical applications. Progresses in deep learning (DL) methods have resulted in significant performance enhancement in tasks relate…

Anatomical Landmark DetectionDiagnostic

Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory

2025-07-26 · Dan Song, Won-Chan Lee, Hong Jiao arxiv

This study investigates the estimation of reliability for large language models (LLMs) in scoring writing tasks from the AP Chinese Language and Culture Exam. Using generalizability theory, the research evaluates and com…

Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach

2026-02-26 · Jodi M. Casabianca, Maggie Beiting-Parrish arxiv

Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psychometric rater models into the AI pipeli…