k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations
Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many natural language processing (NLP) applications that rely on aggregate ratings only report the reliability of individual ratings, which is the incorrect unit of analysis. In these instances, the data reliability is under-reported, and a proposed k-rater reliability (kRR) should be used as the correct data reliability for aggregated datasets. It is a multi-rater generalization of inter-rater reliability (IRR). We conducted two replications of the WordSim-353 benchmark, and present empirical, analytical, and bootstrap-based methods for computing kRR on WordSim-353. These methods produce very similar results. We hope this discussion will nudge researchers to report kRR in addition to IRR.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations
Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many NLP datasets that rely on aggregate r…
Word SimilarityReliability Gaps Between Groups in COMPAS Dataset
This paper investigates the inter-rater reliability of risk assessment instruments (RAIs). The main question is whether different, socially salient groups are affected differently by a lack of inter-rater reliability of …
Reliability of deep learning models for anatomical landmark detection: The role of inter-rater variability
Automated detection of anatomical landmarks plays a crucial role in many diagnostic and surgical applications. Progresses in deep learning (DL) methods have resulted in significant performance enhancement in tasks relate…
Anatomical Landmark DetectionDiagnosticExploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory
This study investigates the estimation of reliability for large language models (LLMs) in scoring writing tasks from the AP Chinese Language and Culture Exam. Using generalizability theory, the research evaluates and com…
Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach
Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psychometric rater models into the AI pipeli…