paper-with-me

Papers

Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater Reliability

2021-08-01 · ACL 2021 5 · Ka Wong, Praveen Paritosh, Lora Aroyo

When collecting annotations and labeled data from humans, a standard practice is to use inter-rater reliability (IRR) as a measure of data goodness (Hallgren, 2012). Metrics such as Krippendorff{'}s alpha or Cohen{'}s kappa are typically required to be above a threshold of 0.6 (Landis and Koch, 1977). These absolute thresholds are unreasonable for crowdsourced data from annotators with high cultural and training variances, especially on subjective topics. We present a new alternative to interpreting IRR that is more empirical and contextualized. It is based upon benchmarking IRR against baseline measures in a replication, one of which is a novel cross-replication reliability (xRR) measure based on Cohen{'}s (1960) kappa. We call this approach the xRR framework. We opensource a replication dataset of 4 million human judgements of facial expressions and analyze it with the proposed framework. We argue this framework can be used to measure the quality of crowdsourced datasets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Cross-replication Reliability -- An Empirical Approach to Interpreting Inter-rater Reliability

2021-06-11 · Ka Wong, Praveen Paritosh, Lora Aroyo

We present a new approach to interpreting IRR that is empirical and contextualized. It is based upon benchmarking IRR against baseline measures in a replication, one of which is a novel cross-replication reliability (xRR…

Benchmarking

k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many NLP datasets that rely on aggregate r…

Word Similarity

k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations

2022-03-24 · ACL 2022 5 · Ka Wong, Praveen Paritosh

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many natural language processing (NLP) app…

RL-TIME: Reinforcement Learning-based Task Replication in Multicore Embedded Systems

2025-03-16 · Roozbeh Siyadatzadeh, Mohsen Ansari, Muhammad Shafique, Alireza Ejlali

Embedded systems power many modern applications and must often meet strict reliability, real-time, thermal, and power requirements. Task replication can improve reliability by duplicating a task's execution to handle tra…

Interpreting the dependence of mutation rates on age and time

2015-07-24

Mutations can arise from the chance misincorporation of nucleotides during DNA replication or from DNA lesions that are not repaired correctly. We introduce a model that relates the source of mutations to their accumulat…