paper-with-me

홈 › Papers

Quantifying and Predicting Disagreement in Graded Human Ratings

2026-05-01 · Leixin Zhang, Çağrı Çöltekin arxiv

It is increasingly recognized that human annotators do not always agree, and such disagreement is inherent in many annotation tasks. However, not all instances in a given task elicit the same degree of opinion divergence. In this paper, we investigate annotation variation patterns in graded human ratings for inappropriate languages, including offensive language, hate speech, and toxic language perception. We examine whether the degree of annotation disagreement can be predicted from textual features. We further propose the Opposition Index, a metric that quantifies perspective opposition among annotators on a given item, and investigate the predictability of instances with potentially opposing human opinions. Our results show a moderate positive correlation between estimated and observed annotation variance. We find that two approaches achieve comparable performance in variance prediction: directly predicting the variance value and estimating it from predicted annotation distributions. Our results on opposition perspective prediction show that items with high opposition index values are more difficult to predict and are often underestimated by models.

📄 PDF Abstract BibTeX arXiv:2605.01168

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Predicting Relevance based on Assessor Disagreement: Analysis and Practical Applications for Search Evaluation

2015-11-23 · Demeester Thomas, Aly Robin, Hiemstra Djoerd, Nguyen Dong 외

Evaluation of search engines relies on assessments of search results for selected test queries, from which we would ideally like to draw conclusions in terms of relevance of the results for general (e.g., future, unknown…

Information RetrievalRetrieval

Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals

2026-05-12 · Yo Ehara arxiv

Automatic generation of educational materials using large language models (LLMs) is becoming increasingly common, but assigning difficulty levels to such materials still requires substantial human effort. LLM-as-a-Judge …

When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks

2023-05-11 · Eve Fleisig, Rediet Abebe, Dan Klein

Though majority vote among annotators is typically used for ground truth labels in natural language processing, annotator disagreement in tasks such as hate speech detection may reflect differences in opinion across grou…

Hate Speech Detection

Investigating the Nature of Disagreements on Mid-Scale Ratings: A Case Study on the Abstractness-Concreteness Continuum

2023-11-08 · Urban Knupleš, Diego Frassinelli, Sabine Schulte im Walde

Humans tend to strongly agree on ratings on a scale for extreme cases (e.g., a CAT is judged as very concrete), but judgements on mid-scale words exhibit more disagreement. Yet, collected rating norms are heavily exploit…

Clustering

Inherent Disagreements in Human Textual Inferences

2019-03-01 · TACL 2019 3 · Ellie Pavlick, Tom Kwiatkowski

We analyze human{'}s disagreements about the validity of natural language inferences. We show that, very often, disagreements are not dismissible as annotation {``}noise{''}, but rather persist as we collect more ratings…

Natural Language InferenceRTE