paper-with-me

Papers

Investigating the Nature of Disagreements on Mid-Scale Ratings: A Case Study on the Abstractness-Concreteness Continuum

2023-11-08 · Urban Knupleš, Diego Frassinelli, Sabine Schulte im Walde

Humans tend to strongly agree on ratings on a scale for extreme cases (e.g., a CAT is judged as very concrete), but judgements on mid-scale words exhibit more disagreement. Yet, collected rating norms are heavily exploited across disciplines. Our study focuses on concreteness ratings and (i) implements correlations and supervised classification to identify salient multi-modal characteristics of mid-scale words, and (ii) applies a hard clustering to identify patterns of systematic disagreement across raters. Our results suggest to either fine-tune or filter mid-scale target words before utilising them.

📄 PDF Abstract BibTeX arXiv:2311.04563

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

Collective Human Opinions in Semantic Textual Similarity

2023-08-08 · Yuxia Wang, Shimin Tao, Ning Xie, Hao Yang 외

Despite the subjective nature of semantic textual similarity (STS) and pervasive disagreements in STS annotation, existing benchmarks have used averaged human ratings as the gold standard. Averaging masks the true distri…

Semantic Textual SimilaritySentenceSTS

Inherent Disagreements in Human Textual Inferences

2019-03-01 · TACL 2019 3 · Ellie Pavlick, Tom Kwiatkowski

We analyze human{'}s disagreements about the validity of natural language inferences. We show that, very often, disagreements are not dismissible as annotation {``}noise{''}, but rather persist as we collect more ratings…

Natural Language InferenceRTE

Direct Uncertainty Prediction for Medical Second Opinions

2018-07-04 · Maithra Raghu, Katy Blumer, Rory Sayres, Ziad Obermeyer 외

The issue of disagreements amongst human experts is a ubiquitous one in both machine learning and medicine. In medicine, this often corresponds to doctor disagreements on a patient diagnosis. In this work, we show that m…

BIG-bench Machine LearningGeneral ClassificationPrediction

How to disagree well: Investigating the dispute tactics used on Wikipedia

2022-12-16 · Christine de Kock, Tom Stafford, Andreas Vlachos

Disagreements are frequently studied from the perspective of either detecting toxicity or analysing argument structure. We propose a framework of dispute tactics that unifies these two perspectives, as well as other dial…

The Past, Present, and Future of Typological Databases in NLP

2023-10-20 · Emi Baylor, Esther Ploeger, Johannes Bjerva

Typological information has the potential to be beneficial in the development of NLP models, particularly for low-resource languages. Unfortunately, current large-scale typological databases, notably WALS and Grambank, a…

Language ModelingLanguage Modelling