paper-with-me

홈 › Papers

Inherent Disagreements in Human Textual Inferences

2019-03-01 · TACL 2019 3 · Ellie Pavlick, Tom Kwiatkowski

We analyze human{'}s disagreements about the validity of natural language inferences. We show that, very often, disagreements are not dismissible as annotation {``}noise{''}, but rather persist as we collect more ratings and as we vary the amount of context provided to raters. We further show that the type of uncertainty captured by current state-of-the-art models for natural language inference is not reflective of the type of uncertainty present in human disagreements. We discuss implications of our results in relation to the recognizing textual entailment (RTE)/natural language inference (NLI) task. We argue for a refined evaluation objective that requires models to explicitly capture the full distribution of plausible human judgments.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language InferenceRTE

Similar Papers 제목 키워드 기반

Did they answer? Subjective acts and intents in conversational discourse

2021-04-09 · NAACL 2021 4 · Elisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk

Discourse signals are often implicit, leaving it up to the interpreter to draw the required inferences. At the same time, discourse is embedded in a social context, meaning that interpreters apply their own assumptions a…

valid

Agree to Disagree: Analysis of Inter-Annotator Disagreements in Human Evaluation of Machine Translation Output

2021-11-01 · CoNLL (EMNLP) 2021 11 · Maja Popović

This work describes an analysis of inter-annotator disagreements in human evaluation of machine translation output. The errors in the analysed texts were marked by multiple annotators under guidance of different quality …

Machine TranslationNegationTranslation

Ambiguity Helps: Classification With Disagreements in Crowdsourced Annotations

2016-06-01 · CVPR 2016 6 · Viktoriia Sharmanska, Daniel Hernandez-Lobato, Jose Miguel Hernandez-Lobato, Novi Quadrianto

Imagine we show an image to a person and ask her/him to decide whether the scene in the image is warm or not warm, and whether it is easy or not to spot a squirrel in the image. For exactly the same image, the answers to…

ClassificationGeneral Classification

Stop Measuring Calibration When Humans Disagree

2022-10-28 · Joris Baan, Wilker Aziz, Barbara Plank, Raquel Fernández

Calibration is a popular framework to evaluate whether a classifier knows when it does not know - i.e., its predictive probabilities are a good indication of how likely a prediction is to be correct. Correctness is commo…

Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data

2025-05-15 · Poli Apollinaire Nemkova, Solomon Ubani, Mark V. Albert

In the era of increasingly sophisticated natural language processing (NLP) systems, large language models (LLMs) have demonstrated remarkable potential for diverse applications, including tasks requiring nuanced textual …

Binary Classificationtext annotation