Collective Human Opinions in Semantic Textual Similarity
Despite the subjective nature of semantic textual similarity (STS) and pervasive disagreements in STS annotation, existing benchmarks have used averaged human ratings as the gold standard. Averaging masks the true distribution of human opinions on examples of low agreement, and prevents models from capturing the semantic vagueness that the individual ratings represent. In this work, we introduce USTS, the first Uncertainty-aware STS dataset with ~15,000 Chinese sentence pairs and 150,000 labels, to study collective human opinions in STS. Analysis reveals that neither a scalar nor a single Gaussian fits a set of observed judgements adequately. We further show that current STS models cannot capture the variance caused by human disagreement on individual instances, but rather reflect the predictive confidence over the aggregate dataset.
Code (1)
Tasks
Semantic Textual SimilaritySentenceSTSSimilar Papers 제목 키워드 기반
Rethinking STS and NLI in Large Language Models
Recent years have seen the rise of large language models (LLMs), where practitioners use task-specific prompts; this was shown to be effective for a variety of tasks. However, when applied to semantic textual similarity …
Natural Language InferenceSemantic Textual SimilaritySTSSemantic Properties of Customer Sentiment in Tweets
An increasing number of people are using online social networking services (SNSs), and a significant amount of information related to experiences in consumption is shared in this new media form. Text mining is an emergin…
ClusteringSentiment AnalysisA model to support collective reasoning: Formalization, analysis and computational assessment
Inspired by e-participation systems, in this paper we propose a new model to represent human debates and methods to obtain collective conclusions from them. This model overcomes drawbacks of existing approaches by allowi…
What Can We Learn from Collective Human Opinions on Natural Language Inference Data?
Despite the subjective nature of many NLP tasks, most NLU evaluations have focused on using the majority label with presumably high agreement as the ground truth. Less attention has been paid to the distribution of human…
Natural Language InferenceLearning Human-Object Interaction as Groups
Human-Object Interaction Detection (HOI-DET) aims to localize human-object pairs and identify their interactive relationships. To aggregate contextual cues, existing methods typically propagate information across all det…
Human-Object Interaction DetectionSemantic Similarity