Improving Reliability of Word Similarity Evaluation by Redesigning Annotation Task and Performance Measure
We suggest a new method for creating and using gold-standard datasets for word similarity evaluation. Our goal is to improve the reliability of the evaluation, and we do this by redesigning the annotation task to achieve higher inter-rater agreement, and by defining a performance measure which takes the reliability of each annotation decision in the dataset into account.
Code (1)
Tasks
Word SimilaritySimilar Papers 제목 키워드 기반
k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations
Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many NLP datasets that rely on aggregate r…
Word SimilarityWord Similarity Datasets for Indian Languages: Annotation and Baseline Systems
With the advent of word representations, word similarity tasks are becoming increasing popular as an evaluation metric for the quality of the representations. In this paper, we present manually annotated monolingual word…
Dependency ParsingMachine TranslationNamed Entity Recognition (NER)Question Answering+4Czech Dataset for Semantic Similarity and Relatedness
This paper introduces a Czech dataset for semantic similarity and semantic relatedness. The dataset contains word pairs with hand annotated scores that indicate the semantic similarity and semantic relatedness of the wor…
Semantic SimilaritySemantic Textual SimilarityApproaching Peak Ground Truth
Machine learning models are typically evaluated by computing similarity with reference annotations and trained by maximizing similarity with such. Especially in the biomedical domain, annotations are subjective and suffe…
Subword Tokenization Strategies for Kurdish Word Embeddings
We investigate tokenization strategies for Kurdish word embeddings by comparing word-level, morpheme-based, and BPE approaches on morphological similarity preservation tasks. We develop a BiLSTM-CRF morphological segment…