paper-with-me

홈 › Papers

Improving Reliability of Word Similarity Evaluation by Redesigning Annotation Task and Performance Measure

2016-11-11 · WS 2016 8 · Oded Avraham, Yoav Goldberg

We suggest a new method for creating and using gold-standard datasets for word similarity evaluation. Our goal is to improve the reliability of the evaluation, and we do this by redesigning the annotation task to achieve higher inter-rater agreement, and by defining a performance measure which takes the reliability of each annotation decision in the dataset into account.

📄 PDF Abstract BibTeX arXiv:1611.03641

Code (1)

oavraham1/ag-evaluation 공식 구현

Tasks

Word Similarity

Similar Papers 제목 키워드 기반

k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human Annotations

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many NLP datasets that rely on aggregate r…

Word Similarity

Word Similarity Datasets for Indian Languages: Annotation and Baseline Systems

2017-04-01 · WS 2017 4 · Syed Sarfaraz Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava 외

With the advent of word representations, word similarity tasks are becoming increasing popular as an evaluation metric for the quality of the representations. In this paper, we present manually annotated monolingual word…

Dependency ParsingMachine TranslationNamed Entity Recognition (NER)Question Answering+4

Czech Dataset for Semantic Similarity and Relatedness

2017-09-01 · RANLP 2017 9 · Miloslav Konop{\'\i}k, Ond{\v{r}}ej Pra{\v{z}}{\'a}k, David Steinberger

This paper introduces a Czech dataset for semantic similarity and semantic relatedness. The dataset contains word pairs with hand annotated scores that indicate the semantic similarity and semantic relatedness of the wor…

Semantic SimilaritySemantic Textual Similarity

Approaching Peak Ground Truth

2022-12-31 · Florian Kofler, Johannes Wahle, Ivan Ezhov, Sophia Wagner 외

Machine learning models are typically evaluated by computing similarity with reference annotations and trained by maximizing similarity with such. Especially in the biomedical domain, annotations are subjective and suffe…

Subword Tokenization Strategies for Kurdish Word Embeddings

2025-11-18 · Ali Salehi, Cassandra L. Jacobs arxiv

We investigate tokenization strategies for Kurdish word embeddings by comparing word-level, morpheme-based, and BPE approaches on morphological similarity preservation tasks. We develop a BiLSTM-CRF morphological segment…