paper-with-me

홈 › Papers

A Multi-word Expression Dataset for Swedish

2020-05-01 · LREC 2020 5 · Murathan Kurfal{\i}, Robert {\"O}stling, Johan Sjons, Mats Wir{\'e}n

We present a new set of 96 Swedish multi-word expressions annotated with degree of (non-)compositionality. In contrast to most previous compositionality datasets we also consider syntactically complex constructions and publish a formal specification of each expression. This allows evaluation of computational models beyond word bigrams, which have so far been the norm. Finally, we use the annotations to evaluate a system for automatic compositionality estimation based on distributional semantics. Our analysis of the disagreements between human annotators and the distributional model reveal interesting questions related to the perception of compositionality, and should be informative to future work in the area.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Crowdsourcing Relative Rankings of Multi-Word Expressions: Experts versus Non-Experts

2022-06-17 · David Alfter, Therese Lindström Tiedemann, Elena Volodina

In this study we investigate to which degree experts and non-experts agree on questions of difficulty in a crowdsourcing experiment. We ask non-experts (second language learners of Swedish) and two groups of experts (tea…

Distributional properties of political dogwhistle representations in Swedish BERT

2022-07-01 · NAACL (WOAH) 2022 7 · Niclas Hertzberg, Robin Cooper, Elina Lindgren, Björn Rönnerstrand 외

“Dogwhistles” are expressions intended by the speaker have two messages: a socially-unacceptable “in-group” message understood by a subset of listeners, and a benign message intended for the out-group. We take the result…

SentenceSurvey

SuperSim: a test set for word similarity and relatedness in Swedish

2021-04-12 · NoDaLiDa 2021 5 · Simon Hengchen, Nina Tahmasebi

Language models are notoriously difficult to evaluate. We release SuperSim, a large-scale similarity and relatedness test set for Swedish built with expert human judgments. The test set is composed of 1,360 word-pairs in…

Word Similarity

SVALex: a CEFR-graded Lexical Resource for Swedish Foreign and Second Language Learners

2016-05-01 · LREC 2016 5 · Thomas Fran{\c{c}}ois, Elena Volodina, Ildik{\'o} Pil{\'a}n, Ana{\"\i}s Tack

The paper introduces SVALex, a lexical resource primarily aimed at learners and teachers of Swedish as a foreign and second language that describes the distribution of 15,681 words and expressions across the Common Europ…

Why Not Simply Translate? A First Swedish Evaluation Benchmark for Semantic Similarity

2020-09-07 · Tim Isbister, Magnus Sahlgren

This paper presents the first Swedish evaluation benchmark for textual semantic similarity. The benchmark is compiled by simply running the English STS-B dataset through the Google machine translation API. This paper dis…

Machine TranslationSemantic SimilaritySemantic Textual SimilaritySTS+2