paper-with-me

Papers

The SemDaX Corpus ― Sense Annotations with Scalable Sense Inventories

2016-05-01 · LREC 2016 5 · Bolette Pedersen, Anna Braasch, Anders Johannsen, H{\'e}ctor Mart{\'\i}nez Alonso, Sanni Nimb, Sussi Olsen, Anders S{\o}gaard, Nicolai Hartvig S{\o}rensen

We launch the SemDaX corpus which is a recently completed Danish human-annotated corpus available through a CLARIN academic license. The corpus includes approx. 90,000 words, comprises six textual domains, and is annotated with sense inventories of different granularity. The aim of the developed corpus is twofold: i) to assess the reliability of the different sense annotation schemes for Danish measured by qualitative analyses and annotation agreement scores, and ii) to serve as training and test data for machine learning algorithms with the practical purpose of developing sense taggers for Danish. To these aims, we take a new approach to human-annotated corpus resources by double annotating a much larger part of the corpus than what is normally seen: for the all-words task we double annotated 60{\%} of the material and for the lexical sample task 100{\%}. We include in the corpus not only the adjucated files, but also the diverging annotations. In other words, we consider not all disagreement to be noise, but rather to contain valuable linguistic information that can help us improve our annotation schemes and our learning algorithms.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SALMA: Arabic Sense-Annotated Corpus and WSD Benchmarks

2023-10-29 · Mustafa Jarrar, Sanad Malaysha, Tymaa Hammouda, Mohammed Khalilia

SALMA, the first Arabic sense-annotated corpus, consists of ~34K tokens, which are all sense-annotated. The corpus is annotated using two different sense inventories simultaneously (Modern and Ghani). SALMA novelty lies …

Word Sense Disambiguation

The MASC Word Sense Corpus

2012-05-01 · LREC 2012 5 · Rebecca J. Passonneau, Collin F. Baker, Christiane Fellbaum, Nancy Ide

The MASC project has produced a multi-genre corpus with multiple layers of linguistic annotation, together with a sentence corpus containing WordNet 3.1 sense tags for 1000 occurrences of each of 100 words produced by mu…

Sentence

EuroSense: Automatic Harvesting of Multilingual Sense Annotations from Parallel Text

2017-07-01 · ACL 2017 7 · Claudio Delli Bovi, Jose Camacho-Collados, Aless Raganato, ro 외

Parallel corpora are widely used in a variety of Natural Language Processing tasks, from Machine Translation to cross-lingual Word Sense Disambiguation, where parallel sentences can be exploited to automatically generate…

Entity LinkingMachine TranslationTranslationWord Sense Disambiguation

Annotating the MASC Corpus with BabelNet

2014-05-01 · LREC 2014 5 · Andrea Moro, Roberto Navigli, Francesco Maria Tucci, Rebecca J. Passonneau

In this paper we tackle the problem of automatically annotating, with both word senses and named entities, the MASC 3.0 corpus, a large English corpus covering a wide range of genres of written and spoken text. We use Ba…

Entity LinkingReading ComprehensionRelation ExtractionWord Sense Disambiguation

Semi-Supervised and Unsupervised Sense Annotation via Translations

2021-06-11 · RANLP 2021 9 · Bradley Hauer, Grzegorz Kondrak, Yixing Luan, Arnob Mallik 외

Acquisition of multilingual training data continues to be a challenge in word sense disambiguation (WSD). To address this problem, unsupervised approaches have been proposed to automatically generate sense annotations fo…

Machine TranslationTranslationWord Sense Disambiguation