paper-with-me

Papers

EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus

2020-05-01 · LREC 2020 5 · Thomas Proisl, Natalie Dykes, Philipp Heinrich, Besim Kabashi, Andreas Blombach, Stefan Evert

The EmpiriST corpus (Bei{\ss}wenger et al., 2016) is a manually tokenized and part-of-speech tagged corpus of approximately 23,000 tokens of German Web and CMC (computer-mediated communication) data. We extend the corpus with manually created annotation layers for word form normalization, lemmatization and lexical semantics. All annotations have been independently performed by multiple human annotators. We report inter-annotator agreements and results of baseline systems and state-of-the-art off-the-shelf tools.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Lemmatization

Similar Papers 제목 키워드 기반

Revisiting NMT for Normalization of Early English Letters

2019-06-01 · WS 2019 6 · Mika H{\"a}m{\"a}l{\"a}inen, Tanja S{\"a}ily, Jack Rueter, J{\"o}rg Tiedemann 외

This paper studies the use of NMT (neural machine translation) as a normalization method for an early English letter corpus. The corpus has previously been normalized so that only less frequent deviant forms are left out…

LemmatizationMachine TranslationNMTTranslation

FOLK-Gold ― A Gold Standard for Part-of-Speech-Tagging of Spoken German

2016-05-01 · LREC 2016 5 · Swantje Westpfahl, Thomas Schmidt

In this paper, we present a GOLD standard of part-of-speech tagged transcripts of spoken German. The GOLD standard data consists of four annotation layers ― transcription (modified orthography), normalization (standard o…

LemmatizationPart-Of-Speech TaggingPOS

Morphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models

2019-12-02 · Nelda Kote, Marenglen Biba, Jenna Kanerva, Samuel Rönnqvist 외

In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…

LemmatizationMorphological TaggingPart-Of-Speech Tagging

Urdu Summary Corpus

2016-05-01 · LREC 2016 5

Language resources, such as corpora, are important for various natural language processing tasks. Urdu has millions of speakers around the world but it is under-resourced in terms of standard evaluation resources. This p…

ArticlesDocument SummarizationLemmatizationMorphological Analysis+1

ZAEBUC: An Annotated Arabic-English Bilingual Writer Corpus

2022-06-01 · LREC 2022 6 · Nizar Habash, David Palfreyman

We present ZAEBUC, an annotated Arabic-English bilingual writer corpus comprising short essays by first-year university students at Zayed University in the United Arab Emirates. We describe and discuss the various guidel…

LemmatizationPart-Of-Speech TaggingPOSPOS Tagging