EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus
The EmpiriST corpus (Bei{\ss}wenger et al., 2016) is a manually tokenized and part-of-speech tagged corpus of approximately 23,000 tokens of German Web and CMC (computer-mediated communication) data. We extend the corpus with manually created annotation layers for word form normalization, lemmatization and lexical semantics. All annotations have been independently performed by multiple human annotators. We report inter-annotator agreements and results of baseline systems and state-of-the-art off-the-shelf tools.
Code (0)
등록된 구현이 없습니다.
Tasks
LemmatizationSimilar Papers 제목 키워드 기반
Revisiting NMT for Normalization of Early English Letters
This paper studies the use of NMT (neural machine translation) as a normalization method for an early English letter corpus. The corpus has previously been normalized so that only less frequent deviant forms are left out…
LemmatizationMachine TranslationNMTTranslationFOLK-Gold ― A Gold Standard for Part-of-Speech-Tagging of Spoken German
In this paper, we present a GOLD standard of part-of-speech tagged transcripts of spoken German. The GOLD standard data consists of four annotation layers ― transcription (modified orthography), normalization (standard o…
LemmatizationPart-Of-Speech TaggingPOSMorphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models
In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…
LemmatizationMorphological TaggingPart-Of-Speech TaggingUrdu Summary Corpus
Language resources, such as corpora, are important for various natural language processing tasks. Urdu has millions of speakers around the world but it is under-resourced in terms of standard evaluation resources. This p…
ArticlesDocument SummarizationLemmatizationMorphological Analysis+1ZAEBUC: An Annotated Arabic-English Bilingual Writer Corpus
We present ZAEBUC, an annotated Arabic-English bilingual writer corpus comprising short essays by first-year university students at Zayed University in the United Arab Emirates. We describe and discuss the various guidel…
LemmatizationPart-Of-Speech TaggingPOSPOS Tagging