An Annotated Social Media Corpus for German
This paper presents the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research, discussing the compilation, pseudonymization and grammatical annotation of the corpus, as well as special linguistic features and peculiarities encountered in the data. Among other things, compounding, accidental and intentional orthographic variation, gendering and the use of emoticons/emojis are addressed in a genre-specific fashion. We present the different layers of linguistic annotation (morphosyntactic, dependencies and semantic types) and explain how a general parser (GerGram) can be made to work on Social Media data, pointing out necessary adaptations and extensions. In an evaluation run on a random cross-section of tweets, the modified parser achieved F-scores of 97{\%} for morphology (fine-grained POS) and 92{\%} for syntax (labeled attachment score). Predictably, performance was twice as good in tweets with standard orthography than in tweets with spelling/casing irregularities or lack of sentence separation, the effect being more marked for morphology than for syntax.
Code (0)
등록된 구현이 없습니다.
Tasks
POSSentenceSimilar Papers 제목 키워드 기반
Cross-lingual Approaches for the Detection of Adverse Drug Reactions in German from a Patient's Perspective
In this work, we present the first corpus for German Adverse Drug Reaction (ADR) detection in patient-generated content. The data consists of 4,169 binary annotated documents from a German patient forum, where users talk…
Binary ClassificationFew-Shot LearningCross-lingual Approaches for the Detection of Adverse Drug Reactions in German from a Patient’s Perspective
In this work, we present the first corpus for German Adverse Drug Reaction (ADR) detection in patient-generated content. The data consists of 4,169 binary annotated documents from a German patient forum, where users talk…
Binary ClassificationFew-Shot LearningFine-grained German Sentiment Analysis on Social Media
Expressing opinions and emotions on social media becomes a frequent activity in daily life. People express their opinions about various targets via social media and they are also interested to know about other opinions o…
Sentiment AnalysisGenerating Sentiment Lexicons for German Twitter
Despite a substantial progress made in developing new sentiment lexicon generation (SLG) methods for English, the task of transferring these approaches to other languages and domains in a sound way still remains open. In…
Detecting de minimis Code-Switching in Historical German Books
Code-switching has long interested linguists, with computational work in particular focusing on speech and social media data (Sitaram et al., 2019). This paper contrasts these informal instances of code-switching to its …
Optical Character RecognitionOptical Character Recognition (OCR)