Adapting a part-of-speech tagset to non-standard text: The case of STTS
The Stuttgart-T{\"u}bingen TagSet (STTS) is a de-facto standard for the part-of-speech tagging of German texts. Since its first publication in 1995, STTS has been used in a variety of annotation projects, some of which have adapted the tagset slightly for their specific needs. Recently, the focus of many projects has shifted from the analysis of newspaper text to that of non-standard varieties such as user-generated content, historical texts, and learner language. These text types contain linguistic phenomena that are missing from or are only suboptimally covered by STTS; in a community effort, German NLP researchers have therefore proposed additions to and modifications of the tagset that will handle these phenomena more appropriately. In addition, they have discussed alternative ways of tag assignment in terms of bipartite tags (stem, token) for historical texts and tripartite tags (lexicon, morphology, distribution) for learner texts. In this article, we report on this ongoing activity, addressing methodological issues and discussing selected phenomena and their treatment in the tagset adaptation process.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain AdaptationMachine TranslationOpinion MiningPart-Of-Speech TaggingSemantic Role LabelingTAGSimilar Papers 제목 키워드 기반
A Universal Part-of-Speech Tagset
To facilitate future research in unsupervised induction of syntactic structure and to standardize best-practices, we propose a tagset that consists of twelve universal part-of-speech categories. In addition to the tagset…
A Fully Expanded Dependency Treebank for Telugu
Treebanks are an essential resource for syntactic parsing. The available Paninian dependency treebank(s) for Telugu is annotated only with inter-chunk dependency relations and not all words of a sentence are part of the …
SentenceA Proposal for a Part-of-Speech Tagset for the Albanian Language
Part-of-speech tagging is a basic step in Natural Language Processing that is often essential. Labeling the word forms of a text with fine-grained word-class information adds new value to it and can be a prerequisite for…
Part-Of-Speech TaggingPart-of-Speech Tagging of Odia Language Using statistical and Deep Learning-Based Approaches
Automatic Part-of-speech (POS) tagging is a preprocessing step of many natural language processing (NLP) tasks such as name entity recognition (NER), speech processing, information extraction, word sense disambiguation, …
Machine TranslationNERPart-Of-Speech TaggingPOS+2Part-of-speech tagging of Swedish texts in the neural era
We train and test five open-source taggers, which use different methods, on three Swedish corpora, which are of comparable size but use different tagsets. The KB-Bert tagger achieves the highest accuracy for part-of-spee…
Morphological TaggingPart-Of-Speech Tagging