FOLK-Gold ― A Gold Standard for Part-of-Speech-Tagging of Spoken German
In this paper, we present a GOLD standard of part-of-speech tagged transcripts of spoken German. The GOLD standard data consists of four annotation layers ― transcription (modified orthography), normalization (standard orthography), lemmatization and POS tags ― all of which have undergone careful manual quality control. It comes with guidelines for the manual POS annotation of transcripts of German spoken data and an extended version of the STTS (Stuttgart T{\"u}bingen Tagset) which accounts for phenomena typically found in spontaneous spoken German. The GOLD standard was developed on the basis of the Research and Teaching Corpus of Spoken German, FOLK, and is, to our knowledge, the first such dataset based on a wide variety of spontaneous and authentic interaction types. It can be used as a basis for further development of language technology and corpus linguistic applications for German spoken language.
Code (0)
등록된 구현이 없습니다.
Tasks
LemmatizationPart-Of-Speech TaggingPOSSimilar Papers 제목 키워드 기반
Cross-lingual Annotation Projection Is Effective for Neural Part-of-Speech Tagging
We tackle the important task of part-of-speech tagging using a neural model in the zero-resource scenario, where we have no access to gold-standard POS training data. We compare this scenario with the low-resource scenar…
Part-Of-Speech TaggingPOSPOS TaggingAlbanian Part-of-Speech Tagging: Gold Standard and Evaluation
Towards High Accuracy Named Entity Recognition for Icelandic
We report on work in progress which consists of annotating an Icelandic corpus for named entities (NEs) and using it for training a named entity recognizer based on a Bidirectional Long Short-Term Memory model. Currently…
Miscellaneousnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3Turning silver into gold: error-focused corpus reannotation with active learning
While high quality gold standard annotated corpora are crucial for most tasks in natural language processing, many annotated corpora published in recent years, created by annotators or tools, contains noisy annotations. …
Active LearningDocument ClassificationPart-Of-Speech TaggingZero-Shot Text-to-Speech as Golden Speech Generator: A Systematic Framework and its Applicability in Automatic Pronunciation Assessment
Second language (L2) learners can improve their pronunciation by imitating golden speech, especially when the speech that aligns with their respective speech characteristics. This study explores the hypothesis that learn…
text-to-speechText to Speech