paper-with-me

홈 › Papers

FOLK-Gold ― A Gold Standard for Part-of-Speech-Tagging of Spoken German

2016-05-01 · LREC 2016 5 · Swantje Westpfahl, Thomas Schmidt

In this paper, we present a GOLD standard of part-of-speech tagged transcripts of spoken German. The GOLD standard data consists of four annotation layers ― transcription (modified orthography), normalization (standard orthography), lemmatization and POS tags ― all of which have undergone careful manual quality control. It comes with guidelines for the manual POS annotation of transcripts of German spoken data and an extended version of the STTS (Stuttgart T{\"u}bingen Tagset) which accounts for phenomena typically found in spontaneous spoken German. The GOLD standard was developed on the basis of the Research and Teaching Corpus of Spoken German, FOLK, and is, to our knowledge, the first such dataset based on a wide variety of spontaneous and authentic interaction types. It can be used as a basis for further development of language technology and corpus linguistic applications for German spoken language.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

LemmatizationPart-Of-Speech TaggingPOS

Similar Papers 제목 키워드 기반

Cross-lingual Annotation Projection Is Effective for Neural Part-of-Speech Tagging

2019-06-01 · WS 2019 6 · Matthias Huck, Diana Dutka, Alex Fraser, er

We tackle the important task of part-of-speech tagging using a neural model in the zero-resource scenario, where we have no access to gold-standard POS training data. We compare this scenario with the low-resource scenar…

Part-Of-Speech TaggingPOSPOS Tagging

Albanian Part-of-Speech Tagging: Gold Standard and Evaluation

2018-05-01 · LREC 2018 5 · Besim Kabashi, Thomas Proisl
Morphological AnalysisPart-Of-Speech Tagging

Towards High Accuracy Named Entity Recognition for Icelandic

2019-09-01 · WS (NoDaLiDa) 2019 9 · Svanhvít Lilja Ingólfsdóttir, Sigurjón Þorsteinsson, Hrafn Loftsson

We report on work in progress which consists of annotating an Icelandic corpus for named entities (NEs) and using it for training a named entity recognizer based on a Bidirectional Long Short-Term Memory model. Currently…

Miscellaneousnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3

Turning silver into gold: error-focused corpus reannotation with active learning

2019-09-01 · RANLP 2019 9 · Pierre Andr{\'e} M{\'e}nard, Antoine Mougeot

While high quality gold standard annotated corpora are crucial for most tasks in natural language processing, many annotated corpora published in recent years, created by annotators or tools, contains noisy annotations. …

Active LearningDocument ClassificationPart-Of-Speech Tagging

Zero-Shot Text-to-Speech as Golden Speech Generator: A Systematic Framework and its Applicability in Automatic Pronunciation Assessment

2024-09-11 · Tien-Hong Lo, Meng-Ting Tsai, Berlin Chen

Second language (L2) learners can improve their pronunciation by imitating golden speech, especially when the speech that aligns with their respective speech characteristics. This study explores the hypothesis that learn…

text-to-speechText to Speech