The making of the Litkey Corpus, a richly annotated longitudinal corpus of German texts written by primary school children
To date, corpus and computational linguistic work on written language acquisition has mostly dealt with second language learners who have usually already mastered orthography acquisition in their first language. In this paper, we present the Litkey Corpus, a richly-annotated longitudinal corpus of written texts produced by primary school children in Germany from grades 2 to 4. The paper focuses on the (semi-)automatic annotation procedure at various linguistic levels, which include POS tags, features of the word-internal structure (phonemes, syllables, morphemes) and key orthographic features of the target words as well as a categorization of spelling errors. Comprehensive evaluations show that high accuracy was achieved on all levels, making the Litkey Corpus a useful resource for corpus-based research on literacy acquisition of German primary school children and for developing NLP tools for educational purposes. The corpus is freely available under https://www.linguistics.rub.de/litkeycorpus/.
Code (0)
등록된 구현이 없습니다.
Tasks
Language AcquisitionPOSSimilar Papers 제목 키워드 기반
ROMBAC: The Romanian Balanced Annotated Corpus
This article describes the collecting, processing and validation of a large balanced corpus for Romanian. The annotation types and structure of the corpus are briefly reviewed. It was constructed at the Research Institut…
ChunkingLemmatizationPOSPOS TaggingA Richly Annotated, Multilingual Parallel Corpus for Hybrid Machine Translation
In recent years, machine translation (MT) research has focused on investigating how hybrid machine translation as well as system combination approaches can be designed so that the resulting hybrid translations show an im…
Machine TranslationTranslationA Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature
We present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, …
ArticlesParticipant Intervention Comparison Outcome ExtractionPICORedHOT: A Corpus of Annotated Medical Questions, Experiences, and Claims on Social Media
We present Reddit Health Online Talk (RedHOT), a corpus of 22,000 richly annotated social media posts from Reddit spanning 24 health conditions. Annotations include demarcations of spans corresponding to medical claims, …
RetrievalPrague Dependency Treebank -- Consolidated 2.0: Enriching a Complex Annotation Scheme
The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, espe…