paper-with-me

홈 › Papers

The making of the Litkey Corpus, a richly annotated longitudinal corpus of German texts written by primary school children

2019-08-01 · WS 2019 8 · Ronja Laarmann-Quante, Stefanie Dipper, Eva Belke

To date, corpus and computational linguistic work on written language acquisition has mostly dealt with second language learners who have usually already mastered orthography acquisition in their first language. In this paper, we present the Litkey Corpus, a richly-annotated longitudinal corpus of written texts produced by primary school children in Germany from grades 2 to 4. The paper focuses on the (semi-)automatic annotation procedure at various linguistic levels, which include POS tags, features of the word-internal structure (phonemes, syllables, morphemes) and key orthographic features of the target words as well as a categorization of spelling errors. Comprehensive evaluations show that high accuracy was achieved on all levels, making the Litkey Corpus a useful resource for corpus-based research on literacy acquisition of German primary school children and for developing NLP tools for educational purposes. The corpus is freely available under https://www.linguistics.rub.de/litkeycorpus/.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language AcquisitionPOS

Similar Papers 제목 키워드 기반

ROMBAC: The Romanian Balanced Annotated Corpus

2012-05-01 · LREC 2012 5 · Radu Ion, Elena Irimia, Dan {\c{S}}tef{\u{a}}nescu, Dan Tufi{\textcommabelow{s}}

This article describes the collecting, processing and validation of a large balanced corpus for Romanian. The annotation types and structure of the corpus are briefly reviewed. It was constructed at the Research Institut…

ChunkingLemmatizationPOSPOS Tagging

A Richly Annotated, Multilingual Parallel Corpus for Hybrid Machine Translation

2012-05-01 · LREC 2012 5 · Eleftherios Avramidis, Marta R. Costa-juss{\`a}, Christian Federmann, Josef van Genabith 외

In recent years, machine translation (MT) research has focused on investigating how hybrid machine translation as well as system combination approaches can be designed so that the resulting hybrid translations show an im…

Machine TranslationTranslation

A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature

2018-06-11 · ACL 2018 7 · Benjamin Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang 외

We present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, …

ArticlesParticipant Intervention Comparison Outcome ExtractionPICO

RedHOT: A Corpus of Annotated Medical Questions, Experiences, and Claims on Social Media

2022-10-12 · Somin Wadhwa, Vivek Khetan, Silvio Amir, Byron Wallace

We present Reddit Health Online Talk (RedHOT), a corpus of 22,000 richly annotated social media posts from Reddit spanning 24 health conditions. Annotations include demarcations of spans corresponding to medical claims, …

Retrieval

Prague Dependency Treebank -- Consolidated 2.0: Enriching a Complex Annotation Scheme

2026-06-23 · Marie Mikulová, Jiří Mírovský, Milan Straka, Pavlína Synková 외 arxiv

The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, espe…