paper-with-me

홈 › Papers

Building a learner corpus

2012-05-01 · LREC 2012 5 · Jirka Hana, Alex Rosen, R, Barbora {\v{S}}tindlov{\'a}, Petr J{\"a}ger

The paper describes a corpus of texts produced by non-native speakers of Czech. We discuss its annotation scheme, consisting of three interlinked levels to cope with a wide range of error types present in the input. Each level corrects different types of errors; links between the levels allow capturing errors in word order and complex discontinuous expressions. Errors are not only corrected, but also classified. The annotation scheme is tested on a doubly-annotated sample of approx. 10,000 words with fair inter-annotator agreement results. We also explore options of application of automated linguistic annotation tools (taggers, spell checkers and grammar checkers) on the learner text to support or even substitute manual annotation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language Acquisition

Similar Papers 제목 키워드 기반

Semi-automatically Annotated Learner Corpus for Russian

2022-06-01 · LREC 2022 6 · Anisia Katinskaia, Maria Lebedeva, Jue Hou, Roman Yangarber

We present ReLCo— the Revita Learner Corpus—a new semi-automatically annotated learner corpus for Russian. The corpus was collected while several thousand L2 learners were performing exercises using the Revita language-l…

Grammatical Error CorrectionGrammatical Error Detection

Building a Large Annotated Corpus of Learner English: The NUS Corpus of Learner English

2013-06-01 · WS 2013 6 · Daniel Dahlmeier, Hwee Tou Ng, Siew Mei Wu
Grammatical Error Correction

Building a learner corpus for Russian

2016-11-01 · WS 2016 11 · Ekaterina Rakhilina, Anastasia Vyrenkova, Elmira Mustakimova, Alina Ladygina 외
Language AcquisitionLanguage Identification

Building a TOCFL Learner Corpus for Chinese Grammatical Error Diagnosis

2018-05-01 · LREC 2018 5 · Lung-Hao Lee, Yuen-Hsien Tseng, Li-Ping Chang
Grammatical Error DetectionLanguage AcquisitionLanguage Identification

From Labels to Facets: Building a Taxonomically Enriched Turkish Learner Corpus

2026-01-30 · Elif Sayar, Tolgahan Türker, Anna Golynskaia Knezhevich, Bihter Dereli 외 arxiv

In terms of annotation structure, most learner corpora rely on holistic flat label inventories which, even when extensive, do not explicitly separate multiple linguistic dimensions. This makes linguistically deep annotat…