Automatic Classification of Russian Learner Errors
Grammatical Error Correction systems are typically evaluated overall, without taking into consideration performance on individual error types because system output is not annotated with respect to error type. We introduce a tool that automatically classifies errors in Russian learner texts. The tool takes an edit pair consisting of the original token(s) and the corresponding replacement and provides a grammatical error category. Manual evaluation of the output reveals that in more than 93% of cases the error categories are judged as correct or acceptable. We apply the tool to carry out a fine-grained evaluation on the performance of two error correction systems for Russian.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationGrammatical Error CorrectionSimilar Papers 제목 키워드 기반
Semi-automatically Annotated Learner Corpus for Russian
We present ReLCo— the Revita Learner Corpus—a new semi-automatically annotated learner corpus for Russian. The corpus was collected while several thousand L2 learners were performing exercises using the Revita language-l…
Grammatical Error CorrectionGrammatical Error DetectionToward a Paradigm Shift in Collection of Learner Corpora
We present the first version of the longitudinal Revita Learner Corpus (ReLCo), for Russian. In contrast to traditional learner corpora, ReLCo is collected and annotated fully automatically, while students perform exerci…
Classifying Syntactic Errors in Learner Language
We present a method for classifying syntactic errors in learner language, namely errors whose correction alters the morphosyntactic structure of a sentence. The methodology builds on the established Universal Dependencie…
ClassificationGeneral ClassificationGrammatical Error CorrectionSentenceRILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts
Many errors in student essays can be explained by influence from the native language (L1). L1 interference refers to errors influenced by a speaker's first language, such as using stadion instead of stadium, reflecting l…
Multiple Admissibility: Judging Grammaticality using Unlabeled Data in Language Learning
We present our work on the problem of Multiple Admissibility (MA) in language learning. Multiple Admissibility occurs in many languages when more than one grammatical form of a word fits syntactically and semantically in…