paper-with-me

Papers

Grammatical Error Correction Using Pseudo Learner Corpus Considering Learner's Error Tendency

2020-07-01 · ACL 2020 6 · Yujin Takahashi, Satoru Katsumata, Mamoru Komachi

Recently, several studies have focused on improving the performance of grammatical error correction (GEC) tasks using pseudo data. However, a large amount of pseudo data are required to train an accurate GEC model. To address the limitations of language and computational resources, we assume that introducing pseudo errors into sentences similar to those written by the language learners is more efficient, rather than incorporating random pseudo errors into monolingual data. In this regard, we study the effect of pseudo data on GEC task performance using two approaches. First, we extract sentences that are similar to the learners{'} sentences from monolingual data. Second, we generate realistic pseudo errors by considering error types that learners often make. Based on our comparative results, we observe that F0.5 scores for the Russian GEC task are significantly improved.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Grammatical Error Correction

Similar Papers 제목 키워드 기반

(Almost) Unsupervised Grammatical Error Correction using Synthetic Comparable Corpus

2019-08-01 · WS 2019 8 · Satoru Katsumata, Mamoru Komachi

We introduce unsupervised techniques based on phrase-based statistical machine translation for grammatical error correction (GEC) trained on a pseudo learner corpus created by Google Translation. We verified our GEC syst…

Grammatical Error CorrectionMachine TranslationTranslation

Towards Unsupervised Grammatical Error Correction using Statistical Machine Translation with Synthetic Comparable Corpus

2019-07-23 · Satoru Katsumata, Mamoru Komachi

We introduce unsupervised techniques based on phrase-based statistical machine translation for grammatical error correction (GEC) trained on a pseudo learner corpus created by Google Translation. We verified our GEC syst…

Grammatical Error CorrectionMachine TranslationTranslation

Construction of an Evaluation Corpus for Grammatical Error Correction for Learners of Japanese as a Second Language

2020-05-01 · LREC 2020 5 · Aomi Koyama, Tomoshige Kiyuna, Kenji Kobayashi, Mio Arai 외

The NAIST Lang-8 Learner Corpora (Lang-8 corpus) is one of the largest second-language learner corpora. The Lang-8 corpus is suitable as a training dataset for machine translation-based grammatical error correction syste…

Grammatical Error CorrectionMachine TranslationNMTTranslation

YACLC: A Chinese Learner Corpus with Multidimensional Annotation

2021-12-30 · Yingying Wang, Cunliang Kong, Liner Yang, Yijun Wang 외

Learner corpus collects language data produced by L2 learners, that is second or foreign-language learners. This resource is of great relevance for second language acquisition research, foreign-language teaching, and aut…

Grammatical Error CorrectionLanguage AcquisitionSentence

Semi-automatically Annotated Learner Corpus for Russian

2022-06-01 · LREC 2022 6 · Anisia Katinskaia, Maria Lebedeva, Jue Hou, Roman Yangarber

We present ReLCo— the Revita Learner Corpus—a new semi-automatically annotated learner corpus for Russian. The corpus was collected while several thousand L2 learners were performing exercises using the Revita language-l…

Grammatical Error CorrectionGrammatical Error Detection