paper-with-me

홈 › Papers

MixEdit: Revisiting Data Augmentation and Beyond for Grammatical Error Correction

2023-10-18 · Jingheng Ye, Yinghui Li, Yangning Li, Hai-Tao Zheng

Data Augmentation through generating pseudo data has been proven effective in mitigating the challenge of data scarcity in the field of Grammatical Error Correction (GEC). Various augmentation strategies have been widely explored, most of which are motivated by two heuristics, i.e., increasing the distribution similarity and diversity of pseudo data. However, the underlying mechanism responsible for the effectiveness of these strategies remains poorly understood. In this paper, we aim to clarify how data augmentation improves GEC models. To this end, we introduce two interpretable and computationally efficient measures: Affinity and Diversity. Our findings indicate that an excellent GEC data augmentation strategy characterized by high Affinity and appropriate Diversity can better improve the performance of GEC models. Based on this observation, we propose MixEdit, a data augmentation approach that strategically and dynamically augments realistic data, without requiring extra monolingual corpora. To verify the correctness of our findings and the effectiveness of the proposed MixEdit, we conduct experiments on mainstream English and Chinese GEC datasets. The results show that MixEdit substantially improves GEC models and is complementary to traditional data augmentation methods.

📄 PDF Abstract BibTeX arXiv:2310.11671

Code (1)

thukelab/mixedit 공식 구현 pytorch

Tasks

Data AugmentationDiversityGrammatical Error Correction

Similar Papers 제목 키워드 기반

Chinese Grammatical Error Correction Based on Hybrid Models with Data Augmentation

2020-12-01 · AACL (NLP-TEA) 2020 12 · Yi Wang, Ruibin Yuan, Yan‘gen Luo, Yufang Qin 외

A better Chinese Grammatical Error Diagnosis (CGED) system for automatic Grammatical Error Correction (GEC) can benefit foreign Chinese learners and lower Chinese learning barriers. In this paper, we introduce our soluti…

Data AugmentationGrammatical Error Correction

Improving Grammatical Error Correction with Data Augmentation by Editing Latent Representation

2020-12-01 · COLING 2020 8 · Zhaohong Wan, Xiaojun Wan, Wenguang Wang

The incorporation of data augmentation method in grammatical error correction task has attracted much attention. However, existing data augmentation methods mainly apply noise to tokens, which leads to the lack of divers…

Data AugmentationDiversityGrammatical Error Correction

Revisiting Classification Taxonomy for Grammatical Errors

2025-02-17 · Deqing Zou, Jingheng Ye, Yulu Liu, Yu Wu 외

Grammatical error classification plays a crucial role in language learning systems, but existing classification taxonomies often lack rigorous validation, leading to inconsistencies and unreliable feedback. In this paper…

Classification

Revisiting Grammatical Error Correction Evaluation and Beyond

2022-11-03 · Peiyuan Gong, Xuebo Liu, Heyan Huang, Min Zhang

Pretraining-based (PT-based) automatic evaluation metrics (e.g., BERTScore and BARTScore) have been widely used in several sentence generation tasks (e.g., machine translation and text summarization) due to their better …

Grammatical Error CorrectionMachine TranslationSentenceText Summarization

Grammaticality Judgments in Humans and Language Models: Revisiting Generative Grammar with LLMs

2025-12-11 · Lars G. B. Johnsen arxiv

What counts as evidence for syntactic structure? In traditional generative grammar, systematic contrasts in grammaticality such as subject-auxiliary inversion and the licensing of parasitic gaps are taken as evidence for…