paper-with-me

Papers

Data Augmentation for Spoken Grammatical Error Correction

2025-07-25 · Penny Karanasou, Mengjie Qian, Stefano Bannò, Mark J. F. Gales, Kate M. Knill arxiv

While there exist strong benchmark datasets for grammatical error correction (GEC), high-quality annotated spoken datasets for Spoken GEC (SGEC) are still under-resourced. In this paper, we propose a fully automated method to generate audio-text pairs with grammatical errors and disfluencies. Moreover, we propose a series of objective metrics that can be used to evaluate the generated data and choose the more suitable dataset for SGEC. The goal is to generate an augmented dataset that maintains the textual and acoustic characteristics of the original data while providing new types of errors. This augmented dataset should augment and enrich the original corpus without altering the language assessment scores of the second language (L2) learners. We evaluate the use of the augmented corpus both for written GEC (the text part) and for SGEC (the audio-text pairs). Our experiments are conducted on the S\&I Corpus, the first publicly available speech dataset with grammar error annotations.

📄 PDF Abstract BibTeX arXiv:2507.19374

Code (0)

등록된 구현이 없습니다.

Tasks

Grammatical Error CorrectionData Augmentation

Similar Papers 제목 키워드 기반

Chinese Grammatical Error Correction Based on Hybrid Models with Data Augmentation

2020-12-01 · AACL (NLP-TEA) 2020 12 · Yi Wang, Ruibin Yuan, Yan‘gen Luo, Yufang Qin 외

A better Chinese Grammatical Error Diagnosis (CGED) system for automatic Grammatical Error Correction (GEC) can benefit foreign Chinese learners and lower Chinese learning barriers. In this paper, we introduce our soluti…

Data AugmentationGrammatical Error Correction

Improving Grammatical Error Correction with Data Augmentation by Editing Latent Representation

2020-12-01 · COLING 2020 8 · Zhaohong Wan, Xiaojun Wan, Wenguang Wang

The incorporation of data augmentation method in grammatical error correction task has attracted much attention. However, existing data augmentation methods mainly apply noise to tokens, which leads to the lack of divers…

Data AugmentationDiversityGrammatical Error Correction

Grammatical error detection in transcriptions of spoken English

2020-12-01 · COLING 2020 8 · Andrew Caines, Christian Bentz, Kate Knill, Marek Rei 외

We describe the collection of transcription corrections and grammatical error annotations for the CrowdED Corpus of spoken English monologues on business topics. The corpus recordings were crowdsourced from native speake…

Grammatical Error CorrectionGrammatical Error Detection

Towards End-to-End Spoken Grammatical Error Correction

2023-11-09 · Stefano Bannò, Rao Ma, Mengjie Qian, Kate M. Knill 외

Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies …

Grammatical Error Correctionspeech-recognitionSpeech Recognition

Collecting fluency corrections for spoken learner English

2017-09-01 · WS 2017 9 · Andrew Caines, Emma Flint, Paula Buttery

We present crowdsourced collection of error annotations for transcriptions of spoken learner English. Our emphasis in data collection is on fluency corrections, a more complete correction than has traditionally been aime…

Grammatical Error CorrectionGrammatical Error DetectionMachine Translation