On Assessing and Developing Spoken ’Grammatical Error Correction’ Systems
Spoken ‘grammatical error correction’ (SGEC) is an important process to provide feedback for second language learning. Due to a lack of end-to-end training data, SGEC is often implemented as a cascaded, modular system, consisting of speech recognition, disfluency removal, and grammatical error correction (GEC). This cascaded structure enables efficient use of training data for each module. It is, however, difficult to compare and evaluate the performance of individual modules as preceeding modules may introduce errors. For example the GEC module input depends on the output of non-native speech recognition and disfluency detection, both challenging tasks for learner data.This paper focuses on the assessment and development of SGEC systems. We first discuss metrics for evaluating SGEC, both individual modules and the overall system. The system-level metrics enable tuning for optimal system performance. A known issue in cascaded systems is error propagation between modules.To mitigate this problem semi-supervised approaches and self-distillation are investigated. Lastly, when SGEC system gets deployed it is important to give accurate feedback to users. Thus, we apply filtering to remove edits with low-confidence, aiming to improve overall feedback precision. The performance metrics are examined on a Linguaskill multi-level data set, which includes the original non-native speech, manual transcriptions and reference grammatical error corrections, to enable system analysis and development.
Code (0)
등록된 구현이 없습니다.
Tasks
Grammatical Error Correctionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback
We introduce the Speak & Improve Corpus 2025, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning plat…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Grammatical Error Correctionspeech-recognition+1IMPARA: Impact-Based Metric for GEC Using Parallel Data
Automatic evaluation of grammatical error correction (GEC) is essential in developing useful GEC systems. Existing methods for automatic evaluation require multiple reference sentences or manual scores. However, such res…
Grammatical Error CorrectionData Augmentation for Spoken Grammatical Error Correction
While there exist strong benchmark datasets for grammatical error correction (GEC), high-quality annotated spoken datasets for Spoken GEC (SGEC) are still under-resourced. In this paper, we propose a fully automated meth…
Grammatical Error CorrectionData AugmentationGrammatical error detection in transcriptions of spoken English
We describe the collection of transcription corrections and grammatical error annotations for the CrowdED Corpus of spoken English monologues on business topics. The corpus recordings were crowdsourced from native speake…
Grammatical Error CorrectionGrammatical Error DetectionTowards End-to-End Spoken Grammatical Error Correction
Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies …
Grammatical Error Correctionspeech-recognitionSpeech Recognition