Correct Me If You Can: Learning from Error Corrections and Markings
Sequence-to-sequence learning involves a trade-off between signal strength and annotation cost of training data. For example, machine translation data range from costly expert-generated translations that enable supervised learning, to weak quality-judgment feedback that facilitate reinforcement learning. We present the first user study on annotation cost and machine learnability for the less popular annotation mode of error markings. We show that error markings for translations of TED talks from English to German allow precise credit assignment while requiring significantly less human effort than correcting/post-editing, and that error-marked data can be used successfully to fine-tune neural machine translation models.
Code (1)
Tasks
Machine Translationreinforcement-learningReinforcement LearningReinforcement Learning (RL)TranslationSimilar Papers 제목 키워드 기반
Enhancing Supervised Learning with Contrastive Markings in Neural Machine Translation Training
Supervised learning in Neural Machine Translation (NMT) typically follows a teacher forcing paradigm where reference tokens constitute the conditioning context in the model's prediction, instead of its own previous predi…
Machine TranslationNMTTranslationPrompting Large Language Models with Human Error Markings for Self-Correcting Machine Translation
While large language models (LLMs) pre-trained on massive amounts of unpaired language data have reached the state-of-the-art in machine translation (MT) of general domain texts, post-editing (PE) is still required to co…
Machine TranslationTranslationType-Driven Multi-Turn Corrections for Grammatical Error Correction
Grammatical Error Correction (GEC) aims to automatically detect and correct grammatical errors. In this aspect, dominant models are trained by one-iteration learning while performing multiple iterations of corrections du…
Data AugmentationGrammatical Error CorrectionVocal Bursts Type PredictionCollecting fluency corrections for spoken learner English
We present crowdsourced collection of error annotations for transcriptions of spoken learner English. Our emphasis in data collection is on fluency corrections, a more complete correction than has traditionally been aime…
Grammatical Error CorrectionGrammatical Error DetectionMachine TranslationEfficient and Interpretable Grammatical Error Correction with Mixture of Experts
Error type information has been widely used to improve the performance of grammatical error correction (GEC) models, whether for generating corrections, re-ranking them, or combining GEC models. Combining GEC models that…
Grammatical Error CorrectionMixture-of-ExpertsRe-Ranking