Data Augmentation Techniques for Machine Translation of Code-Switched Texts: A Comparative Study
Code-switching (CSW) text generation has been receiving increasing attention as a solution to address data scarcity. In light of this growing interest, we need more comprehensive studies comparing different augmentation approaches. In this work, we compare three popular approaches: lexical replacements, linguistic theories, and back-translation (BT), in the context of Egyptian Arabic-English CSW. We assess the effectiveness of the approaches on machine translation and the quality of augmentations through human evaluation. We show that BT and CSW predictive-based lexical replacement, being trained on CSW parallel data, perform best on both tasks. Linguistic theories and random lexical replacement prove to be effective in the lack of CSW parallel data, where both approaches achieve similar results.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationMachine TranslationText GenerationTranslationSimilar Papers 제목 키워드 기반
Textual Augmentation Techniques Applied to Low Resource Machine Translation: Case of Swahili
In this work we investigate the impact of applying textual data augmentation tasks to low resource machine translation. There has been recent interest in investigating approaches for training systems for languages with l…
Data AugmentationMachine TranslationNMTtext-classification+2From Scarcity to Efficiency: Investigating the Effects of Data Augmentation on African Machine Translation
The linguistic diversity across the African continent presents different challenges and opportunities for machine translation. This study explores the effects of data augmentation techniques in improving translation syst…
Machine TranslationData AugmentationEnhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation
Direct speech-to-speech translation (S2ST) models suffer from data scarcity issues as there exists little parallel S2ST data, compared to the amount of data available for conventional cascaded systems that consist of aut…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDecoder+9The Impact of Code-switched Synthetic Data Quality is Task Dependent: Insights from MT and ASR
Code-switching, the act of alternating between languages, emerged as a prevalent global phenomenon that needs to be addressed for building user-friendly language technologies. A main bottleneck in this pursuit is data sc…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMachine Translation+3Low-resource neural machine translation with morphological modeling
Morphological modeling in neural machine translation (NMT) is a promising approach to achieving open-vocabulary machine translation for morphologically-rich languages. However, existing methods such as sub-word tokenizat…
Data AugmentationDecoderLow Resource Neural Machine TranslationLow-Resource Neural Machine Translation+4