Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing
Low-resource indigenous languages often lack the parallel corpora required for effective neural machine translation (NMT). Synthetic data generation offers a practical strategy for mitigating this limitation in data-scarce settings. In this work, we augment curated parallel datasets for indigenous languages of the Americas with synthetic sentence pairs generated using a high-capacity multilingual translation model. We fine-tune a multilingual mBART model on curated-only and synthetically augmented data and evaluate translation quality using chrF++, the primary metric used in recent AmericasNLP shared tasks for agglutinative languages. We further apply language-specific preprocessing, including orthographic normalization and noise-aware filtering, to reduce corpus artifacts. Experiments on Guarani-Spanish and Quechua-Spanish translation show consistent chrF++ improvements from synthetic data augmentation, while diagnostic experiments on Aymara highlight the limitations of generic preprocessing for highly agglutinative languages.
Code (0)
등록된 구현이 없습니다.
Tasks
Synthetic Data GenerationMachine TranslationData AugmentationSimilar Papers 제목 키워드 기반
Revitalization of Indigenous Languages through Pre-processing and Neural Machine Translation: The case of Inuktitut
Indigenous languages have been very challenging when dealing with NLP tasks and applications because of multiple reasons. These languages, in linguistic typology, are polysynthetic and highly inflected with rich morphoph…
Machine TranslationMorphological AnalysisTranslationEthical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers
In recent years machine translation has become very successful for high-resource language pairs. This has also sparked new interest in research on the automatic translation of low-resource languages, including Indigenous…
Machine TranslationTranslationIndT5: A Text-to-Text Transformer for 10 Indigenous Languages
Transformer language models have become fundamental components of natural language processing based pipelines. Although several Transformer models have been introduced to serve many languages, there is a shortage of mode…
Language ModelingLanguage ModellingMachine TranslationTranslationParallel Corpus for Indigenous Language Translation: Spanish-Mazatec and Spanish-Mixtec
In this paper, we present a parallel Spanish-Mazatec and Spanish-Mixtec corpus for machine translation (MT) tasks, where Mazatec and Mixtec are two indigenous Mexican languages. We evaluated the usability of the collecte…
Few-Shot LearningMachine TranslationTransfer LearningTranslationThe Indigenous Languages Technology project at NRC Canada: An empowerment-oriented approach to developing language software
This paper surveys the first, three-year phase of a project at the National Research Council of Canada that is developing software to assist Indigenous communities in Canada in preserving their languages and extending th…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+2