paper-with-me

Papers

Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing

2026-01-06 · Aashish Dhawan, Christopher Driggers-Ellis, Christan Grant, Daisy Zhe Wang arxiv

Low-resource indigenous languages often lack the parallel corpora required for effective neural machine translation (NMT). Synthetic data generation offers a practical strategy for mitigating this limitation in data-scarce settings. In this work, we augment curated parallel datasets for indigenous languages of the Americas with synthetic sentence pairs generated using a high-capacity multilingual translation model. We fine-tune a multilingual mBART model on curated-only and synthetically augmented data and evaluate translation quality using chrF++, the primary metric used in recent AmericasNLP shared tasks for agglutinative languages. We further apply language-specific preprocessing, including orthographic normalization and noise-aware filtering, to reduce corpus artifacts. Experiments on Guarani-Spanish and Quechua-Spanish translation show consistent chrF++ improvements from synthetic data augmentation, while diagnostic experiments on Aymara highlight the limitations of generic preprocessing for highly agglutinative languages.

📄 PDF Abstract BibTeX arXiv:2601.03135

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data GenerationMachine TranslationData Augmentation

Similar Papers 제목 키워드 기반

Revitalization of Indigenous Languages through Pre-processing and Neural Machine Translation: The case of Inuktitut

2020-12-01 · COLING 2020 8 · Tan Ngoc Le, Fatiha Sadat

Indigenous languages have been very challenging when dealing with NLP tasks and applications because of multiple reasons. These languages, in linguistic typology, are polysynthetic and highly inflected with rich morphoph…

Machine TranslationMorphological AnalysisTranslation

Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers

2023-05-31 · Manuel Mager, Elisabeth Mager, Katharina Kann, Ngoc Thang Vu

In recent years machine translation has become very successful for high-resource language pairs. This has also sparked new interest in research on the automatic translation of low-resource languages, including Indigenous…

Machine TranslationTranslation

IndT5: A Text-to-Text Transformer for 10 Indigenous Languages

2021-04-04 · NAACL (AmericasNLP) 2021 6 · El Moatez Billah Nagoudi, Wei-Rui Chen, Muhammad Abdul-Mageed, Hasan Cavusogl

Transformer language models have become fundamental components of natural language processing based pipelines. Although several Transformer models have been introduced to serve many languages, there is a shortage of mode…

Language ModelingLanguage ModellingMachine TranslationTranslation

Parallel Corpus for Indigenous Language Translation: Spanish-Mazatec and Spanish-Mixtec

2023-05-27 · Atnafu Lambebo Tonja, Christian Maldonado-Sifuentes, David Alejandro Mendoza Castillo, Olga Kolesnikova 외

In this paper, we present a parallel Spanish-Mazatec and Spanish-Mixtec corpus for machine translation (MT) tasks, where Mazatec and Mixtec are two indigenous Mexican languages. We evaluated the usability of the collecte…

Few-Shot LearningMachine TranslationTransfer LearningTranslation

The Indigenous Languages Technology project at NRC Canada: An empowerment-oriented approach to developing language software

2020-12-01 · COLING 2020 8 · Roland Kuhn, Fineen Davis, Alain D{\'e}silets, Eric Joanis 외

This paper surveys the first, three-year phase of a project at the National Research Council of Canada that is developing software to assist Indigenous communities in Canada in preserving their languages and extending th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+2