paper-with-me

Papers

Parallel Corpus for Indigenous Language Translation: Spanish-Mazatec and Spanish-Mixtec

2023-05-27 · Atnafu Lambebo Tonja, Christian Maldonado-Sifuentes, David Alejandro Mendoza Castillo, Olga Kolesnikova, Noé Castro-Sánchez, Grigori Sidorov, Alexander Gelbukh

In this paper, we present a parallel Spanish-Mazatec and Spanish-Mixtec corpus for machine translation (MT) tasks, where Mazatec and Mixtec are two indigenous Mexican languages. We evaluated the usability of the collected corpus using three different approaches: transformer, transfer learning, and fine-tuning pre-trained multilingual MT models. Fine-tuning the Facebook M2M100-48 model outperformed the other approaches, with BLEU scores of 12.09 and 22.25 for Mazatec-Spanish and Spanish-Mazatec translations, respectively, and 16.75 and 22.15 for Mixtec-Spanish and Spanish-Mixtec translations, respectively. The findings show that the dataset size (9,799 sentences in Mazatec and 13,235 sentences in Mixtec) affects translation performance and that indigenous languages work better when used as target languages. The findings emphasize the importance of creating parallel corpora for indigenous languages and fine-tuning models for low-resource translation tasks. Future research will investigate zero-shot and few-shot learning approaches to further improve translation performance in low-resource settings. The dataset and scripts are available at \url{https://github.com/atnafuatx/Machine-Translation-Resources}

📄 PDF Abstract BibTeX arXiv:2305.17404

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot LearningMachine TranslationTransfer LearningTranslation

Similar Papers 제목 키워드 기반

Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing

2026-01-06 · Aashish Dhawan, Christopher Driggers-Ellis, Christan Grant, Daisy Zhe Wang arxiv

Low-resource indigenous languages often lack the parallel corpora required for effective neural machine translation (NMT). Synthetic data generation offers a practical strategy for mitigating this limitation in data-scar…

Synthetic Data GenerationMachine TranslationData Augmentation

IndT5: A Text-to-Text Transformer for 10 Indigenous Languages

2021-04-04 · NAACL (AmericasNLP) 2021 6 · El Moatez Billah Nagoudi, Wei-Rui Chen, Muhammad Abdul-Mageed, Hasan Cavusogl

Transformer language models have become fundamental components of natural language processing based pipelines. Although several Transformer models have been introduced to serve many languages, there is a shortage of mode…

Language ModelingLanguage ModellingMachine TranslationTranslation

Improving Neural Machine Translation of Indigenous Languages with Multilingual Transfer Learning

2022-05-14 · Wei-Rui Chen, Muhammad Abdul-Mageed

Machine translation (MT) involving Indigenous languages, including those possibly endangered, is challenging due to lack of sufficient parallel data. We describe an approach exploiting bilingual and multilingual pretrain…

Data AugmentationMachine TranslationTransfer LearningTranslation

Peru is Multilingual, Its Machine Translation Should Be Too?

2021-06-01 · NAACL (AmericasNLP) 2021 6 · Arturo Oncevay

Peru is a multilingual country with a long history of contact between the indigenous languages and Spanish. Taking advantage of this context for machine translation is possible with multilingual approaches for learning b…

Machine TranslationTranslation

How does discourse affect Spanish-Chinese Translation? A case study based on a Spanish-Chinese parallel corpus

2020-11-01 · EMNLP (CODI) 2020 11 · Shuyuan Cao

With their huge speaking populations in the world, Spanish and Chinese occupy important positions in linguistic studies. Since the two languages come from different language systems, the translation between Spanish and C…

Translation