paper-with-me

홈 › Papers

Improving Neural Machine Translation of Indigenous Languages with Multilingual Transfer Learning

2022-05-14 · Wei-Rui Chen, Muhammad Abdul-Mageed

Machine translation (MT) involving Indigenous languages, including those possibly endangered, is challenging due to lack of sufficient parallel data. We describe an approach exploiting bilingual and multilingual pretrained MT models in a transfer learning setting to translate from Spanish to ten South American Indigenous languages. Our models set new SOTA on five out of the ten language pairs we consider, even doubling performance on one of these five pairs. Unlike previous SOTA that perform data augmentation to enlarge the train sets, we retain the low-resource setting to test the effectiveness of our models under such a constraint. In spite of the rarity of linguistic information available about the Indigenous languages, we offer a number of quantitative and qualitative analyses (e.g., as to morphology, tokenization, and orthography) to contextualize our results.

📄 PDF Abstract BibTeX arXiv:2205.06993

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationMachine TranslationTransfer LearningTranslation

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Enhancing Translation for Indigenous Languages: Experiments with Multilingual Models

2023-05-27 · Atnafu Lambebo Tonja, Hellina Hailu Nigatu, Olga Kolesnikova, Grigori Sidorov 외

This paper describes CIC NLP's submission to the AmericasNLP 2023 Shared Task on machine translation systems for indigenous languages of the Americas. We present the system descriptions for three methods. We used two mul…

Machine TranslationTransfer LearningTranslation

Parallel Corpus for Indigenous Language Translation: Spanish-Mazatec and Spanish-Mixtec

2023-05-27 · Atnafu Lambebo Tonja, Christian Maldonado-Sifuentes, David Alejandro Mendoza Castillo, Olga Kolesnikova 외

In this paper, we present a parallel Spanish-Mazatec and Spanish-Mixtec corpus for machine translation (MT) tasks, where Mazatec and Mixtec are two indigenous Mexican languages. We evaluated the usability of the collecte…

Few-Shot LearningMachine TranslationTransfer LearningTranslation

Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing

2026-01-06 · Aashish Dhawan, Christopher Driggers-Ellis, Christan Grant, Daisy Zhe Wang arxiv

Low-resource indigenous languages often lack the parallel corpora required for effective neural machine translation (NMT). Synthetic data generation offers a practical strategy for mitigating this limitation in data-scar…

Synthetic Data GenerationMachine TranslationData Augmentation

Peru is Multilingual, Its Machine Translation Should Be Too?

2021-06-01 · NAACL (AmericasNLP) 2021 6 · Arturo Oncevay

Peru is a multilingual country with a long history of contact between the indigenous languages and Spanish. Taking advantage of this context for machine translation is possible with multilingual approaches for learning b…

Machine TranslationTranslation

IndT5: A Text-to-Text Transformer for 10 Indigenous Languages

2021-04-04 · NAACL (AmericasNLP) 2021 6 · El Moatez Billah Nagoudi, Wei-Rui Chen, Muhammad Abdul-Mageed, Hasan Cavusogl

Transformer language models have become fundamental components of natural language processing based pipelines. Although several Transformer models have been introduced to serve many languages, there is a shortage of mode…

Language ModelingLanguage ModellingMachine TranslationTranslation