Surprise Language Challenge: Developing a Neural Machine Translation System between Pashto and English in Two Months
In the media industry and the focus of global reporting can shift overnight. There is a compelling need to be able to develop new machine translation systems in a short period of time and in order to more efficiently cover quickly developing stories. As part of the EU project GoURMET and which focusses on low-resource machine translation and our media partners selected a surprise language for which a machine translation system had to be built and evaluated in two months(February and March 2021). The language selected was Pashto and an Indo-Iranian language spoken in Afghanistan and Pakistan and India. In this period we completed the full pipeline of development of a neural machine translation system: data crawling and cleaning and aligning and creating test sets and developing and testing models and and delivering them to the user partners. In this paperwe describe rapid data creation and experiments with transfer learning and pretraining for this low-resource language pair. We find that starting from an existing large model pre-trained on 50languages leads to far better BLEU scores than pretraining on one high-resource language pair with a smaller model. We also present human evaluation of our systems and which indicates that the resulting systems perform better than a freely available commercial system when translating from English into Pashto direction and and similarly when translating from Pashto into English.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTransfer LearningTranslationSimilar Papers 제목 키워드 기반
Language Resource Building and English-to-Mizo Neural Machine Translation Encountering Tonal Words
Multilingual country like India has an enormous linguistic diversity and has an increasing demand towards developing language resources such that it will outreach in various natural language processing applications like …
DiversityMachine TranslationTranslationQRev: Machine Translation of User Reviews: What Influences the Translation Quality?
This project aims to identify the important aspects of translation quality of user reviews which will represent a starting point for developing better automatic MT metrics and challenge test sets, and will be also helpfu…
Machine TranslationTranslationBhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
This paper focuses on developing translation models and related applications for 36 Indian languages, including Assamese, Awadhi, Bengali, Bhojpuri, Braj, Bodo, Dogri, English, Konkani, Gondi, Gujarati, Hindi, Hinglish, …
Automatic Post-EditingData AugmentationMachine TranslationTranslationBenchmarking Neural and Statistical Machine Translation on Low-Resource African Languages
Research in machine translation (MT) is developing at a rapid pace. However, most work in the community has focused on languages where large amounts of digital resources are available. In this study, we benchmark state o…
BenchmarkingMachine TranslationNMTTranslationManipuri-English Machine Translation using Comparable Corpus
Unsupervised Machine Translation (MT) model, which has the ability to perform MT without parallel sentences using comparable corpora, is becoming a promising approach for developing MT in low-resource languages. However,…
Machine TranslationTranslationUnsupervised Machine Translation