paper-with-me

홈 › Papers

Exploring Diversity in Back Translation for Low-Resource Machine Translation

2022-06-01 · DeepLo 2022 7 · Laurie Burchell, Alexandra Birch, Kenneth Heafield

Back translation is one of the most widely used methods for improving the performance of neural machine translation systems. Recent research has sought to enhance the effectiveness of this method by increasing the 'diversity' of the generated translations. We argue that the definitions and metrics used to quantify 'diversity' in previous work have been insufficient. This work puts forward a more nuanced framework for understanding diversity in training data, splitting it into lexical diversity and syntactic diversity. We present novel metrics for measuring these different aspects of diversity and carry out empirical analysis into the effect of these types of diversity on final neural machine translation model performance for low-resource English$\leftrightarrow$Turkish and mid-resource English$\leftrightarrow$Icelandic. Our findings show that generating back translation using nucleus sampling results in higher final model performance, and that this method of generation has high levels of both lexical and syntactic diversity. We also find evidence that lexical diversity is more important than syntactic for back translation performance.

📄 PDF Abstract BibTeX arXiv:2206.00564

Code (1)

laurieburchell/exploring-diversity-bt 공식 구현

Tasks

DiversityMachine TranslationTranslation

Similar Papers 제목 키워드 기반

Language Resource Building and English-to-Mizo Neural Machine Translation Encountering Tonal Words

2022-06-01 · WILDRE (LREC) 2022 6 · Vanlalmuansangi Khenglawt, Sahinur Rahman Laskar, Santanu Pal, Partha Pakray 외

Multilingual country like India has an enormous linguistic diversity and has an increasing demand towards developing language resources such that it will outreach in various natural language processing applications like …

DiversityMachine TranslationTranslation

From Scarcity to Efficiency: Investigating the Effects of Data Augmentation on African Machine Translation

2025-09-09 · Mardiyyah Oduwole, Oluwatosin Olajide, Jamiu Suleiman, Faith Hunja 외 arxiv

The linguistic diversity across the African continent presents different challenges and opportunities for machine translation. This study explores the effects of data augmentation techniques in improving translation syst…

Machine TranslationData Augmentation

Exploring Pair-Wise NMT for Indian Languages

2020-12-10 · ICON 2020 12 · Kartheek Akella, Sai Himal Allu, Sridhar Suresh Ragupathi, Aman Singhal 외

In this paper, we address the task of improving pair-wise machine translation for specific low resource Indian languages. Multilingual NMT models have demonstrated a reasonable amount of effectiveness on resource-poor la…

Machine TranslationNMTTranslation

The University of Edinburgh’s English-Tamil and English-Inuktitut Submissions to the WMT20 News Translation Task

2020-11-01 · WMT (EMNLP) 2020 11 · Rachel Bawden, Alexandra Birch, Radina Dobreva, Arturo Oncevay 외

We describe the University of Edinburgh’s submissions to the WMT20 news translation shared task for the low resource language pair English-Tamil and the mid-resource language pair English-Inuktitut. We use the neural mac…

Language ModelingLanguage ModellingMachine TranslationTranslation

Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine Translation

2020-05-01 · ACL 2020 6 · Xabier Soto, Dimitar Shterionov, Alberto Poncelas, Andy Way

Machine translation (MT) has benefited from using synthetic training data originating from translating monolingual corpora, a technique known as backtranslation. Combining backtranslated data from different sources has l…

DiversityMachine TranslationTranslation