paper-with-me

홈 › Papers

Improving a Multi-Source Neural Machine Translation Model with Corpus Extension for Low-Resource Languages

2017-09-26 · LREC 2018 5 · Gyu-Hyeon Choi, Jong-Hun Shin, Young-Kil Kim

In machine translation, we often try to collect resources to improve performance. However, most of the language pairs, such as Korean-Arabic and Korean-Vietnamese, do not have enough resources to train machine translation systems. In this paper, we propose the use of synthetic methods for extending a low-resource corpus and apply it to a multi-source neural machine translation model. We showed the improvement of machine translation performance through corpus extension using the synthetic method. We specifically focused on how to create source sentences that can make better target sentences, including the use of synthetic methods. We found that the corpus extension could also improve the performance of multi-source neural machine translation. We showed the corpus extension and multi-source model to be efficient methods for a low-resource language pair. Furthermore, when both methods were used together, we found better machine translation performance.

📄 PDF Abstract BibTeX arXiv:1709.08898

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

ParCorFull2.0: a Parallel Corpus Annotated with Full Coreference

2022-06-01 · LREC 2022 6 · Ekaterina Lapshinova-Koltunski, Pedro Augusto Ferreira, Elina Lartaud, Christian Hardmeier

In this paper, we describe ParCorFull2.0, a parallel corpus annotated with full coreference chains for multiple languages, which is an extension of the existing corpus ParCorFull (Lapshinova-Koltunski et al., 2018). Simi…

coreference-resolutionCoreference ResolutionMachine TranslationTranslation

Statistical Machine Translation without Source-side Parallel Corpus Using Word Lattice and Phrase Extension

2012-05-01 · LREC 2012 5 · Takanori Kusumoto, Tomoyosi Akiba

Statistical machine translation (SMT) requires a parallel corpus between the source and target languages. Although a pivot-translation approach can be applied to a language pair that does not have a parallel corpus direc…

Machine TranslationTranslation

ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic - English

2022-11-22 · Injy Hamed, Nizar Habash, Slim Abdennadher, Ngoc Thang Vu

We present our work on collecting ArzEn-ST, a code-switched Egyptian Arabic - English Speech Translation Corpus. This corpus is an extension of the ArzEn speech corpus, which was collected through informal interviews wit…

Machine TranslationTranslation

The Multilingual TEDx Corpus for Speech Recognition and Translation

2021-02-02 · Elizabeth Salesky, Matthew Wiesner, Jacob Bremerman, Roldano Cattoni 외

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx t…

speech-recognitionSpeech RecognitionTranslation

Latin-Spanish Neural Machine Translation: from the Bible to Saint Augustine

2020-05-01 · LREC 2020 5 · Eva Mart{\'\i}nez Garcia, {\'A}lvaro Garc{\'\i}a Tejedor

Although there are several sources where to find historical texts, they usually are available in the original language that makes them generally inaccessible. This paper presents the development of state-of-the-art Neura…

Domain AdaptationMachine TranslationTranslation