Corpus Augmentation by Sentence Segmentation for Low-Resource Neural Machine Translation
Neural Machine Translation (NMT) has been proven to achieve impressive results. The NMT system translation results depend strongly on the size and quality of parallel corpora. Nevertheless, for many language pairs, no rich-resource parallel corpora exist. As described in this paper, we propose a corpus augmentation method by segmenting long sentences in a corpus using back-translation and generating pseudo-parallel sentence pairs. The experiment results of the Japanese-Chinese and Chinese-Japanese translation with Japanese-Chinese scientific paper excerpt corpus (ASPEC-JC) show that the method improves translation performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMTSentenceSentence segmentationTranslationSimilar Papers 제목 키워드 기반
Parallel resources for Tunisian Arabic Dialect Translation
The difficulty of processing dialects is clearly observed in the high cost of building representative corpus, in particular for machine translation. Indeed, all machine translation systems require a huge amount and good …
Data AugmentationMachine TranslationManagementSentence+1Data centric approach to Chinese Medical Speech Recognition
Concerning the development of Chinese medical speech recognition technology, this study re-addresses earlier encountered issues in accordance with the process of Machine Learning Engineering for Production (MLOps) from a…
Data Augmentationspeech-recognitionSpeech RecognitionNot Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine Translation
Despite being the seventh most widely spoken language in the world, Bengali has received much less attention in machine translation literature due to being low in resources. Most publicly available parallel corpora for B…
Machine TranslationSentenceSentence segmentationTranslationLanguage Resource Addition: Dictionary or Corpus?
In this paper, we investigate the relative effect of two strategies of language resource additions to the word segmentation problem and part-of-speech tagging problem in Japanese. The first strategy is adding entries to …
Active LearningDomain AdaptationMorphological AnalysisPart-Of-Speech Tagging+2Data Augmentation for Machine Translation via Dependency Subtree Swapping
We present a generic framework for data augmentation via dependency subtree swapping that is applicable to machine translation. We extract corresponding subtrees from the dependency parse trees of the source and target s…
Data AugmentationMachine TranslationTranslation