Sentence Concatenation Approach to Data Augmentation for Neural Machine Translation
Neural machine translation (NMT) has recently gained widespread attention because of its high translation accuracy. However, it shows poor performance in the translation of long sentences, which is a major issue in low-resource languages. It is assumed that this issue is caused by insufficient number of long sentences in the training data. Therefore, this study proposes a simple data augmentation method to handle long sentences. In this method, we use only the given parallel corpora as the training data and generate long sentences by concatenating two sentences. Based on the experimental results, we confirm improvements in long sentence translation by the proposed data augmentation method, despite its simplicity. Moreover, the translation quality is further improved by the proposed method, when combined with back-translation.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationMachine TranslationNMTSentenceTranslationSimilar Papers 제목 키워드 기반
From Scarcity to Efficiency: Investigating the Effects of Data Augmentation on African Machine Translation
The linguistic diversity across the African continent presents different challenges and opportunities for machine translation. This study explores the effects of data augmentation techniques in improving translation syst…
Machine TranslationData AugmentationFocused Concatenation for Context-Aware Neural Machine Translation
A straightforward approach to context-aware neural machine translation consists in feeding the standard encoder-decoder architecture with a window of consecutive sentences, formed by the current sentence and a number of …
DecoderMachine TranslationSentenceTranslationEncoding Sentence Position in Context-Aware Neural Machine Translation with Concatenation
Context-aware translation can be achieved by processing a concatenation of consecutive sentences with the standard Transformer architecture. This paper investigates the intuitive idea of providing the model with explicit…
Machine TranslationPositionSentenceTranslationData Augmentation by Concatenation for Low-Resource Translation: A Mystery and a Solution
In this paper, we investigate the driving factors behind concatenation, a simple but effective data augmentation method for low-resource neural machine translation. Our experiments suggest that discourse context is unlik…
Data AugmentationDiversityLow Resource Neural Machine TranslationLow-Resource Neural Machine Translation+3Coreference and Coherence in Neural Machine Translation: A Study Using Oracle Experiments
Cross-sentence context can provide valuable information in Machine Translation and is critical for translation of anaphoric pronouns and for providing consistent translations. In this paper, we devise simple oracle exper…
Coreference ResolutionLanguage ModelingLanguage ModellingMachine Translation+3