BPE and CharCNNs for Translation of Morphology: A Cross-Lingual Comparison and Analysis
Neural Machine Translation (NMT) in low-resource settings and of morphologically rich languages is made difficult in part by data sparsity of vocabulary words. Several methods have been used to help reduce this sparsity, notably Byte-Pair Encoding (BPE) and a character-based CNN layer (charCNN). However, the charCNN has largely been neglected, possibly because it has only been compared to BPE rather than combined with it. We argue for a reconsideration of the charCNN, based on cross-lingual improvements on low-resource data. We translate from 8 languages into English, using a multi-way parallel collection of TED transcripts. We find that in most cases, using both BPE and a charCNN performs best, while in Hebrew, using a charCNN over words is best.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationNMTTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Comparison of Transformer and Recurrent Neural Networks on Multilingual Neural Machine Translation
Recently, neural machine translation (NMT) has been extended to multilinguality, that is to handle more than one translation direction with a single system. Multilingual NMT showed competitive performance against pure bi…
Machine TranslationNMTTransfer LearningTranslationSlovene SuperGLUE Benchmark: Translation and Evaluation
We present a Slovene combined machine-human translated SuperGLUE benchmark. We describe the translation process and problems arising due to differences in morphology and grammar. We evaluate the translated datasets in se…
TranslationMorphologically Aware Word-Level Translation
We propose a novel morphologically aware probability model for bilingual lexicon induction, which jointly models lexeme translation and inflectional morphology in a structured way. Our model exploits the basic linguistic…
Bilingual Lexicon InductionTranslationMAAM: A Morphology-Aware Alignment Model for Unsupervised Bilingual Lexicon Induction
The task of unsupervised bilingual lexicon induction (UBLI) aims to induce word translations from monolingual corpora in two languages. Previous work has shown that morphological variation is an intractable challenge for…
Bilingual Lexicon InductionDenoisingLanguage ModelingLanguage Modelling+1Cross-lingual Name Tagging and Linking for 282 Languages
The ambitious goal of this work is to develop a cross-lingual name tagging and linking framework for 282 languages that exist in Wikipedia. Given a document in any of these languages, our framework is able to identify na…
TranslationWord Translation