Latin-Spanish Neural Machine Translation: from the Bible to Saint Augustine
Although there are several sources where to find historical texts, they usually are available in the original language that makes them generally inaccessible. This paper presents the development of state-of-the-art Neural Machine Systems for the low-resourced Latin-Spanish language pair. First, we build a Transformer-based Machine Translation system on the Bible parallel corpus. Then, we build a comparable corpus from Saint Augustine texts and their translations. We use this corpus to study the domain adaptation case from the Bible texts to Saint Augustine{'}s works. Results show the difficulties of handling a low-resourced language as Latin. First, we noticed the importance of having enough data, since the systems do not achieve high BLEU scores. Regarding domain adaptation, results show how using in-domain data helps systems to achieve a better quality translation. Also, we observed that it is needed a higher amount of data to perform an effective vocabulary extension that includes in-domain vocabulary.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain AdaptationMachine TranslationTranslationSimilar Papers 제목 키워드 기반
The eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource Languages
Efficiently and accurately translating a corpus into a low-resource language remains a challenge, regardless of the strategies employed, whether manual, automated, or a combination of the two. Many Christian organization…
BenchmarkingMachine TranslationNMTTranslationEvaluating Indirect Strategies for Chinese-Spanish Statistical Machine Translation
Although, Chinese and Spanish are two of the most spoken languages in the world, not much research has been done in machine translation for this language pair. This paper focuses on investigating the state-of-the-art of …
Machine TranslationTranslationTranslating Spanish into Spanish Sign Language: Combining Rules and Data-driven Approaches
This paper presents a series of experiments on translating between spoken Spanish and Spanish Sign Language glosses (LSE), including enriching Neural Machine Translation (NMT) systems with linguistic features, and creati…
Machine TranslationNMTTranslationCo-reference Resolution of Elided Subjects and Possessive Pronouns in Spanish-English Statistical Machine Translation
This paper presents a straightforward method to integrate co-reference information into phrase-based machine translation to address the problems of i) elided subjects and ii) morphological underspecification of pronouns …
Coreference ResolutionMachine TranslationTranslation