Better OOV Translation with Bilingual Terminology Mining
Unseen words, also called out-of-vocabulary words (OOVs), are difficult for machine translation. In neural machine translation, byte-pair encoding can be used to represent OOVs, but they are still often incorrectly translated. We improve the translation of OOVs in NMT using easy-to-obtain monolingual data. We look for OOVs in the text to be translated and translate them using simple-to-construct bilingual word embeddings (BWEs). In our MT experiments we take the 5-best candidates, which is motivated by intrinsic mining experiments. Using all five of the proposed target language words as queries we mine target-language sentences. We then back-translate, forcing the back-translation of each of the five proposed target-language OOV-translation-candidates to be the original source-language OOV. We show that by using this synthetic data to fine-tune our system the translation of OOVs can be dramatically improved. In our experiments we use a system trained on Europarl and mine sentences containing medical terms from monolingual data.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationNMTTranslationWord EmbeddingsSimilar Papers 제목 키워드 기반
Bilingual Terminology Extraction from Comparable E-Commerce Corpora
Bilingual terminologies are important machine translation resources in the field of e-commerce, which are usually either manually translated or automatically extracted from parallel data. The human translation is costly …
Machine TranslationSentenceTranslationA Method of Augmenting Bilingual Terminology by Taking Advantage of the Conceptual Systematicity of Terminologies
In this paper, we propose a method of augmenting existing bilingual terminologies. Our method belongs to a {``}generate and validate{''} framework rather than extraction from corpora. Although many studies have proposed …
Transfer LearningUsing WordNet and Semantic Similarity for Bilingual Terminology Mining from Comparable Corpora
Domain Terminology Integration into Machine Translation: Leveraging Large Language Models
This paper discusses the methods that we used for our submissions to the WMT 2023 Terminology Shared Task for German-to-English (DE-EN), English-to-Czech (EN-CS), and Chinese-to-English (ZH-EN) language pairs. The task a…
Automatic Post-Editingde-enMachine TranslationTranslationTermMind: Alibaba’s WMT21 Machine Translation Using Terminologies Task Submission
This paper describes our work in the WMT 2021 Machine Translation using Terminologies Shared Task. We participate in the shared translation terminologies task in English to Chinese language pair. To satisfy terminology c…
Data AugmentationMachine TranslationTranslation