paper-with-me

Papers

On the Integration of LinguisticFeatures into Statistical and Neural Machine Translation

2020-03-31 · Eva Vanmassenhove

New machine translations (MT) technologies are emerging rapidly and with them, bold claims of achieving human parity such as: (i) the results produced approach "accuracy achieved by average bilingual human translators" (Wu et al., 2017b) or (ii) the "translation quality is at human parity when compared to professional human translators" (Hassan et al., 2018) have seen the light of day (Laubli et al., 2018). Aside from the fact that many of these papers craft their own definition of human parity, these sensational claims are often not supported by a complete analysis of all aspects involved in translation. Establishing the discrepancies between the strengths of statistical approaches to MT and the way humans translate has been the starting point of our research. By looking at MT output and linguistic theory, we were able to identify some remaining issues. The problems range from simple number and gender agreement errors to more complex phenomena such as the correct translation of aspectual values and tenses. Our experiments confirm, along with other studies (Bentivogli et al., 2016), that neural MT has surpassed statistical MT in many aspects. However, some problems remain and others have emerged. We cover a series of problems related to the integration of specific linguistic features into statistical and neural MT, aiming to analyse and provide a solution to some of them. Our work focuses on addressing three main research questions that revolve around the complex relationship between linguistics and MT in general. We identify linguistic information that is lacking in order for automatic translation systems to produce more accurate translations and integrate additional features into the existing pipelines. We identify overgeneralization or 'algorithmic bias' as a potential drawback of neural MT and link it to many of the remaining linguistic issues.

📄 PDF Abstract BibTeX arXiv:2003.14324

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Dynamic Terminology Integration Methods in Statistical Machine Translation

2015-05-01 · WS 2015 5 · M{\=a}rcis Pinnis
Domain AdaptationLanguage ModellingMachine TranslationTranslation

Bootstrapping Phrase-based Statistical Machine Translation via WSD Integration

2013-10-01 · IJCNLP 2013 10 · Hien Vu Huy, Phuong-Thai Nguyen, Tung-Lam Nguyen, M.L Nguyen
Active LearningLanguage ModellingMachine TranslationTranslation+1

Data Augmentation and Terminology Integration for Domain-Specific Sinhala-English-Tamil Statistical Machine Translation

2020-11-05 · Aloka Fernando, Surangika Ranathunga, Gihan Dias

Out of vocabulary (OOV) is a problem in the context of Machine Translation (MT) in low-resourced languages. When source and/or target languages are morphologically rich, it becomes even worse. Bilingual list integration …

Data AugmentationMachine TranslationTranslation

HABLex: Human Annotated Bilingual Lexicons for Experiments in Machine Translation

2019-11-01 · IJCNLP 2019 11 · Brian Thompson, Rebecca Knowles, Xuan Zhang, Huda Khayrallah 외

Bilingual lexicons are valuable resources used by professional human translators. While these resources can be easily incorporated in statistical machine translation, it is unclear how to best do so in the neural framewo…

Machine TranslationTranslation

MTUOC: easy and free integration of NMT systems in professional translation environments

2020-11-01 · EAMT 2020 11 · Antoni Oliver

In this paper the MTUOC project, aiming to provide an easy integration of neural and statistical machine translation systems, is presented. Almost all the required software to train and use neural and statistical MT syst…

Machine TranslationNMTTranslation