paper-with-me

홈 › Papers

Robust Neural Machine Translation: Modeling Orthographic and Interpunctual Variation

2020-09-11 · Toms Bergmanis, Artūrs Stafanovičs, Mārcis Pinnis

Neural machine translation systems typically are trained on curated corpora and break when faced with non-standard orthography or punctuation. Resilience to spelling mistakes and typos, however, is crucial as machine translation systems are used to translate texts of informal origins, such as chat conversations, social media posts and web pages. We propose a simple generative noise model to generate adversarial examples of ten different types. We use these to augment machine translation systems' training data and show that, when tested on noisy data, systems trained using adversarial examples perform almost as well as when translating clean data, while baseline systems' performance drops by 2-3 BLEU points. To measure the robustness and noise invariance of machine translation systems' outputs, we use the average translation edit rate between the translation of the original sentence and its noised variants. Using this measure, we show that systems trained on adversarial examples on average yield 50% consistency improvements when compared to baselines trained on clean data.

📄 PDF Abstract BibTeX arXiv:2009.05460

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin

2024-04-28 · Pin-Jie Lin, Merel Scholman, Muhammed Saeed, Vera Demberg

Nigerian Pidgin is an English-derived contact language and is traditionally an oral language, spoken by approximately 100 million people. No orthographic standard has yet been adopted, and thus the few available Pidgin d…

Data AugmentationMachine TranslationSentiment AnalysisTranslation

A Survey of Orthographic Information in Machine Translation

2020-08-04 · Bharathi Raja Chakravarthi, Priya Rani, Mihael Arcan, John P. McCrae

Machine translation is one of the applications of natural language processing which has been explored in different languages. Recently researchers started paying attention towards machine translation for resource-poor la…

Bilingual Lexicon InductionMachine TranslationSurveyTranslation

Learning variable length units for SMT between related languages via Byte Pair Encoding

2016-10-20 · WS 2017 9 · Anoop Kunchukuttan, Pushpak Bhattacharyya

We explore the use of segments learnt using Byte Pair Encoding (referred to as BPE units) as basic units for statistical machine translation between related languages and compare it with orthographic syllables, which are…

Machine TranslationTranslation

Naver Labs Europe's Systems for the WMT19 Machine Translation Robustness Task

2019-07-15 · WS 2019 8 · Alexandre Bérard, Ioan Calapodescu, Claude Roux

This paper describes the systems that we submitted to the WMT19 Machine Translation robustness task. This task aims to improve MT's robustness to noise found on social media, like informal language, spelling mistakes and…

Domain AdaptationMachine TranslationTranslation

Graphemic Normalization of the Perso-Arabic Script

2022-10-21 · Raiomond Doctor, Alexander Gutkin, Cibu Johny, Brian Roark 외

Since its original appearance in 1991, the Perso-Arabic script representation in Unicode has grown from 169 to over 440 atomic isolated characters spread over several code pages representing standard letters, various dia…

Language ModelingLanguage ModellingMachine Translation