paper-with-me

홈 › Papers

OCR Improves Machine Translation for Low-Resource Languages

2022-02-27 · Findings (ACL) 2022 5 · Oana Ignat, Jean Maillard, Vishrav Chaudhary, Francisco Guzmán

We aim to investigate the performance of current OCR systems on low resource languages and low resource scripts. We introduce and make publicly available a novel benchmark, OCR4MT, consisting of real and synthetic data, enriched with noise, for 60 low-resource languages in low resource scripts. We evaluate state-of-the-art OCR systems on our benchmark and analyse most common errors. We show that OCR monolingual data is a valuable resource that can increase performance of Machine Translation models, when used in backtranslation. We then perform an ablation study to investigate how OCR errors impact Machine Translation performance and determine what is the minimum level of OCR quality needed for the monolingual data to be useful for Machine Translation.

📄 PDF Abstract BibTeX arXiv:2202.13274

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationOptical Character Recognition (OCR)Translation

Similar Papers 제목 키워드 기반

A Survey of Orthographic Information in Machine Translation

2020-08-04 · Bharathi Raja Chakravarthi, Priya Rani, Mihael Arcan, John P. McCrae

Machine translation is one of the applications of natural language processing which has been explored in different languages. Recently researchers started paying attention towards machine translation for resource-poor la…

Bilingual Lexicon InductionMachine TranslationSurveyTranslation

Love Thy Neighbor: Combining Two Neighboring Low-Resource Languages for Translation

2021-08-01 · MTSummit 2021 8 · John E. Ortega, Richard Alexander Castro Mamani, Jaime Rafael Montoya Samame

Low-resource languages sometimes take on similar morphological and syntactic characteristics due to their geographic nearness and shared history. Two low-resource neighboring languages found in Peru, Quechua and Ashanink…

Machine TranslationTransfer LearningTranslationVocal Bursts Valence Prediction

Low-Resource Machine Translation Using Cross-Lingual Language Model Pretraining

2021-06-01 · NAACL (AmericasNLP) 2021 6 · Francis Zheng, Machel Reid, Edison Marrese-Taylor, Yutaka Matsuo

This paper describes UTokyo’s submission to the AmericasNLP 2021 Shared Task on machine translation systems for indigenous languages of the Americas. We present a low-resource machine translation system that improves tra…

Language ModelingLanguage ModellingMachine TranslationTranslation

Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data

2021-05-31 · ACL 2021 5 · Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary 외

The scarcity of parallel data is a major obstacle for training high-quality machine translation systems for low-resource languages. Fortunately, some low-resource languages are linguistically related or similar to high-r…

DenoisingMachine TranslationNMTTranslation

Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation

2020-05-11 · ACL 2020 6 · Aditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat 외

Over the last few years two promising research directions in low-resource neural machine translation (NMT) have emerged. The first focuses on utilizing high-resource languages to improve the quality of low-resource langu…

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMT+1