paper-with-me

홈 › Papers

Translation Errors Significantly Impact Low-Resource Languages in Cross-Lingual Learning

2024-02-03 · Ashish Sunil Agrawal, Barah Fazili, Preethi Jyothi

Popular benchmarks (e.g., XNLI) used to evaluate cross-lingual language understanding consist of parallel versions of English evaluation sets in multiple target languages created with the help of professional translators. When creating such parallel data, it is critical to ensure high-quality translations for all target languages for an accurate characterization of cross-lingual transfer. In this work, we find that translation inconsistencies do exist and interestingly they disproportionally impact low-resource languages in XNLI. To identify such inconsistencies, we propose measuring the gap in performance between zero-shot evaluations on the human-translated and machine-translated target text across multiple target languages; relatively large gaps are indicative of translation errors. We also corroborate that translation errors exist for two target languages, namely Hindi and Urdu, by doing a manual reannotation of human-translated test instances in these two languages and finding poor agreement with the original English labels these instances were supposed to inherit.

📄 PDF Abstract BibTeX arXiv:2402.02080

Code (1)

csalt-research/translation-errors-crosslingual-learning 공식 구현 pytorch

Tasks

Cross-Lingual TransferTranslation

Similar Papers 제목 키워드 기반

OCR Improves Machine Translation for Low-Resource Languages

2022-02-27 · Findings (ACL) 2022 5 · Oana Ignat, Jean Maillard, Vishrav Chaudhary, Francisco Guzmán

We aim to investigate the performance of current OCR systems on low resource languages and low resource scripts. We introduce and make publicly available a novel benchmark, OCR4MT, consisting of real and synthetic data, …

Machine TranslationOptical Character Recognition (OCR)Translation

Misgendering and Assuming Gender in Machine Translation when Working with Low-Resource Languages

2024-01-24 · Sourojit Ghosh, Srishti Chatterjee

This chapter focuses on gender-related errors in machine translation (MT) in the context of low-resource languages. We begin by explaining what low-resource languages are, examining the inseparable social and computation…

Machine Translation

Cross-Lingual Conversational Speech Summarization with Large Language Models

2024-08-12 · Max Nelson, Shannon Wotherspoon, Francis Keith, William Hartmann 외

Cross-lingual conversational speech summarization is an important problem, but suffers from a dearth of resources. While transcriptions exist for a number of languages, translated conversational speech is rare and datase…

Machine Translationspeech-recognitionSpeech RecognitionTranslation

Impact of Domain-Adapted Multilingual Neural Machine Translation in the Medical Domain

2022-12-05 · Miguel Rios, Raluca-Maria Chereji, Alina Secara, Dragos Ciobanu

Multilingual Neural Machine Translation (MNMT) models leverage many language pairs during training to improve translation quality for low-resource languages by transferring knowledge from high-resource languages. We stud…

Machine TranslationTranslation

Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine

2023-01-20 · Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang 외

This report provides a preliminary evaluation of ChatGPT for machine translation, including translation prompt, multilingual translation, and translation robustness. We adopt the prompts advised by ChatGPT to trigger its…

Machine TranslationSentenceTranslation