Machine Translation Evaluation for Arabic using Morphologically-enriched Embeddings
Evaluation of machine translation (MT) into morphologically rich languages (MRL) has not been well studied despite posing many challenges. In this paper, we explore the use of embeddings obtained from different levels of lexical and morpho-syntactic linguistic analysis and show that they improve MT evaluation into an MRL. Specifically we report on Arabic, a language with complex and rich morphology. Our results show that using a neural-network model with different input representations produces results that clearly outperform the state-of-the-art for MT evaluation into Arabic, by almost over 75{\%} increase in correlation with human judgments on pairwise MT evaluation quality task. More importantly, we demonstrate the usefulness of morpho-syntactic representations to model sentence similarity for MT evaluation and address complex linguistic phenomena of Arabic.
Code (0)
등록된 구현이 없습니다.
Tasks
Community Question AnsweringMachine TranslationMorphological AnalysisMorphological InflectionQuestion AnsweringSentenceSentence SimilarityTranslationWord EmbeddingsSimilar Papers 제목 키워드 기반
Translating Between Morphologically Rich Languages: An Arabic-to-Turkish Machine Translation System
This paper introduces the work on building a machine translation system for Arabic-to-Turkish in the news domain. Our work includes collecting parallel datasets in several ways for a new and low-resourced language pair, …
Machine TranslationTranslationThe Arabic Parallel Gender Corpus 2.0: Extensions and Analyses
Gender bias in natural language processing (NLP) applications, particularly machine translation, has been receiving increasing attention. Much of the research on this issue has focused on mitigating gender bias in Englis…
Machine TranslationText GenerationTranslationJoint Segmentation and POS Tagging for Arabic Using a CRF-based Classifier
Arabic is a morphologically rich language, and Arabic texts abound of complex word forms built by concatenation of multiple subparts, corresponding for instance to prepositions, articles, roots prefixes, or suffixes. The…
ArticlesBIG-bench Machine LearningMachine TranslationMorphological Analysis+4Morphological Word Embeddings for Arabic Neural Machine Translation in Low-Resource Settings
Neural machine translation has achieved impressive results in the last few years, but its success has been limited to settings with large amounts of parallel data. One way to improve NMT for lower-resource settings is to…
Low Resource NMTMachine TranslationNMTTranslation+2AraBench: Benchmarking Dialectal Arabic-English Machine Translation
Low-resource machine translation suffers from the scarcity of training data and the unavailability of standard evaluation sets. While a number of research efforts target the former, the unavailability of evaluation bench…
BenchmarkingData AugmentationMachine TranslationTranslation