Modifications of Machine Translation Evaluation Metrics by Using Word Embeddings
Traditional machine translation evaluation metrics such as BLEU and WER have been widely used, but these metrics have poor correlations with human judgements because they badly represent word similarity and impose strict identity matching. In this paper, we propose some modifications to the traditional measures based on word embeddings for these two metrics. The evaluation results show that our modifications significantly improve their correlation with human judgements.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSemantic Textual SimilarityTranslationWord EmbeddingsWord SimilaritySimilar Papers 제목 키워드 기반
DEMETR: Diagnosing Evaluation Metrics for Translation
While machine translation evaluation metrics based on string overlap (e.g., BLEU) have their limitations, their computations are transparent: the BLEU score assigned to a particular candidate translation can be traced ba…
DiagnosticMachine TranslationSensitivityTranslationMetrics for Evaluation of Word-level Machine Translation Quality Estimation
WMDO: Fluency-based Word Mover's Distance for Machine Translation Evaluation
We propose WMDO, a metric based on distance between distributions in the semantic vector space. Matching in the semantic space has been investigated for translation evaluation, but the constraints of a translation{'}s wo…
Machine TranslationTranslationWord EmbeddingsIs This Translation Error Critical?: Classification-Based Human and Automatic Machine Translation Evaluation Focusing on Critical Errors
This paper discusses a classification-based approach to machine translation evaluation, as opposed to a common regression-based approach in the WMT Metrics task. Recent machine translation usually works well but sometime…
Machine TranslationTranslationPutting Evaluation in Context: Contextual Embeddings Improve Machine Translation Evaluation
Accurate, automatic evaluation of machine translation is critical for system tuning, and evaluating progress in the field. We proposed a simple unsupervised metric, and additional supervised metrics which rely on context…
Machine TranslationSentenceTranslationWord Embeddings