paper-with-me

홈 › Papers

Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains

2024-02-28 · Vilém Zouhar, Shuoyang Ding, Anna Currey, Tatyana Badeka, Jenyuan Wang, Brian Thompson

We introduce a new, extensive multidimensional quality metrics (MQM) annotated dataset covering 11 language pairs in the biomedical domain. We use this dataset to investigate whether machine translation (MT) metrics which are fine-tuned on human-generated MT quality judgements are robust to domain shifts between training and inference. We find that fine-tuned metrics exhibit a substantial performance drop in the unseen domain scenario relative to metrics that rely on the surface form, as well as pre-trained metrics which are not fine-tuned on MT quality judgments.

📄 PDF Abstract BibTeX arXiv:2402.18747

Code (1)

amazon-science/bio-mqm-dataset 공식 구현 pytorch

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Instruction-tuned Large Language Models for Machine Translation in the Medical Domain

2024-08-29 · Miguel Rios

Large Language Models (LLMs) have shown promising results on machine translation for high resource language pairs and domains. However, in specialised domains (e.g. medical) LLMs have shown lower performance compared to …

Machine TranslationTranslation

Multilingual Machine Translation Evaluation Metrics Fine-tuned on Pseudo-Negative Examples for WMT 2021 Metrics Task

2021-11-01 · WMT (EMNLP) 2021 11 · Kosuke Takahashi, Yoichi Ishibashi, Katsuhito Sudoh, Satoshi Nakamura

This paper describes our submission to the WMT2021 shared metrics task. Our metric is operative to segment-level and system-level translations. Our belief toward a better metric is to detect a significant error that cann…

AttributeMachine Translation

Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling

2025-04-18 · Shaomu Tan, Christof Monz

A key challenge in MT evaluation is the inherent noise and inconsistency of human ratings. Regression-based neural metrics struggle with this noise, while prompting LLMs shows promise at system-level evaluation but perfo…

Machine TranslationTranslation

Are Large Language Models State-of-the-art Quality Estimators for Machine Translation of User-generated Content?

2024-10-08 · Shenbin Qian, Constantin Orăsan, Diptesh Kanojia, Félix do Carmo

This paper investigates whether large language models (LLMs) are state-of-the-art quality estimators for machine translation of user-generated content (UGC) that contains emotional expressions, without the use of referen…

In-Context LearningMachine Translationparameter-efficient fine-tuningTranslation

Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation

2024-10-28 · Yirong Sun, Dawei Zhu, Yanjun Chen, Erjia Xiao 외

Large language models (LLMs) have excelled in various NLP tasks, including machine translation (MT), yet most studies focus on sentence-level translation. This work investigates the inherent capability of instruction-tun…

Document Level Machine TranslationMachine TranslationSentenceTranslation