paper-with-me

Papers

MQM-Chat: Multidimensional Quality Metrics for Chat Translation

2024-08-29 · Yunmeng Li, Jun Suzuki, Makoto Morishita, Kaori Abe, Kentaro Inui

The complexities of chats pose significant challenges for machine translation models. Recognizing the need for a precise evaluation metric to address the issues of chat translation, this study introduces Multidimensional Quality Metrics for Chat Translation (MQM-Chat). Through the experiments of five models using MQM-Chat, we observed that all models generated certain fundamental errors, while each of them has different shortcomings, such as omission, overly correcting ambiguous source content, and buzzword issues, resulting in the loss of stylized information. Our findings underscore the effectiveness of MQM-Chat in evaluating chat translation, emphasizing the importance of stylized content and dialogue consistency for future studies.

📄 PDF Abstract BibTeX arXiv:2408.16390

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Multidimensional Evaluation for Text Style Transfer Using ChatGPT

2023-04-26 · Huiyuan Lai, Antonio Toral, Malvina Nissim

We investigate the potential of ChatGPT as a multidimensional evaluator for the task of \emph{Text Style Transfer}, alongside, and in comparison to, existing automatic metrics as well as human judgements. We focus on a z…

Style TransferText GenerationText Style Transfer

Is Context Helpful for Chat Translation Evaluation?

2024-03-13 · Sweta Agrawal, Amin Farajian, Patrick Fernandes, Ricardo Rei 외

Despite the recent success of automatic metrics for assessing translation quality, their application in evaluating the quality of machine-translated chats has been limited. Unlike more structured texts like news, chat co…

Language ModelingLanguage ModellingLarge Language ModelSentence+1

An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation

2024-12-24 · Andre Rusli, Makoto Shishido

This paper analyses how traditional baseline metrics, such as BLEU and TER, and neural-based methods, such as BERTScore and COMET, score several NMT models performance on chat translation and how these metrics perform wh…

NMTTranslation

Evaluating LLM-Based Translation of a Low-Resource Technical Language: The Medical and Philosophical Greek of Galen

2026-02-27 · James L. Zainaldin, Cameron Pattison, Manuela Marai, Jacob Wu 외 arxiv

Purpose: This study evaluates the quality of commercial large language model (LLM) machine translation (MT) for Ancient Greek technical prose and benchmarks standard automated MT evaluation metrics against expert human j…

Machine Translation

Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation

2024-01-10 · Zhaokun Jiang, Qianxi Lv, Ziyin Zhang, Lei Lei

Large language models have demonstrated parallel and even superior translation performance compared to neural machine translation (NMT) systems. However, existing comparative studies between them mainly rely on automated…

Machine TranslationNMTPrompt EngineeringTranslation