paper-with-me

홈 › Papers

Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling

2025-04-18 · Shaomu Tan, Christof Monz

A key challenge in MT evaluation is the inherent noise and inconsistency of human ratings. Regression-based neural metrics struggle with this noise, while prompting LLMs shows promise at system-level evaluation but performs poorly at segment level. In this work, we propose ReMedy, a novel MT metric framework that reformulates translation evaluation as a reward modeling task. Instead of regressing on imperfect human ratings directly, ReMedy learns relative translation quality using pairwise preference data, resulting in a more reliable evaluation. In extensive experiments across WMT22-24 shared tasks (39 language pairs, 111 MT systems), ReMedy achieves state-of-the-art performance at both segment- and system-level evaluation. Specifically, ReMedy-9B surpasses larger WMT winners and massive closed LLMs such as MetricX-13B, XCOMET-Ensemble, GEMBA-GPT-4, PaLM-540B, and finetuned PaLM2. Further analyses demonstrate that ReMedy delivers superior capability in detecting translation errors and evaluating low-quality translations.

📄 PDF Abstract BibTeX arXiv:2504.13630

Code (1)

Smu-Tan/Remedy 공식 구현 pytorch

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations

2025-12-21 · Shaomu Tan, Ryosuke Mitani, Ritvik Choudhary, Qiyu Wu 외 arxiv

Over the years, automatic MT metrics have hillclimbed benchmarks and presented strong and sometimes human-level agreement with human ratings. Yet they remain black-box, offering little insight into their decision-making …

Reinforcement LearningMachine Translation

PMMT: Preference Alignment in Multilingual Machine Translation via LLM Distillation

2024-10-15 · Shuqiao Sun, Yutong Yao, Peiwen Wu, Feijun Jiang 외

Translation is important for cross-language communication, and many efforts have been made to improve its accuracy. However, less investment is conducted in aligning translations with human preferences, such as translati…

Machine TranslationTranslation

Reinforced Self-Training (ReST) for Language Modeling

2023-08-17 · Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova 외

Reinforcement learning from human feedback (RLHF) can improve the quality of large language model's (LLM) outputs by aligning them with human preferences. We propose a simple algorithm for aligning LLMs with human prefer…

Language ModelingLanguage ModellingMachine TranslationOffline RL+4

Automatic Post-Editing for Vietnamese

2021-04-25 · ALTA 2021 12 · Thanh Vu, Dai Quoc Nguyen

Automatic post-editing (APE) is an important remedy for reducing errors of raw translated texts that are produced by machine translation (MT) systems or software-aided translation. In this paper, we present a systematic …

Automatic Post-EditingMachine TranslationSentenceTranslation

Cross-lingual Human-Preference Alignment for Neural Machine Translation with Direct Quality Optimization

2024-09-26 · Kaden Uhlig, Joern Wuebker, Raphael Reinauer, John DeNero

Reinforcement Learning from Human Feedback (RLHF) and derivative techniques like Direct Preference Optimization (DPO) are task-alignment algorithms used to repurpose general, foundational models for specific tasks. We sh…

Machine TranslationNMTTranslation