paper-with-me

홈 › Papers

Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings

2024-11-08 · Miguel Moura Ramos, Tomás Almeida, Daniel Vareta, Filipe Azevedo, Sweta Agrawal, Patrick Fernandes, André F. T. Martins

Reinforcement learning (RL) has been proven to be an effective and robust method for training neural machine translation systems, especially when paired with powerful reward models that accurately assess translation quality. However, most research has focused on RL methods that use sentence-level feedback, leading to inefficient learning signals due to the reward sparsity problem -- the model receives a single score for the entire sentence. To address this, we propose a novel approach that leverages fine-grained, token-level quality assessments along with error severity levels using RL methods. Specifically, we use xCOMET, a state-of-the-art quality estimation system, as our token-level reward model. We conduct experiments on small and large translation datasets with standard encoder-decoder and large language models-based machine translation systems, comparing the impact of sentence-level versus fine-grained reward signals on translation quality. Our results show that training with token-level rewards improves translation quality across language pairs over baselines according to both automatic and human evaluation. Furthermore, token-level reward optimization improves training stability, evidenced by a steady increase in mean rewards over training epochs.

📄 PDF Abstract BibTeX arXiv:2411.05986

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderMachine TranslationReinforcement Learning (RL)SentenceTranslation

Similar Papers 제목 키워드 기반

From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference Optimization

2026-02-01 · Chaoqun Cui, Shijing Wang, Liangbin Huang, Qingqing Gu 외 arxiv

The rapid development of Large Language Models (LLMs) has significantly enhanced the general capabilities of machine translation. However, as application scenarios become more complex, the limitations of LLMs in vertical…

Machine Translation

GRRM: Group Relative Reward Modeling for Machine Translation

2026-02-15 · Sen Yang, Shanbo Cheng, Lu Xu, Jianbing Zhang 외 arxiv

While Group Relative Policy Optimization (GRPO) offers a powerful framework for LLM post-training, its effectiveness in open-ended domains like Machine Translation hinges on accurate intra-group ranking. We identify that…

Machine Translation

$M^2PO$: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation

2025-10-15 · Hao Wang, Linlong Xu, Heng Liu, Yangyang Liu 외 arxiv

Aligning Large Language Models (LLMs) with human preferences is pivotal for Machine Translation (MT), yet current approaches are often hindered by misleading reward signals. Our analysis reveals that prevailing Quality E…

Machine Translation

Structured Document Translation via Format Reinforcement Learning

2025-12-04 · Haiyue Song, Johannes Eschbach-Dymanus, Hour Kaing, Sumire Honda 외 arxiv

Recent works on structured text translation remain limited to the sentence level, as they struggle to effectively handle the complex document-level XML or HTML structures. To address this, we propose \textbf{Format Reinf…

Reinforcement Learning

Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization

2025-08-12 · Chaoqun Cui, Liangbin Huang, Shijing Wang, Zhe Tong 외 arxiv

Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information …

Machine Translation