paper-with-me

홈 › Papers

$M^2PO$: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation

2025-10-15 · Hao Wang, Linlong Xu, Heng Liu, Yangyang Liu, Xiaohu Zhao, Bo Zeng, Liangying Shao, Longyue Wang, Weihua Luo, Kaifu Zhang arxiv

Aligning Large Language Models (LLMs) with human preferences is pivotal for Machine Translation (MT), yet current approaches are often hindered by misleading reward signals. Our analysis reveals that prevailing Quality Estimation (QE) models exhibit a systematic blind spot toward partial errors, specifically partial hallucinations and omissions, often favoring superficially fluent but unfaithful translations. To address this issue, we propose $M^2PO$ (Multi-Perspective Multi-Pair Preference Optimization), a data-centric framework for preference optimization in machine translation. First, to correct the bias toward fluency, $M^2PO$ uses a dual-perspective mechanism that decouples semantic fidelity from fluency and prioritizes faithfulness through a curriculum strategy. Second, after correcting this bias, partial errors fall between perfect and severely incorrect translations, making them difficult to learn through standard best-versus-worst comparisons. We therefore introduce a multi-pair objective that leverages the full candidate list to capture these fine-grained error signals. Experiments on WMT23, WMT24, and FLORES-200 show that $M^2PO$ enables a 9B model to outperform leading open-source baselines and achieve parity with proprietary models such as GPT-4o and Gemini-2.0-Flash, demonstrating strong potential for efficient and high-fidelity LLM-based translation. Our code and dataset will be released.

📄 PDF Abstract BibTeX arXiv:2510.13434

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Preference learning along multiple criteria: A game-theoretic perspective

2021-05-05 · NeurIPS 2020 12 · Kush Bhatia, Ashwin Pananjady, Peter L. Bartlett, Anca D. Dragan 외

The literature on ranking from ordinal data is vast, and there are several ways to aggregate overall preferences from pairwise comparisons between objects. In particular, it is well known that any Nash equilibrium of the…

Autonomous Driving

CAPO: Confidence Aware Preference Optimization Learning for Multilingual Preferences

2025-11-10 · Rhitabrat Pokharel, Yufei Tao, Ameeta Agrawal arxiv

Preference optimization is a critical post-training technique used to align large language models (LLMs) with human preferences, typically by fine-tuning on ranked response pairs. While methods like Direct Preference Opt…

Calibrated Multi-Preference Optimization for Aligning Diffusion Models

2025-02-04 · CVPR 2025 1 · Kyungmin Lee, Xiaohang Li, Qifei Wang, Junfeng He 외

Aligning text-to-image (T2I) diffusion models with preference optimization is valuable for human-annotated datasets, but the heavy cost of manual data collection limits scalability. Using reward models offers an alternat…

Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization

2024-03-31 · Hritik Bansal, Ashima Suvarna, Gantavya Bhatt, Nanyun Peng 외

A common technique for aligning large language models (LLMs) relies on acquiring human preferences by comparing multiple generations conditioned on a fixed context. This method, however, relies solely on pairwise compari…

Preference Ranking Optimization for Human Alignment

2023-06-30 · Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu 외

Large language models (LLMs) often contain misleading content, emphasizing the need to align them with human values to ensure secure AI systems. Reinforcement learning from human feedback (RLHF) has been employed to achi…