paper-with-me

홈 › Papers

PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation

2026-08-04 · Yongshi Ye, Biao Fu, Chongxuan Huang, Yidong Chen, Xiaodong Shi arxiv

Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambiguation, terminology control, and stylistic adaptation. Large reasoning models (LRMs) make such decisions explicit through intermediate translation steps, but our analysis across 15 domains and four translation directions shows that this explicit reasoning is double-edged: it improves long-form and high-difficulty translation, yet often drifts in terminology-intensive and stylistically constrained settings. We trace this failure to a credit-assignment bottleneck: existing methods optimize final outputs or coarse trajectories, but cannot identify which translation steps actually help the final translation. To address this, we propose PAMT, a process-aligned training framework that combines cold-start domain-aware Long-CoT supervision with reinforcement learning. PAMT uses sequence-level format and outcome rewards for the final translation, together with a step-level process reward that measures how much each explicit translation step increases the likelihood of the reference translation. Across two backbones, PAMT improves over base models, outperforms MT-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in-domain, OOD, and multilingual settings.

📄 PDF Abstract BibTeX arXiv:2608.03077

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMachine Translation

Similar Papers 제목 키워드 기반

PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data

2020-05-02 · ICCV 2019 10 · Zheng Tang, Milind Naphade, Stan Birchfield, Jonathan Tremblay 외

In comparison with person re-identification (ReID), which has been widely studied in the research community, vehicle ReID has received less attention. Vehicle ReID is challenging due to 1) high intra-class variability (c…

AttributeMulti-Task LearningPerson Re-IdentificationPose Estimation+1

Propagation with Adaptive Mask then Training for Node Classification on Attributed Networks

2022-06-21 · Jinsong Chen, Boyu Li, Qiuting He, Kun He

Node classification on attributed networks is a semi-supervised task that is crucial for network analysis. By decoupling two critical operations in Graph Convolutional Networks (GCNs), namely feature transformation and n…

AttributeNode Classification

Prompt-Guided Adaptive Model Transformation for Whole Slide Image Classification

2024-03-19 · Yi Lin, Zhengjie ZHU, Kwang-Ting Cheng, Hao Chen

Multiple instance learning (MIL) has emerged as a popular method for classifying histopathology whole slide images (WSIs). Existing approaches typically rely on frozen pre-trained models to extract instance features, neg…

image-classificationImage ClassificationMultiple Instance Learningwhole slide images

A Parametric Memory Head for Continual Generative Retrieval

2026-04-25 · Kidist Amde Mekonnen, Yubao Tang, Maarten de Rijke arxiv

Generative information retrieval (GenIR) consolidates retrieval into a single neural model that decodes document identifiers (docids) directly from queries. While this model-as-index paradigm offers architectural simplic…

parameter-efficient fine-tuningInformation RetrievalNatural Questions

Optimal Subpattern Assignment Metric for Multiple Tracks (OSPAMT Metric)

2019-04-16

In this paper, we propose a new metric which measures the distance between two finite sets of tracks (a track is a path of either a real or estimated target). This metric is based on the same principle as the Optimal Sub…