PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation
Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambiguation, terminology control, and stylistic adaptation. Large reasoning models (LRMs) make such decisions explicit through intermediate translation steps, but our analysis across 15 domains and four translation directions shows that this explicit reasoning is double-edged: it improves long-form and high-difficulty translation, yet often drifts in terminology-intensive and stylistically constrained settings. We trace this failure to a credit-assignment bottleneck: existing methods optimize final outputs or coarse trajectories, but cannot identify which translation steps actually help the final translation. To address this, we propose PAMT, a process-aligned training framework that combines cold-start domain-aware Long-CoT supervision with reinforcement learning. PAMT uses sequence-level format and outcome rewards for the final translation, together with a step-level process reward that measures how much each explicit translation step increases the likelihood of the reference translation. Across two backbones, PAMT improves over base models, outperforms MT-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in-domain, OOD, and multilingual settings.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningMachine TranslationSimilar Papers 제목 키워드 기반
PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data
In comparison with person re-identification (ReID), which has been widely studied in the research community, vehicle ReID has received less attention. Vehicle ReID is challenging due to 1) high intra-class variability (c…
AttributeMulti-Task LearningPerson Re-IdentificationPose Estimation+1Propagation with Adaptive Mask then Training for Node Classification on Attributed Networks
Node classification on attributed networks is a semi-supervised task that is crucial for network analysis. By decoupling two critical operations in Graph Convolutional Networks (GCNs), namely feature transformation and n…
AttributeNode ClassificationPrompt-Guided Adaptive Model Transformation for Whole Slide Image Classification
Multiple instance learning (MIL) has emerged as a popular method for classifying histopathology whole slide images (WSIs). Existing approaches typically rely on frozen pre-trained models to extract instance features, neg…
image-classificationImage ClassificationMultiple Instance Learningwhole slide imagesA Parametric Memory Head for Continual Generative Retrieval
Generative information retrieval (GenIR) consolidates retrieval into a single neural model that decodes document identifiers (docids) directly from queries. While this model-as-index paradigm offers architectural simplic…
parameter-efficient fine-tuningInformation RetrievalNatural QuestionsOptimal Subpattern Assignment Metric for Multiple Tracks (OSPAMT Metric)
In this paper, we propose a new metric which measures the distance between two finite sets of tracks (a track is a path of either a real or estimated target). This metric is based on the same principle as the Optimal Sub…