paper-with-me

Papers

LangMark: A Multilingual Dataset for Automatic Post-Editing

2025-11-21 · Diego Velazquez, Mikaela Grace, Konstantinos Karageorgos, Lawrence Carin, Aaron Schliem, Dimitrios Zaikis, Roger Wechsler arxiv

Automatic post-editing (APE) aims to correct errors in machine-translated text, enhancing translation quality, while reducing the need for human intervention. Despite advances in neural machine translation (NMT), the development of effective APE systems has been hindered by the lack of large-scale multilingual datasets specifically tailored to NMT outputs. To address this gap, we present and release LangMark, a new human-annotated multilingual APE dataset for English translation to seven languages: Brazilian Portuguese, French, German, Italian, Japanese, Russian, and Spanish. The dataset has 206,983 triplets, with each triplet consisting of a source segment, its NMT output, and a human post-edited translation. Annotated by expert human linguists, our dataset offers both linguistic diversity and scale. Leveraging this dataset, we empirically show that Large Language Models (LLMs) with few-shot prompting can effectively perform APE, improving upon leading commercial and even proprietary machine translation systems. We believe that this new resource will facilitate the future development and evaluation of APE systems.

📄 PDF Abstract BibTeX arXiv:2511.17153

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

MLQE-PE: A Multilingual Quality Estimation and Post-Editing Dataset

2020-10-09 · LREC 2022 6 · Marina Fomicheva, Shuo Sun, Erick Fonseca, Chrysoula Zerva 외

We present MLQE-PE, a new dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE). The dataset contains eleven language pairs, with human labels for up to 10,000 translations per lan…

ArticlesAutomatic Post-EditingMachine TranslationSentence+1

X-RiSAWOZ: High-Quality End-to-End Multilingual Dialogue Datasets and Few-shot Agents

2023-06-30 · Mehrad Moradshahi, Tianhao Shen, Kalika Bali, Monojit Choudhury 외

Task-oriented dialogue research has mainly focused on a few popular languages like English and Chinese, due to the high dataset creation cost for a new language. To reduce the cost, we apply manual editing to automatical…

Entity AlignmentMachine TranslationTranslation

Together We Can: Multilingual Automatic Post-Editing for Low-Resource Languages

2024-10-23 · Sourabh Deoghare, Diptesh Kanojia, Pushpak Bhattacharyya

This exploratory study investigates the potential of multilingual Automatic Post-Editing (APE) systems to enhance the quality of machine translations for low-resource Indo-Aryan languages. Focusing on two closely related…

Automatic Post-EditingData AugmentationDomain AdaptationMulti-Task Learning

AlphaMWE: Construction of Multilingual Parallel Corpora with MWE Annotations

2020-11-07 · COLING (MWE) 2020 12 · Lifeng Han, Gareth Jones, Alan Smeaton

In this work, we present the construction of multilingual parallel corpora with annotation of multiword expressions (MWEs). MWEs include verbal MWEs (vMWEs) defined in the PARSEME shared task that have a verb as the head…

Machine TranslationSentenceTranslation

DivEMT: Neural Machine Translation Post-Editing Effort Across Typologically Diverse Languages

2022-05-24 · Gabriele Sarti, Arianna Bisazza, Ana Guerberof Arenas, Antonio Toral

We introduce DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages. Using a strictly controlled setup, 18 professional translator…

Machine TranslationNMTTranslation