paper-with-me

Papers

Quality Estimation for Synthetic Parallel Data Generation

2014-05-01 · LREC 2014 5 · Raphael Rubino, Antonio Toral, Nikola Ljube{\v{s}}i{\'c}, Gema Ram{\'\i}rez-S{\'a}nchez

This paper presents a novel approach for parallel data generation using machine translation and quality estimation. Our study focuses on pivot-based machine translation from English to Croatian through Slovene. We generate an English―Croatian version of the Europarl parallel corpus based on the English―Slovene Europarl corpus and the Apertium rule-based translation system for Slovene―Croatian. These experiments are to be considered as a first step towards the generation of reliable synthetic parallel data for under-resourced languages. We first collect small amounts of aligned parallel data for the Slovene―Croatian language pair in order to build a quality estimation system for sentence-level Translation Edit Rate (TER) estimation. We then infer TER scores on automatically translated Slovene to Croatian sentences and use the best translations to build an English―Croatian statistical MT system. We show significant improvement in terms of automatic metrics obtained on two test sets using our approach compared to a random selection of synthetic parallel data.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

AR: Auto-Repair the Synthetic Data for Neural Machine Translation

2020-04-05 · Shanbo Cheng, Shaohui Kuang, Rongxiang Weng, Heng Yu 외

Compared with only using limited authentic parallel data as training corpus, many studies have proved that incorporating synthetic parallel data, which generated by back translation (BT) or forward translation (FT, or se…

de-enMachine TranslationNMTSentence+1

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

2026-08-19 · Milan Gritta, Patrik Lambert, Jihye Back, Amril Nazir arxiv

The rapid progress in Artificial Intelligence has largely bypassed African languages, creating a digital divide that limits AI adoption on the continent. Recent open-source LLMs systematically underperform on African lan…

Machine Translation

Evaluation of large-scale synthetic data for Grammar Error Correction

2022-10-31 · Vanya Bannihatti Kumar

Grammar Error Correction(GEC) mainly relies on the availability of high quality of large amount of synthetic parallel data of grammatically correct and erroneous sentence pairs. The quality of the synthetic data is evalu…

DiversitySentence

A density ratio framework for evaluating the utility of synthetic data

2024-08-23 · Thom Benjamin Volker, Peter-Paul de Wolf, Erik-Jan van Kesteren

Synthetic data generation is a promising technique to facilitate the use of sensitive data while mitigating the risk of privacy breaches. However, for synthetic data to be useful in downstream analysis tasks, it needs to…

Density Ratio EstimationSynthetic Data Generation

Semi-Synthetic Parallel Data for Translation Quality Estimation: A Case Study of Dataset Building for an Under-Resourced Language Pair

2026-03-12 · Assaf Siani, Anna Kernerman, Ilan Kernerman arxiv

Quality estimation (QE) plays a crucial role in machine translation (MT) workflows, as it serves to evaluate generated outputs that have no reference translations and to determine whether human post-editing or full retra…

Machine Translation