paper-with-me

Papers

Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine Translation

2020-05-01 · ACL 2020 6 · Xabier Soto, Dimitar Shterionov, Alberto Poncelas, Andy Way

Machine translation (MT) has benefited from using synthetic training data originating from translating monolingual corpora, a technique known as backtranslation. Combining backtranslated data from different sources has led to better results than when using such data in isolation. In this work we analyse the impact that data translated with rule-based, phrase-based statistical and neural MT systems has on new MT systems. We use a real-world low-resource use-case (Basque-to-Spanish in the clinical domain) as well as a high-resource language pair (German-to-English) to test different scenarios with backtranslation and employ data selection to optimise the synthetic corpora. We exploit different data selection strategies in order to reduce the amount of data used, while at the same time maintaining high-quality MT systems. We further tune the data selection method by taking into account the quality of the MT systems used for backtranslation and lexical diversity of the resulting corpora. Our experiments show that incorporating backtranslated data from different sources can be beneficial, and that availing of data selection can yield improved performance.

📄 PDF Abstract BibTeX arXiv:2005.00308

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityMachine TranslationTranslation

Similar Papers 제목 키워드 기반

Better Alignment with Instruction Back-and-Forth Translation

2024-08-08 · Thao Nguyen, Jeffrey Li, Sewoong Oh, Ludwig Schmidt 외

We propose a new method, instruction back-and-forth translation, to construct high-quality synthetic data grounded in world knowledge for aligning large language models (LLMs). Given documents from a web corpus, we gener…

DiversityTranslationWorld Knowledge

Defending LLMs against Jailbreaking Attacks via Backtranslation

2024-02-26 · Yihan Wang, Zhouxing Shi, Andrew Bai, Cho-Jui Hsieh

Although many large language models (LLMs) have been trained to refuse harmful requests, they are still vulnerable to jailbreaking attacks which rewrite the original prompt to conceal its harmful intent. In this paper, w…

Language Modelling

Reliable Feature Selection for Adversarially Robust Cyber-Attack Detection

2024-04-05 · João Vitorino, Miguel Silva, Eva Maia, Isabel Praça

The growing cybersecurity threats make it essential to use high-quality data to train Machine Learning (ML) models for network traffic analysis, without noisy or missing data. By selecting the most relevant features for …

Computational EfficiencyCyber Attack Detectionfeature selection

Multilingual Document-Level Translation Enables Zero-Shot Transfer From Sentences to Documents

2021-09-21 · ACL 2022 5 · Biao Zhang, Ankur Bapna, Melvin Johnson, Ali Dabirmoghaddam 외

Document-level neural machine translation (DocNMT) achieves coherent translations by incorporating cross-sentence context. However, for most language pairs there's a shortage of parallel documents, although parallel sent…

Machine TranslationSentenceTransfer LearningTranslation

CUNI Submission for Low-Resource Languages in WMT News 2019

2019-08-01 · WS 2019 8 · Tom Kocmi, Ond{\v{r}}ej Bojar

This paper describes the CUNI submission to the WMT 2019 News Translation Shared Task for the low-resource languages: Gujarati-English and Kazakh-English. We participated in both language pairs in both translation direct…

Transfer LearningTranslation