paper-with-me

Papers

Selecting Artificially-Generated Sentences for Fine-Tuning Neural Machine Translation

2019-09-26 · WS 2019 10 · Alberto Poncelas, Andy Way

Neural Machine Translation (NMT) models tend to achieve best performance when larger sets of parallel sentences are provided for training. For this reason, augmenting the training set with artificially-generated sentence pairs can boost performance. Nonetheless, the performance can also be improved with a small number of sentences if they are in the same domain as the test set. Accordingly, we want to explore the use of artificially-generated sentences along with data-selection algorithms to improve German-to-English NMT models trained solely with authentic data. In this work, we show how artificially-generated sentences can be more beneficial than authentic pairs, and demonstrate their advantages when used in combination with data-selection algorithms.

📄 PDF Abstract BibTeX arXiv:1909.12016

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTSentenceTranslation

Similar Papers 제목 키워드 기반

Not Enough Data? Deep Learning to the Rescue!

2019-11-08 · Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 외

Based on recent advances in natural language modeling and those in text generation capabilities, we propose a novel data augmentation method for text classification tasks. We use a powerful pre-trained neural network mod…

Data AugmentationDeep LearningGeneral ClassificationLAMBADA+5

Neural Grammatical Error Correction for Romanian

2026-04-26 · Teodor-Mihai Cotet, Stefan Ruseti, Mihai Dascalu arxiv

Resources for Grammatical Error Correction (GEC) in non-English languages are scarce, while available spellcheckers in these languages are mostly limited to simple corrections and rules. In this paper we introduce a firs…

Grammatical Error Correction

BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification

2025-11-03 · Ayesha Afroza Mohsin, Mashrur Ahsan, Nafisa Maliyat, Shanta Maria 외 arxiv

Toxic language in Bengali remains prevalent, especially in online environments, with few effective precautions against it. Although text detoxification has seen progress in high-resource languages, Bengali remains undere…

Artificial Error Generation with Machine Translation and Syntactic Patterns

2017-07-17 · WS 2017 9 · Marek Rei, Mariano Felice, Zheng Yuan, Ted Briscoe

Shortage of available training data is holding back progress in the area of automated error detection. This paper investigates two alternative methods for artificially generating writing errors, in order to create additi…

Grammatical Error DetectionMachine TranslationTranslation

SemEval-2022 Task 3: PreTENS-Evaluating Neural Networks on Presuppositional Semantic Knowledge

2022-07-01 · SemEval (NAACL) 2022 7 · Roberto Zamparelli, Shammur Chowdhury, Dominique Brunato, Cristiano Chesi 외

We report the results of the SemEval 2022 Task 3, PreTENS, on evaluation the acceptability of simple sentences containing constructions whose two arguments are presupposed to be or not to be in an ordered taxonomic relat…

Data Augmentation