Translating away Translationese without Parallel Data
Translated texts exhibit systematic linguistic differences compared to original texts in the same language, and these differences are referred to as translationese. Translationese has effects on various cross-lingual natural language processing tasks, potentially leading to biased results. In this paper, we explore a novel approach to reduce translationese in translated texts: translation-based style transfer. As there are no parallel human-translated and original data in the same language, we use a self-supervised approach that can learn from comparable (rather than parallel) mono-lingual original and translated data. However, even this self-supervised approach requires some parallel data for validation. We show how we can eliminate the need for parallel validation data by combining the self-supervised loss with an unsupervised loss. This unsupervised loss leverages the original language model loss over the style-transferred output and a semantic similarity loss between the input and style-transferred output. We evaluate our approach in terms of original vs. translationese binary classification in addition to measuring content preservation and target-style fluency. The results show that our approach is able to reduce translationese classifier accuracy to a level of a random classifier after style transfer while adequately preserving the content and fluency in the target original style.
Code (0)
등록된 구현이 없습니다.
Tasks
Binary ClassificationLanguage ModellingSemantic SimilaritySemantic Textual SimilarityStyle TransferSimilar Papers 제목 키워드 기반
Translating Translationese: A Two-Step Approach to Unsupervised Machine Translation
Given a rough, word-by-word gloss of a source language sentence, target language natives can uncover the latent, fully-fluent rendering of the translation. In this work we explore this intuition by breaking translation i…
DecoderMachine TranslationSentenceTranslation+2Lexicogrammatic translationese across two targets and competence levels
This research employs genre-comparable data from a number of parallel and comparable corpora to explore the specificity of translations from English into German and Russian produced by students and professional translato…
SpecificityTranslationVocal Bursts Valence PredictionTowards Debiasing Translation Artifacts
Cross-lingual natural language processing relies on translation, either by humans or machines, at different levels, from translating training data to translating test sets. However, compared to original texts in the same…
Natural Language InferenceSentenceTranslationLost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation
Translated texts bear several hallmarks distinct from texts originating in the language. Though individual translated texts are often fluent and preserve meaning, at a large scale, translated texts have statistical tende…
Abstract Meaning RepresentationMachine TranslationParaphrase GenerationTranslationA Parallel Corpus of Translationese
We describe a set of bilingual English--French and English--German parallel corpora in which the direction of translation is accurately and reliably annotated. The corpora are diverse, consisting of parliamentary proceed…
Machine TranslationTranslation