Investigating and Explaining Feature and Representation Learning in Translationese Classification
Recent work has shown that neural feature- and representation-learning approaches, and specifically the BERT model, demonstrates superior performance over traditional manual feature engineering and an SVM classifier for the task of translationese classification for various source and target languages. However, to date it is unclear whether the performance differences are due to better representations, better classifiers or both. Moreover, it remains unclear whether the features learnt by BERT overlap with commonly used manual features. To answer these, we exchange features between BERT-based and SVM classifiers, and show that, an SVM fed with BERT representations performs at the level of the best BERT classifiers, and BERT learning and using hand-crafted features performs at the level of traditional classifiers using hand-crafted features. Our experiments indicate that our hand-crafted feature set does not provide any additional information that BERT has not learnt already, and is likely to be a subset of features automatically learnt by BERT. Finally, we apply Integrated Gradients to examine token importance for the BERT model, and find that part of its top performance results are due to just topic differences and spurious correlations with translationese.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationFeature EngineeringRepresentation LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Explaining Translationese: why are Neural Classifiers Better and what do they Learn?
Recent work has shown that neural feature- and representation-learning, e.g. BERT, achieves superior performance over traditional manual feature engineering based approaches, with e.g. SVMs, in translationese classificat…
Feature EngineeringRepresentation LearningAn Investigation of Translationese in the Generations of Multilingual Large Language Models
Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) …
Fiction in Russian Translation: A Translationese Study
This paper presents a translationese study based on the parallel data from the Russian National Corpus (RNC). We explored differences between literary texts originally authored in Russian and fiction translated into Russ…
Binary ClassificationLanguage IdentificationTranslationInformation Density and Quality Estimation Features as Translationese Indicators for Human Translation Classification
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
Translationese refers to linguistic properties that usually occur in translated texts. Previous works study translationese by framing it as a binary classification between original texts and translated texts. In this pap…
Binary ClassificationMachine Translation