paper-with-me

홈 › Papers

Investigating and Explaining Feature and Representation Learning in Translationese Classification

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Recent work has shown that neural feature- and representation-learning approaches, and specifically the BERT model, demonstrates superior performance over traditional manual feature engineering and an SVM classifier for the task of translationese classification for various source and target languages. However, to date it is unclear whether the performance differences are due to better representations, better classifiers or both. Moreover, it remains unclear whether the features learnt by BERT overlap with commonly used manual features. To answer these, we exchange features between BERT-based and SVM classifiers, and show that, an SVM fed with BERT representations performs at the level of the best BERT classifiers, and BERT learning and using hand-crafted features performs at the level of traditional classifiers using hand-crafted features. Our experiments indicate that our hand-crafted feature set does not provide any additional information that BERT has not learnt already, and is likely to be a subset of features automatically learnt by BERT. Finally, we apply Integrated Gradients to examine token importance for the BERT model, and find that part of its top performance results are due to just topic differences and spurious correlations with translationese.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationFeature EngineeringRepresentation Learning

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Adam 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

Explaining Translationese: why are Neural Classifiers Better and what do they Learn?

2022-10-24 · Kwabena Amponsah-Kaakyire, Daria Pylypenko, Josef van Genabith, Cristina España-Bonet

Recent work has shown that neural feature- and representation-learning, e.g. BERT, achieves superior performance over traditional manual feature engineering based approaches, with e.g. SVMs, in translationese classificat…

Feature EngineeringRepresentation Learning

An Investigation of Translationese in the Generations of Multilingual Large Language Models

2026-08-18 · Maria Valentini, Téa Wright, Julisa Granados, Eliana Colunga 외 arxiv

Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) …

Fiction in Russian Translation: A Translationese Study

2021-09-01 · RANLP 2021 9 · Maria Kunilovskaya, Ekaterina Lapshinova-Koltunski, Ruslan Mitkov

This paper presents a translationese study based on the parallel data from the Russian National Corpus (RNC). We explored differences between literary texts originally authored in Russian and fiction translated into Russ…

Binary ClassificationLanguage IdentificationTranslation

Information Density and Quality Estimation Features as Translationese Indicators for Human Translation Classification

2016-06-01 · NAACL 2016 6 · Raphael Rubino, Ekaterina Lapshinova-Koltunski, Josef van Genabith
General ClassificationMachine TranslationTranslation

Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese

2025-07-16 · Yikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang 외 arxiv

Translationese refers to linguistic properties that usually occur in translated texts. Previous works study translationese by framing it as a binary classification between original texts and translated texts. In this pap…

Binary ClassificationMachine Translation