paper-with-me

홈 › Papers

Small-Scale Cross-Language Authorship Attribution on Social Media Comments

2021-08-01 · MTSummit 2021 8 · Benjamin Murauer, Gunther Specht

Cross-language authorship attribution is the challenging task of classifying documents by bilingual authors where the training documents are written in a different language than the evaluation documents. Traditional solutions rely on either translation to enable the use of single-language features, or language-independent feature extraction methods. More recently, transformer-based language models like BERT can also be pre-trained on multiple languages, making them intuitive candidates for cross-language classifiers which have not been used for this task yet. We perform extensive experiments to benchmark the performance of three different approaches to a smallscale cross-language authorship attribution experiment: (1) using language-independent features with traditional classification models, (2) using multilingual pre-trained language models, and (3) using machine translation to allow single-language classification. For the language-independent features, we utilize universal syntactic features like part-of-speech tags and dependency graphs, and multilingual BERT as a pre-trained language model. We use a small-scale social media comments dataset, closely reflecting practical scenarios. We show that applying machine translation drastically increases the performance of almost all approaches, and that the syntactic features in combination with the translation step achieve the best overall classification performance. In particular, we demonstrate that pre-trained language models are outperformed by traditional models in small scale authorship attribution problems for every language combination analyzed in this paper.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Authorship AttributionLanguage ModellingMachine TranslationTranslation

Similar Papers 제목 키워드 기반

I Can Find You in Seconds! Leveraging Large Language Models for Code Authorship Attribution

2025-01-14 · Soohyeon Choi, Yong Kiam Tan, Mark Huasong Meng, Mohamed Ragab 외

Source code authorship attribution is important in software forensics, plagiarism detection, and protecting software patch integrity. Existing techniques often rely on supervised machine learning, which struggles with ge…

Adversarial RobustnessAttributeAuthorship AttributionFew-Shot Learning+1

Authorship Attribution Using Word Network Features

2013-11-12 · Shibamouli Lahiri, Rada Mihalcea

In this paper, we explore a set of novel features for authorship attribution of documents. These features are derived from a word network representation of natural language text. As has been noted in previous studies, na…

Authorship AttributionBIG-bench Machine Learning

DT-grams: Structured Dependency Grammar Stylometry for Cross-Language Authorship Attribution

2021-06-10 · Benjamin Murauer, Günther Specht

Cross-language authorship attribution problems rely on either translation to enable the use of single-language features, or language-independent feature extraction methods. Until recently, the lack of datasets for this p…

Authorship AttributionTranslation

Cross-Language Authorship Attribution

2014-05-01 · LREC 2014 5 · Dasha Bogdanova, Angeliki Lazaridou

This paper presents a novel task of cross-language authorship attribution (CLAA), an extension of authorship attribution task to multilingual settings: given data labelled with authors in language X, the objective is to …

Authorship AttributionInformation RetrievalMachine TranslationText Classification+1

A Supervised Authorship Attribution Framework for Bengali Language

2016-07-13 · Shanta Phani, Shibamouli Lahiri, Arindam Biswas

Authorship Attribution is a long-standing problem in Natural Language Processing. Several statistical and computational methods have been used to find a solution to this problem. In this paper, we have proposed methods t…

Authorship Attribution