paper-with-me

홈 › Papers

Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains

2025-09-20 · Junghwan Kim, Haotian Zhang, David Jurgens arxiv

Authorship representation (AR) learning, which models an author's unique writing style, has demonstrated strong performance in authorship attribution tasks. However, prior research has primarily focused on monolingual settings-mostly in English-leaving the potential benefits of multilingual AR models underexplored. We introduce a novel method for multilingual AR learning that incorporates two key innovations: probabilistic content masking, which encourages the model to focus on stylistically indicative words rather than content-specific words, and language-aware batching, which improves contrastive learning by reducing cross-lingual interference. Our model is trained on over 4.5 million authors across 36 languages and 13 domains. It consistently outperforms monolingual baselines in 21 out of 22 non-English languages, achieving an average Recall@8 improvement of 4.85%, with a maximum gain of 15.91% in a single language. Furthermore, it exhibits stronger cross-lingual and cross-domain generalization compared to a monolingual model trained solely on English. Our analysis confirms the effectiveness of both proposed techniques, highlighting their critical roles in the model's improved performance.

📄 PDF Abstract BibTeX arXiv:2509.16531

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationContrastive Learning

Similar Papers 제목 키워드 기반

Enhancing Representation Generalization in Authorship Identification

2023-09-30 · Haining Wang

Authorship identification ascertains the authorship of texts whose origins remain undisclosed. That authorship identification techniques work as reliably as they do has been attributed to the fact that authorial style is…

Domain Generalization

Authorship Attribution in Multilingual Machine-Generated Texts

2025-08-03 · Lucio La Cava, Dominik Macko, Róbert Móro, Ivan Srba 외 arxiv

As Large Language Models (LLMs) have reached human-like fluency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly difficult. While early efforts in MGT detection h…

Binary Classification

Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI

2026-04-23 · Hieu Man, Van-Cuong Pham, Nghia Trung Ngo, Franck Dernoncourt 외 arxiv

Learning robust representations of authorial style is crucial for authorship attribution and AI-generated text detection. However, existing methods often struggle with content-style entanglement, where models learn spuri…

Representation LearningContrastive LearningFew-Shot LearningText Detection

Small-Scale Cross-Language Authorship Attribution on Social Media Comments

2021-08-01 · MTSummit 2021 8 · Benjamin Murauer, Gunther Specht

Cross-language authorship attribution is the challenging task of classifying documents by bilingual authors where the training documents are written in a different language than the evaluation documents. Traditional solu…

Authorship AttributionLanguage ModellingMachine TranslationTranslation

Trends and Challenges in Authorship Analysis: A Review of ML, DL, and LLM Approaches

2025-05-21 · Nudrat Habib, Tosin Adewumi, Marcus Liwicki, Elisa Barney

Authorship analysis plays an important role in diverse domains, including forensic linguistics, academia, cybersecurity, and digital content authentication. This paper presents a systematic literature review on two key s…

Author AttributionDomain GeneralizationSystematic Literature ReviewText Detection