paper-with-me

홈 › Papers

Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI

2026-04-23 · Hieu Man, Van-Cuong Pham, Nghia Trung Ngo, Franck Dernoncourt, Thien Huu Nguyen arxiv

Learning robust representations of authorial style is crucial for authorship attribution and AI-generated text detection. However, existing methods often struggle with content-style entanglement, where models learn spurious correlations between authors' writing styles and topics, leading to poor generalization across domains. To address this challenge, we propose Explainable Authorship Variational Autoencoder (EAVAE), a novel framework that explicitly disentangles style from content through architectural separation-by-design. EAVAE first pretrains style encoders using supervised contrastive learning on diverse authorship data, then finetunes with a Variational Autoencoder (VEA) architecture using separate encoders for style and content representations. Disentanglement is enforced through a novel discriminator that not only distinguishes whether pairs of style/content representations belong to the same or different authors/content sources, but also generates natural language explanation for their decision, simultaneously mitigating confounding information and enhancing interpretability. Extensive experiments demonstrate the effectiveness of EAVAE. On authorship attribution, we achieve state-of-the-art performance on various datasets, including Amazon Reviews, PAN21, and HRS. For AI-generated text detection, EAVAE excels in few-shot learning over the M4 dataset. Code and data repositories are available online\footnote{https://github.com/hieum98/avae} \footnote{https://huggingface.co/collections/Hieuman/document-level-authorship-datasets}.

📄 PDF Abstract BibTeX arXiv:2604.21300

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningContrastive LearningFew-Shot LearningText Detection

Similar Papers 제목 키워드 기반

Layered Insights: Generalizable Analysis of Authorial Style by Leveraging All Transformer Layers

2025-03-02 · Milad Alshomary, Nikhil Reddy Varimalla, Vishal Anand, Kathleen McKeown

We propose a new approach for the authorship attribution task that leverages the various linguistic representations learned at different layers of pre-trained transformer-based models. We evaluate our approach on three d…

AllAuthorship Attribution

Sui Generis: Large Language Models for Authorship Attribution and Verification in Latin

2024-10-11 · Gleb Schmidt, Svetlana Gorovaia, Ivan P. Yamshchikov

This paper evaluates the performance of Large Language Models (LLMs) in authorship attribution and authorship verification tasks for Latin texts of the Patristic Era. The study showcases that LLMs can be robust in zero-s…

Authorship AttributionAuthorship VerificationDecision MakingFeature Engineering

Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution

2024-09-11 · Milad Alshomary, Narutatsu Ri, Marianna Apidianaki, Ajay Patel 외

Recent state-of-the-art authorship attribution methods learn authorship representations of texts in a latent, non-interpretable space, hindering their usability in real-world applications. Our work proposes a novel appro…

Authorship Attribution

Improving Explainability of Disentangled Representations using Multipath-Attribution Mappings

2023-06-15 · Lukas Klein, João B. S. Carvalho, Mennatallah El-Assady, Paolo Penna 외

Explainable AI aims to render model behavior understandable by humans, which can be seen as an intermediate step in extracting causal relations from correlative patterns. Due to the high risk of possible fatal decisions …

Relation Extraction

Learning Text Styles: A Study on Transfer, Attribution, and Verification

2025-07-22 · Zhiqiang Hu arxiv

This thesis advances the computational understanding and manipulation of text styles through three interconnected pillars: (1) Text Style Transfer (TST), which alters stylistic properties (e.g., sentiment, formality) whi…

Text Style Transfer