paper-with-me

홈 › Papers

Can Authorship Representation Learning Capture Stylistic Features?

2023-08-22 · Andrew Wang, Cristina Aggazzotti, Rebecca Kotula, Rafael Rivera Soto, Marcus Bishop, Nicholas Andrews

Automatically disentangling an author's style from the content of their writing is a longstanding and possibly insurmountable problem in computational linguistics. At the same time, the availability of large text corpora furnished with author labels has recently enabled learning authorship representations in a purely data-driven manner for authorship attribution, a task that ostensibly depends to a greater extent on encoding writing style than encoding content. However, success on this surrogate task does not ensure that such representations capture writing style since authorship could also be correlated with other latent variables, such as topic. In an effort to better understand the nature of the information these representations convey, and specifically to validate the hypothesis that they chiefly encode writing style, we systematically probe these representations through a series of targeted experiments. The results of these experiments suggest that representations learned for the surrogate authorship prediction task are indeed sensitive to writing style. As a consequence, authorship representations may be expected to be robust to certain kinds of data shift, such as topic drift over time. Additionally, our findings may open the door to downstream applications that require stylistic representations, such as style transfer.

📄 PDF Abstract BibTeX arXiv:2308.11490

Code (1)

llnl/luar 공식 구현 pytorch

Tasks

Authorship AttributionRepresentation LearningStyle Transfer

Similar Papers 제목 키워드 기반

Enhancing Representation Generalization in Authorship Identification

2023-09-30 · Haining Wang

Authorship identification ascertains the authorship of texts whose origins remain undisclosed. That authorship identification techniques work as reliably as they do has been attributed to the fact that authorial style is…

Domain Generalization

Capturing Style in Author and Document Representation

2024-07-18 · Enzo Terreau, Antoine Gourru, Julien Velcin

A wide range of Deep Natural Language Processing (NLP) models integrates continuous and low dimensional representations of words and documents. Surprisingly, very few models study representation learning for authors. The…

Authorship AttributionRecommendation SystemsRepresentation Learning

Learning Text Styles: A Study on Transfer, Attribution, and Verification

2025-07-22 · Zhiqiang Hu arxiv

This thesis advances the computational understanding and manipulation of text styles through three interconnected pillars: (1) Text Style Transfer (TST), which alters stylistic properties (e.g., sentiment, formality) whi…

Text Style Transfer

Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings

2026-05-11 · Benjamin Icard, Lila Sainero, Alice Breton, Evangelia Zve 외 arxiv

Large language models (LLMs) can convincingly imitate human writing styles, yet it remains unclear how much stylistic information is encoded in embeddings from any language model and retained after LLM rewriting. We inve…

Distinguishing Fictional Voices: a Study of Authorship Verification Models for Quotation Attribution

2024-01-30 · Gaspard Michel, Elena V. Epure, Romain Hennequin, Christophe Cerisara

Recent approaches to automatically detect the speaker of an utterance of direct speech often disregard general information about characters in favor of local information found in the context, such as surrounding mentions…

Authorship Verification