paper-with-me

홈 › Papers

IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation

2026-04-06 · Anjali Kantharuban, Aarohi Srivastava, Fahim Faisal, Orevaoghene Ahia, Antonios Anastasopoulos, David Chiang, Yulia Tsvetkov, Graham Neubig arxiv

Existing sentence representations primarily encode what a sentence says, rather than how it is expressed, even though the latter is important for many applications. In contrast, we develop sentence representations that capture style and dialect, decoupled from semantic content. We call this the task of idiolectal representation learning. We introduce IDIOLEX, a framework for training models that combines supervision from a sentence's provenance with linguistic features of a sentence's content, to learn a continuous representation of each sentence's style and dialect. We evaluate the approach on dialects of both Arabic and Spanish. The learned representations capture meaningful variation and transfer across domains for analysis and classification. We further explore the use of these representations as training objectives for stylistically aligning language models. Our results suggest that jointly modeling individual and community-level variation provides a useful perspective for studying idiolect and supports downstream applications requiring sensitivity to stylistic differences, such as developing diverse and accessible LLMs.

📄 PDF Abstract BibTeX arXiv:2604.04704

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Idiosyncratic but not Arbitrary: Learning Idiolects in Online Registers Reveals Distinctive yet Consistent Individual Styles

2021-09-07 · EMNLP 2021 11 · Jian Zhu, David Jurgens

An individual's variation in writing style is often a function of both social and personal attributes. While structured social variation has been extensively studied, e.g., gender based variation, far less is known about…

Chasing the Ghosts of Ibsen: A computational stylistic analysis of drama in translation

2015-01-05 · Gerard Lynch, Carl Vogel

Research into the stylistic properties of translations is an issue which has received some attention in computational stylistics. Previous work by Rybicki (2006) on the distinguishing of character idiolects in the work o…

Translation

StyleDecipher: Robust and Explainable Detection of LLM-Generated Texts with Stylistic Analysis

2025-10-14 · Siyuan Li, Aodu Wulianghai, Xi Lin, Guangyan Li 외 arxiv

With the increasing integration of large language models (LLMs) into open-domain writing, detecting machine-generated text has become a critical task for ensuring content authenticity and trust. Existing approaches rely …

Text Detection

Capturing Style in Author and Document Representation

2024-07-18 · Enzo Terreau, Antoine Gourru, Julien Velcin

A wide range of Deep Natural Language Processing (NLP) models integrates continuous and low dimensional representations of words and documents. Surprisingly, very few models study representation learning for authors. The…

Authorship AttributionRecommendation SystemsRepresentation Learning

Stylistic Attribute Control in Latent Diffusion Models

2026-05-04 · Max Reimann, Benito Buchheim, Jürgen Döllner arxiv

Text-to-image diffusion models have revolutionized image synthesis and editing, but precise control over stylistic attributes remains a challenge, often causing unintended content modifications. We propose an approach fo…

Image Editing