Idiosyncratic but not Arbitrary: Learning Idiolects in Online Registers Reveals Distinctive yet Consistent Individual Styles
An individual's variation in writing style is often a function of both social and personal attributes. While structured social variation has been extensively studied, e.g., gender based variation, far less is known about how to characterize individual styles due to their idiosyncratic nature. We introduce a new approach to studying idiolects through a massive cross-author comparison to identify and encode stylistic features. The neural model achieves strong performance at authorship identification on short texts and through an analogy-based probing task, showing that the learned representations exhibit surprising regularities that encode qualitative and quantitative shifts of idiolectal styles. Through text perturbation, we quantify the relative contributions of different linguistic elements to idiolectal variation. Furthermore, we provide a description of idiolects through measuring inter- and intra-author variation, showing that variation in idiolects is often distinctive yet consistent.
Code (1)
Similar Papers 제목 키워드 기반
Using idiolects and sociolects to improve word prediction
Toward Multilingual Identification of Online Registers
We consider cross- and multilingual text classification approaches to the identification of online registers (genres), i.e. text varieties with specific situational characteristics. Register is the most important predict…
Multilingual text classificationMultilingual Word Embeddingstext-classificationText Classification+1Liquidity Provision with Adverse Selection and Inventory Costs
We study one-shot Nash competition between an arbitrary number of identical dealers that compete for the order flow of a client. The client trades either because of proprietary information, exposure to idiosyncratic risk…
Decomposing Global Bank Network Connectedness: What is Common, Idiosyncratic and When?
We propose a novel approach to estimate high-dimensional global bank network connectedness in both the time and frequency domains. By employing a factor model with sparse VAR idiosyncratic components, we decompose system…
Density EstimationAutomatic register identification for the open web using multilingual deep learning
This article investigates how well deep learning models can identify web registers -- text varieties such as news reports and discussion forums -- across 16 languages. We introduce the Multilingual CORE corpora, which co…
AttributeMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION