A Lexical, Syntactic, and Semantic Perspective for Understanding Style in Text
With a growing interest in modeling inherent subjectivity in natural language, we present a linguistically-motivated process to understand and analyze the writing style of individuals from three perspectives: lexical, syntactic, and semantic. We discuss the stylistically expressive elements within each of these levels and use existing methods to quantify the linguistic intuitions related to some of these elements. We show that such a multi-level analysis is useful for developing a well-knit understanding of style - which is independent of the natural language task at hand, and also demonstrate its value in solving three downstream tasks: authors' style analysis, authorship attribution, and emotion prediction. We conduct experiments on a variety of datasets, comprising texts from social networking sites, user reviews, legal documents, literary books, and newswire. The results on the aforementioned tasks and datasets illustrate that such a multi-level understanding of style, which has been largely ignored in recent works, models style-related subjectivity in text and can be leveraged to improve performance on multiple downstream tasks both qualitatively and quantitatively.
Code (0)
등록된 구현이 없습니다.
Tasks
Authorship AttributionSimilar Papers 제목 키워드 기반
Examining Scientific Writing Styles from the Perspective of Linguistic Complexity
Publishing articles in high-impact English journals is difficult for scholars around the world, especially for non-native English-speaking scholars (NNESs), most of whom struggle with proficiency in English. In order to …
ArticlesDiversitySentenceStyle-aware Neural Model with Application in Authorship Attribution
Writing style is a combination of consistent decisions associated with a specific author at different levels of language production, including lexical, syntactic, and structural. In this paper, we introduce a style-aware…
Authorship AttributionLexical Features Are More Vulnerable, Syntactic Features Have More Predictive Power
Understanding the vulnerability of linguistic features extracted from noisy text is important for both developing better health text classification models and for interpreting vulnerabilities of natural language models. …
ClassificationGeneral Classificationtext-classificationText ClassificationEncoding a syntactic dictionary into a super granular unification grammar
We show how to turn a large-scale syntactic dictionary into a dependency-based unification grammar where each piece of lexical information calls a separate rule, yielding a super granular grammar. Subcategorization, rais…
Evaluating language models for the retrieval and categorization of lexical collocations
Lexical collocations are idiosyncratic combinations of two syntactically bound lexical items (e.g., {``}heavy rain{''} or {``}take a step{''}). Understanding their degree of compositionality and idiosyncrasy, as well the…
Retrievalvalid