paper-with-me

Papers

Quasi Error-free Text Classification and Authorship Recognition in a large Corpus of English Literature based on a Novel Feature Set

2020-10-21 · Arthur M. Jacobs, Annette Kinder

The Gutenberg Literary English Corpus (GLEC) provides a rich source of textual data for research in digital humanities, computational linguistics or neurocognitive poetics. However, so far only a small subcorpus, the Gutenberg English Poetry Corpus, has been submitted to quantitative text analyses providing predictions for scientific studies of literature. Here we show that in the entire GLEC quasi error-free text classification and authorship recognition is possible with a method using the same set of five style and five content features, computed via style and sentiment analysis, in both tasks. Our results identify two standard and two novel features (i.e., type-token ratio, frequency, sonority score, surprise) as most diagnostic in these tasks. By providing a simple tool applicable to both short poems and long novels generating quantitative predictions about features that co-determe the cognitive and affective processing of specific text categories or authors, our data pave the way for many future computational and empirical studies of literature or experiments in reading psychology.

📄 PDF Abstract BibTeX arXiv:2010.10801

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticSentiment Analysistext-classificationText Classification

Similar Papers 제목 키워드 기반

Robust Authorship Verification with Transfer Learning

2019-02-27 · Anonymous

We address the problem of open-set authorship verification, a classification task that consists of attributing texts of unknown authorship to a given author when the unknown documents in the test set are excluded from th…

Authorship VerificationGenerative Adversarial NetworkLanguage ModelingLanguage Modelling+2

Authorship Classification in a Resource Constraint Language Using Convolutional Neural Networks

2021-07-09 · IEEE Access 2021 7 · Md. Rajib Hossain, Mohammed Moshiul Hoque, M. Ali Akber Dewan, Nazmul Siddique 외

Authorship classification is a method of automatically determining the appropriate author of an unknown linguistic text. Although research on authorship classification has significantly progressed in high-resource langua…

Classification

Uncertainty Estimation for the Open-Set Text Classification systems

2026-03-17 · Leonid Erlygin, Alexey Zaytsev arxiv

Accurate uncertainty estimation is essential for building robust and trustworthy recognition systems. In this paper, we consider the open-set text classification (OSTC) task - and uncertainty estimation for it. For OSTC …

Intent ClassificationText Classification

Text Classification For Authorship Attribution Analysis

2013-10-18 · M. Sudheep Elayidom, Chinchu Jose, Anitta Puthussery, Neenu K Sasi

Authorship attribution mainly deals with undecided authorship of literary texts. Authorship attribution is useful in resolving issues like uncertain authorship, recognize authorship of unknown texts, spot plagiarism so o…

Authorship AttributionClassificationGeneral ClassificationSentence+2

Domain Specific Author Attribution Based on Feedforward Neural Network Language Models

2016-02-24 · Zhenhao Ge, Yufang Sun

Authorship attribution refers to the task of automatically determining the author based on a given sample of text. It is a problem with a long history and has a wide range of application. Building author profiles using l…

Author AttributionAuthorship AttributionLanguage ModelingLanguage Modelling