Annotating Characters in Literary Corpora: A Scheme, the CHARLES Tool, and an Annotated Novel
Characters form the focus of various studies of literary works, including social network analysis, archetype induction, and plot comparison. The recent rise in the computational modelling of literary works has produced a proportional rise in the demand for character-annotated literary corpora. However, automatically identifying characters is an open problem and there is low availability of literary texts with manually labelled characters. To address the latter problem, this work presents three contributions: (1) a comprehensive scheme for manually resolving mentions to characters in texts. (2) A novel collaborative annotation tool, CHARLES (CHAracter Resolution Label-Entry System) for character annotation and similiar cross-document tagging tasks. (3) The character annotations resulting from a pilot study on the novel Pride and Prejudice, demonstrating the scheme and tool facilitate the efficient production of high-quality annotations. We expect this work to motivate the further production of annotated literary corpora to help meet the demand of the community.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Modelling Narrative Elements in a Short Story: A Study on Annotation Schemes and Guidelines
Text-processing algorithms that annotate main components of a story-line are presently in great need of corpora and well-agreed annotation schemes. The Text World Theory of cognitive linguistics offers a model that gener…
Annotating Character Relationships in Literary Texts
We present a dataset of manually annotated relationships between characters in literary texts, in order to support the training and evaluation of automatic methods for relation type prediction in this domain (Makazhanov …
Type prediction“Politeness, you simpleton!” retorted [MASK]: Masked prediction of literary characters
What is the best way to learn embeddings for entities, and what can be learned from them? We consider this question for the case of literary characters. We address the highly challenging task of guessing, from a sentence…
Entity EmbeddingsSentenceAn Annotated Corpus of Direct Speech
We propose a scheme for annotating direct speech in literary texts, based on the Text Encoding Initiative (TEI) and the coreference annotation guidelines from the Message Understanding Conference (MUC). The scheme encode…
Finding a Character's Voice: Stylome Classification on Literary Characters
We investigate in this paper the problem of classifying the stylome of characters in a literary work. Previous research in the field of authorship attribution has shown that the writing style of an author can be characte…
Authorship AttributionClassificationGeneral ClassificationText Categorization