Automatic Alignment and Annotation Projection for Literary Texts
This paper presents a modular NLP pipeline for the creation of a parallel literature corpus, followed by annotation transfer from the source to the target language. The test case we use to evaluate our pipeline is the automatic transfer of quote and speaker mention annotations from English to German. We evaluate the different components of the pipeline and discuss challenges specific to literary texts. Our experiments show that after applying a reasonable amount of semi-automatic postprocessing we can obtain high-quality aligned and annotated resources for a new language.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Towards annotation of text worlds in a literary work
Literary texts are usually rich in meanings and their interpretation complicates corpus studies and automatic processing. There have been several attempts to create collections of literary texts with annotation of litera…
TAGAnnotating Characters in Literary Corpora: A Scheme, the CHARLES Tool, and an Annotated Novel
Characters form the focus of various studies of literary works, including social network analysis, archetype induction, and plot comparison. The recent rise in the computational modelling of literary works has produced a…
Towards Coreference for Literary Text: Analyzing Domain-Specific Phenomena
Coreference resolution is the task of grouping together references to the same discourse entity. Resolving coreference in literary texts could benefit a number of Digital Humanities (DH) tasks, such as analyzing the depi…
coreference-resolutionCoreference ResolutionNatural Language InferenceSentiment AnalysisAutomatic Translation Alignment Pipeline for Multilingual Digital Editions of Literary Works
This paper investigates the application of translation alignment algorithms in the creation of a Multilingual Digital Edition (MDE) of Alessandro Manzoni's Italian novel "I promessi sposi" ("The Betrothed"), with transla…
TranslationAnnotating Character Relationships in Literary Texts
We present a dataset of manually annotated relationships between characters in literary texts, in order to support the training and evaluation of automatic methods for relation type prediction in this domain (Makazhanov …
Type prediction