paper-with-me

홈 › Papers

BookSum: A Collection of Datasets for Long-form Narrative Summarization

2021-05-18 · Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev

The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases. While relevant, such datasets will offer limited challenges for future generations of text summarization systems. We address these issues by introducing BookSum, a collection of datasets for long-form narrative summarization. Our dataset covers source documents from the literature domain, such as novels, plays and stories, and includes highly abstractive, human written summaries on three levels of granularity of increasing difficulty: paragraph-, chapter-, and book-level. The domain and structure of our dataset poses a unique set of challenges for summarization systems, which include: processing very long documents, non-trivial causal and temporal dependencies, and rich discourse structures. To facilitate future work, we trained and evaluated multiple extractive and abstractive summarization models as baselines for our dataset.

📄 PDF Abstract BibTeX arXiv:2105.08209

Code (2)

salesforce/booksum 공식 구현
fladhak/creative-summ-data

Tasks

Abstractive Text SummarizationFormLong-Form Narrative SummarizationText Summarization

Similar Papers 제목 키워드 기반

BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization

2022-01-16 · ACL ARR January 2022 1 · Anonymous

The majority of existing text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases. While relevant, such …

Abstractive Text SummarizationFormLong-Form Narrative SummarizationText Summarization

Codified Foreshadowing-Payoff Text Generation

2026-01-11 · Longfei Yun, Kun Zhou, Yupeng Hou, Letian Peng 외 arxiv

Foreshadowing and payoff are ubiquitous narrative devices through which authors introduce commitments early in a story and resolve them through concrete, observable outcomes. However, despite advances in story generation…

Story GenerationText Generation

S^2tory: Story Spine Distillation for Movie Script Summarization

2026-05-05 · Mingzhe Lu, Yanbing Liu, Qihao Wang, Jiarui Zhang 외 arxiv

Movie scripts pose a fundamental challenge for automatic summarization due to their non-linear, cross-cut narrative structure, which makes surface-level saliency methods ineffective at preserving core story progression. …

Domain Generalization

COLING 2022 Shared Task: LED Finteuning and Recursive Summary Generation for Automatic Summarization of Chapters from Novels

2022-10-01 · COLING (CreativeSumm) 2022 10 · Prerna Kashyap

We present the results of the Workshop on Automatic Summarization for Creative Writing 2022 Shared Task on summarization of chapters from novels. In this task, we finetune a pretrained transformer model for long document…

16k

EMNS /Imz/ Corpus: An emotive single-speaker dataset for narrative storytelling in games, television and graphic novels

2023-05-22 · Kari Ali Noriy, Xiaosong Yang, Jian Jun Zhang

The increasing adoption of text-to-speech technologies has led to a growing demand for natural and emotive voices that adapt to a conversation's context and emotional tone. The Emotive Narrative Storytelling (EMNS) corpu…

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech