paper-with-me

Papers

Summarization Beyond News: The Automatically Acquired Fandom Corpora

2020-05-01 · LREC 2020 5 · Benjamin H{\"a}ttasch, Nadja Geisler, Christian M. Meyer, Carsten Binnig

Large state-of-the-art corpora for training neural networks to create abstractive summaries are mostly limited to the news genre, as it is expensive to acquire human-written summaries for other types of text at a large scale. In this paper, we present a novel automatic corpus construction approach to tackle this issue as well as three new large open-licensed summarization corpora based on our approach that can be used for training abstractive summarization models. Our constructed corpora contain fictional narratives, descriptive texts, and summaries about movies, television, and book series from different domains. All sources use a creative commons (CC) license, hence we can provide the corpora for download. In addition, we also provide a ready-to-use framework that implements our automatic construction approach to create custom corpora with desired parameters like the length of the target summary and the number of source documents from which to create the summary. The main idea behind our automatic construction approach is to use existing large text collections (e.g., thematic wikis) and automatically classify whether the texts can be used as (query-focused) multi-document summaries and align them with potential source texts. As a final contribution, we show the usefulness of our automatic construction approach by running state-of-the-art summarizers on the corpora and through a manual evaluation with human annotators.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Abstractive Text SummarizationDescriptive

Similar Papers 제목 키워드 기반

Shaping Political Discourse using multi-source News Summarization

2023-12-18 · Charles Rajan, Nishit Asnani, Shreya Singh

Multi-document summarization is the process of automatically generating a concise summary of multiple documents related to the same topic. This summary can help users quickly understand the key information from a large c…

Document SummarizationMulti-Document SummarizationNews Summarization

Overview of the VLSP 2022 -- Abmusu Shared Task: A Data Challenge for Vietnamese Abstractive Multi-document Summarization

2023-11-27 · Mai-Vu Tran, Hoang-Quynh Le, Duy-Cat Can, Quoc-An Nguyen

This paper reports the overview of the VLSP 2022 - Vietnamese abstractive multi-document summarization (Abmusu) shared task for Vietnamese News. This task is hosted at the 9$^{th}$ annual workshop on Vietnamese Language …

Document SummarizationMulti-Document SummarizationNews Summarization

NewsQs: Multi-Source Question Generation for the Inquiring Mind

2024-02-28 · Alyssa Hwang, Kalpit Dixit, Miguel Ballesteros, Yassine Benajiba 외

We present NewsQs (news-cues), a dataset that provides question-answer pairs for multiple news documents. To create NewsQs, we augment a traditional multi-document summarization dataset with questions automatically gener…

ArticlesDocument SummarizationMulti-Document SummarizationQNLI+2

Automatically Discarding Straplines to Improve Data Quality for Abstractive News Summarization

2022-05-01 · nlppower (ACL) 2022 5 · Amr Keleg, Matthias Lindemann, Danyang Liu, Wanqiu Long 외

Recent improvements in automatic news summarization fundamentally rely on large corpora of news articles and their summaries. These corpora are often constructed by scraping news websites, which results in including not …

ArticlesNews Summarization

Towards Automatic Construction of News Overview Articles by News Synthesis

2017-09-01 · EMNLP 2017 9 · Jianmin Zhang, Xiaojun Wan

In this paper we investigate a new task of automatically constructing an overview article from a given set of news articles about a news event. We propose a news synthesis approach to address this task based on passage s…

ArticlesDocument SummarizationMulti-Document Summarization