Unsupervised Multi-document Summarization with Holistic Inference
Multi-document summarization aims to obtain core information from a collection of documents written on the same topic. This paper proposes a new holistic framework for unsupervised multi-document extractive summarization. Our method incorporates the holistic beam search inference method associated with the holistic measurements, named Subset Representative Index (SRI). SRI balances the importance and diversity of a subset of sentences from the source documents and can be calculated in unsupervised and adaptive manners. To demonstrate the effectiveness of our method, we conduct extensive experiments on both small and large-scale multi-document summarization datasets under both unsupervised and adaptive settings. The proposed method outperforms strong baselines by a significant margin, as indicated by the resulting ROUGE scores and diversity measures. Our findings also suggest that diversity is essential for improving multi-document summary performance.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityDocument SummarizationExtractive SummarizationMulti-Document SummarizationSimilar Papers 제목 키워드 기반
An Unsupervised Masking Objective for Abstractive Multi-Document News Summarization
We show that a simple unsupervised masking objective can approach near supervised performance on abstractive multi-document news summarization. Our method trains a state-of-the-art neural summarization model to predict t…
Extractive SummarizationNews SummarizationQBSUM: a Large-Scale Query-Based Document Summarization Dataset from Real-world Applications
Query-based document summarization aims to extract or generate a summary of a document which directly answers or is relevant to the search query. It is an important technique that can be beneficial to a variety of applic…
Document SummarizationMachine Reading ComprehensionReading ComprehensionGLIMMER: Incorporating Graph and Lexical Features in Unsupervised Multi-Document Summarization
Pre-trained language models are increasingly being used in multi-document summarization tasks. However, these models need large-scale corpora for pre-training and are domain-dependent. Other non-neural unsupervised summa…
Document SummarizationInformativenessMulti-Document SummarizationSentenceUnsupervised Multi-Granularity Summarization
Text summarization is a user-preference based task, i.e., for one document, users often have different priorities for summary. As a key aspect of customization in summarization, granularity is used to measure the semanti…
Abstractive Text SummarizationText SummarizationSummPip: Unsupervised Multi-Document Summarization with Sentence Graph Compression
Obtaining training data for multi-document summarization (MDS) is time consuming and resource-intensive, so recent neural models can only be trained for limited domains. In this paper, we propose SummPip: an unsupervised…
ClusteringDocument SummarizationMulti-Document SummarizationSentence