Unsupervised Multi-document Summarization for News Corpus with Key Synonyms and Contextual Embeddings
Information overload has been one of the challenges regarding information from the Internet. It is not a matter of information access, instead, the focus had shifted towards the quality of the retrieved data. Particularly in the news domain, multiple outlets report on the same news events but may differ in details. This work considers that different news outlets are more likely to differ in their writing styles and the choice of words, and proposes a method to extract sentences based on their key information by focusing on the shared synonyms in each sentence. Our method also attempts to reduce redundancy through hierarchical clustering and arrange selected sentences on the proposed orderBERT. The results show that the proposed unsupervised framework successfully improves the coverage, coherence, and, meanwhile, reduces the redundancy for a generated summary. Moreover, due to the process of obtaining the dataset, we also propose a data refinement method to alleviate the problems of undesirable texts, which result from the process of automatic scraping.
Code (0)
등록된 구현이 없습니다.
Tasks
Document SummarizationMulti-Document SummarizationSentenceSimilar Papers 제목 키워드 기반
An Unsupervised Masking Objective for Abstractive Multi-Document News Summarization
We show that a simple unsupervised masking objective can approach near supervised performance on abstractive multi-document news summarization. Our method trains a state-of-the-art neural summarization model to predict t…
Extractive SummarizationNews SummarizationThe Next Step for Multi-Document Summarization: A Heterogeneous Multi-Genre Corpus Built with a Novel Construction Approach
Research in multi-document summarization has focused on newswire corpora since the early beginnings. However, the newswire genre provides genre-specific features such as sentence position which are easy to exploit in sum…
Document SummarizationMulti-Document SummarizationSentencePre-training Meets Clustering: A Hybrid Extractive Multi-document Summarization Model
In this era where a large amount of information has flooded the Internet, manual extraction and consumption of relevant information is very difficult and time-consuming. Therefore, an automated document summarization too…
ClusteringDocument Summarizationdocument understandingExtractive Text Summarization+4Priberam Compressive Summarization Corpus: A New Multi-Document Summarization Corpus for European Portuguese
In this paper, we introduce the Priberam Compressive Summarization Corpus, a new multi-document summarization corpus for European Portuguese. The corpus follows the format of the summarization corpora for English in rece…
Document SummarizationInformation RetrievalMulti-Document SummarizationSentence+2NewsQs: Multi-Source Question Generation for the Inquiring Mind
We present NewsQs (news-cues), a dataset that provides question-answer pairs for multiple news documents. To create NewsQs, we augment a traditional multi-document summarization dataset with questions automatically gener…
ArticlesDocument SummarizationMulti-Document SummarizationQNLI+2