HowSumm: A Multi-Document Summarization Dataset Derived from WikiHow Articles
We present HowSumm, a novel large-scale dataset for the task of query-focused multi-document summarization (qMDS), which targets the use-case of generating actionable instructions from a set of sources. This use-case is different from the use-cases covered in existing multi-document summarization (MDS) datasets and is applicable to educational and industrial scenarios. We employed automatic methods, and leveraged statistics from existing human-crafted qMDS datasets, to create HowSumm from wikiHow website articles and the sources they cite. We describe the creation of the dataset and discuss the unique features that distinguish it from other summarization corpora. Automatic and human evaluations of both extractive and abstractive summarization models on the dataset reveal that there is room for improvement.
Code (1)
Tasks
Abstractive Text SummarizationArticlesDocument SummarizationMulti-Document SummarizationSimilar Papers 제목 키워드 기반
MSˆ2: Multi-Document Summarization of Medical Studies
To assess the effectiveness of any medical intervention, researchers must conduct a time-intensive and manual literature review. NLP systems can help to automate or assist in parts of this expensive process. In support o…
Document SummarizationMulti-Document SummarizationMS2: Multi-Document Summarization of Medical Studies
To assess the effectiveness of any medical intervention, researchers must conduct a time-intensive and highly manual literature review. NLP systems can help to automate or assist in parts of this expensive process. In su…
Document SummarizationMulti-Document SummarizationMiRANews: Dataset and Benchmarks for Multi-Resource-Assisted News Summarization
One of the most challenging aspects of current single-document news summarization is that the summary often contains 'extrinsic hallucinations', i.e., facts that are not present in the source document, which are often de…
ArticlesDocument SummarizationMulti-Document SummarizationNews Summarization+1OpenAsp: A Benchmark for Multi-document Open Aspect-based Summarization
The performance of automatic summarization models has improved dramatically in recent years. Yet, there is still a gap in meeting specific information needs of users in real-world scenarios, particularly when a targeted …
Document SummarizationMulti-Document SummarizationNeural Abstractive Summarization with Structural Attention
Attentional, RNN-based encoder-decoder architectures have achieved impressive performance on abstractive summarization of news articles. However, these methods fail to account for long term dependencies within the senten…
Abstractive Text SummarizationArticlesCommunity Question AnsweringDecoder+4