NEWTS: A Corpus for News Topic-Focused Summarization
Text summarization models are approaching human levels of fidelity. Existing benchmarking corpora provide concordant pairs of full and abridged versions of Web, news or, professional content. To date, all summarization datasets operate under a one-size-fits-all paradigm that may not reflect the full range of organic summarization needs. Several recently proposed models (e.g., plug and play language models) have the capacity to condition the generated summaries on a desired range of themes. These capacities remain largely unused and unevaluated as there is no dedicated dataset that would support the task of topic-focused summarization. This paper introduces the first topical summarization corpus NEWTS, based on the well-known CNN/Dailymail dataset, and annotated via online crowd-sourcing. Each source article is paired with two reference summaries, each focusing on a different theme of the source document. We evaluate a representative range of existing techniques and analyze the effectiveness of different prompting methods.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingText SummarizationSimilar Papers 제목 키워드 기반
Controllable Topic-Focused Abstractive Summarization
Controlled abstractive summarization focuses on producing condensed versions of a source article to cover specific aspects by shifting the distribution of generated text towards a desired style, e.g., a set of topics. Su…
Abstractive Text SummarizationTopic-Selective Graph Network for Topic-Focused Summarization
Due to the success of the pre-trained language model (PLM), existing PLM-based summarization models show their powerful generative capability. However, these models are trained on general-purpose summarization datasets, …
ARCLanguage ModelingLanguage ModellingLogit Reweighting for Topic-Focused Summarization
Generating abstractive summaries that adhere to a specific topic remains a significant challenge for language models. While standard approaches, such as fine-tuning, are resource-intensive, simpler methods like prompt en…
Prompt EngineeringPriberam Compressive Summarization Corpus: A New Multi-Document Summarization Corpus for European Portuguese
In this paper, we introduce the Priberam Compressive Summarization Corpus, a new multi-document summarization corpus for European Portuguese. The corpus follows the format of the summarization corpora for English in rece…
Document SummarizationInformation RetrievalMulti-Document SummarizationSentence+2NEWSFARM: the Largest Chinese Corpus for Long News Summarization
Recently, driven by a large number of datasets, the field of natural language processing(NLP) has developed rapidly. However, the lack of large-scale and high-quality Chinese datasets is still a critical bottleneck for f…
News SummarizationSemantic SimilaritySemantic Textual SimilarityText Summarization