QBSUM: a Large-Scale Query-Based Document Summarization Dataset from Real-world Applications
Query-based document summarization aims to extract or generate a summary of a document which directly answers or is relevant to the search query. It is an important technique that can be beneficial to a variety of applications such as search engines, document-level machine reading comprehension, and chatbots. Currently, datasets designed for query-based summarization are short in numbers and existing datasets are also limited in both scale and quality. Moreover, to the best of our knowledge, there is no publicly available dataset for Chinese query-based document summarization. In this paper, we present QBSUM, a high-quality large-scale dataset consisting of 49,000+ data samples for the task of Chinese query-based document summarization. We also propose multiple unsupervised and supervised solutions to the task and demonstrate their high-speed inference and superior performance via both offline experiments and online A/B tests. The QBSUM dataset is released in order to facilitate future advancement of this research field.
Code (0)
등록된 구현이 없습니다.
Tasks
Document SummarizationMachine Reading ComprehensionReading ComprehensionSimilar Papers 제목 키워드 기반
LMGQS: A Large-scale Dataset for Query-focused Summarization
Query-focused summarization (QFS) aims to extract or generate a summary of an input document that directly answers or is relevant to a given query. The lack of large-scale datasets in the form of documents, queries, and …
DiversityLanguage ModelingLanguage ModellingQuery-focused Summarization+1AQuaMuSe: Automatically Generating Datasets for Query-Based Multi-Document Summarization
Summarization is the task of compressing source document(s) into coherent and succinct passages. This is a valuable tool to present users with concise and accurate sketch of the top ranked documents related to their quer…
Document SummarizationMulti-Document SummarizationQuestion AnsweringText Summarization with Latent Queries
The availability of large-scale datasets has driven the development of neural models that create summaries from single documents, for generic purposes. When using a summarization system, users often have specific intents…
Abstractive Text SummarizationLanguage ModelingLanguage ModellingQuery-focused Summarization+1Beyond Relevant Documents: A Knowledge-Intensive Approach for Query-Focused Summarization using Large Language Models
Query-focused summarization (QFS) is a fundamental task in natural language processing with broad applications, including search engines and report generation. However, traditional approaches assume the availability of r…
Language ModelingLanguage ModellingLarge Language ModelQuery-focused Summarization+1Data Augmentation for Abstractive Query-Focused Multi-Document Summarization
The progress in Query-focused Multi-Document Summarization (QMDS) has been limited by the lack of sufficient largescale high-quality training datasets. We present two QMDS training datasets, which we construct using two …
Data AugmentationDocument SummarizationMulti-Document Summarization