SKT5SciSumm -- Revisiting Extractive-Generative Approach for Multi-Document Scientific Summarization
Summarization for scientific text has shown significant benefits both for the research community and human society. Given the fact that the nature of scientific text is distinctive and the input of the multi-document summarization task is substantially long, the task requires sufficient embedding generation and text truncation without losing important information. To tackle these issues, in this paper, we propose SKT5SciSumm - a hybrid framework for multi-document scientific summarization (MDSS). We leverage the Sentence-Transformer version of Scientific Paper Embeddings using Citation-Informed Transformers (SPECTER) to encode and represent textual sentences, allowing for efficient extractive summarization using k-means clustering. We employ the T5 family of models to generate abstractive summaries using extracted sentences. SKT5SciSumm achieves state-of-the-art performance on the Multi-XScience dataset. Through extensive experiments and evaluation, we showcase the benefits of our model by using less complicated models to achieve remarkable results, thereby highlighting its potential in advancing the field of multi-document summarization for scientific text.
Code (0)
등록된 구현이 없습니다.
Tasks
Document SummarizationExtractive SummarizationMulti-Document SummarizationSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
CIST@CL-SciSumm 2020, LongSumm 2020: Automatic Scientific Document Summarization
Our system participates in two shared tasks, CL-SciSumm 2020 and LongSumm 2020. In the CL-SciSumm shared task, based on our previous work, we apply more machine learning methods on position features and content features …
Abstractive Text SummarizationDocument SummarizationExtractive SummarizationPosition+1Overview and Results: CL-SciSumm Shared Task 2019
The CL-SciSumm Shared Task is the first medium-scale shared task on scientific document summarization in the computational linguistics~(CL) domain. In 2019, it comprised three tasks: (1A) identifying relationships betwee…
Document SummarizationInformation RetrievalRetrievalScientific Document Summarization1A-Team / Martin-Luther-Universität Halle-Wittenberg@CLSciSumm 20
This document demonstrates our groups approach to the CL-SciSumm shared task 2020. There are three tasks in CL-SciSumm 2020. In Task 1a, we apply a Siamese neural network to identify the spans of text in the reference pa…
SciSummPip: An Unsupervised Scientific Paper Summarization Pipeline
The Scholarly Document Processing (SDP) workshop is to encourage more efforts on natural language understanding of scientific task. It contains three shared tasks and we participate in the LongSumm shared task. In this p…
ClusteringGraph Clusteringgraph constructionLanguage Modeling+5Monash-Summ@LongSumm 20 SciSummPip: An Unsupervised Scientific Paper Summarization Pipeline
The Scholarly Document Processing (SDP) workshop is to encourage more efforts on natural language understanding of scientific task. It contains three shared tasks and we participate in the LongSumm shared task. In this p…
Graph Clusteringgraph constructionLanguage ModelingLanguage Modelling+4