paper-with-me

홈 › Papers

On Generating Extended Summaries of Long Documents

2020-12-28 · Sajad Sotudeh, Arman Cohan, Nazli Goharian

Prior work in document summarization has mainly focused on generating short summaries of a document. While this type of summary helps get a high-level view of a given document, it is desirable in some cases to know more detailed information about its salient points that can't fit in a short summary. This is typically the case for longer documents such as a research paper, legal document, or a book. In this paper, we present a new method for generating extended summaries of long papers. Our method exploits hierarchical structure of the documents and incorporates it into an extractive summarization model through a multi-task learning approach. We then present our results on three long summarization datasets, arXiv-Long, PubMed-Long, and Longsumm. Our method outperforms or matches the performance of strong baselines. Furthermore, we perform a comprehensive analysis over the generated results, shedding insights on future research for long-form summary generation task. Our analysis shows that our multi-tasking approach can adjust extraction probability distribution to the favor of summary-worthy sentences across diverse sections. Our datasets, and codes are publicly available at https://github.com/Georgetown-IR-Lab/ExtendedSumm

📄 PDF Abstract BibTeX arXiv:2012.14136

Code (1)

Georgetown-IR-Lab/ExtendedSumm 공식 구현 pytorch

Tasks

Document SummarizationExtended SummarizationExtractive SummarizationMulti-Task Learning

Similar Papers 제목 키워드 기반

TSTR: Too Short to Represent, Summarize with Details! Intro-Guided Extended Summary Generation

2022-06-02 · NAACL 2022 7 · Sajad Sotudeh, Nazli Goharian

Many scientific papers such as those in arXiv and PubMed data collections have abstracts with varying lengths of 50-1000 words and average length of approximately 200 words, where longer abstracts typically convey more i…

Extended Summarization

Generating Extended and Multilingual Summaries with Pre-trained Transformers

2022-06-01 · LREC 2022 6 · Rémi Calizzano, Malte Ostendorff, Qian Ruan, Georg Rehm

Almost all summarisation methods and datasets focus on a single language and short summaries. We introduce a new dataset called WikinewsSum for English, German, French, Spanish, Portuguese, Polish, and Italian summarisat…

Articles

CNLP-NITS @ LongSumm 2021: TextRank Variant for Generating Long Summaries

2021-06-01 · NAACL (sdp) 2021 6 · Darsh Kaushik, Abdullah Faiz Ur Rahman Khilji, Utkarsh Sinha, Partha Pakray

The huge influx of published papers in the field of machine learning makes the task of summarization of scholarly documents vital, not just to eliminate the redundancy but also to provide a complete and satisfying crux o…

Extractive Summarization

GUIR @ LongSumm 2020: Learning to Generate Long Summaries from Scientific Documents

2020-11-01 · EMNLP (sdp) 2020 11 · Sajad Sotudeh Gharebagh, Arman Cohan, Nazli Goharian

This paper presents our methods for the LongSumm 2020: Shared Task on Generating Long Summaries for Scientific Documents, where the task is to generatelong summaries given a set of scientific papers provided by the organ…

Leveraging Graph to Improve Abstractive Multi-Document Summarization

2020-05-20 · ACL 2020 6 · Wei Li, Xinyan Xiao, Jiachen Liu, Hua Wu 외

Graphs that capture relations between textual units have great benefits for detecting salient information from multiple documents and generating overall coherent summaries. In this paper, we develop a neural abstractive …

Document SummarizationMulti-Document Summarization