ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks
Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article's impacts on research community. This paper provides novel solutions to these two challenges. We 1) develop and release the first large-scale manually-annotated corpus for scientific papers (on computational linguistics) by enabling faster annotation, and 2) propose summarization methods that integrate the authors' original highlights (abstract) and the article's actual impacts on the community (citations), to create comprehensive, hybrid summaries. We conduct experiments to demonstrate the efficacy of our corpus in training data-driven models for scientific paper summarization and the advantage of our hybrid summaries over abstracts and traditional citation-based summaries. Our large annotated corpus and hybrid methods provide a new framework for scientific paper summarization research.
Code (1)
Tasks
Scientific Document SummarizationText SummarizationSimilar Papers 제목 키워드 기반
Overview and Results: CL-SciSumm Shared Task 2019
The CL-SciSumm Shared Task is the first medium-scale shared task on scientific document summarization in the computational linguistics~(CL) domain. In 2019, it comprised three tasks: (1A) identifying relationships betwee…
Document SummarizationInformation RetrievalRetrievalScientific Document SummarizationWhite Paper: Challenges and Considerations for the Creation of a Large Labelled Repository of Online Videos with Questionable Content
This white paper presents a summary of the discussions regarding critical considerations to develop an extensive repository of online videos annotated with labels indicating questionable content. The main discussion poin…
A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries
Interdisciplinary research is accelerating, yet scientific papers remain difficult to understand outside their home fields. We study large language model (LLM)-based simplification of scientific texts and present a human…
Pro-TEXT: an Annotated Corpus of Keystroke Logs
Pro-TEXT is a corpus of keystroke logs written in French. Keystroke logs are recordings of the writing process executed through a keyboard, which keep track of all actions taken by the writer (character additions, deleti…
Ubuntu-fr: A Large and Open Corpus for Multi-modal Analysis of Online Written Conversations
We present a large, free, French corpus of online written conversations extracted from the Ubuntu platform{'}s forums, mailing lists and IRC channels. The corpus is meant to support multi-modality and diachronic studies …