paper-with-me

Papers

ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks

2019-09-04 · Michihiro Yasunaga, Jungo Kasai, Rui Zhang, Alexander R. Fabbri, Irene Li, Dan Friedman, Dragomir R. Radev

Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article's impacts on research community. This paper provides novel solutions to these two challenges. We 1) develop and release the first large-scale manually-annotated corpus for scientific papers (on computational linguistics) by enabling faster annotation, and 2) propose summarization methods that integrate the authors' original highlights (abstract) and the article's actual impacts on the community (citations), to create comprehensive, hybrid summaries. We conduct experiments to demonstrate the efficacy of our corpus in training data-driven models for scientific paper summarization and the advantage of our hybrid summaries over abstracts and traditional citation-based summaries. Our large annotated corpus and hybrid methods provide a new framework for scientific paper summarization research.

📄 PDF Abstract BibTeX arXiv:1909.01716

Code (1)

WING-NUS/scisumm-corpus 공식 구현

Tasks

Scientific Document SummarizationText Summarization

Similar Papers 제목 키워드 기반

Overview and Results: CL-SciSumm Shared Task 2019

2019-07-23 · Muthu Kumar Chandrasekaran, Michihiro Yasunaga, Dragomir Radev, Dayne Freitag 외

The CL-SciSumm Shared Task is the first medium-scale shared task on scientific document summarization in the computational linguistics~(CL) domain. In 2019, it comprised three tasks: (1A) identifying relationships betwee…

Document SummarizationInformation RetrievalRetrievalScientific Document Summarization

White Paper: Challenges and Considerations for the Creation of a Large Labelled Repository of Online Videos with Questionable Content

2021-01-25 · Thamar Solorio, Mahsa Shafaei, Christos Smailis, Mona Diab 외

This white paper presents a summary of the discussions regarding critical considerations to develop an extensive repository of online videos annotated with labels indicating questionable content. The main discussion poin…

A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries

2026-07-28 · Kyuri Im, Michael Färber arxiv

Interdisciplinary research is accelerating, yet scientific papers remain difficult to understand outside their home fields. We study large language model (LLM)-based simplification of scientific texts and present a human…

Pro-TEXT: an Annotated Corpus of Keystroke Logs

2022-06-01 · LREC 2022 6 · Aleksandra Miletic, Christophe Benzitoun, Georgeta Cislaru, Santiago Herrera-Yanez

Pro-TEXT is a corpus of keystroke logs written in French. Keystroke logs are recordings of the writing process executed through a keyboard, which keep track of all actions taken by the writer (character additions, deleti…

Ubuntu-fr: A Large and Open Corpus for Multi-modal Analysis of Online Written Conversations

2016-05-01 · LREC 2016 5 · Hern, Nicolas ez, Soufian Salim, Elizaveta Loginova Clouet

We present a large, free, French corpus of online written conversations extracted from the Ubuntu platform{'}s forums, mailing lists and IRC channels. The corpus is meant to support multi-modality and diachronic studies …