paper-with-me

Papers

S2ORC: The Semantic Scholar Open Research Corpus

2019-11-07 · ACL 2020 6 · Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, Dan S. Weld

We introduce S2ORC, a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for 8.1M open access papers. Full text is annotated with automatically-detected inline mentions of citations, figures, and tables, each linked to their corresponding paper objects. In S2ORC, we aggregate papers from hundreds of academic publishers and digital archives into a unified source, and create the largest publicly-available collection of machine-readable academic text to date. We hope this resource will facilitate research and development of tools and tasks for text mining over academic text.

📄 PDF Abstract BibTeX arXiv:1911.02782

Code (2)

allenai/s2-gorc 공식 구현
allenai/s2orc 공식 구현

Tasks

Language Modelling

Similar Papers 제목 키워드 기반

The ACL OCL Corpus: Advancing Open Science in Computational Linguistics

2023-05-24 · Shaurya Rohatgi, Yanxia Qin, Benjamin Aw, Niranjana Unnithan 외

We present ACL OCL, a scholarly corpus derived from the ACL Anthology to assist Open scientific research in the Computational Linguistics domain. Integrating and enhancing the previous versions of the ACL Anthology, the …

ChunkingText Generation

Finding Pragmatic Differences Between Disciplines

2023-09-30 · NAACL (sdp) 2021 6 · Lee Kezar, Jay Pujara

Scholarly documents have a great degree of variation, both in terms of content (semantics) and structure (pragmatics). Prior work in scholarly document understanding emphasizes semantics through document summarization an…

DiversityDocument Summarizationdocument understandingLanguage Modeling+2

Change Summarization of Diachronic Scholarly Paper Collections by Semantic Evolution Analysis

2021-12-07 · Naman Paharia, Muhammad Syafiq Mohd Pozi, Adam Jatowt

The amount of scholarly data has been increasing dramatically over the last years. For newcomers to a particular science domain (e.g., IR, physics, NLP) it is often difficult to spot larger trends and to position the lat…

Articles

SChuBERT: Scholarly Document Chunks with BERT-encoding boost Citation Count Prediction

2020-12-21 · Thomas van Dongen, Gideon Maillette de Buy Wenniger, Lambert Schomaker

Predicting the number of citations of scholarly documents is an upcoming task in scholarly document processing. Besides the intrinsic merit of this information, it also has a wider use as an imperfect proxy for quality w…

Citation PredictionPrediction

SChuBERT: Scholarly Document Chunks with BERT-encoding boost Citation Count Prediction.

2020-11-01 · EMNLP (sdp) 2020 11 · Thomas van Dongen, Gideon Maillette de Buy Wenniger, Lambert Schomaker

Predicting the number of citations of scholarly documents is an upcoming task in scholarly document processing. Besides the intrinsic merit of this information, it also has a wider use as an imperfect proxy for quality w…

Citation PredictionPrediction